The Delhi High Court on Friday refused to restrain OpenAI from using news agency ANI’s copyrighted content, holding at the interim stage that storing publicly available literary work to train large language models (LLMs) could qualify as “fair dealing” for private use and research under Indian copyright law.
Justice Amit Bansal dismissed ANI Media’s application for an interim injunction against the ChatGPT maker, but clarified that the findings were prima facie and would not affect the final outcome of the copyright infringement suit. Citing ANI’s October 2024 offer to license its content to OpenAI for $7.5 million, the court said this indicated that its claim was quantifiable, and that monetary compensation could be awarded if the news agency succeeded in the suit.
ANI has accused OpenAI of copyright infringement by copying and storing its content to train LLMs underlying ChatGPT and reproducing its articles in responses generated for users.
The court held that ANI failed to establish a prima facie case on either claim. It found no evidence at this stage that ChatGPT memorised and regurgitated ANI’s articles or that its responses substantially reproduced the news agency’s original expression.
While storing a copyrighted literary work ordinarily amounts to reproduction under the Copyright Act, the court held that it should be read with Section 52 of the Act.
Under Section 52(1)(a), a fair dealing with any work, except for computer programs, for private or personal use, including research, does not constitute copyright infringement. The court held that “private use” could extend to companies as well.
The training data was used within a closed system, was not available to the public and was accessible only to the models, the judgment said. The process of analysing data and converting it into machine-readable inputs to improve a model could also be considered research aimed at generating knowledge and advancing AI systems.
The court used the “doctrine of updating construction” to interpret the Copyright Act in light of technological developments that lawmakers could not have anticipated when the provision was last amended in 2012.
“Research/learning is no longer confined to humans. It is now being done through Artificial Intelligence,” the court said, adding that such research was undertaken at the behest of and for the benefit of humans.
The fact that OpenAI was a commercial entity did not by itself take its activity outside the protection of Section 52, the court held.
Lack of proof:
On ANI’s allegation that ChatGPT reproduced its news reports, the court said copyright in news extended to its original form and manner of expression, not to the underlying facts. The threshold for proving substantial similarity would be higher in the case of news because its primary purpose was to report events that had occurred, it said.
The court said the articles cited to demonstrate copying were published in August and September 2024. The training cut-off dates for the models concerned were April 2022 for GPT-4 and April 2024 for GPT-4o.
The articles could consequently not have formed part of the models’ training data, ruling out ANI’s claim that the responses resulted from memorisation, it said.
Comparing the ChatGPT responses with ANI’s articles, the court said they conveyed similar facts but used different expressions and, in some instances, added their own commentary. It also said ANI had used prompts seeking to elicit the “exact” words spoken in an interview, describing this as adversarial prompting.
Even after such prompts, ANI had not obtained an output that could be characterised as a substantial reproduction of its copyrighted content, the court said.
The court, however, accepted prima facie that ANI owned copyright in the original works created by its professionals and published on its website. It also held that making a copyrighted work freely accessible online did not strip it of copyright protection.
The court said OpenAI’s use was limited to training its models and that its activities did not result in ANI losing subscribers or suffering a loss from its news syndication business.
It also rejected OpenAI’s preliminary objection that Indian courts lacked jurisdiction because its training data was stored and processed on servers in the US.
ANI has its registered and principal office in Delhi, while OpenAI offers services to users and subscribers in India, the court said. The location of servers abroad could not be used to sever a chain of events that began with accessing copyrighted material in India and resulted in outputs being generated here, it said.
While refusing an injunction, the court cited the public benefits of generative AI. An injunction could affect millions of Indian users and impede the development of domestic AI models by making it necessary to obtain licences from numerous sources, the court said.
The main suit will continue, with questions requiring evidence to be determined at trial.