AI copyright enters a new era
Guest Column: The ANI and Anthropic cases are beginning to define when training language models is lawful and when it crosses into infringement, writes M. Gautham Machaiah
by
Published: Jul 29, 2026 11:40 AM | 5 min read
- Recent court decisions in India and the U.S. are redefining the legal boundaries of copyright as it pertains to artificial intelligence, particularly regarding the development of language models.
- The Delhi High Court ruled that storing publicly available literary works for AI model training may qualify as "fair dealing," while the U.S. court found that using lawfully acquired books for AI training constitutes transformative fair use.
- Both cases emphasize the distinction between learning from copyrighted works and reproducing them, suggesting that AI models can learn without infringing copyright if they do not produce substantial portions of protected expression.
- The outcomes signal a shift in copyright jurisprudence, indicating that technology companies must ensure proper sourcing of data and prevent unauthorized reproduction, while content creators may face challenges in proving copyright violations without evidence of market harm.
For nearly three years, the debate over artificial intelligence and copyright has been framed as a battle between technology companies and content creators. News organisations, publishers and authors have argued that companies built powerful language models by exploiting copyrighted material without permission. Developers, on the other hand, have maintained that machine learning is fundamentally different from copying. Two recent court decisions—one in India and the other in the United States—have begun to redefine the legal boundaries.
The Delhi High Court’s interim order in ANI Media Pvt. Ltd. v. OpenAI and the settlement in the US Bartz v. Anthropic litigation represent important milestones in the evolution of AI law. Although they arise from different legal systems, they converge on a principle that is likely to shape the future of generative systems: developing a language model is legally distinct from reproducing copyrighted content.
ANI alleged that OpenAI scraped and stored its news reports without authorisation to build ChatGPT. It also claimed that the chatbot generated fabricated quotations attributed to the news agency.
The litigation extends beyond model training. ANI sought to restrain ChatGPT from generating responses based on its content; other publishers and media bodies later joined the proceedings, reflecting wider concerns within the news industry over the use of journalistic content by automated platforms. In the US, a group of authors sued Anthropic, alleging that it had used millions of copyrighted books, including material obtained from pirate libraries, to develop its Claude language model. What makes these cases significant is not merely their outcomes but the legal reasoning behind them.
Training is Not Copying
Justice Amit Bansal’s order is India’s first substantive judicial interpretation of how copyright law applies to generative AI. The court held, at least prima facie, that storing publicly available literary works to develop a large language model may qualify as “fair dealing” for research under Section 52(1)(a) of the Copyright Act. Importantly, the court adopted an “updating construction” of the 1957 legislation, recognising that research and learning today are no longer activities confined to humans. Artificial intelligence, the court observed, undertakes learning at the behest of and for the benefit of humans.
The commercial nature of OpenAI’s business, by itself, did not negate the fair dealing defence. Nor did the mere act of storing copyrighted works automatically amount to infringement. The court focused instead on what the model ultimately produces. ANI was unable to establish that ChatGPT memorised or regurgitated substantial portions of its reports. The court distinguished between unprotected facts and protected expression, observing that copyright safeguards the author’s original presentation rather than the underlying information. Even prompts designed to elicit verbatim passages largely produced summaries or differently expressed responses.
Learning Is Fair Use, Piracy Is Not
A similar distinction emerged in the Anthropic litigation. Judge William Alsup held that using lawfully acquired books to train an AI model constitutes transformative fair use because the model analyses patterns in language rather than reproducing the books themselves. The court likened the process to a student reading thousands of books to improve language and writing techniques.
But the judgment drew an equally firm line: obtaining material from pirated "shadow libraries" amounted to copyright infringement. Faced with potentially enormous statutory damages, Anthropic agreed to pay US$1.5 billion, destroy the illicitly acquired files and compensate authors and publishers. On July 20, 2026, a federal judge granted final approval to the settlement, bringing the litigation to a close. Because the case ended through a negotiated settlement rather than an appellate review, Judge William Alsup's fair use ruling does not constitute binding precedent. Nevertheless, the reasoning is likely to carry considerable persuasive value as courts around the world grapple with similar issues.
The Law Is Still Evolving
Taken together, these decisions signal the direction in which copyright jurisprudence is evolving. The central question is no longer whether these systems “use” copyrighted works, but how they acquire them, learn from them and ultimately reproduce them. If data is obtained from publicly available sources or legitimately purchased material, courts are unlikely to treat the learning process as infringing. But if the material itself is acquired illegally, as with pirated repositories, liability can arise even before a model is trained. And regardless of how a model is built, developers remain responsible if their systems reproduce substantial portions of protected expression instead of generating genuinely original outputs.
For the media industry, these developments present both challenges and opportunities. Publishers may find it harder to argue that model training alone violates copyright unless they can demonstrate actual market harm or verbatim reproduction. At the same time, technology companies cannot treat copyrighted works as a free resource. They must ensure rightful sourcing of data, implement safeguards against memorisation and provide mechanisms to prevent infringing outputs.
The Emerging Consensus
These decisions are unlikely to be the final word. The ANI case remains at the interim stage, while the US decision, though now final as a settlement, leaves the underlying legal reasoning untested by an appellate court.
The emerging consensus is that large language models, much like human learners, may read and learn from existing works. They cannot, however, become machines for reproducing them. That distinction between learning and copying may well become the cornerstone of copyright law in the age of artificial intelligence.
(The author is a certified independent director who has held senior leadership positions across print, broadcast and digital platforms.)
Disclaimer: The views expressed here are solely those of the author and do not in any way represent the views of exchange4media.com.
Read more news about Digital Media, Internet Advertising, Marketing News, Television Media, Radio Media
For more updates, be socially connected with us onInstagram, LinkedIn, Twitter, Facebook, YouTube & Google News
