GLIPA India is a proud Community Partner at the 17th Annual GIPC, January 11–12, 2025. Register Now

AI Training and Indian Copyright Law: Navigating Data Scraping and Infringement

ABSTRACT-

This paper aims to examine the legal issues that have emerged from the process of generative AI (GAI) training and the Indian copyright  act 1957. With regards to the implications of GAI training, it highlights the use of automated data scraping at a large scale, especially in the context of Indian copyright law. The study explores whether feeding copyrighted works to (machine) learning models is an unauthorized “reproduction” or “adaptation” under s 14 of the Copyright Act, 1957 and if the commercial use is covered by the narrow protective exceptions of the fair dealing doctrine under s 52. It also looks into the articulation of the concept-expression dichotomy, as well as the mechanics of using a digital token as a non-expressive use, to examine how the boundaries in the law are changing for domestic AI builders.   

It examines  the ongoing litigation Asian News International (ANI) v. OpenAI before the Delhi High Court  which forms a very crucial test on content protection v. technological innovation in India. In addition, it explores the scope of limitation of AI authorship under Section 2(d) of the Act by interpreting the “RAGHAV” benchmark further which places emphasis on the subjective intent of its use of Indian copyright registration whereby the registration is well grounded in human creativity. It concludes with a call to action: advocates a hybrid model that strikes a balance and recommends implementing a ‘One Nation, One License’ system – both through Collective Management Organisations (CMOs) and legally binding opt-out standards – to guarantee fair and equitable remuneration for creators while encouraging technological innovation. 

Keywords: Generative AI, Copyright Act 1957, Data Scraping, Fair Dealing, ANI v. OpenAI, AI Authorship, RAGHAV Standard, One Nation One License  

KEY TAKEAWAYS-

  • The large-scale commercial practice of data scraping for the training of Generative AI technologies poses substantial copyright infringement concerns,  
  • The Copyright Act, 1957, fails to include specific references to digital tokenization, and courts have then had to determine whether temporary machine-readable copies are in violation of the section pertaining to “reproduction” under section 14 of the act. 
  • The commercialization of AI-driven data extraction extends beyond a limited scope of ‘fair dealing’ within Section 52, and could potentially pose a threat to both creators and tech firms.  
  • Indian Jurisprudence as opposed to that of the United States does not offer a broader ‘transformative use’ defense, making it challenging to defend commercial use of artificial intelligence for training that infringes upon the original work’s value.  
  • The case of ANI vs OpenAI before the Delhi High Court will be the final word in India on whether a “Bharat standard” is protected for creators or a Silicon Valley “laissez-faire” approach is warranted.  
  • The Indian copyright framework (as per Section 2(d) and RAGHAV standard) does not allow autonomous AI systems to have authorship.  

 

A COMPARATIVE JURISDICTIONAL ANALYSIS

Jurisdiction 

Statutory Law 

AI Authorship & Ownership  

AI Training and Text and Data Mining Exceptions 

India 

Copyright Act, 1957 & Designs Act, 2000;  

Strictly Human-Centric approach and it requires human intellect and creativity 

The recent case of RAGHAV AI confirmed authorship cannot be granted to AI. 

Indian IPR laws currently lack explicit text/data mining rules.  

Commercial data scraping faces strict scrutiny under narrow fair dealing exemptions (e.g., the ongoing litigation ANI v. OpenAI). 

United States (USA) 

Copyright Act of 1976,  

Based on Strict Human Authorship  

Rejects pure AI authorship (for instance the case of Thaler v. Perlmutter). (Zarya of the Dawn). 

Tech firms rely on the judicial Fair Use doctrine. The United States opts for an inclusive fair use exception and transformative doctrine 

(NYT v. OpenAI.) 

European Union (EU) 

EU AI Act (2024) r/w 

CDSM Directive (2019). 

Divergent Regimes.  

Copyright protection strictly requires human creativity  

On the contrary the EU Design Law protects pure AI outputs if they possess visual novelty (e.g., the Midjourney Plate). 

 

 Enacts mandatory exemptions for scientific research, (Text and Data mining exception clause)  

while commercial entities are legally permitted to mine data unless rightsholders structurally “opt out.” 

United Kingdom (UK) 

Copyright, Designs and Patents Act 1988 (CDPA). 

Section 9(3) of the Act bypasses the non-human barrier by granting authorship to “the person by whom the arrangements necessary for the creation… are undertaken.” 

Strict Fair Dealing. Does not recognize broad transformative use defenses. Commercial training datasets are highly legally vulnerable (e.g., Getty Images v. Stability AI). 

The constant development of generative artificial intelligence (AI), its widespread use in today’s digital processes, and the challenge of “web-scraped” data have become central topics in debates about IP. The data AI systems draw upon are vast, unorganized, and often composed of copyrighted articles, academic journals, creative blogs, published works, and high-quality photographs.1 This data includes an indiscriminate mix of information that is protected by copyright. Strictly speaking, data scraping refers to the automated, unauthorized extraction of proprietary information from digital databases to feed into the machine learning model training process. The legal implications of such an extraction process carry significant threats of infringement of intellectual property rights, with questions regarding the legalities of using such content for “non-consumptive” machine learning purposes remaining one of the most contentious and litigious legal areas.2 As the generative models become more complex and reach more customers, the amount of data needed to train such models is increasing at an exponential rate, and it is becoming harder to see how this can be protected by the copyright system, which is traditionally based on human creativity. What is at stake is the question of whether the use of these expressive works, whether for their expressive content or as statistical fuel for a machine, is an infringement of the author’s monopoly on his/her work.3 

The Indian Legal Framework and Web Scraping

The Copyright Act, 1957, was drafted with the idea of protecting the creative expression of literary, dramatic, musical and artistic work but the Indian jurisprudence is very unclear on the legality of high-volume data scraping. Section 14 of the Copyright Act gives the copyright owner the exclusive right to reproduce, store, and distribute their work to the public. The primary legal conflict is figuring out if the feeding of copyrighted works into a “latent” (mathematical) training algorithm of an AI model constitutes “reproduction” under current law. 

Indian law does not say much about this particular technical application, but the law does provide that the unauthorized copying as well as downstream republishing could lead to significant liability under Section 51, while Section 52 of the Act codifies the doctrine of fair dealing, which allows use of the copyrighted work for specific, enumerated purposes, including for research, private study, criticism, and news review.4 Large-scale commercial data scraping for the purpose of training into profit-driven AI engines, is no longer within the historical scope of these narrow, protective exceptions, and thus is increasingly likely to become the subject of legal challenges. As a result, the current notion of fair dealing for domestic AI developers is jeopardized. There is also the grey area of “fair dealing” when it comes to commercial AI use, which raises a dangerous uncertainty for both those providing the content as well as the tech companies that will use the output5. 

The Mechanics of Copying: Reproduction and Adaptation

One of the key technical and legal concerns is the process of “copying” used in the AI training phase. As digital data are absorbed and analysed by modern machine learning systems, they make high-speed copies of the information. Lacking a specific statutory exception, these temporary copies may be considered “reproduction”, which could mean that the process of “AI training” is actually a form of copyright infringement. In the absence of a specific statutory carve-out, these temporary copies may be considered ‘reproduction’, meaning the process of ‘AI training’ could be a form of copyright infringement. 

There was no specific definition of reproduction or copying in the Indian copyright law to specifically address digital tokenization, and courts are now faced with the question if even storage of a work’s expression, for machine-readable purpose, amounts to infringement. During the training of generative AI the system breaks down data into tokens to learn the statistical patterns without really “reading” the work from  human perspective. According to this process, the AI is learning about the structure of the data and not copying the expression, hence the term “non-expressive” use of data. But the judiciary must now determine whether to apply the broadest reading of reproduction rights on which the interests of both creators and users of digital tools could be based, or to limit the scope of the rights to traditional uses of expression. 

The Reproduction Right in Section 14 is based on the concept of “substantial similarity” that could replace the original in its main commercial market. After the precedent in R.G. Anand v. Deluxe Films, infringement is accepted only if there is an unmistakeable similarity of the final output to the original work as a whole.6 Generative AI outputs will not come close to this, unless they regurgitate the training data, which can be done intentionally, through “prompt injections. Additionally, the Adaptation Right under Indian law is not likely to be violated if the AI-generated output closely mirrors the original in a different form or material, or if it results in minor, insignificant changes of the original. The courts in India, specifically in the case of Barbara Taylor Bradford vs Sahara Media have given a narrow definition to the term alteration. The “ideas vs. expression” distinction will be the last major line of defence against over-protection, with AI models that learn styles and patterns of authorial expression not necessarily the specific content potentially falling under the fair-use exceptions.  

International Lawsuits and Legal Discussions

The legal landscape surrounding AI training data is in flux worldwide, with several key court cases, including the blockbuster lawsuit by The New York Times against OpenAI, shaping the current situation.7 This litigation focuses on the question of whether the use of newspaper articles in training generative models is infringing use or a “transformative technical purpose” protected by fair use. The debates that have been proposed in this section usually fall into two conflicting positions: defenders of the copyright claim will say that the AI is trying to make a substitute product for their copyrighted material, while advocates for AI will respond that the AI is simply trying to learn from a general pattern that does not violate copyright. But this conflict is being fought in courts around the world, as copyright was penned long before the modern era of sophisticated generative technology, leaving gaps in the current statutory language.8 This is a growing conflict between the undeniable necessity for technological progress and the need to ensure the economic worth and dignity of human-produced material. However, with the various rulings coming from international courts, there is a mounting need on the part of the Indian judiciary to hold the technology sector on a leash, that is, to develop a unified and locally relevant strategy.  

Landmark Litigation in India: ANI versus OpenAI

 Asian News International (ANI) v. OpenAI9 in the Delhi High Court is the first Indian case of an organization taking on a businessman for employing training data used in the creation of AI algorithms. The plaintiff claims that OpenAI illegally scraped its articles of the news to use in the training of its large language models (LLMs), such as ChatGPT. This case will set an important precedent for the billion-dollar AI industry in India. Presiding Judge Amit Bansal has remarked on creating a Bharat standard where creators and publishers are protected over a “laissez-faire” stance that is predominant in Silicon Valley. The case was filed by the news agency ANI on 19 November 2024 in the Delhi High Court, and the judgment is presently reserved vide order dated 27 March 2026. 

According to ANI, hacking their paywalls to access breaking news content and use it in Retrieval Augmented Generation (RAG) systems inflict irreparable damage.10 But, according to OpenAI, this is a self-contained process that can’t harm the organic numbers of the news agency, and the company has not been liable in other key geographies for copyright infringements, and the procedure does not create a copy and paste reroll but rather combines answers from a wide range of publicly available sources of data. 11  

When the case reaches the Hon’ble Supreme Court, Sections 14 and 51 will provide much clarity on whether the automated compilation of proprietary news content for training AI constitutes a breach of the copyright owner’s exclusive reproduction rights. The outcomes are far-reaching, with their findings and reasoning potentially shaping the future of other creative sectors, such as fashion and textile design, where generative systems infringe intellectual property rights. It’s a kind of litmus test to see how Indian copyright law would develop, towards the interests of content holders or the efficiency of AI innovation. 

Transformative Use

Indian law does not recognise the “transformative use” doctrine, which is contrary to the United States copyright laws. Delhi High Court in University of Cambridge vs. B.D. Bhandari recognised another purpose for a work but at the same time emphasised that the reproduction right is limited to the original expressive purpose of the author of a work. This implies that Indian courts don’t appear to be going to take the wide-ranging reading of transformative use used by American courts as a general defence for AI training. Any alteration or “improvement” of the original content would have to be done in a way that gave the original something new rather than taking away from the vitality and perceived success of its communication. The court’s hesitation to come around to this doctrine demonstrates that Indian legal thinking views copyright as something more exclusive than the protection of the original expression and not something more flexible that will allow a wide range of commercial changes. 

The Doctrine of Extractive Use

The extractive use theory assumes that facts, information, or knowledge that is contained in a copyrighted work should not itself be protected by copyright. In the case of Akuate Internet Services v.Star India12, the court pointed out that to monopolise facts would hamper the constitutional pursuit of information dissemination. OpenAI has contended that its models learn facts, but not the expression: this, however, suggests that the underlying facts will have to be extracted from the entirety of the expressed work, raising other potential reproduction questions under the current law. The court has thus to tread a tightrope to strip away information not protected by the copyright and retain the expression in which it is packaged.13 

OpenAI is big on the concept and the expression dichotomy, a concept that they say they’re only using unprotected facts to share the information democratically. This is grounded in the right to freedom of information and speech in the Constitution, Article 19(1)(a). Consequently, the claim is that an LLM is a public good since it is a mechanism that indexes information, not the narrative-based content that is generally thought of as being protected by copyright. In this argument, the debate about AI is framed as a new iteration of a “library” or a “search engine”, yet publishers have signalled that synthesising information is a new type of disruptive and fundamentally different activity from the one they have traditionally been selling.14 

The existing technical solutions are “Robots.txt” protocols and opt-out mechanisms that developers/publishers use. These mechanisms can be used to indicate to users that they don’t wish to be scraped (all three are currently being discussed), but their effectiveness is open to debate; they may be disregarded by illegal or violating users. 15 In addition, publishers are urged to be careful about unauthorised mirrors of their content to not unknowingly include publisher information in training sets. Whether the bars can be legally enforced is another hot topic, with publishers saying that if they are not respected, it is interpreted as “wilful infringement”.16 

Jurisprudential Perspectives on AI Authorship and the RAGHAV Standard

Whether Artificial Intelligence can qualify as an “author” of a copyright work is a key dispute in IP law. This debate tests the fundamental limits of traditional copyright frameworks, which were historically designed around the concept of human creativity as the essential bedrock of authorship. In India, the legal framework is governed by the Copyright Act, 1957. The definition of an author is given in Section 2(d), which deems the author to be the person who created the work.17 In particular, Section 2(d)(vi) applies to computer-generated works, which refers to literary works, dramatic works, musical works, and artistic works where “author” means the person causing the work to be generated. This is reaffirmation of statutory language that says that the author must be a human or legal person. 

The incident is known as RAGHAV (Robust Artificially Intelligent Graphics and Art Visualizer) and has become a legal battle over the use of artificial intelligence.18 The RAGHAV incident comes from the RAGHAV (Robust Artificially Intelligent Graphics and Art Visualizer) and is being fought as a legal wrangle in terms of artificial intelligence usage. The artist, Ankit Sahni, an Indian, designed a work called ‘Suryast’ with RAGHAV AI and submitted it for registration to the Indian Copyright Office. 

For the first time, they also double-list the person and the AI tool as co-authors. The registration was accepted for a short time, but eventually the Copyright Office withdrew it, raising significant legal anxieties about the comprehension of a computer-generated work as the authorship of AI.19  This was a nod to the fact that the Indian copyright regime is firmly rooted in the creative contributions of humans. Currently, Indian law does not have any specific provision granting an AI system authorship in its own right, and the judiciary has been reiterating that the presence of intellectual creativity and subjective intent is essential for authorship, and not something that an AI can have. 

The international position on AI as an author (or a co-author) is consistent; the majority are steadfastly against it. The courts in the United States clearly have required that human authorship be one of the essential elements for the copyrightability of any content. At the same time, in Thaler vs. Perlmutter, it was similarly agreed by the courts that the Copyright Act assumes human authorship of copyrightable works, establishes a copyright term based on an author’s life, and requires human intent in order to qualify a group of two people as joint authors.20 In one case, Stephen Thaler tried to copyright a work that was authored by an AI system, but the court stated that it didn’t agree because copyright law presumed that humans have a legal life and this is how the duration of the copyright is measured. 

Although mainly AI-created content is not subject to copyright, there is a trend for some jurisdictions to differentiate between works created solely by AI and works produced with the help of AI. If human input is substantial, either as detailed direction or specific elements or significant post-generation refinement, then the human can be noted as the only author of the final piece. In these cases, the AI becomes a powerful instrument akin to a camera or computer program, but it is no creative partner. Leaving the polished, formal name out of who created an AI-generated work has some strategic legal implications for creators. First, to be protected, a human has to show a degree of creative control and that the AI was a tool for human intellectual expression. Second, as AI do not possess legal personality and legal intentions, we can’t set up a legal joint authorship. 

Given this, the way forward to a balance between innovation and the protection of copyright owners is a balanced hybrid system. This method resolves the binary litigation scheme by adopting a structured licensing scheme instead. A promising concept would be to develop a ‘One Nation, One License’ system through a Collective Management Organisation (CMO).21 With this model, the owners of AI models would dedicate a portion of their earnings and their usage towards a centralised pool, and we would see the rightsholders’ share going from the collection central according to any measurable metrics of scraping. This would move the burden from copyright owners to a “standardised” approach that would make compensation proportional to the usefulness of the data applied.22 

There would be an opportunity to check data usage on the fly with auditing and having to submit annual transparency reports. In addition, establishing a legally binding regime of ‘Opt-Out Standard’ (OS) would shift the industry from ‘scrape by-default’ to ‘respect by-default’. The publishers could prevent unauthorised training easily by marking content with unambiguous tags, with the result of the violation being statutory damages. This could minimise the cost burden of individual litigation, be more cost-effective and make sure that the economic value of the AI ecosystem is distributed fairly to those who create the ‘raw material’ for technological development. This approach acknowledges the creative industry’s incentive structure and the enormous technological potential of the AI industry, as the only way of guaranteeing a stable direction of development in an era of rapid changes. 

As the boundaries of generative AI with copyright law continue to evolve, they also present a landscape of intricate dynamics in litigation proceedings and nascent regulatory initiatives, both domestically and internationally. Extractive use and fair dealing doctrines could afford some protection; however, their utility for big-picture AI training is undecided. Publishers want protection for their work, technically and legally, whilst the AI industry is championing the idea of ubiquitous information. Indian legal community will need to navigate with care to ensure that the principles of authorship and creativity are not trampled down to advance the speed of algorithms, and that the traditional role and importance of the creative class within the cultural and economic life of the Digital Age are not sacrificed. 

PROPOSED POLICY FRAMEWORK FOR AI TRAINING AND COPYRIGHT

The paper suggests a hybrid approach based on the interdependence of tech companies and creative authors in the context of AI training and copyright. The main recommendation that suggests a way forward is the establishment of a One Nation, One License system with authority resting with the Collective Management Organisations that collect a share of the AI revenue, allocating it to creators depending on the amount of their content being scraped. This is paired with legally binding opt-out standards (similar to EU, AI act 2024) allowing publishers to tag their content to block scraping by default along with mandatory data audits and annual transparency reports to track usage. 

PROS:
  • Protects the rights of creators and ensures they are fairly compensated for the AI systems  trained on their content.   
  • Offers a well-defined legal framework for creators of artificial intelligence, replacing the perils of copyright litigation with a viable and orderly licensing process.  
  • Provides copyright owners with means of restoring autonomy and enforceable opt-out options for publishers to prevent unauthorized scraping of content they don’t want scraped. 
  • Minimizes transaction costs systemically, in both industries, by removing need for individual and costly court proceedings and substituting a common, centralised system of collection.   
CONS:
  • Opt-Out regulations can be ignored by infringers, or by an A.I. developer from another country, thus making digital risks a possibility.  
  • Severe technical enforcement issues around accurately auditing and measuring the exact amount of data affecting a specific AI model.  
  • Creates substantial financial and regulatory requirements on tech startups, necessitating new AI companies to forfeit share in their earnings and handle comprehensive compliance regulations. 
  • Creates complex administrative overhead to manage a national licensing pool, which requires continuous real-time data auditing and complex royalty distribution metrics to function properly 

Authored by: Lamia Kidwai 

Leave a Reply

Your email address will not be published. Required fields are marked *