Training AI on Unlicensed Copyrighted Data: Copyright Infringement and the Legal Liability of Companies


Introduction: Who Is Liable When Artificial Intelligence Is Trained on Copyrighted Material?

Generative artificial intelligence has created one of the most significant copyright disputes of the digital era.

Large language models, image generators, and other generative systems rely on vast datasets during development which frequently incorporating copyrighted text, images, code and media scraped from the web.

Yet, public accessiblility on the internet does not equate to a legal license for AI training.

This distinction has created a major legal question:

Can a company train an AI system on copyrighted works without obtaining a licence from the copyright holder?

An equally important question follows:

If the training process infringes copyright, who bears legal responsibility—the model developer, the company supplying the dataset, the company deploying the AI system, its directors, its employees or multiple parties simultaneously?

There is currently no single global answer.

The United States, European Union and Türkiye approach AI training through substantially different copyright frameworks. Recent litigation has also demonstrated that courts may distinguish between the use of copyrighted works for machine learning and the manner in which those works were originally acquired.

For technology companies, startups, investors and businesses integrating generative AI into commercial operations, copyright compliance is therefore becoming an essential part of AI legal due diligence.


Why Does AI Training Raise Copyright Issues?

AI models generally learn patterns from very large datasets.

During this process, copyrighted materials may need to be downloaded, copied, converted, stored, tokenised or otherwise processed.

Copyright law traditionally grants the copyright owner exclusive rights over certain uses of a protected work, including reproduction.

The key legal issue is therefore not simply whether an AI model eventually reproduces an entire copyrighted book, photograph or song.

Copyright concerns may arise at an earlier stage: when the copyrighted work itself is copied or stored for inclusion in a training dataset.

This distinction is crucial.

An AI developer may argue that the model does not store a conventional copy of a work after training and merely learns statistical relationships from the dataset. Copyright owners, however, may argue that the initial ingestion and reproduction of their works already constitutes an exercise of an exclusive copyright right.

This conflict lies at the centre of modern AI copyright litigation.


Publicly Available Does Not Mean Copyright-Free

One of the most dangerous assumptions for AI developers is that material available on the open internet can automatically be used for machine learning.

That assumption is legally risky.

A newspaper article available without a paywall may still be protected by copyright.

A photograph appearing in Google search results may still belong to a photographer.

A book available through an unauthorised online archive may still be protected.

Source code published on GitHub may be governed by specific licence conditions.

Music, images and videos available through websites or social media platforms may similarly remain subject to intellectual property rights.

The Turkish Ministry of Culture and Tourism likewise expressly states that publication of a photograph or another work on the internet does not mean the work may be freely used; copyright continues unless an applicable permission or exception exists.

Consequently, companies designing AI training pipelines should distinguish between:

  • publicly accessible content;
  • public-domain content;
  • licensed content;
  • open-source or Creative Commons content subject to licence conditions;
  • copyrighted content for which an exception may apply; and
  • illegally distributed or pirated content.

These categories do not carry the same legal risk.


Is Training an AI Model on Copyrighted Material Automatically Copyright Infringement?

Not necessarily.

The legal answer depends heavily on the applicable jurisdiction.

For example, United States copyright law contains the flexible fair use doctrine, while European Union copyright law provides specific text and data mining exceptions.

Türkiye currently follows a different framework.

Therefore, a training practice that may arguably qualify as lawful in one country may create significantly higher infringement risk in another.

Companies operating internationally should not assume that the legal status of their training dataset can be determined under the law of the country where their headquarters are located.

AI systems are developed, hosted, licensed, supplied and commercially deployed across borders. Copyright disputes may therefore involve complex questions of jurisdiction and applicable law.


AI Training and Copyright Law in the United States

The United States has become one of the most important jurisdictions for AI copyright litigation.

The central defence relied upon by many AI developers is fair use under Section 107 of the U.S. Copyright Act.

Fair use requires a case-specific analysis involving several factors, including the purpose and character of the use, the nature of the copyrighted work, the amount used and the effect of the use on the potential market for the copyrighted work.

There is therefore no general statutory rule stating that all AI training is lawful or unlawful.

Recent court decisions demonstrate why this distinction matters.


The Anthropic Case: AI Training and the Source of the Dataset May Be Separate Questions

In Bartz v. Anthropic, a U.S. federal court considered the use of copyrighted books in connection with training Anthropic’s generative AI models.

In 2025, the court concluded that the use of books for AI training could constitute fair use under the circumstances considered by the court. However, it separately distinguished the company’s acquisition and retention of books obtained from pirate sources.

This distinction is highly significant for AI companies.

It suggests that two different legal questions may need to be asked:

First: Was the use of the work for machine learning legally permissible?

Second: Was the copy used in the training process lawfully obtained in the first place?

A company may therefore have a potentially stronger argument regarding the transformative nature of machine learning while still facing liability arising from the unlawful acquisition or storage of copyrighted material.

The consequences can be substantial. In July 2026, a U.S. federal judge approved a $1.5 billion settlement in litigation involving authors and Anthropic concerning pirated books used in the company’s library.

For corporate legal departments, this provides a critical compliance lesson:

training-data provenance matters.

Knowing where the data came from can be as important as knowing how the AI model uses it.


Meta and the Continuing Uncertainty Around Fair Use

Another significant dispute concerns Meta’s Llama models.

In a 2025 ruling involving authors who claimed their books were copied for AI training, the court granted summary judgment in Meta´s favor.

However, the decision did not create a universal rule that training generative AI on copyrighted works is always fair use.

The court specifically emphasised the importance of potential market harm and explained that a different evidentiary record concerning market dilution could produce a different outcome.

The dispute has also continued to evolve. In May 2026, several major publishers—including Elsevier, Cengage, Hachette, Macmillan and McGraw Hill—filed further claims accusing Meta of using copyrighted books and journal materials to train Llama.

The lesson for companies is straightforward:

U.S. fair use is a defence, not an automatic licence to scrape copyrighted material.


Thomson Reuters v. Ross Intelligence: Competitive Use May Increase Copyright Risk

The Thomson Reuters v. Ross Intelligence litigation provides another important example.

The dispute concerned the use of material from Westlaw to develop a competing AI-powered legal research product.

In 2025, the court rejected Ross Intelligence’s fair-use defence in relation to copyrighted material used to create a competing product.

The decision is particularly relevant to commercial AI companies because the relationship between the original work and the market served by the resulting AI system can significantly affect the legal analysis.

If copyrighted material is used to build a product that directly competes with the copyright owner’s existing or reasonably foreseeable market, infringement risk may be greater.


The U.S. Copyright Debate Remains Unresolved

The broader American legal position remains contested.

The U.S. Copyright Office has conducted an extensive study of copyright and artificial intelligence, including a dedicated report concerning generative AI training.

At the same time, litigation involving publishers, musicians, writers and AI developers continues.

As recently as September 2, 2026, the U.S. government filed a brief supporting OpenAI’s fair-use position in the copyright litigation brought by The New York Times and other publishers. The publishers dispute that position, and the litigation remains important to the developing legal framework.

Days earlier, Sony Music Publishing and Warner Chappell brought further claims against Anthropic concerning copyrighted songs allegedly used in AI training.

Accordingly, companies should be cautious about statements such as:

“AI training is fair use in the United States.”

A more accurate legal position is:

Certain forms of AI training may qualify as fair use depending on the specific facts, but neither the scope of the doctrine nor its application to all types of generative AI training has been definitively settled.


AI Training and Copyright Law in the European Union

The European Union follows a substantially different structure.

The Directive (EU) 2019/790 on Copyright in the Digital Single Market, commonly referred to as the DSM Directive, introduced express rules concerning text and data mining (TDM).

Article 4 provides a copyright exception for reproductions and extractions of lawfully accessible works for text and data mining.

However, there is an essential limitation.

The exception applies where the copyright holder has not expressly reserved its rights in an appropriate manner.

For publicly available online content, such reservations may be expressed through machine-readable means.

This opt-out mechanism is becoming increasingly important for AI developers and publishers.


The EU AI Act Introduces Specific Copyright Compliance Obligations

The legal environment changed further with the EU Artificial Intelligence Act.

Article 53 imposes specific obligations on providers of general-purpose AI models.

Among other things, providers must:

  • maintain technical documentation;
  • establish a policy to comply with EU copyright and related-rights law;
  • identify and respect rights reservations made by copyright holders; and
  • publish a sufficiently detailed summary of the content used to train the model.

The general-purpose AI obligations became applicable from 2 August 2025.

The European Commission has also introduced a General-Purpose AI Code of Practice containing dedicated transparency and copyright chapters to assist providers with compliance.

This development is significant because copyright compliance is no longer solely a matter of defending private infringement claims.

For companies supplying general-purpose AI models to the European market, training-data governance has become a regulatory compliance issue as well.


Robots.txt, Machine-Readable Opt-Outs and Rights Reservations

The EU framework is moving toward a technological compliance model in which rights holders can communicate that their content should not be used for text and data mining.

The European Commission’s GPAI Code of Practice addresses machine-readable methods for identifying rights reservations, including respect for robots.txt and developing technical protocols.

For AI developers, this means that a crawler designed solely to gather as much internet content as technically possible may no longer be sufficient.

A legally compliant crawler may need to distinguish between:

  • accessible content;
  • authorised content;
  • opt-out signals;
  • robots.txt instructions;
  • licensing conditions; and
  • websites whose terms prohibit scraping or AI training.

This creates a new intersection between software architecture and legal compliance.

Copyright rules increasingly need to be built directly into data acquisition systems.


What Is the Position Under Turkish Copyright Law?

Türkiye presents a particularly important issue for domestic AI companies and foreign companies developing AI systems using Turkish content.

Copyright protection in Türkiye is primarily regulated by Law No. 5846 on Intellectual and Artistic Works (Fikir ve Sanat Eserleri Kanunu – FSEK).

Article 22 grants the copyright holder the exclusive reproduction right.

The provision defines reproduction broadly, covering complete or partial, direct or indirect, temporary or permanent copying by any means or method.

AI training may involve technical processes that fall within this broad concept of reproduction, including downloading, storing or converting protected content.

Unlike the European Union, Türkiye does not currently have a specific statutory text-and-data-mining exception comparable to Articles 3 and 4 of the DSM Directive.

Recent Turkish legal scholarship has specifically identified the absence of a dedicated TDM exception in Turkish copyright legislation and the resulting legal uncertainty for AI training.

This distinction is critical.

Türkiye also does not apply the same broad fair use doctrine found in U.S. copyright law.

Therefore, businesses should be particularly cautious before assuming that large-scale commercial AI training using copyrighted Turkish works without authorisation is automatically lawful.


Can Existing FSEK Exceptions Protect AI Training?

FSEK contains several exceptions and limitations concerning the use of copyrighted works.

However, applying traditional exceptions designed for quotation, education, research or personal reproduction to large-scale commercial AI training can be legally difficult.

Modern AI training may involve millions or billions of automated reproductions undertaken for commercial purposes.

This is materially different from an individual quoting a limited portion of a work for scientific analysis or making a copy for personal use.

As of September 2026, Turkish law does not contain a dedicated AI-training or general commercial TDM exception equivalent to the EU model. Turkish academic analysis published in 2026 has likewise described the existing framework as insufficient to provide full legal certainty for modern text and data mining activities.

Consequently, licensing and data provenance may play an especially important role for companies developing AI models in Türkiye.


Can a Company Be Liable for Copyright Infringement Caused During AI Development?

Yes.

Corporate liability is one of the most important issues arising from AI training.

Under Article 66 FSEK, where infringement occurs through representatives or employees of an enterprise while performing their duties, an action for removal of infringement may also be brought against the owner of the enterprise. The provision expressly states that fault is not required for the action contemplated by that rule.

This provision is highly relevant to AI development.

Consider the following example:

A technology company’s data engineering team downloads millions of books, photographs and articles to build a training dataset. Senior management may never manually examine the individual files.

That does not necessarily mean the company is insulated from copyright claims.

Companies cannot generally eliminate intellectual-property risk merely by delegating dataset construction to developers, employees or contractors.

Corporate AI compliance therefore requires institutional controls.


What Remedies May Copyright Holders Seek in Türkiye?

Copyright infringement under Turkish law may create significant financial exposure.

Under Article 68 FSEK, a copyright holder whose work has been used without the required written permission may, under the conditions of that provision, claim up to three times the contractual or market value that could have been requested if permission had been properly obtained.

Article 70 also provides mechanisms for material and moral damages and, under relevant conditions, recovery connected with profits obtained from infringement.

The Ministry of Culture and Tourism identifies among the available civil remedies:

  • claims under Article 68;
  • prevention of infringement;
  • removal of infringement;
  • material damages;
  • moral damages; and
  • claims relating to profits obtained through infringement.

For AI companies relying on enormous training datasets, this creates potentially substantial exposure.

If thousands or millions of individual protected works are implicated, the financial scale of a dispute can become commercially significant even before final judgment.


Can AI Copyright Infringement Also Create Criminal Risk in Türkiye?

Potentially, depending on the conduct and the individuals involved.

The Turkish Ministry of Culture and Tourism confirms that certain intentional violations of economic and moral rights protected under FSEK can lead to criminal proceedings, including unauthorised reproduction and distribution of protected works.

The criminal-law analysis must be undertaken separately from corporate civil liability and depends on the specific facts, the conduct of individuals and the relevant statutory elements.

Companies developing AI technologies should therefore avoid treating copyright compliance exclusively as a contractual or civil-litigation issue.


The Risk Is Not Limited to the Company That Trains the Model

Another major issue is the allocation of responsibility throughout the AI supply chain.

An AI ecosystem may involve:

Dataset provider → model developer → API provider → software company → enterprise customer → end user

Different intellectual-property risks may arise at each stage.

For example, a software company may purchase access to a third-party foundation model without having participated in its original training.

Its legal exposure may differ substantially from the entity that created the training dataset.

However, risk can increase where the company:

  • fine-tunes the model on its own unlicensed dataset;
  • uploads copyrighted databases to a model;
  • generates and commercially distributes infringing outputs;
  • intentionally designs the system to reproduce copyrighted works;
  • ignores known copyright infringement;
  • provides datasets to another AI developer; or
  • gives contractual warranties concerning AI rights that it cannot substantiate.

Therefore, businesses should distinguish foundation-model risk from deployment risk and fine-tuning risk.


Fine-Tuning Can Create a Separate Copyright Problem

Many companies do not develop foundation models from scratch.

Instead, they fine-tune existing models using industry-specific data.

For example:

  • a legal AI company may use court commentary, textbooks and subscription databases;
  • a healthcare AI developer may use medical publications;
  • a financial technology company may use proprietary financial reports;
  • an e-commerce business may use product photographs and descriptions;
  • a media company may use newspaper archives;
  • a music startup may use commercially released recordings.

The fact that the underlying foundation model was lawfully obtained does not automatically make the company’s fine-tuning dataset lawful.

The company must have an independent legal basis for using its own training materials.


Licensing Is Becoming a Core AI Business Issue

The growing copyright dispute has also created a new market: AI training licences.

Publishers, media organisations, image libraries and other content owners increasingly view their archives as valuable training assets.

This means that training data should no longer be treated as a free technical input.

For many AI businesses, data may become one of their largest intellectual-property costs.

Companies negotiating AI training licences should consider provisions dealing with:

  • permitted machine-learning uses;
  • foundation training versus fine-tuning;
  • model ownership;
  • permitted outputs;
  • sublicensing;
  • derivative models;
  • territorial scope;
  • duration;
  • deletion obligations;
  • audit rights;
  • attribution;
  • confidentiality;
  • indemnification; and
  • liability for infringement.

A poorly drafted data licence can create serious problems in future investment or acquisition transactions.


AI Training Data Is Becoming an M&A and Investment Due Diligence Issue

Investors acquiring or funding AI companies should investigate how the company’s models were trained.

A technically impressive AI system may carry hidden copyright liabilities.

During legal due diligence, investors should ask:

Where did the training data originate?

Does the company have records identifying the datasets?

Were any datasets downloaded from torrent networks, shadow libraries or unauthorised repositories?

Were copyright licences obtained?

Were website opt-outs respected?

Does the company rely on fair use or another statutory exception?

In which jurisdictions was training performed?

Can the company provide evidence supporting its legal position?

Do customer contracts contain IP indemnities?

Could an injunction prevent continued use of the model?

A lack of reliable answers can affect company valuation, transaction warranties, indemnity arrangements and investment decisions.


Why “We Do Not Know What Was in the Training Data” Is Becoming an Unacceptable Corporate Position

Historically, some AI developers treated training datasets primarily as an engineering issue.

That approach is becoming increasingly difficult to defend.

The EU AI Act now expressly requires general-purpose AI providers to publish sufficiently detailed summaries of training content.

Investors and corporate customers are also becoming more concerned about intellectual-property provenance.

A company that cannot explain the origin of its training data may face difficulties involving:

  • regulatory compliance;
  • litigation;
  • investment;
  • insurance;
  • corporate acquisitions;
  • enterprise procurement;
  • contractual warranties; and
  • reputation.

Training-data records should therefore form part of the company’s permanent compliance documentation.


What Should AI Companies Do to Reduce Copyright Risk?

A strong AI copyright compliance programme should begin before training begins.

Companies developing or fine-tuning AI systems should consider implementing a training-data governance framework.

This may include:

  1. identifying every significant dataset used for training or fine-tuning;
  2. recording the source and date of acquisition;
  3. identifying the licence governing the dataset;
  4. verifying whether access to the underlying content was lawful;
  5. detecting copyright and machine-readable opt-out signals;
  6. excluding known pirate sources;
  7. documenting the legal basis relied upon for each category of content;
  8. separating licensed, public-domain and exception-based datasets;
  9. retaining evidence supporting copyright permissions;
  10. implementing output testing for memorisation and reproduction;
  11. reviewing third-party model and dataset contracts;
  12. establishing takedown and rightsholder complaint procedures; and
  13. obtaining legal review before releasing commercially significant models.

For larger organisations, copyright compliance should also form part of the company’s wider AI governance policy.


Contracts With AI Vendors Should Address Training Data

Companies purchasing AI solutions should not focus exclusively on functionality, security and price.

Vendor agreements should also address intellectual-property risk.

Important contractual questions include:

  • Does the provider warrant that it has sufficient rights to the training data?
  • Does the provider disclose the categories of training sources?
  • Who is responsible for third-party copyright claims?
  • Does the vendor provide an IP indemnity?
  • Is liability capped?
  • Does the indemnity cover training-data infringement or only infringing outputs?
  • Can customer data be used to train future models?
  • Does the customer retain ownership of uploaded materials?
  • Can the provider use confidential information for model improvement?

These provisions can materially affect a company’s exposure if a model later becomes the subject of copyright litigation.


Copyright Risk Can Affect the Entire Value of an AI Company

For traditional software companies, a copyright dispute may affect a specific module, image or codebase.

For an AI company, training-data litigation can potentially challenge the foundation of the company’s principal asset: the model itself.

A claimant may seek damages, but financial compensation may not be the only concern.

Depending on the jurisdiction and circumstances, disputes may involve requests for:

  • injunctions;
  • deletion of infringing datasets;
  • restrictions on further training;
  • restrictions on distribution;
  • disclosure concerning training sources; or
  • other remedies affecting continued commercial exploitation.

For that reason, intellectual-property provenance should be considered part of an AI company’s core corporate infrastructure.


AI Developers, Startups and Directors Should Treat Copyright as a Governance Issue

Copyright compliance should not remain solely within an engineering department.

Management should know whether the company’s model has been trained on:

  • licensed data;
  • proprietary company data;
  • customer data;
  • scraped internet data;
  • public-domain material;
  • copyrighted databases;
  • academic publications;
  • books;
  • news content;
  • photographs;
  • music;
  • audiovisual content; or
  • source code.

The board and senior management should also understand the legal basis relied upon for commercially significant datasets.

As AI copyright litigation becomes increasingly expensive, failure to create appropriate compliance procedures may become relevant not only to copyright disputes but also to broader corporate governance questions.


The Future of AI Training May Depend on Data Provenance

The early phase of generative AI development focused primarily on the quantity of data available for training.

The next phase is increasingly focused on the legal quality of that data.

The question for AI companies is gradually changing from:

“Can we technically collect this data?”

to:

“Do we have the legal right to use this data to build a commercial AI model?”

That shift will likely increase the importance of licensed datasets, transparent data sources, rights-management protocols and AI-specific intellectual-property agreements.


Frequently Asked Questions About AI Training and Copyright

Is it legal to train AI on copyrighted material?

There is no universal answer. The legality depends on the applicable jurisdiction, how the material was obtained, the purpose of the training, applicable copyright exceptions and the commercial impact of the AI system.

Does publicly available internet content automatically become free AI training data?

No. Public accessibility does not automatically terminate copyright protection.

Is AI training always fair use in the United States?

No. Fair use requires a fact-specific analysis. Recent cases have produced important decisions favouring AI developers in certain circumstances, but the doctrine does not create a universal exemption for AI training.

Does the European Union allow text and data mining?

Yes, the DSM Directive contains text and data mining exceptions, but commercial TDM under Article 4 is subject to conditions, including rights reservations by copyright owners.

Does the EU AI Act regulate copyright?

The AI Act does not replace copyright law, but it requires providers of general-purpose AI models to adopt copyright-compliance policies and publish information concerning training content.

Does Turkish law have a specific AI training exception?

As of September 2026, Turkish copyright legislation does not contain a dedicated general AI-training or text-and-data-mining exception comparable to the EU framework.

Can a Turkish company be liable when employees use copyrighted data for AI development?

Potentially yes. Article 66 FSEK expressly permits certain infringement claims against an enterprise owner where infringement occurs through representatives or employees performing their duties.

Can copyright infringement create financial liability under Turkish law?

Yes. Depending on the circumstances, FSEK provides claims relating to prevention and removal of infringement, damages and claims under Article 68 that may reach up to three times the relevant contractual or market remuneration.


Conclusion: AI Training Data Is Now a Legal Asset, Not Merely a Technical Resource

Artificial intelligence companies can no longer treat training data simply as raw material available for unrestricted technological use.

Copyright law increasingly requires organisations to understand what data they use, where it came from, whether they have permission to use it and which jurisdiction’s law governs that use.

The United States continues to develop its approach through fair-use litigation.

The European Union has adopted a structured text-and-data-mining regime combined with new transparency and copyright-compliance requirements under the AI Act.

Türkiye currently lacks a dedicated AI-training or commercial TDM exception, making copyright licences, dataset provenance and careful legal analysis especially important for AI developers operating under Turkish law.

For companies, the legal risk does not end with the initial model developer.

Dataset suppliers, companies fine-tuning models, AI service providers and businesses commercially deploying AI systems may each face different forms of intellectual-property exposure.

For this reason, AI copyright compliance should be integrated into software development, corporate governance, investment due diligence and commercial contracting from the earliest stages of an AI project.

Technology companies, startups, investors and international businesses developing or deploying artificial intelligence in Türkiye may require legal assistance concerning AI training datasets, copyright licensing, intellectual-property due diligence, AI development agreements, software contracts, EU AI Act compliance and copyright infringement disputes.

This article is intended for general informational purposes only and does not constitute legal advice. Artificial intelligence and copyright law are developing rapidly, and the applicable legal position should be assessed according to the relevant jurisdiction, technology, dataset and circumstances of each individual case.

Categories:

No Responses

Leave a Reply

Your email address will not be published. Required fields are marked *

Our Client

We provide a wide range of Turkish legal services to businesses and individuals throughout the world. Our services include comprehensive, updated legal information, professional legal consultation and representation

Our Team

.Our team includes business and trial lawyers experienced in a wide range of legal services across a broad spectrum of industries.

Why Choose Us

We will hold your hand. We will make every effort to ensure that you understand and are comfortable with each step of the legal process.

Call Now Button