Training an AI on copyrighted books and news articles is legally contested, not clearly legal or illegal — it depends on the country, on whether a fair-use or text-and-data-mining exception applies, and on how the training data was obtained.
In the United States, AI developers generally argue that training is "fair use" under copyright law, while authors, publishers, and news organizations argue it is unauthorized copying on an industrial scale.
Courts have not settled the question, and several high-profile lawsuits are still unresolved. So the honest answer to "is it legal?" is: sometimes, in some places, under some conditions — and nobody can yet give you a blanket yes or no.
The mechanism behind the dispute is that AI training involves two separate legal steps, and people often blur them together. First, the developer has to acquire the material — by buying books, scraping websites, or using a dataset someone else assembled. Second, the model processes that text to learn statistical patterns.
Copying is the part copyright law cares about most, because making a copy is normally the copyright owner's exclusive right. The defense usually raised is fair use in the US, or text-and-data-mining exceptions in places like the European Union and the United Kingdom. Fair use is decided case by case using factors such as the purpose of the use, the nature of the work, how much was used, and the effect on the market for the original.
That is why two similar-sounding cases can come out differently: a model trained on lawfully purchased books and one trained on pirated copies are not in the same legal position, even if the output looks identical.
A concrete example helps here. Suppose a company buys thousands of ebooks through normal retail channels and trains a language model on them. The copyright owner's strongest argument is that the company made copies without a license, and that the model now competes with the books by generating similar text.
The company's strongest argument is that the training is transformative — it does not reproduce the books for readers — and that the public benefits from the resulting tool. Now change one fact: the company downloaded the same books from a pirate site instead. The fair-use argument gets much weaker, because the acquisition itself was unlawful, and courts tend to look unfavorably on a defense built on stolen inputs.
Same model, same output, very different legal exposure. News articles add another wrinkle, because publishers argue that AI summaries and chatbot answers substitute for visiting the original article, which hits the market-harm factor directly.
Where this advice breaks down is important. First, the law is genuinely unsettled: there is no final ruling that covers all books and news articles, and appellate courts in different countries can reach opposite conclusions. Second, jurisdiction matters enormously — a practice that fits inside the EU's text-and-data-mining rules may not fit US fair use, and vice versa.
Third, the rules differ by content type: a public-domain novel, a licensed news archive, and a scraped paywalled article are three different situations. Fourth, licensing is increasingly the practical answer. Some publishers now sell training rights directly, which sidesteps the court fight entirely.
If you are a creator wondering whether your book was used, the realistic path is to check whether the developer published a dataset list, look for an opt-out mechanism, and consult a lawyer rather than rely on a general article. If you are building something small, the safest route is to use licensed, public-domain, or explicitly permitted data.
For broader guidance on keeping AI use clean and defensible, see our guide on how to stop AI from spreading misinformation in your content.
One insight worth holding onto: the legality of training and the legality of output are two different questions. A model can be trained in a way a court might excuse, yet still produce output that infringes if it reproduces protected text too closely. That is why developers add filters and why some tools refuse to generate long passages from known works.
Our AI tool database tracks a wide range of tools with a pricing and capability snapshot recorded at verification time, but it does not track the licensing status of every training dataset — and no directory can, because those disclosures are voluntary and change constantly. Treat any confident claim that AI training is "definitely legal" or "definitely theft" as a sign the speaker has picked a side, not read the case law.