Ask a language model about a well documented historical event and it answers smoothly, in its own words, without pointing anywhere. Ask the same model a question that triggers a live search and it often comes back with a tight citation and a link. The difference is not mood or inconsistency. It comes from two structurally different ways a model can hold information, plus a set of legal and business pressures that have grown up around each one.

This document walks through both sides of that split. It looks at how training and retrieval work as separate technical pipelines, what copyright rulings from 2025 and settlements from later that year say about training on copyrighted material, how a fast growing market for publisher licensing deals is shaping which sources get named and linked, and what all of this means in practice for readers, publishers and the people building these systems.

How AI Memory Works

PARAMETRIC MEMORY

During training, a model reads an enormous amount of text, commonly hundreds of billions to trillions of words drawn from books, articles, code and web pages, and repeatedly adjusts billions of internal numerical weights so that patterns of grammar, fact and style become part of the network itself. Researchers call this parametric memory, because the knowledge lives in the model's parameters rather than in any retrievable document. Nothing learned this way keeps a pointer back to the specific article, book or webpage it came from. The model has not stored a copy of any single source, it has absorbed a statistical impression spread across millions of documents, so the only honest way to answer from this memory is a synthesis in its own words.

RETRIEVED MEMORY

A different pipeline runs when a model performs a live web search, reads a file someone uploaded, or pulls from a connected database. In that process, a query is converted into a numerical representation and compared against an index of documents, the most relevant passages are pulled into the model's working context alongside the user's question, and the model then generates its answer with that exact text visible to it. Researchers describe this overall pattern as retrieval augmented generation, an approach formalized in 2020 research that combined a model's trained parameters with an external, searchable index the model could consult at answer time. Because the retrieved passage is a discrete, addressable piece of text with a known origin, the system can quote a short piece of it, name where it came from, and link straight to it.

This is the real fork in the road. Parametric memory has no addresses to hand out, so it produces paraphrase. Retrieved memory has an address for nearly everything it uses, so it produces citations and links. The same model can run both pipelines within a single conversation, choosing based on what the question in front of it actually requires.

Training compresses huge collections of documents into a single set of weights, which is why a model's trained knowledge rarely comes with a pointer back to one source. Photo via Wikimedia Commons, CC BY 2.0.

Why Retrieved Answers Are Not Always Reliable Either

Having a link attached to an answer is not the same as the answer being correct. Researchers who tested a set of eight publicly available AI chatbots on bibliographic reference retrieval found that a substantial share of the citations produced, more than a third of the entries checked, were incomplete or contained inaccurate details, most often mixing information from different editions of the same book, such as blending a publication year from one edition with an author list from another. Accuracy also varied noticeably between systems, with some chatbots performing clearly better than others on the same test set.

The same research raised a related concern: because many of the books these systems referenced have long publishing histories, some of the underlying text may have reached the model through sources of questionable legality rather than through a licensed edition, which connects this reliability question back to the copyright issues discussed in the next section. The practical takeaway is that a citation should be treated as a starting point for verification, not as proof on its own, whether it came from a model's memory or from a live retrieval step.

What Copyright Law Says About AI Training

Whether an AI company can train on copyrighted material without permission is being tested directly in court, and the rules are still being written case by case. Under United States copyright law, a use of someone else's work can be excused as fair use if it passes a four part test: the purpose and character of the use, including whether it is transformative, the nature of the original work, how much of it was used, and whether the use harms the market for the original.

In 2025, two federal rulings, Bartz v. Anthropic and Kadrey v. Meta, both found that training a language model on copyrighted books can qualify as a transformative, and therefore fair, use. The two courts split on a narrower point, whether flooding a market with AI generated material can still count as harm to that market, so the legal picture is not fully settled even on this question.

The Anthropic case took a further turn after that ruling. The same judge who found that training on legally acquired books was fair use also found that Anthropic's use of pirated copies, downloaded from shadow libraries including Library Genesis and Pirate Library Mirror, was not covered by fair use, and set the piracy question for trial. Rather than proceed to trial, Anthropic agreed in September 2025 to a settlement covering roughly 500,000 books at about 3,000 dollars per work, for a total of 1.5 billion dollars, described by the parties and the court as the largest reported copyright class action recovery on record. The settlement resolved the piracy claim without establishing new case law on it, since the parties settled rather than let a jury decide.

Litigation continues elsewhere on related questions. The Chicago Tribune sued the AI search company Perplexity in December 2025, alleging that its answers substitute for the newspaper's original reporting and bypass its subscription paywall. A regional court in Munich, Germany reportedly found that OpenAI's training on copyrighted song lyrics violated German copyright law, in a case brought by the music rights organization GEMA. The New York Times' lawsuit against OpenAI remains active rather than settled, and continues to shape public debate over how far training can go without a license.

The United States Copyright Office's own analysis of AI training identified the fourth fair use factor, market effect, as the one courts tend to weigh most heavily, and raised a concept sometimes called market dilution, the idea that flooding a market with AI generated material in a similar style could undercut demand for original works even without any direct copying of specific text.

Legal scholars studying this area have separately argued that a chatbot's output sits on firmer fair use ground when it stays close to a brief summary, cites its source precisely, and links back to the original, since that combination lets the copyrighted work keep circulating instead of being displaced by the AI's version. That argument lines up with what these systems already tend to do in practice: a long, uncredited quotation risks looking like a substitute for the original article, while a short paraphrase paired with a visible link sends the reader onward instead of replacing the trip.

Training corpora are built from enormous, mixed collections of copyrighted and public material. What courts decide about that mix is still being argued case by case. Photo via Wikimedia Commons, CC BY 2.0.

How Publisher Licensing Deals Work

Alongside the courtroom fight, a separate and faster moving trend has taken shape: a market for licensing deals between publishers and AI companies. Industry trackers following this market count roughly zero deals in 2022, 12 in 2023, 28 in 2024, a dip in 2025, and a projected 36 for 2026. News and journalism content dominates the deal count, with around 48 tracked agreements, well ahead of music and audio at 16 and images or video at 12. Deals structured around ongoing, attributed access, rather than a one time training data purchase, have grown fastest of all, from about 2 in 2023 to a projected 34 in 2026.

The individual figures vary widely by company. OpenAI has struck roughly two dozen disclosed publisher agreements, including one reported at 250 million dollars spread over five years with News Corp. Google signed the Associated Press to feed real time news into its AI products. Meta reached multi year agreements covering outlets including CNN, Fox News and USA Today. Amazon licensed the New York Times, along with Conde Nast and Hearst, for use in its shopping assistant.

Academic and scholarly publishers report similar arrangements. Wiley disclosed 49 million dollars in AI licensing revenue for its 2026 fiscal year, up 23 percent from the year before, with lifetime AI related revenue surpassing 110 million dollars since it began signing these deals. Taylor and Francis, part of Informa, struck a non exclusive content and data licensing agreement with Microsoft reported at roughly 10 million dollars in its first year. Reddit disclosed 203 million dollars in aggregate data licensing contract value in its initial public offering filing.

Trade groups have also stepped in directly. In 2026 the News or Media Alliance, representing about 2,200 member publishers, introduced an opt in licensing arrangement with the AI company Bria built around a revenue share close to an even split, tied to how much a given AI product actually draws on a publisher's content.

Even so, the leverage in this market is uneven. Analysts tracking the deals note that even the largest individual publishers have signed only single digit numbers of agreements, and that real negotiating power tends to sit with sources that are genuinely hard to substitute, such as a major wire service or a single flagship newspaper. The broader conclusion drawn across several 2026 industry analyses is that the long tail of small and mid size publishers is unlikely to see meaningful direct licensing revenue, even as the headline deals grow larger.

This matters for the paraphrase and link question because a licensing deal is usually paying for exactly the thing that makes linking possible: an ongoing feed of fresh content, plus a contractual requirement to credit and link the source. A publisher without a deal, or a source the model only ever saw once during a training run years earlier, has no such live pipeline feeding it into the model's answers, so it is far more likely to show up folded into paraphrase, if it shows up in an identifiable way at all.

A growing share of what a chatbot cites by name traces back to a formal licensing arrangement, not to an accident of what happened to be on the open web. Photo via Wikimedia Commons, CC BY-SA 3.0.

When Each Mode Shows Up

Put together, the mechanism, the law and the market tend to sort answers into a few recurring situations, summarized below.

MODESITUATIONWHAT THE MODEL DRAWS ONWHAT YOU SEE IN THE ANSWER
TrainedA general knowledge question about a well documented topicPatterns absorbed across many sources during trainingA paraphrased answer in the model's own words, with no link attached
RetrievedA question about breaking news or a fast changing figureA live web search or a connected news feed pulled in during the conversationA short summary or excerpt with a named source and a link to that specific page
Session contextA question about a document the user has sharedThe exact text of that document, held in context for the conversationA close paraphrase or a short quoted phrase, referenced back to the uploaded file
DeclinedA request to reproduce an entire article, song or poemEither training data or a retrieved copy of the complete workA decline to reproduce the full work, usually paired with a short summary or a link to read it in full

What This Means in Practice

For a reader, this split is a useful signal. An answer with no link is the model's own synthesis of whatever it absorbed during training, worth double checking if the details matter, since it can be outdated or quietly blended from more than one source. An answer with a link is traceable to one specific page fetched during that conversation, and it can be opened and verified directly, though as the earlier section on reliability shows, a citation is a starting point for checking, not a guarantee of accuracy.

For a publisher or writer, visibility increasingly depends on being retrievable at the moment someone asks, not just on having been part of a training run years earlier. That is pushing some site owners toward a file called llms.txt, a plain markdown file proposed in 2024 that can be placed at the root of a website to give AI systems a curated map of its most important pages, similar in spirit to the older robots.txt file used for search engines. Unlike robots.txt, it is not an enforceable access control, it is a curated signal that a compliant AI system may or may not choose to use, since adoption among AI companies is still uneven and the standard remains informal. Combined with attribution focused licensing arrangements such as the News or Media Alliance's 2026 arrangement, it represents an attempt by publishers to be retrievable and creditable going forward, rather than relying on having been remembered from training.

For the people building these systems, the underlying architecture, described in the research literature as combining parametric and non parametric memory, exists specifically to address this gap: giving a model a way to ground a claim in evidence a person can inspect, rather than asking anyone to simply trust a fluent paragraph.

Retrieval infrastructure exists to give a model something specific to point at. It is the technical counterpart to a legal and commercial push toward attribution. Photo via Wikimedia Commons.

Conclusion

Paraphrase and link trace back to two different pipelines inside these systems. One turns patterns absorbed during training into new sentences with no address attached. The other holds a specific, fetched document in view and points directly at it. Court rulings, including the 2025 decisions in Bartz v. Anthropic and Kadrey v. Meta and Anthropic's subsequent 1.5 billion dollar settlement, are still drawing the legal line between the two. Publishers, through an expanding set of licensing and attribution deals, are still negotiating the commercial line. The basic mechanism, trained weights on one side, retrieved documents on the other, is what actually decides which sentence a reader gets.