Content licensing became the AI industry's quiet annexation: model developers, needing text and images beyond the open web, signed deals with publishers — and the documented record now runs to dozens of agreements, from the Associated Press in July 2023 through 2025's archive sales by Condé Nast and Time. What follows is the documented map: who signed, for what, at what reported prices, and where the record is dark. It is information, not legal advice.
What are the landmark deals on the record?
The documented sequence. July 2023: OpenAI and the Associated Press — the first major news licensing deal, terms undisclosed. December 2023: Axel Springer, the German publisher, reported at tens of millions of euros annually over multiple years. May 2024: News Corp, the largest disclosed deal — reported by the Wall Street Journal at more than $250 million over five years, covering Dow Jones, the New York Post, the Times and Sun of London. Through 2024: the Financial Times, Le Monde, Prisa, Associated Newspapers, and others, mostly undisclosed. Into 2025: the second wave — Perplexity's content-partnership program with revenue sharing (launched mid-2024 after publisher complaints), Amazon reportedly negotiating with news outlets, and Google extending its own licensing programs. Alongside them, the archive sales: Condé Nast reportedly licensed its photographic archive for more than $100 million, and Time's print-archive license was reported in the tens of millions.
What are the actual terms?
Mostly dark, by design. The documented pattern from publishers' public statements and reporting: multi-year terms, per-year pricing in the millions to tens of millions for the largest portfolios, training rights plus display rights in AI answers (the valuable and contested part), and attribution requirements of varying strength. The two structural questions every deal answers differently: whether the license covers future model training or only current products, and whether the publisher can withdraw when the term ends — rights that matter because a model trained on the content keeps the learning. The deals that leaked in detail — News Corp's among them — show escalating payments structured to rise as AI products monetize, an equity-like participation without the equity.
How does litigation set the floor?
The parallel track is the courtroom. The New York Times sued OpenAI and Microsoft in December 2023, the landmark case of the category; other suits followed from the Chicago Tribune's owner, the Center for Investigative Reporting, and authors including the class actions consolidated in New York. The litigation's documented effects: it established the price of refusal — publishers with strong copyright positions negotiated from a credible threat — and it slowed nothing else, since most defendants kept training while the cases proceeded through 2025. A related event: in late 2025, OpenAI and News Corp's dispute over whether chat products counted under their license landed in arbitration, the first documented intradeal fight over what 'display' means — a preview of the definitional battles every contract now faces. The courtroom's unresolved core, fair use for training, remained undecided as of early 2026.
What did the record look like for the sell-side?
The publishers' documented splits. The licensors' case: recurring high-margin revenue in a collapsing ad market, and attribution that drives some traffic. The refusers' case, printed in the Times' suit and others: licensing at current prices funds the very substitution — AI answers ending the click-through economy — that destroys publishers' long-term position; the reported per-year figures, even $50 million-plus, are a rounding error against the labs' capital and against the traffic declines already documented at search-dependent publishers. Both positions are documented; the market has mostly chosen the money.
What should startups take from the map?
Three operating notes. Data licensing is now a funded startup category — the market for licensed training and evaluation data (specialty corpora, expert data, human feedback) repriced upward with every litigation milestone, and suppliers with clean provenance command premiums. Provenance is the product: post-settlement and post-arbitration, buyers increasingly require documented rights chains, and startups that can prove theirs win deals. And the definitional risk is real: every license's words — training, display, derivative models — are being contested in arbitration and courts, so the drafting, not the headline price, is where the value actually moves.
The record shows an industry converting lawsuits into invoices: prices rising, terms secret, and the underlying question — what training is worth — still unanswered by any court. The deals are the market's interim answer, renewed annually under threat.
For more context, read Open Source in the AI Era: How Projects Now Get Funded.
For more context, read genius act stablecoin law.
For more context, read Datacenter Energy: The Wall AI Startups Are Heading Toward.

