All updates

Update · July 2026

What the Reddit–Google Deal Suggests Creative Work Is Worth to AI

A correction. An earlier version of this article was published and withdrawn. It reported that the deal implied a market valuation of creative work at 9 to 12% of the AI revenue it enables. That result depended on a figure that Reddit supplies about 45% of the community text used in AI training, which was cited incorrectly. This version derives that input from Reddit's measured token count and treats the split of the fee between training and live access as unknown.

Google reportedly pays Reddit $60 million a year for access to Reddit's data, both to pre-train Gemini and to ground its live answers and AI summaries (CBS News). That agreement is now up for renewal. On 22 July the Wall Street Journal reported that Reddit has discussed shutting Google out of its content altogether, and Reddit's shares fell 8% on the day (CNBC).

That fee is one of the few public prices a frontier AI developer has put on a large body of human writing.

If we apply a top-down method, of the kind emerging from the field of data economics (Open Data Labs, 2025, and see Related work below), to estimate what creative work contributes to the value an AI system produces, we can use the reported Reddit–Google deal to estimate the share of AI revenue implied by that negotiated price.

This method is not an exact science. More precise measures, such as estimating the influence of specific data on a model's output (Grosse et al., 2023; Ghorbani and Zou, 2019), are being developed. But a top-down method provides a workable way to estimate a candidate pool of value attributable to creators' contributions to the development of AI.

For this exploratory model, we assume that data's contribution to model performance provides a reasonable proxy for its contribution to attributable revenue.

Applied to the Reddit deal, the method puts the contribution of creative work at a range of 7 to 24% of the AI revenue it enables. At a conservative 10%, the implied creative-work pool across OpenAI and Anthropic alone is valued at about $7 billion a year.

A top-down method for valuing creative work

The top-down method establishes the value an AI product generates and narrows that value, in stages, to the share attributable to a single body of work or corpus of data, in this case Reddit data.

Each stage narrows Gemini's estimated $10 billion of annual revenue to the part attributable to Reddit's archive.

Stage 1. The revenue pool. Gemini generates roughly $10 billion a year in attributable revenue, our estimate within a range of $8 to 13 billion, extrapolated from disclosed anchors. Gemini subscription revenue was estimated at approximately $1.2 billion in 2025 (Business of Apps, 2026), Google Cloud grew 63% year on year to $20.0 billion in the first quarter of 2026 (Alphabet Q1 2026 results), with revenue from products built on generative AI models up nearly 800% (Pichai, Q1 2026 earnings call), and a growing share of search advertising is served through Gemini-powered AI Overviews. Subscriptions are the only one of those lines estimated separately. Alphabet does not break out what Gemini contributes to cloud or search, so the $8 to 13 billion band is a judgement built on those anchors rather than a figure derived directly from them.

Stage 2. The data share. Analysis of historical language-model progress attributes 60 to 95% of performance gains to compute and training data together (Ho et al., 2024). Erdil and Besiroglu (2022) separately estimate that data scaling accounted for approximately 10 to 30% of performance gains in image classification. No equivalent decomposition is currently available for frontier language models. We therefore treat 10 to 30% as an exploratory range and use 10%, its conservative end, as the anchor. The 10% anchor is also broadly consistent with comparable content markets, from BMI's blanket terrestrial-radio licence at 2.2% of gross revenue (Music Business Worldwide) to the 10 to 15% of hardcover price that a trade publisher may pay an author (The Bindery Agency).

Stage 3. The class share. Community writing, meaning forum posts, discussion threads and reviews, makes up 8.4% of training text. We measured this directly by classifying an open training corpus of 200 billion tokens (Wettig et al., 2025; Roche, Creative Value Monitor composition research, forthcoming). Where inputs are varied below, we treat the class share as a 7 to 10% band.

Stage 4. The corpus share. Reddit's archive holds over one billion posts and more than sixteen billion comments (Reddit S-1), which filters to roughly 200 to 400 billion training-ready tokens of Reddit text. A corpus can be usefully repeated about four times in training before returns fade (Muennighoff et al., 2023), so the archive can supply at most 0.8 to 1.6 trillion effective tokens. A frontier run uses 15 to 30 trillion tokens, and the community class gets 1.26 to 2.52 trillion of them. Pairing the smallest archive against the largest run puts Reddit's feasible capacity at just under a third of the community-text allocation, which we round down to 30%. The upper end of the arithmetic exceeds the whole class, which is not credible for a single site, so we impose a conservative modelling cap of 50%. Reddit's feasible capacity share is therefore 30 to 50% of the community class, assuming substantial utilisation of the archive.

Forward, the chain prices what the corpus could be worth for training. Backwards, it returns the valuation of creative work implied by whatever share of the $60 million fee pays for training.

A simulation across 200,000 runs puts the median at 10.5%

Every figure above is an estimate built on imperfect data. A Monte Carlo simulation takes imperfect data and, instead of fixing each input to a single best guess, gives each one a range, then runs the whole calculation many thousands of times, each run drawing a different combination of values from those ranges. The result is not one number but a spread, showing where the answer tends to land and how often.

For the Reddit deal, across 200,000 runs, and treating the split of the fee between training and live access as unknown, the implied share of creative work in Gemini's revenue has a median of 10.5%. Nine in ten runs fall between 4.6% and 19.6%. Fewer than one run in two thousand lands above 30%. The four inputs are drawn uniformly: the revenue pool at $8 to 13 billion, the class share at 7 to 10%, Reddit's corpus share at 30 to 50%, and the training share of the fee at 25 to 100%.

The calculation is:

Implied content-value share = training portion of the Reddit fee ÷ (Gemini-attributable revenue × community-text share × Reddit share of community text).

Histogram of creative work's implied share of AI revenue across 200,000 Monte Carlo runs. A solid line at 3% marks what creators are paid today, a dotted line marks the normative anchor at 10%, and a dashed line marks the deal-implied median at 10.5%. A bracket spans the middle nine-tenths of runs, from 4.6% to 19.6%, and a marker at 30% notes that fewer than one run in two thousand lands above it. Footnotes list the input ranges drawn uniformly and note that payments run sector-wide against two firms' revenue, so 3% is an upper bound.
Each of 200,000 runs draws a different combination of values from the input ranges. The median implied share is 10.5%. Nine runs in ten fall between 4.6% and 19.6%. Fewer than one run in two thousand lands above 30%, the ceiling set at Stage 2. The deal inversion and the scaling-research estimate use separate inputs.

What the deal cannot tell us

The fee buys a bundle: part of it pays for data to train on, and the rest pays for live access to current conversations, which Gemini retrieves at inference to ground and cite its answers. Deals structured around attribution and live access grew from 2 in 2023 to 18 in 2025, with 34 projected for 2026 (Media & the Machine, June 2026). The contractual split between the two is not public.

Training-only archive deals command real fees elsewhere in the market, and Reddit's archive is the largest of its class. On those grounds, a reading of half to all of the fee as training is a plausible zone. If all of the fee pays for training, creative work's implied share of model value is 14 to 24%. If half, it is 7 to 12%. If a quarter, it is 4 to 6%.

At the 10% anchor, the forward chain prices training use, the archive plus a year of new content, at $25 to 42 million a year, about 40 to 70% of the reported fee.

The Reddit deal places creative work's pool at about $7 billion a year

In Google's case, the 10% anchor applies to Gemini's estimated $10 billion a year, which values creative work and content's contribution at roughly $1 billion.

For the rest of the AI sector, revenues are large and growing. OpenAI and Anthropic are annualising around $72 billion between them. OpenAI topped a $25 billion run rate at the end of February 2026 (The Information, via Reuters). Anthropic reported a $47 billion run rate towards the end of May 2026 (CNBC). Run rates annualise a single month, so these are snapshots of current scale rather than forecasts of booked revenue. Both figures are months old. That total excludes Google, Microsoft and Meta.

At the 10% value share, the creative-work pool across just these two firms, OpenAI and Anthropic, is $7.2 billion a year.

What content owners currently receive is a fraction of that $7.2 billion. The best available public estimate put annual content-licensing payments at approximately $800 million in 2024, although it was based on incomplete deal disclosures and should be treated as indicative rather than definitive (Media & the Machine, January 2025). The number of publicly reported agreements has since risen from 34 to 91 (Media & the Machine, June 2026). Applying the same average deal size to 91 deals gives roughly $2.2 billion a year, or about 3% of OpenAI and Anthropic's combined annualised revenue.

The market signals that creative content contributes around a tenth of what AI earns, while current payments appear to amount to only around 3%.

The next problem is dividing the pool

Knowing that around 10% of an AI system's value may be attributable to the creative work and content behind it sets a candidate size for the pool, but it does not determine how that pool should be divided among the many who contributed to it. Dividing it will need attribution methods at scale.

This work sits alongside a small but growing field. In June, writing in Harvard Business Review, Glen Weyl and Raul Castro Fernandez argued that AI companies already produce the two ingredients needed to price content whenever they train a model: scaling laws, which indicate how much value the training data creates, and the training-data mixture, which indicates how any value pool should be divided.

References

  1. Google and Reddit licensing deal, reported $60 million a year — CBS News, February 2024: cbsnews.com
  2. Reddit reconsidering renewal of the Google AI content deal — CNBC, 22 July 2026: cnbc.com
  3. Reddit Q2 2026 results, other revenue up 24% to $43 million — CNBC, 30 July 2026: cnbc.com
  4. Reddit, Inc. Form S-1, post and comment counts — SEC, 2024: sec.gov
  5. Algorithmic Progress in Language Models — Ho et al., Epoch AI, 2024: epoch.ai
  6. Algorithmic Progress in Computer Vision — Erdil and Besiroglu, 2022: arxiv.org/abs/2212.05153
  7. Training Compute-Optimal Large Language Models — Hoffmann et al., 2022: arxiv.org/abs/2203.15556
  8. Scaling Data-Constrained Language Models — Muennighoff et al., 2023: arxiv.org/abs/2305.16264
  9. Organize the Web: Constructing Domains Enhances Pre-Training Data Curation — Wettig et al., 2025: arxiv.org/abs/2502.10341
  10. Creative Value Monitor composition research — Roche, forthcoming
  11. Google Gemini revenue and usage statistics — Business of Apps, 2026: businessofapps.com
  12. Alphabet Q1 2026 earnings release — SEC, Exhibit 99.1: sec.gov
  13. Q1 2026 earnings call remarks — Sundar Pichai, Google, 2026: blog.google
  14. OpenAI annualised revenue tops $25 billion — The Information, via Reuters, March 2026: finance.yahoo.com
  15. Anthropic annualised revenue around $47 billion — CNBC, June 2026: cnbc.com
  16. AI content-licensing market estimate — Media & the Machine, January 2025: mediaandthemachine.substack.com
  17. AI content-licensing deal tracker — Media & the Machine, June 2026: mediaandthemachine.substack.com
  18. BMI blanket licence rate from US terrestrial radio — Music Business Worldwide: musicbusinessworldwide.com
  19. How publishers pay authors: hardcover royalty rates — The Bindery Agency: thebinderyagency.com
  20. The Economics of AI Training Data: A Research Agenda — Oderinwale and Kazlauskas, Open Data Labs, 2025: arxiv.org/abs/2510.24990
  21. Studying Large Language Model Generalization with Influence Functions — Grosse et al., 2023: arxiv.org/abs/2308.03296
  22. Data Shapley: Equitable Valuation of Data for Machine Learning — Ghorbani and Zou, 2019: arxiv.org/abs/1904.02868
  23. How AI Companies Can Pay Fair Rates for the Content They Need — Weyl and Castro Fernandez, Harvard Business Review, June 2026: hbr.org
  24. Creative Work: the $7 Billion Missing Line Item — earlier article: creativevaluemonitor.ai/updates

Anthropic's Claude was used for research, reference sourcing, calculations and drafting support. The assumptions and calculations in this article are open to scrutiny and revision as better evidence becomes available.