A correction. An earlier version of this article was published and withdrawn. It reported that the deal implied a market valuation of creative work at 9 to 12% of the AI revenue it enables. That result depended on a figure that Reddit supplies about 45% of the community text used in AI training, which was cited incorrectly. This version derives that input from Reddit's measured token count and treats the split of the fee between training and live access as unknown.
Google reportedly pays Reddit $60 million a year for access to Reddit's data, both to pre-train Gemini and to ground its live answers and AI summaries (CBS News). That agreement is now up for renewal. On 22 July the Wall Street Journal reported that Reddit has discussed shutting Google out of its content altogether, and Reddit's shares fell 8% on the day (CNBC).
That fee is one of the few public prices a frontier AI developer has put on a large body of human writing.
If we apply a top-down method, of the kind emerging from the field of data economics (Open Data Labs, 2025, and see Related work below), to estimate what creative work contributes to the value an AI system produces, we can use the reported Reddit–Google deal to estimate the share of AI revenue implied by that negotiated price.
This method is not an exact science. More precise measures, such as estimating the influence of specific data on a model's output (Grosse et al., 2023; Ghorbani and Zou, 2019), are being developed. But a top-down method provides a workable way to estimate a candidate pool of value attributable to creators' contributions to the development of AI.
For this exploratory model, we assume that data's contribution to model performance provides a reasonable proxy for its contribution to attributable revenue.
Applied to the Reddit deal, the method puts the contribution of creative work at a range of 7 to 24% of the AI revenue it enables. At a conservative 10%, the implied creative-work pool across OpenAI and Anthropic alone is valued at about $7 billion a year.
A top-down method for valuing creative work
The top-down method establishes the value an AI product generates and narrows that value, in stages, to the share attributable to a single body of work or corpus of data, in this case Reddit data.
The revenue pool
Gemini's attributable annual revenue
~$10bn
range tested $8–13bn
Extrapolated from disclosed anchors: Gemini subscriptions, Google Cloud growth and AI Overviews (Business of Apps, 2026; Alphabet Q1 2026 results).
The data share
Training data's share of the model's value
10%
exploratory range 10–30%
Anchored at the conservative end of the scaling-research range (Ho et al., 2024; Erdil and Besiroglu, 2022), consistent with comparable content markets.
The class share
Community writing — forum posts, discussion threads and reviews — as a share of training text
8.4%
range tested 7–10%
Measured by classifying an open training corpus of 200 billion tokens (Wettig et al., 2025; Creative Value Monitor composition research, forthcoming).
The corpus share
Reddit's feasible share of the community class
30–50%
assuming substantial utilisation
From Reddit's measured token count and training repetition limits (Reddit S-1; Muennighoff et al., 2023).
Stage 1. The revenue pool. Gemini generates roughly $10 billion a year in attributable revenue, our estimate within a range of $8 to 13 billion, extrapolated from disclosed anchors. Gemini subscription revenue was estimated at approximately $1.2 billion in 2025 (Business of Apps, 2026), Google Cloud grew 63% year on year to $20.0 billion in the first quarter of 2026 (Alphabet Q1 2026 results), with revenue from products built on generative AI models up nearly 800% (Pichai, Q1 2026 earnings call), and a growing share of search advertising is served through Gemini-powered AI Overviews. Subscriptions are the only one of those lines estimated separately. Alphabet does not break out what Gemini contributes to cloud or search, so the $8 to 13 billion band is a judgement built on those anchors rather than a figure derived directly from them.
Stage 2. The data share. Analysis of historical language-model progress attributes 60 to 95% of performance gains to compute and training data together (Ho et al., 2024). Erdil and Besiroglu (2022) separately estimate that data scaling accounted for approximately 10 to 30% of performance gains in image classification. No equivalent decomposition is currently available for frontier language models. We therefore treat 10 to 30% as an exploratory range and use 10%, its conservative end, as the anchor. The 10% anchor is also broadly consistent with comparable content markets, from BMI's blanket terrestrial-radio licence at 2.2% of gross revenue (Music Business Worldwide) to the 10 to 15% of hardcover price that a trade publisher may pay an author (The Bindery Agency).
Stage 3. The class share. Community writing, meaning forum posts, discussion threads and reviews, makes up 8.4% of training text. We measured this directly by classifying an open training corpus of 200 billion tokens (Wettig et al., 2025; Roche, Creative Value Monitor composition research, forthcoming). Where inputs are varied below, we treat the class share as a 7 to 10% band.
Stage 4. The corpus share. Reddit's archive holds over one billion posts and more than sixteen billion comments (Reddit S-1), which filters to roughly 200 to 400 billion training-ready tokens of Reddit text. A corpus can be usefully repeated about four times in training before returns fade (Muennighoff et al., 2023), so the archive can supply at most 0.8 to 1.6 trillion effective tokens. A frontier run uses 15 to 30 trillion tokens, and the community class gets 1.26 to 2.52 trillion of them. Pairing the smallest archive against the largest run puts Reddit's feasible capacity at just under a third of the community-text allocation, which we round down to 30%. The upper end of the arithmetic exceeds the whole class, which is not credible for a single site, so we impose a conservative modelling cap of 50%. Reddit's feasible capacity share is therefore 30 to 50% of the community class, assuming substantial utilisation of the archive.
Forward, the chain prices what the corpus could be worth for training. Backwards, it returns the valuation of creative work implied by whatever share of the $60 million fee pays for training.
A simulation across 200,000 runs puts the median at 10.5%
Every figure above is an estimate built on imperfect data. A Monte Carlo simulation takes imperfect data and, instead of fixing each input to a single best guess, gives each one a range, then runs the whole calculation many thousands of times, each run drawing a different combination of values from those ranges. The result is not one number but a spread, showing where the answer tends to land and how often.
For the Reddit deal, across 200,000 runs, and treating the split of the fee between training and live access as unknown, the implied share of creative work in Gemini's revenue has a median of 10.5%. Nine in ten runs fall between 4.6% and 19.6%. Fewer than one run in two thousand lands above 30%. The four inputs are drawn uniformly: the revenue pool at $8 to 13 billion, the class share at 7 to 10%, Reddit's corpus share at 30 to 50%, and the training share of the fee at 25 to 100%.
The calculation is:
Implied content-value share = training portion of the Reddit fee ÷ (Gemini-attributable revenue × community-text share × Reddit share of community text).
What the deal cannot tell us
The fee buys a bundle: part of it pays for data to train on, and the rest pays for live access to current conversations, which Gemini retrieves at inference to ground and cite its answers. Deals structured around attribution and live access grew from 2 in 2023 to 18 in 2025, with 34 projected for 2026 (Media & the Machine, June 2026). The contractual split between the two is not public.
Training-only archive deals command real fees elsewhere in the market, and Reddit's archive is the largest of its class. On those grounds, a reading of half to all of the fee as training is a plausible zone. If all of the fee pays for training, creative work's implied share of model value is 14 to 24%. If half, it is 7 to 12%. If a quarter, it is 4 to 6%.
At the 10% anchor, the forward chain prices training use, the archive plus a year of new content, at $25 to 42 million a year, about 40 to 70% of the reported fee.
The Reddit deal places creative work's pool at about $7 billion a year
In Google's case, the 10% anchor applies to Gemini's estimated $10 billion a year, which values creative work and content's contribution at roughly $1 billion.
For the rest of the AI sector, revenues are large and growing. OpenAI and Anthropic are annualising around $72 billion between them. OpenAI topped a $25 billion run rate at the end of February 2026 (The Information, via Reuters). Anthropic reported a $47 billion run rate towards the end of May 2026 (CNBC). Run rates annualise a single month, so these are snapshots of current scale rather than forecasts of booked revenue. Both figures are months old. That total excludes Google, Microsoft and Meta.
At the 10% value share, the creative-work pool across just these two firms, OpenAI and Anthropic, is $7.2 billion a year.
What content owners currently receive is a fraction of that $7.2 billion. The best available public estimate put annual content-licensing payments at approximately $800 million in 2024, although it was based on incomplete deal disclosures and should be treated as indicative rather than definitive (Media & the Machine, January 2025). The number of publicly reported agreements has since risen from 34 to 91 (Media & the Machine, June 2026). Applying the same average deal size to 91 deals gives roughly $2.2 billion a year, or about 3% of OpenAI and Anthropic's combined annualised revenue.
The market signals that creative content contributes around a tenth of what AI earns, while current payments appear to amount to only around 3%.
The next problem is dividing the pool
Knowing that around 10% of an AI system's value may be attributable to the creative work and content behind it sets a candidate size for the pool, but it does not determine how that pool should be divided among the many who contributed to it. Dividing it will need attribution methods at scale.
Related work
This work sits alongside a small but growing field. In June, writing in Harvard Business Review, Glen Weyl and Raul Castro Fernandez argued that AI companies already produce the two ingredients needed to price content whenever they train a model: scaling laws, which indicate how much value the training data creates, and the training-data mixture, which indicates how any value pool should be divided.