for agents/llms.txtv0.6.18 · 28 Sep 2026

Home / Articles / Sixteen thousand fetches, ten clicks, and a token bill nobody is sending: the case for paying publishers to be easy to read

Sixteen thousand fetches, ten clicks, and a token bill nobody is sending: the case for paying publishers to be easy to read

By · 2026-09-28 · v0.6.18 · aipublishingtokensllms-txtmarkdownprovenancefractal-semantic-graphseconomicsnewsarticle

Abstract: A media analyst posted a week of Cloudflare logs this weekend, showing AI answer engines fetching a small publisher's pages 16,000 times and sending 10 readers back, and called it predatory. The numbers are consistent with everything Cloudflare, TollBit and Wikimedia have published, and the usual reading is a tragedy of the commons, to be fixed by pricing the withdrawal. This article makes a second reading that the debate has missed: those 16,000 fetches are also a cost to the fetcher. Every one is a page of HTML parsed, extracted and turned into tokens, more than half of them re-fetches of pages that have not changed, on a web where ninety per cent of what crawlers process is unique and so defeats every cache. Measured on this site's own 183 pages, the markdown twin of a page is 62% fewer tokens than the HTML; Cloudflare's own example is 81%. Dates, hashes and change signals remove whole fetches; frozen, hashed sources remove the verification round trips; a typed graph lets an agent load the altitude a question needs rather than the page. Every payment rail built so far, pay per crawl, RSL, Microsoft's marketplace, Perplexity's pool, Cloudflare's pay per use, prices the content. None prices the format. The hypothesis is that a publisher who serves structure is saving the provider money the provider is already spending, and that a share of the saving, paid in money or in the provider's own tokens, is a monetisation angle that needs no licensing deal and works for a site with ten clicks a week. The arithmetic for a single site is small and the article says so. It also says what data would settle the question, and notes that this site has already been running the publisher's half of the experiment.

One week on one small publisher's site, as an AI answer engine sees it. Five retrieval bots fetch the same HTML pages 16,000 times, each fetch parsed and tokenised on the provider's side, more than half of them for pages that have not changed since the last time. Ten readers come back. Figures from Thomas Baekdal's Cloudflare dashboard, Cloudflare's own crawl research, and this site's token measurements.

This weekend the media analyst Thomas Baekdal posted a week of Cloudflare's AI tracking for baekdal.com. The category Cloudflare calls AI answer retrievals, requests made by bots that fetch a page to answer somebody's question, came to 16,000. The referrals those answers sent back came to 10. By bot: Claude-SearchBot 4,307 requests and no clicks; ChatGPT-User 3,831 and seven; PerplexityBot 3,446 and none; Applebot 2,414 and none; MistralAI-User 1,815 and none. "They are using my work, but providing me with no conversions in return," the post said. "This entire thing is 100% predatory." And then a line that matters for what follows: "this is not about paying for it either. I'm not looking to get paid by AIs for what they use. I want instead to have people come to my site."

Two things to say before the argument. The post carries its own caveat: an update at the top says an inconsistent Cloudflare block rule was found afterwards and that the author was checking whether the requests actually came through, so the exact figures may move. And the shape does not depend on them. Cloudflare's network-wide data for July 2025 put Anthropic at 38,000 pages crawled for every visitor referred, OpenAI at about 1,100 and Perplexity at about 195, against Google's 5.4. TollBit, which meters bot access for hundreds of publishers, counted one AI bot visit for every 31 human visits by the last quarter of 2025, from one in 200 at the start of that year, and for European publishers 179 bot visits for every human referral an AI app sent back. Baekdal's week is one publisher's instance of a pattern everybody who measures it has found.

The reading everybody makes of that pattern is a tragedy of the commons: the answer engines graze on a shared field and give nothing back to the people who plant it, so the field will be fenced and the grazing priced. That reading is right, and the fencing is well under way. I want to make a second reading, which I have not seen made, and which I think has money in it that nobody is currently invoicing.

In short

The commons reading, and where it has got to

The fencing is a matter of record now, so here it is briefly, because the second reading has to sit alongside it rather than replace it.

Cloudflare, which sits in front of a fifth of the web, began blocking AI crawlers by default in July 2025 and let sites answer them with HTTP 402, Payment Required, and a price per request. On 1 July 2026 it said that was not enough: by June, bots were 57.4% of all traffic to HTML content on its network, and the company moved from charging per fetch towards paying publishers "when their content actually appears inside an answer." From 15 September 2026, crawlers that will not say whether they are search, training or an agent are blocked by default on pages that carry ads. In the year between those two announcements a whole layer of pricing machinery was built: the Really Simple Licensing standard, backed by Reddit, Yahoo, Ziff Davis, O'Reilly and Cloudflare itself, which puts licence terms into robots.txt and supports pay per crawl and pay per inference; the IAB Tech Lab's content monetisation protocols for the same purpose; Microsoft's Publisher Content Marketplace, launched in February 2026 with the Associated Press, Condé Nast, Hearst, USA Today and Vox on one side and Copilot and Yahoo on the other, paying on usage; and Perplexity's Comet Plus, which puts 80% of a $5 subscription into a pool paid out for human visits, citations and agent actions. Above all of it sit the licensing deals, of which the largest is News Corp's, more than $250 million over five years from OpenAI, paid, notably, in cash and credits.

I wrote about the publisher's side of this a week ago, in the future of news is the story vault, not the paywall: the industry sells the one output whose price is going to zero, the article, and discards the one asset that is not, the evidence the article was made from. Everything on the list above is a way of pricing the article, or the fetch of the article, or the citation of the article. All of it prices the content. What I want to look at is the thing none of it prices, which is the form the content arrives in, and what that form costs the party fetching it.

The other side of the ledger

Take the answer engine's point of view for a moment, which almost nobody in this debate does, because the answer engine is the villain.

When ChatGPT-User or Claude-SearchBot fetches one of Baekdal's pages, what arrives is HTML. Navigation, a cookie banner, script tags, a sidebar, the article, related links, a footer. Before any of it can be shown to a model, the provider has to fetch it, parse it, decide which part is the article, strip the rest, and tokenise what is left, because the model is billed by the token whether the tokens are the article or the cookie banner. Anthropic's own pricing page gives the rule of thumb: an average 10 kB web page is about 2,500 tokens, a 100 kB documentation page about 25,000. The proposal that started the markdown-for-agents movement, Jeremy Howard's llms.txt from September 2024, put the reason plainly: "An HTML page wraps its information in navigation, ads, and JavaScript, and converting it back into clean text is difficult and imprecise," and "every wasted token costs time and money."

How much is wasted is measurable, and it has been measured, twice, from opposite ends.

Cloudflare measured it at the edge when it launched Markdown for Agents in February 2026, a feature that converts a page to markdown when the request carries Accept: text/markdown. Its example was one of its own blog posts: 16,180 tokens as HTML, 3,150 as markdown, a saving of 81%. The response even carries a header, x-markdown-tokens, that tells the agent what the page cost. Cloudflare also noted which agents already ask for markdown: Claude Code and OpenCode, both coding agents, neither a search bot.

I measured it at the source, on this site, because every page here has had a markdown twin at the same address since it was built, generated from the same content as the HTML. Counting with a standard tokenizer, the 183 HTML pages come to 1,290,795 tokens and their markdown twins to 488,625: 62% fewer, an average page falling from 7,054 tokens to 2,670. On the fourteen articles alone the saving is 53% in total and 64% at the median, the long pieces saving least because the words dominate and the short ones saving most because the chrome does. Different tokenizers will give different absolute counts. The ratio is the point, and it sits where Cloudflare's does and where everybody who has tried it reports: between two and five times.

Now add the two findings from Cloudflare's own research that turn a saving into a waste. First, more than half of the crawl traffic from bots it classifies as legitimate "goes toward re-fetching pages that have not changed." Second, in a paper with ETH Zurich, Rethinking Web Cache Design for the AI Era, more than 90% of the pages processed by large crawlers such as Common Crawl are unique by content. The web's caches were built for humans, who mostly read the same popular pages, so that a cache in the nearest city serves them. Bots read the long tail, one page each, and every one is a miss that goes back to the origin. Wikimedia said the same thing from the receiving end in April 2025: bots were about 35% of its pageviews and 65% of its most expensive traffic, because "bots performing bulk downloads of varied pages are more likely to get forwarded to the core datacenter." Its line was "our content is free, our infrastructure is not."

Put those together and Baekdal's week looks different. Sixteen thousand fetches of HTML, each one parsed and tokenised on the provider's side, more than half of them for pages the provider had already read and that had not changed, every one a cache miss for both parties. The publisher paid for the bandwidth and got ten clicks. The provider paid for the tokens and got, for the re-fetches, nothing it did not already have.

The arithmetic, honestly

Let me put numbers on it, and then say why the numbers are smaller than they look and why that does not end the argument.

Assume Baekdal's pages are like this site's, 7,054 tokens as HTML and 2,670 as markdown. Sixteen thousand fetches a week of HTML is about 113 million tokens. The same fetches of markdown would be about 43 million. The difference is 70 million tokens a week. At the list price of Anthropic's cheapest current model, $1 per million input tokens, that is $70 a week; at the mid-tier, $2 per million, $140; at an Opus-class model, $5 per million, $350. Perplexity, which sells retrieval as a product, prices a URL fetch at $0.0005 and a web search at $0.0025 on its agent API, so its own list price for Baekdal's 16,000 fetches is about $8, before the tokens the fetched pages then cost to read.

So for one small publisher the saving is tens of dollars a week, and there are three reasons the real figure is lower still. The provider does not feed raw HTML to the frontier model; it runs an extractor first, so the tokens the model actually sees are already fewer than my HTML count. It caches what it can and reads the cheap version of the model where it can. And retrieval often truncates: a page is read until the answer is found, not to the end.

Three reasons, then, why the argument survives its own arithmetic. First, the extractor is a guess, run 16,000 times a week per site, and the site knows the answer the extractor is guessing at. That is why Cloudflare's example is 81% and mine is 62%: the site's own markdown drops what the site knows to be chrome. Second, none of the mitigations touch the re-fetches. A page that has not changed costs the same to fetch, extract and read whether it is HTML or markdown; the only saving there is not fetching it, and the signal that makes that possible, nothing has changed here, is one only the publisher can send. Cloudflare's own words, this July: "A signal that just says 'nothing's changed here' lets a crawler skip the trip. That saves the answer engine compute. More importantly, it saves site owners from serving and paying for requests they never needed to." Third, and this is the one that changes the size of the number, the saving is per site but the spend is across the web. Cloudflare sees bots as 57% of HTML traffic on a network in front of a fifth of the web. The provider's retrieval bill is not Baekdal's $70. It is Baekdal's $70 multiplied by every site its bots read, every week, and whatever fraction of that is waste is a line on a very large invoice that the provider already pays.

That is the money on the table. It is not new money. It is money being spent today, by the answer engines, on reading the web the hard way.

What "easy to read" means, in four rungs

The ladder, and what each rung saves. A markdown twin cuts the tokens of every fetch. A date and a hash remove the fetch altogether when nothing has changed. Frozen, hashed sources remove the round trips an agent makes to check a claim. A typed graph lets the agent load the altitude the question needs and nothing else. Each rung exists today, on this site and elsewhere; the marked figures are measured.

The publisher's half of the exchange is a ladder, and the rungs are in order of how much they save and how much they ask.

A markdown twin, and a map. Every page served as clean markdown at a predictable address, and an llms.txt at the root that says what the site is and where the pages are. This is the llms.txt proposal, and it is what this site does: swap .html for .md on any page here and the twin is there, generated from the same source so the two cannot drift, with links that point at other twins so an agent can walk the site without parsing a tag. Cloudflare does it at the edge for sites on its Pro plan and above, and it works because the mechanism is a plain HTTP content negotiation, Accept: text/markdown, that has been a registered media type for a decade. The honest caveat is adoption. In June 2025 Google's John Mueller said "FWIW no AI system currently uses llms.txt," and pointed at server logs to prove it; a month later a colleague confirmed Google does not either. The search bots in Baekdal's list are not, as far as anyone can tell, asking for markdown. The coding agents are. That is the chicken and the egg, and the point of an exchange is to pay for the egg.

A date and a hash. The second rung asks for almost nothing and saves the most: tell the fetcher when the page last changed, and give it a hash so it can tell for itself. Web servers have had this for thirty years in Last-Modified and ETag, and Cloudflare's finding that half of all bot fetches are of unchanged pages is a finding that the bots do not use them, or that the pages do not send them honestly. A publisher who does, and whose markdown twin carries the date and the hash inside it, has made the strongest possible case for a fetch to be skipped. This site's version log is a coarse version of the same signal: every release stamped with the commit it shipped in, so a reader, human or not, can see whether anything moved.

Frozen, hashed sources. The third rung is the one from the news article. An agent asked whether a claim is true will go and check it, which means fetching the source too, and the source's source. A publisher whose claims already carry the byte range and the SHA-256 of the page they came from, frozen on the day, has done that round trip once for everyone. The Evidence Dispatch, a newsroom vault built by another agent and published on this site, hashes every source and ties every correction to the claim it supersedes, so a claim can be walked to its bytes without a second fetch. Each verification an agent does not have to make is tokens it does not have to spend, and it is also the thing that makes the citation resolvable, which is the road back to Baekdal's ten clicks becoming more.

A typed graph. The top rung is the one this site exists to argue for. A page is one projection of what a publisher knows, cut for one reader at one altitude. An agent answering a question about the price of apples does not need the article about apples; it needs the three claims in it that bear on the question, with their sources, and it needs to know that the same publisher's article on bread contradicts one of them. That is a fractal semantic graph: typed nodes, edges that are verbs, every leaf tied to a hashed byte range, and the ability to load exactly as much of the graph as this question needs and no more. The tokens saved are not a percentage of the page. They are the page, minus the answer. Wikimedia's move in April 2025 is the mainstream version of this rung: rather than have AI developers scrape article HTML, it published structured JSON, "abstracts, short descriptions, infobox-style key-value data, image links, and clearly segmented article sections," so that, in its words, "instead of scraping or parsing raw article text" developers work with the structure directly. Cloudflare's AI Index, announced last September, is the infrastructure version: an llms.txt, a structured search API and a protocol endpoint per site, gated by the site's own access rules and wired to its payment features.

Every rung exists today. What does not exist is anyone paying for them.

The exchange

Here is the proposal, stated so it can be wrong.

A provider that spends fewer tokens reading a site that serves structure has saved money it was already spending, and should hand part of the saving back to the site. Not as a licence fee for the content, which is a separate negotiation that twenty publishers in the world get to have. As a rebate for the format, measurable per fetch, owed to any site of any size that makes itself cheap to read.

The measurement is the easy part, because the provider already has it. Every fetch produces a token count; Cloudflare's markdown header even puts it in the response. A provider knows, per domain, per week, what it spent reading a site as HTML and what it would have spent reading the twin, and it knows how many fetches a change signal let it skip. The saving is a number in its own logs.

The currency is the interesting part, and there are two.

Money, by the rails that now exist for paying in pence with no fixed fee. The x402 protocol puts a payment inside the HTTP 402 response, and this month a search consultant made a website charge one cent a page and watched Claude Code, running on a laptop, pay it, in test currency, and fetch the page. "The payments work," he wrote, "but I'm on both sides of them." The rail is real. The counterparty is the missing piece, and the rebate proposed here is the same rail running the other way: the site is not charging for the fetch, the provider is paying for the discount.

Tokens, and this is the one I think is under-explored. An AI company's own inference is the one thing it can give away at cost. A publisher who has cut a provider's reading bill for the site by 70 million tokens a week could be paid in, say, a fraction of that, in the provider's credits. A small publisher cannot do much with $70. It can do a great deal with a few million tokens a week, because the tokens are the budget that builds the site: the agent that writes the markdown twin, that hashes the sources, that builds the graph, that answers the reader's question on the site itself rather than in somebody else's chat window. The precedent is already in the biggest deal in the industry, where News Corp took part of its $250 million in credits for OpenAI's technology. If credits are good enough currency for News Corp, they are good enough for baekdal.com. And once a publisher holds credits it did not buy, a marketplace that lets it sell them to a publisher who needs them follows, which is the point at which a rebate becomes a revenue line and the small publisher has a monetisation angle that needs no licensing team.

Two rules of engagement, both borrowed from the standards already on the table. The terms are machine-readable and live where the bots already look, in robots.txt and the response headers, which is what RSL and Cloudflare's content signals were built for; a site declares that it serves markdown, dates and hashes, and what it expects in return. And the accounting is per fetch and per domain, reported to the site, which is what Microsoft's marketplace already does for its publishers and Cloudflare already does for its crawl control. Nothing here needs a new protocol. It needs one line on an invoice that already exists, or a discount on the fee the same parties are already negotiating.

Why it might not work

The honest objections, in the order I find them serious.

The per-site number is small. Seventy dollars a week is a rounding error for OpenAI and a coffee a day for Baekdal. It will not save journalism. The reply is that the aggregate is not small, that the scaling is linear in fetches so the sites most read are the sites most rewarded, and that the mechanism costs the provider nothing it is not already spending. A rebate on waste is the easiest money in any negotiation to agree to, because the payer is better off even after paying.

A single site cannot send the invoice. True. Baekdal cannot bill Anthropic for serving markdown, and there is no counter to walk up to. This is why the broker will be whoever sits between the bots and the sites, which today means Cloudflare, TollBit and the marketplaces, and why the natural home for the proposal is a clause in pay per use or in RSL rather than a startup. Cloudflare in particular already holds all three numbers: the fetch count, the token count and the change signal.

The provider has already fixed it its own way. Extractors, caches, cheap models. Some of the saving I count is already taken. The re-fetch half is not, and cannot be, without the publisher's signal, and Cloudflare's number for it was measured this year on traffic from the bots that have all those mitigations.

Nobody reads llms.txt. Mueller was right in June 2025 and may still be. But the objection cuts the other way: the reason nobody serves markdown is that nobody pays for it, and the reason the search bots do not ask for it is that nobody serves it. The coding agents broke the loop because for them the saving is immediate and their own. An exchange that pays the publisher breaks it for everybody else.

Baekdal does not want money. This is the objection I take most seriously, because it is the one the post itself makes. "The value to me and my publication is whether I can have an audience." Tokens are not an audience. My answer is that the two are not in competition. Structure is what makes a citation land on the right paragraph with the publisher's name on it, which is the only route by which an answer engine's user becomes a reader; and tokens are the budget that builds the site that reader arrives at. The Comet Plus design already pays for citations and agent actions alongside visits, which is a recognition that the visit is no longer the only thing of value passing between the parties. But if a publisher's position is that no exchange short of a human reader counts, this proposal does not answer it, and I would rather say so than pretend otherwise.

The numbers in the post may be wrong. They may. The author says so. The Cloudflare, TollBit and Wikimedia numbers are not in dispute, and they say the same thing at scale.

What this site already does, and what would settle it

I should say where this site stands, because the argument would be cheap if it did not.

Every page on sgit.ai has a markdown twin, listed in an llms.txt that an agent can traverse without parsing HTML, and each sibling site in the network publishes its own. Every release is logged and stamped with its commit. The vaults published here hold their sources frozen and hashed, and the graph ones are typed. When I measured the 62%, I was measuring the publisher's half of the exchange as it already runs, for free, for any agent that asks. If you are reading this in markdown, the saving has already happened.

What would settle the hypothesis is the other half, and it lives in the providers' logs. Three numbers, per domain, for a week of real retrieval: the tokens spent reading a site as HTML, the tokens the markdown twin would have cost, and the fetches a change signal would have skipped. Any answer engine has them. Cloudflare has them for a fifth of the web. If the aggregate is what I think it is, the rebate is the cheapest thing on the table for the providers, because it pays for itself, and the first monetisation line for small publishers that does not depend on being one of the twenty that get a licensing call. If the aggregate is smaller than I think, we will have learned that too, and the ladder will still have been worth climbing, because a site that is easy for an agent to read is a site whose citations resolve, whose corrections propagate and whose claims can be checked, and those were the products all along.

The commons reading says the answer engines are taking without paying, and it is right. The second reading says they are also paying, every week, to take the hard way, and that the publisher is the only party who can make it easy. Somebody should invoice for that.

Sources

© 2026 Dinis Cruz. This article's own text is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). You're free to share and adapt it, as long as you give credit. Quoted material and linked sources keep their own licences. The token measurements were made on 28 September 2026 with the cl100k_base tokenizer over the 183 pages of this site that have markdown twins.

← All articles