# The future of news is the story vault, not the paywall, sgit.ai

> The news industry runs on two commercial models, advertising and subscriptions, and both are bad for the reader. One sells the reader to somebody else. The other charges rent on something most people have stopped using. Both are now being dismantled from outside, by a search layer that has stopped sending traffic and by consumer law that arrives in January 2027. This article is about what to build instead, in practical terms. The objective is a commercial model that rewards investigative journalism, so that the expensive, evidenced kind of reporting drives usage, usage drives revenue that depends on neither search nor renewals, and that revenue funds more of the same. The mechanism is to stop selling the article and start selling what the article was made from. The story is a graph, a fractal semantic graph in which meaning comes from connectivity and every claim walks down to hashed evidence, so that trust comes through provenance and provenance comes via evidence. The article is one projection of it. From that one graph a newsroom can sell five things, on demand and in pence, to readers, to firms and to agents, and every payment walks back to the people who made the facts. It is built, in parts, on things we have already published.

*Source: <https://sgit.ai/articles/future-of-news-story-vault-not-paywall.html> · site v0.5.6 · this file is generated from the same content as the page, so the two cannot drift. Every page on this site has a `.md` twin; internal links below point at them.*

---

[Home](../index.md) / [Articles](index.md) / The future of news is the story vault, not the paywall

# The future of news is the story vault, not the paywall

By [Dinis Cruz](../about/index.md) · 2026-09-22 · [v0.5.5](../admin/versions.md) · newspublishingprovenancemicropaymentssemantic-graphsarticle

***Abstract:** The news industry runs on two commercial models, advertising and subscriptions, and both are bad for the reader. One sells the reader to somebody else. The other charges rent on something most people have stopped using. Both are now being dismantled from outside, by a search layer that has stopped sending traffic and by consumer law that arrives in January 2027. This article is about what to build instead, in practical terms. The objective is a commercial model that rewards investigative journalism, so that the expensive, evidenced kind of reporting drives usage, usage drives revenue that depends on neither search nor renewals, and that revenue funds more of the same. The mechanism is to stop selling the article and start selling what the article was made from. The story is a graph, a fractal semantic graph in which meaning comes from connectivity and every claim walks down to hashed evidence, so that trust comes through provenance and provenance comes via evidence. The article is one projection of it. From that one graph a newsroom can sell five things, on demand and in pence, to readers, to firms and to agents, and every payment walks back to the people who made the facts. It is built, in parts, on things we have already published.*

One story, held as a graph in a vault, and five things that can be sold from it. The public article is free and brings the reader to the door. Everything to its right is paid, in pence and on demand, and every payment walks back to the people who made the facts.

Here is the claim, stated so it can be wrong. **The news industry sells the one thing whose price is going to zero, the article, and throws away the one thing nobody else has, the evidence the article was made from.** Fixing that is not a better paywall. It is a different product, and this piece is about what that product is, how each part of it works, who pays for it, and which parts of it already run.

## In short

The whole argument, for the reader who will not get to the end.

- **Both of the industry's models are bad for the reader.** Advertising sells the reader to somebody else, so the reader is the product and the content is bait. A subscription charges rent on something most subscribers have stopped using, so the reader is the hostage and the content is the excuse. Neither pays for what the reader came for.
- **Both models are being dismantled from outside.** Google traffic to publishers fell a third in a year and the crawlers that replaced it are now tolled. The UK's subscription rules, with renewal reminders and easy cancellation, were brought forward in August to 1 January 2027. The industry's answer to both is more subscriptions and more personalisation, which is to say, more of the thing both are running against.
- **The objective is a loop, not a price.** A commercial model that rewards investigative journalism, so the expensive, evidenced kind of reporting drives usage, usage drives revenue that depends on neither search nor renewals, and that revenue funds more reporting. The current models run the opposite loop: cheaper content, less trust, less use.
- **The story is a graph. The article is a projection.** A story being reported is claims, evidence, sources, raw materials, analysis and drafts. The article is one walk through that graph, for one audience, at one moment. Keep the graph, in a vault, and the article becomes a build artefact.
- **The graph is fractal, and meaning comes from connectivity.** A claim means nothing on its own; what it is arrives through its edges: who stated it, what it rests on, what it contradicts, what replaced it. Open any node and it is a graph of its own, a source opens into an identity world, a claim opens into the sentence and the bytes it came from, and the regulator, the company and the newsroom each keep their own vocabulary, joined by edges the newsroom draws. That is what connecting the dots has always meant, made inspectable.
- **Trust through provenance, provenance via evidence.** Trust is not asserted, it is the result of a chain that can be walked: from a claim, along named edges, to evidence that is frozen bytes with a hash. Because nothing on that ladder is asserted, every rung of it can be sold.
- **Five things to sell from one graph.** The public article, free, as the door. The licensed article, in pence, for the reader who wants this one and the firm that has to forward it. The evidence vault, licensed to the analysts, lawyers and newsrooms who already pay for data. The customised projection, one graph cut for a role, a sector or a language. And the verification API, an on-record answer with a warranty, sold per query to companies and to agents.
- **Trust, provenance and customisation are the products.** The words can be regenerated by any model in a second. The frozen evidence, the chain from claim to bytes, the journalist's track record and the on-record confirmation cannot, and those are what the reader, the firm and the agent were trying to buy all along. A customised projection is one altitude of the graph, loaded for one question.
- **The money goes to whoever made the fact.** A split of 60 to the original researcher, 25 to the data organisation, 10 to the journalist and 5 to the outlet, against today's model where nearly all of it stops at the outlet. The rails to pay in pence, with no fixed fee, now exist.
- **Parts of this already run.** Hash-verified regulation graphs, a Portuguese newsroom with 92 frozen sources and an editor-gated pipeline, a penetration test sold as eight projections of one graph, and a published estate of thirty-one vaults costing a third of a gigabyte of storage. What does not run is the billing, and the article says so.

The rest of this piece is the long form of those ten points, with the evidence, and with the parts of the solution spelt out in enough detail to be built.

## Both models are bad for the reader

The problem has been written about a great deal, so I will keep it to what matters for the solution, and start from the reader rather than the publisher, because that is where the fault is clearest.

**Advertising sells the reader.** Under an advertising model the reader is not the customer. The reader is the inventory, and the content exists to put the reader in front of the buyer. Every incentive follows from that. The headline is written to be clicked, not to be right. The page is built to be scrolled, not read. The tracking is there because the reader is what is being measured and sold. A reader who understands this, and most now do, treats the content as bait, because that is what the model makes it.

**A subscription charges rent.** Under a subscription model the reader is the customer, which is an improvement, but the reader is paying for unlimited access to everything, and the value of unlimited access is set by how much of it they use. Somebody in my own household pays roughly £50 a month across news subscriptions and reads, in a good month, a tenth of what that buys. That is the model working as designed: you pay 100% for the 10% you actually wanted, and the business depends on you forgetting to do anything about it. As [subscriptions.sgit.ai](https://subscriptions.sgit.ai/) puts it, a subscription *is a discount for committing to regular use. It is not rent on something you have the right to ignore.* Most news subscriptions are the second thing.

**Neither pays for what the reader came for.** The reader came for a piece of reporting they can rely on. Under advertising, reliability does not move the click and so is not funded. Under subscriptions, reliability does not move the renewal, because the renewal is moved by inertia, and so is not funded either. The most expensive thing a newsroom does, investigative work, is the first cost cut under both models, because it is the most expensive thing that does not move the number the model optimises. Both models are, structurally, a race to the bottom, and the 2026 Reuters Institute data is what the bottom looks like: [trust in news at 37%](https://reutersinstitute.politics.ox.ac.uk/digital-news-report/2026/dnr-executive-summary), the lowest since it was first measured in 2015, and 42% of people saying they sometimes or often avoid the news altogether.

Two loops. Advertising and rent optimise for the click and the renewal, so investigative work is the first cut, trust falls, use falls, and the next turn is worse. The model this article describes runs the other way: evidenced reporting fills a story vault, the vault is used on demand, the usage pays the people who made the facts, and the revenue funds the next investigation.

## The objective, stated once

So the objective is not a better paywall, a cheaper subscription or a new advertising format. **The objective is a commercial model that rewards investigative journalism.** One in which the expensive, evidenced kind of reporting is the thing that drives usage, because it is the thing readers, firms and agents will pay for on demand; in which that usage produces revenue that depends on neither the search engine's traffic nor the subscriber's inertia; and in which that revenue funds more of the same reporting, so the loop runs up instead of down.

Every design decision in the rest of this piece is there to serve that loop. If a proposal does not reward the reporting, it is not part of the model, however well it monetises.

## The two clocks, briefly

Both existing models are also being taken apart from outside, by two clocks the industry does not set. The evidence matters mainly for the timing, so here it is, briefly.

**The deal with search is ending, on the search engine's schedule.** Chartbeat data across more than 2,500 sites, reported by [Press Gazette](https://pressgazette.co.uk/media-audience-and-business-data/google-traffic-down-2025-trends-report-2026/), has organic Google search traffic to publishers down 33% globally in the year to November 2025, and 38% in the United States. Digiday [attributes a 25% referral drop](https://digiday.com/media/google-ai-overviews-linked-to-25-drop-in-publisher-referral-traffic-new-data-shows/) to AI Overviews specifically, and publishers surveyed for the 2026 Digital News Report expect search traffic to fall a further 43% in three years. The crawlers that replaced the traffic are now being tolled: Cloudflare began [blocking AI crawlers by default](https://www.theregister.com/2025/07/01/cloudflare_creates_ai_crawler_toll/) in July 2025, answering them with HTTP 402, Payment Required, and on 1 July 2026 said [that was not enough](https://techcrunch.com/2026/07/01/cloudflares-new-policy-pushes-ai-companies-to-pay-for-publishers-content/): more than half of AI crawler traffic re-fetches unchanged pages, bots now outnumber humans on its network, and from 15 September 2026 any crawler that will not say whether it is search, training or an agent is blocked. Its customers send [more than a billion 402s a day](https://ppc.land/cloudflare-stops-charging-ai-per-crawl-and-starts-paying-per-answer/). The licensing deals that some publishers got instead, News Corp's [$250 million over five years](https://llmpulse.ai/blog/openai-publisher-deals/) from OpenAI and the like, exist for perhaps twenty publishers in the world. For everyone else the words are interchangeable, and a model that has read a thousand articles about an event does not need the thousand and first.

**The law is coming for rent, on the legislature's schedule.** Across the twenty countries the Reuters Institute has tracked for a decade, 17% pay for online news, down from 18%, and the 2026 report describes 10% to 20% as *"a ceiling in most markets."* Of those who do not pay, [71% told the 2025 survey](https://reutersinstitute.politics.ox.ac.uk/digital-news-report/2025/dnr-executive-summary) that nothing on offer would persuade them, 65% in the UK. The thing that made subscriptions work despite that ceiling was inertia, and inertia is what the law is now removing. The UK's Digital Markets, Competition and Consumers Act carries a subscription regime with renewal reminders, a fourteen-day cooling-off period when a trial converts or a long contract renews, and cancellation by a route no harder than the one used to join, with fines of up to 10% of global turnover. On 10 August 2026 the government [brought it forward to 1 January 2027](https://ppa.co.uk/burnham-government-brings-forward-implementation-of-subscription-rules-to-january-2027), and the Professional Publishers Association called that *"genuinely difficult for publishers to operationalise."* It is difficult because the model depends on the thing the rules remove. A renewal reminder to a reader who has not opened the app in three months is a cancellation notice with extra steps. The United States is a step behind and moving the same way: the FTC's click-to-cancel rule was vacated on 8 July 2025 on procedure, the FTC has restarted it, and California's auto-renewal amendments have been in force since 1 July 2025.

The two clocks, with dates. The deal with search is being ended by the search layer, and the subscription is being regulated by the legislature. The industry's answer to both is more of the thing both clocks are running against.

**And the industry's answer to both is the same answer.** Read any commentary from the past year and it is about subscriptions and personalisation: how to convert the reader, how to keep the reader, how to show the reader more of what they already read. That was a reasonable answer to a traffic problem in 2016. In 2026 it prices the one part of the operation whose price is falling fastest, because the same model that summarises an article in a search result summarises it in a chatbot, a briefing and a translation, and one in ten people already get news from a chatbot every week.

## Why paying per item failed before, and why that is about the wrong thing

Anything smaller than a subscription has a well-made case against it, and the solution below has to answer it, so here it is at full strength.

James Ball's [2020 piece for Columbia Journalism Review](https://www.cjr.org/opinion/micropayments-subscription-pay-by-article.php) is the best version. A $100 subscription lost is 500 micropayments to replace. A newspaper is a bundle, and unbundling it breaks the cross-subsidy. You only know whether an article was worth it after you have read it. And any scheme that needs dozens of publishers to co-operate on a shared wallet will fail, because *"the success of any project is inversely correlated with the amount that requires publishers to work together."* Matthew Guay's [history of the failures](https://buttondown.com/blog/why-micropayments-do-not-work), from May 2026, adds the psychology: Nick Szabo in 1996 on transactions that are not worth the brain cycles, Clay Shirky on the anxiety in every decision to buy, and Blendle, which had more than a million users, of whom 150,000 ever paid, and pivoted to subscriptions in 2019.

All of that is true, and all of it is about a reader deciding whether an article is worth twenty pence. I agree that the reader will not stop to decide that. But look at who is standing at the till today and what they are trying to buy, because it is not the article.

**The reader who wants one thing.** I can buy The Guardian or The Telegraph for £3 this morning and I cannot buy one article from either online at any price. The Toronto Star [sells one for 75 cents](https://www.amediaoperator.com/analysis/the-toronto-star-launches-micropayments/), Cornwall Reports for 20p, and the striking thing about that list is how short it is. In the UK, 10% pay for online news, and of those, 66% do it by subscription and 7% by one-off payment. The one-off option barely exists, so almost nobody uses it, and that is then cited as proof nobody wants it. Dominic Young of Axate, quoted in [my April 2025 piece on micro and nano payments](https://docs.diniscruz.ai/2025/04/02/the-future-of-news-monetization__embracing-micro-and-nano-payments.html): *"There are more people in the market willing to pay than willing to subscribe."*

**The firm that has to forward it.** Inside a company an article is not read, it is processed. Somebody summarises it for the people who matter, puts it in the format the board reads, attaches it to the risk register, translates it for the regional office and cites it in a filing. That is transformation, it is where the value is created, and it is where the provenance is lost. That reader does not want a subscription to the newspaper. They want this article, with its evidence, in a form they are allowed to transform and forward, and they will pay a licence price that would look absurd to a consumer, because the alternative is an analyst spending an afternoon reconstructing the sources.

**The agent, paying per query.** Cloudflare's billion 402s a day are not being sent to people. The buyer that now arrives most often is a program with a budget and an instruction to find out whether something is true. It has no anxiety and no brain cycles to conserve. It has a wallet. That buyer did not exist in 2020, and it makes the mental-transaction-cost argument beside the point, because the transaction is not mental.

None of the three wants the article. The reader wants this one thing, the firm wants the thing plus the right to transform it, and the agent wants to know whether the thing is true. Ball's arithmetic assumes a micropayment cannibalises a subscription. It cannot cannibalise a subscription the reader was never going to buy, and 71% of them were never going to buy it.

**The story is a graph. The article is a projection. Sell the graph.**

## The story is a graph, and the article is a projection

Here is the shift everything else follows from, and it is the thesis of [newsroom.sgit.ai](https://newsroom.sgit.ai/thesis): *the story is a graph; the article is a projection.*

A story, while it is being reported, is not an article. It is a set of claims, each resting on evidence, drawn from sources, some of whom cannot be named, held together by analysis that has hypotheses and gaps and a confidence that changes as the reporting goes on. It has raw materials: the interviews, the documents, the dataset somebody pulled. It has drafts, and the decisions about what to leave out. All of that is the story. **The article is a projection of it**: a walk through the graph, at one moment, for one audience, in one language, at one length. The infographic is another projection. The translation is another. The three-paragraph summary a model produces in a search result is another, and that is the whole problem in one sentence: the industry has been selling the cheapest projection and throwing the graph away.

Because it does throw it away. In most newsrooms the graph lives in the journalist's head, a notes app and an email thread, and the moment the article ships the graph starts to decay. A year later, when the story is contested, or somebody wants to build on it, or an agent wants to know whether a claim in it still holds, there is nothing to walk back to. The evidence was never a first-class object. It was a step on the way to the words.

**The dataset is worth more than the article.** The words are what any model can regenerate from the dataset in a second. The dataset is what nobody else has: the frozen copy of the source page before it changed, the interview, the on-record confirmation, the spreadsheet, the chain from a claim to the bytes it rests on. A newsroom that keeps the graph and gives the words away is sitting on an asset. A newsroom that sells the words and loses the graph is selling the only thing that has stopped being scarce.

I have been making this argument since early 2025, and the dates matter, so here they are. [Monetising Trust and Knowledge](https://docs.diniscruz.ai/2025/02/02/monetising-trust-and-knowledge-for-news-providers.html), 2 February 2025, set out the tiers, the APIs and the verification services a structured newsroom could sell, and argued that *"publishers that structure their content for intelligent consumption will define the future of trusted journalism."* [Building Trust Through Fact Provenance](https://docs.diniscruz.ai/2025/02/05/the-future-of-news-building-trust-through-fact-provenance.html), three days later, made provenance the pillar the rest stands on: *"trust in information emerges over time, formed by repeated demonstrations of ethical sourcing and consistent reliability."* [Journalists' Challenges with Digital Content Provenance and Trust](https://docs.diniscruz.ai/2025/03/24/journalists-challenges-with-digital-content-provenance-and-trust.html), 24 March 2025, made a point that is still under-appreciated: *"the value of provenance is probably more for the world of journalism than it is for the world of consumers."* The reader will rarely walk the chain. The editor, the lawyer, the regulator and the agent will, and they are the ones who pay. The [identity graphs piece](https://docs.diniscruz.ai/2025/04/21/strengthening-trust-in-news__implementing-identity-graphs-for-authors-and-sources.html) of 21 April 2025, the [Dan Raywood briefing](https://docs.diniscruz.ai/2025/06/06/personalised-briefing-for-dan-raywood-on-the-future-of-news.html) of 6 June 2025 and the [Cloudflare piece](https://docs.diniscruz.ai/2025/07/04/from-free-scraping-to-fair-compensation-cloudflares-genai-crawler-charges-and-the-future-of-news-monetization.html) of 4 July 2025 each appear below where they are used. All of them were written before the traffic numbers, before the tolls, before the January 2027 date and before the payment rails existed. The predictions have not needed revising. The evidence arrived to meet them.

## The graph is fractal, and meaning comes from connectivity

"Graph" is doing a lot of work in that section, so it is worth saying which kind, because the kind is what makes the rest of this piece buildable rather than a metaphor. The graph is a [Fractal Semantic Graph](../demos/fractal-graphs/index.md), and two of its properties are the ones a newsroom has always run on without a way to write them down.

**Meaning comes from connectivity.** The thesis of [graphs.sgit.ai](https://graphs.sgit.ai/) is one sentence: *a node means nothing on its own; what it is arrives through its edges.* In a semantic graph every edge is a verb with a named inverse, evidenced_by, stated_by, contradicts, supersedes, so a link reads correctly from either end, and an edge with no verb is banned, because two things always relate and an edge that says only that constrains nothing. Now apply that to a fact in a story. The fact on its own is a sentence. What it *means* is who stated it, in answer to what, what document it rests on, which earlier claim it contradicts, which later one replaced it, and when. That is not an abstraction laid over journalism. It is the craft. A reporter who knows that a quote from the regulator on Tuesday contradicts the company's filing from March has done exactly one thing: drawn an edge, with a verb. The Portugal newsroom's rule, *we connect the dots, we do not draw them*, is a description of semantic graph construction, and the graph is what makes each connection inspectable rather than trusted.

A story, zoomed. At the desk's altitude it is a node among stories, entities and events. Open it and it is claims, sources, evidence and versions, every edge a verb. Open a claim and it is the sentence, the words in it, and the byte range in a hashed file it came from. Open a source and it is an identity world of its own. The vocabulary changes at every altitude. The grammar does not, and neither does the provenance.

**And the graph is fractal.** Zoom into any node and you enter a new world with its own node types, its own verbs and its own vocabulary, joined to the one you left by a single named edge. The story is one node in the desk's graph of stories, entities and events. Open it and it is a graph of claims, sources, evidence and versions. Open a claim and it is the sentence it rests on, the words in that sentence, and the byte range in a hashed file where they sit. Open a source and it is an identity world: credentials, past claims, corrections, in the vocabulary of the [identity graphs piece](https://docs.diniscruz.ai/2025/04/21/strengthening-trust-in-news__implementing-identity-graphs-for-authors-and-sources.html) from April 2025, and the flag that says this source is anonymous lives on that node, where every projection has to pass it. Open the regulator and it is the regulator's world, in the regulator's vocabulary; open the company and it is the company's, filings and ownership and funding. None of them was asked to adopt the newsroom's schema. The three layers graphs.sgit.ai describes, *shared facts owned by nobody, per-party formulas, declared bridges between them*, are a description of what a good newsroom already is: a place that draws edges between worlds that will never agree on a vocabulary.

Four things fall out of the fractal property that matter here.

- **Provenance comes free.** When the leaf is a word, and the word is connected to the byte range it came from and the hash of the file that held it, every claim at every altitude above it is traceable to source with no extra machinery. graphs.sgit.ai states the discipline as a build rule: *every fact carries the sentence it came from, and the build fails if that sentence has moved.* A newsroom whose build fails when a source page changes under it has solved the correction problem before it is a problem.
- **A correction propagates instead of being republished.** Mark one node superseded and every path that rested on it becomes a query, *what did we build on this?*, rather than an archaeology project. A document cannot do that. The correction is a new document, and nothing connects it to the thousand that already cite the error.
- **The smallest node is whatever the question needs.** For the reader it is the headline. For the analyst it is the claim and its sources. For the lawyer it is the word, and whether the word in the filing is the same word the regulator used. You load as much of the graph as this question needs and no more, because the rest stays one edge away. That is what a customised projection is: one altitude, one context, and it is why the same graph can be sold at five prices without being copied five times.
- **Every door teaches you something, including the empty ones.** The page on fractal graphs puts it well: even when the answer is *nothing to see here*, you now know that door was empty, and that is knowledge too. Any investigative reporter will recognise the description. Most of an investigation is doors that were empty, and under the current model that work is invisible and unpaid. In the graph it is a node with a verb on it, and it is part of what the evidence vault is worth.

## Trust through provenance, provenance via evidence

The two products this whole piece is about, trust and provenance, are not the same thing, and the order between them matters, so here is the ladder stated once.

**Evidence** is the bottom rung, and it is bytes. A source page fetched on a date, frozen, and hashed with SHA-256. An interview recording. A filing. A dataset. Not a link to any of them, because a link is a promise that the thing will still be there and still say the same, and the web breaks that promise constantly. The [Regulation Graph](../demos/vaults/regulation-graph/index.md) ends every chain in a hash of the retrieved bytes for exactly this reason, and so does the [Portugal newsroom](https://newsroom.sgit.ai/portugal), for 88 sources today.

**Provenance** is the path from a claim, down through named edges, to that evidence. It is what lets anyone, a reader, an editor, a lawyer, an agent, walk from a sentence in the article to the bytes it rests on, and see at each step what the edge says: supports, partially supports, extends beyond, contradicts. Provenance is not a statement that the claim is true. It is a statement of what the claim rests on, which is a narrower thing and a checkable one.

**Trust** is what a reader, or a market, extends to a source whose provenance they have walked before and found to hold. It is earned by consistency over time, which is why the fact provenance piece of February 2025 said *trust emerges over time, formed by repeated demonstrations of ethical sourcing and consistent reliability*, and why the journalist's credibility in this model is computed from the graph rather than asserted: how often their claims were superseded, how quickly they corrected, how independent their sources were.

The reason to be precise about the ladder is commercial. Every rung of it can be sold, because none of it is an assertion. The evidence vault sells the bottom rung. The licensed article and the customised projection sell a walk along the middle one. The verification API sells the top one, an answer on the record from a source with a track record, with a warranty. And none of those sales is possible if the newsroom kept only the article, because the article is the one artefact that contains none of the three.

## What a story vault holds

Take the thesis literally and put the story in a vault. Not a folder, and not a content management system. A [vault](../docs/what-is-sgit.md): encrypted, versioned like a git repository, opened with a single read key, carrying its own app, with the server that stores it unable to read it. Then the story graph is not a metaphor. It is the file.

Inside it, the newsroom design on [newsroom.sgit.ai](https://newsroom.sgit.ai/provenance/articles-as-vaults) puts *"the published text, the evidence it cites, the story graph it belongs to, the source list, the translations, and every prompt and decision that produced it, held together as one addressable unit."* Each of the properties that falls out of that shape is something the industry currently pays for separately or does not have at all.

- **Evidence is frozen bytes, not a link.** Every source page is fetched, saved with a date, and hashed with SHA-256, so a claim walks back to what the source said on the day, not to whatever the URL serves now. [The Portugal newsroom](https://newsroom.sgit.ai/portugal) does this for 88 sources today, and its rule is the right one: *"We connect the dots, we do not draw them."*
- **Claims are typed, dated and never deleted.** A correction does not overwrite. It supersedes, and the superseded claim stays, marked from the date, so the graph can answer a question no archive of articles can: [what did we believe on this date, and why](https://newsroom.sgit.ai/corrections/how-a-graph-answers-it). The edge from a claim to its source carries a type, supports, partially supports, extends beyond, contradicts, so a correction sits where anyone would meet the claim rather than in a box at the bottom of a page a week later.
- **Anonymous sources are marked at the node.** A source the journalist cannot name is still a node, flagged as anonymous, and the flag means every projection built from the graph redacts it automatically: the licensed article, the evidence vault sold to an analyst, the briefing for a sector desk. The protection is structural rather than a matter of remembering to remove a name from a paragraph.
- **The journalist's credibility is computable.** Not by a score assigned by somebody else, but by consistency: how often their claims were later superseded, how quickly they corrected, how independent their sources were. The design's principle is that *"more evidence does not mean more confidence unless the evidence is independent"*, and independence can only be measured when the sources are in the graph. The identity graphs piece is the long argument, and its motivating problem has only got worse: fake experts quoted in dozens of articles because nobody checked.
- **The author is the oracle.** Lifting text into typed claims requires somebody to arbitrate what a sentence meant, and that is the author. Where the extraction and the author disagree, that is information, not failure. It is also the answer to what a journalist does in a newsroom where agents do the fetching: they decide what the story is.
- **Evidence, not truth.** The vault does not claim a story is true. It measures what is evidenced, shows the chain, and lets the reader or the reader's agent walk it. That is a narrower claim than truth, and it is one that can be sold, because it can be checked.

The cost of producing a story this way is not hidden either. The newsroom design publishes a [worked story](https://newsroom.sgit.ai/provenance/a-worked-story): an English and Portuguese piece on the EU AI Act's implementation in Portugal, 12 sources, two expert validations, four drafts, £8.40 and six hours twenty-three minutes of compute, and zero human hours, stated plainly. It is a worked design example rather than a production ledger, and the page says so. But when the research, the fact-checking and the translation are lines on a bill in pounds and pence, the price of a projection stops being a guess.

## Five things to sell, and how each one works

Go back to the diagram at the top. One graph, five projections, and the free one is the door. This is the part of the argument that is new, so it gets the detail.

### 1. The public article

Free, static, findable. The projection that search and agents are allowed to read, and the one that carries the free-to-read version of the story to whoever is looking for it.

The reason to give it away is not generosity. It is the honest objection to everything else on this page: a vault opened in the browser is invisible to a crawler, and *"most of a news site's reach comes from being findable by search and by other agents."* So the public projection stays public, stays static, and links to the vault it was made from. Its job is to be the door. What it does not need to be is the product, and the mistake of the last twenty years was making the door carry the whole business.

### 2. The licensed article

The same projection, bought once, for the two buyers who want this one thing.

**For the reader**, a price in pence, paid on demand, with no account to create and no subscription to forget. A read key is the whole credential: send it, and the reader opens the article with nothing installed. The Portuguese newsroom's [wallet](https://pt.newsroom.sgit.ai/carteira/) is a demonstration of what the reader sees: each page costs a cent, the wallet debits it, and the ledger of what was spent is *"yours, not anyone else's."* The price sits where Axate and the Toronto Star have found it works, somewhere between 20p and a dollar, and the reader who buys three in a week is offered the day.

**For the firm**, a licence rather than a copy. The corporate version carries the right to transform, because that is what companies do with news, and it carries its provenance with it, because the graph is what makes the transformation defensible. When the analyst's summary of the article goes to the board, the board can walk from the summary to the claim to the frozen source, and nobody in the chain had to trust the analyst. That is worth a licence price that would look absurd to a consumer, and it is a price nobody currently charges, because nobody currently has the graph to attach.

### 3. The evidence vault

The graph itself, or the part of it that can be released, with anonymous sources redacted at the node. This is the tier nobody sells today and the one with the highest value, because its buyers are not readers.

They are analysts, lawyers, academics, regulators and other newsrooms, and they are used to paying for data. An investigative journalist who has spent a year on a story has, at the end of it, a dataset with a market of its own: the documents, the interviews, the timeline, the entity graph, the things checked and found false. The industry currently pays them for the words and lets the dataset rot. **Under this model investigative journalism is the most valuable thing a newsroom produces**, not because more people read it but because twenty organisations need what it was made from, and each of them will pay for a read key to it. That is the payment that rewards the reporting rather than the click, and it is the payment that makes the loop run up.

Mechanically it is a vault with a published read key per licensee, a version history so the buyer sees every correction as it lands, and the anonymous-source flag doing the redaction. Nothing runs on the newsroom's side between purchases; [the cost is the storage](../demos/fractal-graphs/performance.md).

### 4. The customised projection

The same graph, cut for a role, a sector, a language or a single reader. Personalisation is the word in every industry deck, and once the graph exists it means this and only this: a different walk through the same evidence. Without the graph, it means showing people more of what they already clicked on.

Two instances exist. The [personalised briefing for Dan Raywood](https://docs.diniscruz.ai/2025/06/06/personalised-briefing-for-dan-raywood-on-the-future-of-news.html), from 6 June 2025, is the future-of-news argument projected for a cybersecurity editor with seventeen years on that beat, produced by hand. [pt.newsroom.sgit.ai](https://pt.newsroom.sgit.ai/) is the other, produced by a pipeline: it is *natively Portuguese, not a translation*, because it is projected from the graph rather than translated from the English words. The buyer here is a desk, a firm or a country, and the price is for the cut, not the content: the sector briefing that a bank's risk team receives every morning is the same graph the public article came from, projected for them, and it can be produced for £8.40 rather than by a research team. In the language of the fractal graph, a projection is one altitude and one context, loaded with exactly as much of the graph as that reader's question needs, and the rest one edge away.

### 5. The verification API

The B2B product, and the one that turns a newsroom's most expensive habit, checking things, into its most valuable service.

The flow is short. A company that is about to cite a claim, or an agent that is about to act on one, sends a query: did you report this, is it still current, and is this use of it sound? The newsroom answers from the graph, on the record, with the chain attached and a warranty on the answer. The design calls it [trust as a service](https://newsroom.sgit.ai/economics/trust-as-a-service), and its two prices are the two things being bought: the fact, and the confirmation that the fact still holds today. The second is the more valuable, because the graph records [freshness](https://newsroom.sgit.ai/corrections/how-a-graph-answers-it), how long since a claim was last checked against its sources and whether any of them have changed, and nobody else can answer that question about a newsroom's reporting except the newsroom.

For the agent, the whole exchange is the 402 that Cloudflare already sends a billion times a day, with the payment inside it. For the company, it is an on-record confirmation they can put in a filing. The design also names the risk honestly: *selling trust makes the seller a target*, and a warranted answer that turns out wrong is a claim against whoever warranted it. That is not a reason to avoid the market. It is the reason the market pays, and it is why the credibility in the graph has to be computed rather than asserted.

## Who gets paid, and how the money moves

The subscription's other quiet property is where the money stops. It stops at the outlet. The newsroom design sets out a [split](https://newsroom.sgit.ai/economics/paying-the-fact-creator) closer to where value is actually created: 60% to the original researcher whose work the story rests on, 25% to the organisation that holds the data, 10% to the journalist who synthesised it and 5% to the outlet that distributed it. *"Today, essentially all of that revenue is captured at the 5% layer."* The numbers are a proposal. What is not negotiable is the mechanism that makes any split possible: you cannot pay the fact creator unless the graph names them. The Portuguese wallet page puts it in one line: *"Uma página que não soubesse nomear a sua fonte não saberia a quem pagar."* A page that could not name its source would not know whom to pay.

This is also what makes the loop a loop rather than a slogan. If the evidence vault of an investigation is licensed twenty times and 60% of each licence goes to the people who did the investigating, the reporting has been paid for by its use, and the next investigation has a budget that no advertiser and no renewal rate can touch.

For a decade the answer to "why not pay per item" ended with the rails. A £1 card top-up returned about 59p of usable credit after fees, and nothing under a pound was worth processing. That ended in the last eighteen months. The [x402 protocol](https://newsroom.sgit.ai/economics/rails), which puts a payment inside the HTTP 402 response, moved under the Linux Foundation in April 2026 with more than twenty founding members, settles in about 200 milliseconds, charges no protocol fee, and carried roughly 169 million transactions in its first year. AWS announced agent payments in May 2026 with Coinbase and Stripe, at ticket sizes from a tenth of a cent to a thousand dollars. Cloudflare's Monetization Gateway charges for any resource behind it with the same protocol. The Raywood briefing of June 2025 said *"the infrastructure for micro/nano payments exists – what's been lacking is the industry will to implement it."* It was slightly early. The infrastructure exists now.

Two honest notes belong here. Nothing on sgit.ai is wired to any of those rails today: no fact-creator payment, no per-query billing and no trust-as-a-service product runs anywhere, and [the newsroom site says so](https://newsroom.sgit.ai/shipped) in its own words: *"most of this site is an argument, not a product."* And the Portuguese wallet keeps its ledger in the reader's own browser, as a demonstration of what the reader would see, not as a payment.

## The parts that already run

What does run is the substrate, in public, which is the only reason to believe the rest.

- **[Regulation Graph](../demos/vaults/regulation-graph/index.md).** The EU AI Act parsed from the official Formex XML, hash-verified against the bytes it came from, so that a risk in somebody's register points at a named obligation in a real instrument instead of asserting one. This is the evidence layer of a story vault, done for a law rather than a news event.
- **[A government toolkit, turned into four connected worlds](../demos/vaults/dsit-ai-risk-toolkit/index.md).** Every claim carries the bytes it came from, and every edge declares whether it was curated by a person or merely found by a lexical match. That distinction is the one a newsroom needs between what a journalist established and what a model suggested.
- **[VoiceDebrief](../demos/vaults/voice-debrief/index.md).** Meaning lifted out of unstructured text, from voice notes to Article 9(2) of the EU AI Act, into typed semantic graphs. This article started as a voice memo. That is the first step of the pipeline.
- **[Penetration Test Report](../demos/vaults/pentest-report/index.md).** One engagement, eight audience-specific views, evidence attached to each finding, and a runnable retest per finding. A different industry and exactly the same shape: one graph, many projections, and the evidence travels with each of them. It is what tier four looks like when it is sold.
- **[The Portuguese newsroom](https://pt.newsroom.sgit.ai/newsroom/).** Ninety-two source files frozen and SHA-256 verified, 312 nodes and 942 edges extracted from them, five articles published, and a six-stage pipeline in which the last stage, publication, is the only one no agent can move. A department that writes outside its own folder fails a gate, and that gate, the site says, *"is what makes this a newsroom rather than a script with role names in the comments."* Its English sibling holds 329 nodes and 1,106 typed edges across 64 speakers and 61 organisations, with a named editor of record and no legal review, both stated.
- **[The board](../demos/vaults/board/index.md) and [the catalogue](../demos/vaults/catalogue/index.md).** The site's own task board is a vault and the site renders a snapshot of it; the index of vaults is itself a vault, listed in itself. Both exist to show that the vault is the source of truth and the page is a projection, which is the whole argument in miniature.
- **The cost.** The entire published estate, thirty-one vaults, is 2,662 files and 295 MB of object storage with nothing running between requests, and [the measurements are printed](../demos/fractal-graphs/performance.md). A story vault costs what its files cost. A local newsroom, one journalist, or a blogger with a beat can run this. The model is scale-free, which matters because the industry's crisis is worst at the bottom, and this is the first economic model I know of that works there.

## What I would do, if I ran a publisher

Suggestions, in order, and each one can be started this quarter.

1. **Store the story as a graph from the first day of reporting.** Claims, evidence, sources, raw materials, drafts. Make the article a build artefact. If the graph is not there at the end, the story was never captured, only its words.
2. **Freeze and hash every source.** Not a link. The bytes, on the day, with a hash. It is the cheapest thing on this list and the one that changes what a correction means.
3. **Keep giving the public projection away, and keep it findable.** Static, crawlable, linked to the vault it came from. The door has to be open for the shop to work, and the Cloudflare tolls are for the crawlers that never come through the door.
4. **Sell the licensed article in pence, on demand, and stop treating it as dilutive.** The people who buy it were not going to subscribe. Price the corporate version as a licence to transform, not as a copy.
5. **Adopt the five clauses of the subscription standard before January makes you.** Tell subscribers how much they used, tell them without being asked, let them leave by the route they joined, warn them before charging for something they have stopped using, and do not charge for outages. [The proposal](https://subscriptions.sgit.ai/standard/proposal) is written, and three of its five clauses go beyond what the UK regime will require. A publisher that does this first is the one readers will trust with recurring billing after the others are fined.
6. **Open the evidence vault to the B2B market that already pays for verification.** Lawyers, analysts, regulators, other newsrooms. Redact at the node. License the graph. This is where investigative work gets paid what it cost, and it is the step that makes the loop run.
7. **Put the correction in the graph, and let credibility be computed.** Supersede, never delete. Type the edge. Publish the record, never the verdict. Then a journalist's track record is a fact rather than a reputation, and the verification API has something to warrant.
8. **Publish the cost per story.** If the research, the checking and the translation are lines on a bill, the price of each projection is a decision rather than a guess, and the 60/25/10/5 question becomes answerable.

## The dataset was always worth more than the article

Both clocks are set by somebody else. The search layer will finish ending the deal on its own timetable, and every publisher's answer has been to ask the search layer for a better deal. The legislature will finish regulating the subscription on its timetable, and every publisher's answer has been to ask for more time. Neither answer touches what is actually wrong, which is that the industry sells the one output whose marginal cost has gone to zero and discards the one asset that has not, and that under both of its models the reporting worth paying for is the first thing cut.

The story was always a graph. The article was always a projection. The reader who wanted one article, the firm that needed to forward it, and now the agent that needs to know whether it is true were always standing at the till trying to buy something the shop did not stock: the evidence, the trail, the on-record confirmation, the view cut for them, paid when delivered, in pence, to the people who made the facts. That is the product. It rewards the reporting, it does not depend on search or on renewals, and it has been buildable for eighteen months and sellable for about six.

The industry has never lacked the material. It has lacked a way to sell it. Now it has one, and the only question that matters is the same one I would ask of any product: [when it is taken away, does anybody miss it?](../articles/the-question-is-whether-they-miss-it.md) Take away the words and, increasingly, nobody does. Take away the graph, and the lawyer, the analyst, the regulator and the agent all notice on the same afternoon.

[← All articles](index.md)


---

*[Site index for agents](../llms.txt) · [HTML version](https://sgit.ai/articles/future-of-news-story-vault-not-paywall.html)*
