for agents/llms.txtv0.6.75 · 6 Oct 2026

Home / Articles / If somebody built a company on code review: how I would do it, and why it is only now possible

If somebody built a company on code review: how I would do it, and why it is only now possible

By · 2026-10-05 · updated 2026-10-06 · v0.6.74 · code-reviewfractal-semantic-graphsgraphswardley-mapsagent-behaviour-policiesexplorer-villager-town-plannerexample-mappingrefactoringblast-radiusfive-whysopen-sourcestartupsai-generated-codearticle

Abstract: A reader of the code review article replied with seven good questions, and a voice memo of mine answered them with a change of frame: if somebody were building a company on code review, this is how I would do it. The answers turn on a distinction the first article did not make clearly enough. What is fractal in a fractal semantic graph is the grammar; which layers exist is decided by each company, and a product that standardises them away loses the thing it was meant to review. Two things make the rest possible only now. One technology can read every layer, from strategy to bytecode, so the graphs can be built at every altitude and built close to reality. And that moves code review from an art of opinion and power to a science of facts, provided the models are used to build, prune and maintain the graphs and then taken out of the line. From there: a projected graph from stories before the code exists and a derived graph from the code, with the review as the join; a refactor as relative to the layer held still, correcting the first article; the deploy as a layer; who reads the code at each stage of evolution, after Wardley; reshaping a change by reach; budgets as the objective good enough and the five whys as the loop; behaviour policies for the agents doing the work; and open source as the only model that fits.

What is fractal and what is not. The constant across every codebase is the grammar: nodes, edges, an ontology, a taxonomy and a way to zoom in and out. Which layers exist, what they are called, which languages fill the bottom layers and what the business counts as a change that matters are decided by the company you are in. The first article's seven altitudes were one Python repository's layers, not the layers.
Where this comes from, and what is behind it. A reader of Code review as a fractal semantic graph replied with seven questions, each of them the kind that improves the thing it questions, and a voice memo of mine on 5 October 2026 answered them with a change of frame: not what the graph is, but what you would build if the graph were the company. Behind both sits a run of earlier writing this article quotes with its dates: Who should be reading the code (March 2026), How vibe coded apps actually ship to production (May 2026), the five whys as a translator between domains (June 2026), the open source strategy briefs (June and July 2026), the serialised pull request brief (August 2026), and the architecture brief that defines fractal as a precise claim (July 2026). All of those live in the public SG/Send team record on GitHub, as markdown. The reader is not named here because the questions, not the questioner, are the subject; they know who they are and they have my thanks. Where this article corrects the first one, it says so.

In short

One technology, every layer

Most of what follows is something good engineers have always done by instinct. A senior reviewer reads a diff and sees the story it breaks; an architect looks at a pull request and sees a boundary crossed; a product owner reads a commit message and knows the feature it does not finish. The move between altitudes, from what the business wants to what the code does, has been made in people's heads for as long as there has been code review, sometimes deliberately, mostly by instinct. What did not scale was doing it explicitly, because no one tool could process every layer. The tools that read syntax trees could not read a strategy document, and the people who read strategy documents did not read syntax trees, and the layers in between were held together by meetings.

That is the thing that changed, and it is worth stopping on how strange it is. One technology can now read a board objective, a user story, an architecture diagram, a class, a method body, a syntax tree and, if you want, the bytecode, in the same path or delegated to sub-agents that are the same technology. It is not a different tool at each altitude with a format conversion between them. It is the same reader all the way down, and it can name a node and the verb to the next one at any altitude for a fraction of a cent. That is what makes a graph at every layer buildable. And because the same reader can also be shown the code, the tests and the users, it is what makes the graph buildable in a good way, correct or close to reality, rather than a drawing somebody made in a workshop.

I say this near the top because a great many people will agree with everything below in principle and then say it is ridiculous, that companies have spent millions trying to connect requirements to code and failed, and they will be right about the history. The attempts failed because they needed a different tool and a different team at each layer, and the joins between them were hand-maintained and died. The approach here has one grammar and one reader, it grows along the paths people actually walk rather than by mapping everything, and it is driven by what the change can reach and what the business values, which is a different proposition from the ones that failed.

From art to science

Code review, like security review, has mostly been an art. It was opinionated, and the opinions were not always about the code. A good deal of it was power: who got to say no, whose taste set the style, whose agenda the sprint served, and sometimes a reviewer's quiet wish that the team were working on something else. Some of those judgements were right. There are changes that should not be made because they make no sense in context, and a reviewer who says so is doing their job. But the reason so much of it was politics is that there was no fact base to go back to. You could not say we are doing this because of that, or more usefully, we are not doing this because of that, and point at something.

The map gives the first fact. We are not adding tests to this because it is an explorer project, and that is fine. We have to add them now because it has customers, and it is a villager project. We are stopping at the boundary of this component because it is a commodity and the behaviour at the boundary is what matters. Each of those is a sentence about where the code is, not about who is in the room, and each can be checked against what the code is doing and who depends on it.

The graph gives the rest. A change reaches these commands and these stories, or it does not. The class shapes held, or they did not. A rule holds everywhere, or here are the nodes where it does not. Those are facts, and the point of this whole article is that the models can get review to them, but only if they are used in a particular way: to build the graphs, to review the graphs, to prune and maintain them against reality, to run the feedback loops that correct them, and then to get out of the line. The destination is graphs, deterministic checks, formulas and flows that run on every change without a model in them, which is what a science of code review would look like. A model in the line on every change is the art again with a different reviewer.

People sometimes hear that as a race to the lowest common denominator, as if science meant fewer judgements and worse code. I have found the opposite, every time. The better the structure, the better the principles, the better the components and the more commoditised the parts, the faster a team can go, because it can react, because it can see, and because it can stop. Budgets are the clearest case. Being able to say to a team that the budget for this is spent was one of the most useful controls I ever had, because teams are very good at keeping going. Engineering has a bad name for over-engineering, and some of it is deserved, but a lot of what gets called over-engineering is just work that should have stopped at the point of diminishing returns and had no objective way to know it had reached it. Code review is the same: technically you can keep reviewing forever. A budget per altitude is an objective good enough, and the five whys are how the good enough moves, one rule at a time, in the right direction.

And it makes code review accountable, which is what it should always have been. Code review when it works is one of the most enjoyable things in software. People like doing it, it adds value you can see, and it gives you the confidence to ship, which is the confidence not to slow down. The test of it is also plain: a review that works is not contradicted when the change hits production. That contradiction, or its absence, is the feedback loop, and it is the thing that separates what can be automated from what cannot. What a check caught, or should have caught, becomes a rule. What only a person saw is the judgement that stays with people, and the review product's job is to deliver that judgement the fifty lines it needs and not the ten thousand it does not.

The grammar is fractal, the layers are yours

The word fractal gets read as "there are many levels", and that reading produces the wrong product. The precise claim is the one the architecture brief for the Send platform made in July 2026: "Fractal is a precise claim, not a decoration. It means four things. Self-similarity: the same node and edge grammar describes an entity, a message, a conversation, a customer, and a tenant. Scale invariance: the same validators, the same query engine, and the same visualisation tools work unchanged at every altitude. Composition: a graph can be a node in another graph, which is the graphs-of-graphs property the corpus already runs on. Recursion: zoom into any node and it expands into a graph obeying the identical rules, with no new format and no special case."

Notice what is in that list and what is not. The grammar is in it: nodes with a type and a name, edges with a verb and a named inverse, an ontology that says what the types and verbs mean at this altitude, a taxonomy that says how the nodes group, and the move in and out. The layers are not in it. Which layers exist, what they are called, which languages fill the bottom layers and what counts as a change that matters are not properties of the graph. They are properties of the company you are in.

That is the answer to the reader's first instinct and to mine. The first article drew seven altitudes for one Python repository: stories, commands, packages, classes, methods and calls, syntax, and the machine below. Those were that repository's layers. A team of three with a Flask app and a bank with four hundred services in six languages do not have the same layers, and should not be made to. One calls them epics and stories, another calls them services and bounded contexts, a third has infrastructure as a layer that the first two do not have at all. The programming languages decide the bottom layers, and the languages people speak decide the top ones. The culture names them. The objectives decide which layer counts.

So if somebody were building a company on this, the first design decision would be to embrace the differences. Every development team has its own logic, its own graphs and its own altitudes, and that is what makes the company what it is. A code review product that tries to standardise the layers away loses the thing it was meant to review. The generic parts are the grammar, the tools that work on any graph in that grammar, and a handful of workflows. The rest is the customer's, and the product's job is to learn it, from the stories the customer writes, the names in their code, the shape of their deploys, and to hold it as the customer's configuration, not the vendor's opinion.

Build the graph from the top, before the code exists

The reader's first question was the best one. Could the graph be projected from the user stories, before any code exists, and later married to the graph derived from the code? And would the naming system then surface the stories nobody wrote?

Yes, and this is how I would run a project from the start. The order I use for anything I build now is: brief, then plan, then architecture, then a review of the architecture, then a simulation of it with synthetic users, and the code last, when the shape has survived everything that can be thrown at it without code. Each of those steps produces a graph in the same grammar. The brief names objectives. The plan names stories, and each story names its rules and the examples that make the rules concrete. The architecture names features, flows, the commands and surfaces the stories will touch, and the components they will need, sketched as far down as design can honestly see and no further. Below that the projection stops, and it should say so, because everything under the lowest honest layer is wishful thinking until a parser can read it.

Two graphs, one grammar. The projected graph is built top down before the code exists: objectives, stories with their rules and examples, features and flows, the commands and surfaces they will touch, and the components as far as design can see. The derived graph is built bottom up from the syntax tree. The review is the join, and the join has three outcomes.

The code, when it exists, derives its own graph upward from the syntax tree: every node of every file fingerprinted, methods and the calls between them resolved, the classes and their fields, the modules and packages as the architecture was actually built, the commands and surfaces the code actually exposes, and at the top a layer of stories that a model proposes from the code and marks as proposed.

The review is the join between those two graphs, and the join has exactly three outcomes, each a review finding of a different kind.

A derived surface that serves a projected story is a match. The change is doing what was meant. Progress on a big feature is the count of projected nodes that now have a derived match, which answers the reader's question about small increments on a big feature: an increment that moves no user-visible needle still moves that count, and the review can say so.

A derived node that no projected story asked for is "derived, not projected". It is a bug, or a missed implication of the design, or a story nobody wrote down because it was obvious to the person who built it. Each is a finding; a person decides which. This is where the naming system surfaces the stories the reader asked about. When the model proposes a story from the code and nothing in the plan matches it, the proposal is the question "did you mean this?", and the answer goes into the plan, as a new story or a deleted feature.

A projected story with no derived node under it is "projected, not derived". Either it is not built yet, which is the plan, or it was forgotten, which is the review. The age of the gap says which.

There is a version of this already running in the way this site and the Send platform are built, and it is worth naming because it is the top half without the bottom half. The Librarian role keeps a set of reality files whose rule is blunt: "If it's not in a domain index, it does not exist. No agent may claim a feature is 'working' or 'shipped' unless it appears in the appropriate domain's EXISTS section." The reality files are the projected graph being reconciled with the code by hand, by an agent reading briefs and checking the repository, because, as the same file explains, "Agents were confusing ideas described in briefs with features that actually exist in code." A code review company would make that reconciliation the product: the derived graph as the EXISTS section, computed, and the projected graph as the briefs, and the diff between them as the review.

The layers the example had and the text did not walk

The reader asked where the commands and modules layers were in the worked example, and whether the review skipped them. The honest answer is that the delta had them and the prose did not visit them.

The commit in the vault changed seven methods in two classes, which live in two modules, the command-line vault module and the automatic transport module, in two packages, the command-line package and the network package. Climbing the call graph from the changed methods reached nine of the seventy-two commands and six of the eleven stories. The delta figure shows every one of those layers in red, including the commands and the modules. The text took the two ends, methods and stories, because the shape of a fix is clearest there, and left the middle layers to the picture.

That is a writing fault and the vault is unchanged by it, but it also hides a point that a product would need. The middle layers are where the review's audience changes. Methods are read by the author and whoever owns the file. Modules and packages are read by whoever owns the architecture, because that is the layer where a change can cross a boundary it should not. Commands and surfaces are read by whoever owns the users, because that is where behaviour becomes visible. A review that visits every layer is also a review that routes each layer to the person who can judge it, and the middle layers are the ones with the fewest readers today, which is one reason architecture drifts.

Stories as rules and examples

The reader also thought the stories were too high level, and that hand-writing them in prose was the weak link, and pointed at Example Mapping. They were right on both.

Example Mapping, as Matt Wynne described it in December 2015, is a conversation that produces four kinds of card: the story, the rules that constrain it, the examples that make each rule concrete, and the questions nobody in the room can answer. It was designed as a way to run a conversation in twenty-five minutes, and it has the property the graph needs, which is that the story is not a paragraph. It is a small tree. The story is the root, the rules are its children, and the examples are the leaves, and a leaf is a thing a test or a run can satisfy.

In the graph that is the top layer done properly. A story is a node whose children are rules, each rule a node whose children are examples, and each example a node with an edge to the test that exercises it or the recorded run that passed it. "Hand-written" then stops being a weakness, because what is written by hand is the smallest unit that can be checked. The question cards are the "derived, not projected" and "projected, not derived" findings before the code exists: things the conversation could not settle, held in the graph as questions rather than lost in a meeting.

The worked example's eleven prose stories were the right first layer and the wrong final one. The right final one is rules and examples that tests can be matched to by execution, not by a model's guess at which story a test belongs to. That is on the vault's own list of what does not exist yet, and it moves up the list.

A refactor is relative to a layer

The reader challenged the first article's sentence that "a refactor should change the method and syntax layers and leave the class shapes, the commands and the stories untouched", with the case of an architecture migration: a monolith split into services, or the reverse, where everything below the client interface is new and only the interface holds. By the article's test that is not a refactor. By any working engineer's test it is.

The reader is right and the sentence was too narrow. It described one refactor at one altitude, the tidy-a-method refactor the vault's example happens to be, and stated it as the rule. The rule is this: a refactor is any change in which some layers move and one chosen layer does not, and the chosen layer is the one with the tests or the contract. Which layer that is depends on what you are refactoring.

Three changes, three held layers. Tidying a method holds the class shapes and is proved by unit tests on them. Migrating an architecture holds the client interface and is proved by contract tests on it, while everything below is new. A change that holds nothing is not a refactor, whatever the commit message says, and the review should treat it as a fix or a feature.

Tidy a method and the held layer is the class shape: same fields, same signatures, bodies rewritten, syntax tree all new, unit tests on the methods prove it. Migrate the architecture and the held layer is the client interface: every call and every response the same, services and classes and everything below them new, contract tests on the interface prove it. Change what the software does and no layer is held; a story now holds that did not, a command's behaviour changed, and however the commit is labelled the review should treat it as a fix or a feature.

This also answers the team's own practice, which the first article's sentence would have contradicted. The brief that created the Villager team in March 2026 lists, under refactoring, "Move code to better locations (file structure, module boundaries)" and "Create classes and abstraction layers", and the Villager's working agreement is "Never change public behavior during a refactoring, only internal structure." Those refactors move the class and module layers and hold the behaviour layer. They are refactors held at a higher altitude, which the general rule allows and the narrow sentence did not.

What the review checks, then, is not whether the refactoring strategy was the right one, which is an argument, but whether the layer the author said they were holding actually held, which is a fact the graph and the tests can state. The blast radius of a refactor is the set of layers that moved, and the claim to check is the one layer that was supposed not to.

Down to the deploy

The reader asked where deploys and infrastructure go, and whether a blast radius at the infrastructure level and a graph of the non-functional requirements belong in the same picture.

They are layers. A deployment pipeline, the environments it deploys to, the services running in each, the configuration they read and the secrets they hold are a graph with its own vocabulary, and the vocabulary changes at that altitude as sharply as it does between types and syntax. The nodes are things like a pipeline stage, an environment, a running service, a bucket, a role; the verbs are things like deploys-to, reads-from, assumes, exposes. A change's blast radius does not stop at the command that can reach it. It continues through the deploy graph to what is running where, and a change to a configuration file with no code in it can have a blast radius wider than any method.

The non-functional requirements live at that altitude and above it. Performance, availability, cost, security posture and recoverability are properties of the deployed system, not of the code alone, and the sibling site nfrs.sgit.ai is this site's attempt to hold its own non-functional requirements as measured claims rather than restated ones. The rule there is "LINK THE MEASUREMENT, NEVER RESTATE IT. GENERATE OR DATE EVERY NUMBER." A review that reads a change at the deploy layer asks which measured claims it can reach, and the Cartographer role in the Send team puts the standard for that layer in one sentence: "If a dependency, data flow, or security boundary exists but is not visible on a map, the Cartographer has failed."

I would add that the deploy graph is where the explorer-to-villager handover becomes visible, because explorer code runs in an explorer environment and villager code runs in a tested one, and the deploy edge from one to the other is a review in itself.

Who reads the code

The reader's questions about depth and cost have one answer underneath them, and it is Simon Wardley's. The review a change deserves depends on where the thing it changes is on the map: being discovered, being made into a product, or running as a commodity. Wardley's names for the people who do each are explorers, villagers and town planners. He had called them pioneers, settlers and town planners in 2015 and asked people to use the new words in January 2023, writing later that "the colonialist overtone of those terms made it problematic for many and with good reason". I have used the new names since a June 2025 piece on the generative AI divide, for the split between the people who find things, the people who make them solid and the people who run them at scale.

In March 2026 I wrote an article titled Who should be reading the code, and it is the frame this company would be built on. It opens with a caveat worth keeping: "Before AI agents started writing code, you could argue that 90% of the code running in production was not being read by anyone." Then it answers by phase. "So who reads Explorer code? Mostly other agents. Agents that do security scans, basic quality checks, dependency audits. These are imperfect reviewers, but they catch the obvious problems early." Villager code is where reading begins: "This is where code needs to be read. Not glanced at. Read. Understood. Questioned. Refactored." And town planners "really read it. They debate, question, and verify at an architectural level."

Code review by stage of evolution. For explorer code, review is for speed and feedback and the person reads the behaviour, not the lines. For villager code it is for quality and reachability, all the way down, because this is where changes land. For town planner code it is for guarantees at the boundary, read by deterministic checks with a person for the exceptions. The company sits in the two moves between columns.

The graph changes what each of those sentences costs. For explorer code, map the stories, the commands and a sketch of the components, enough to see the shape, and stop; the model can propose everywhere here because nothing depends on the code yet. For villager code, map all the way down, because this is where changes reach customers and where the blast radius lives; humans read the contracts, the class shapes and the call paths that changed, and agents read the rest. For town planner code, stop at the boundary. Map a commodity's behaviour, not its insides; the checks are graph queries and they have to be deterministic, and a person reads the exceptions. A model has no place in line on a mission-critical path, and the skill lifecycle brief from June 2026 says why in the voice of the memo it came from: "Anybody who spends a lot of money on tokens has an engineering problem. They are using explorer-type code and solutions in a commodity environment."

This is where the company sits. In the two moves. From explorer to villager, when a guess becomes a product and needs a review it never had, and the graph has to be built for the first time from code that was written to be thrown away. From villager to town planner, when a product becomes a commodity and the review becomes a set of checks that run without anyone, and the graph has to be good enough to be the checks. The trap on both sides is the same, and the May 2026 article on how vibe coded apps actually ship names it: teams discover that a model's mistakes are too risky and swing to the other extreme, reviewing everything by hand, which the volume of generated code makes impractical, and both extremes come from not knowing which stage the code is at.

Review is not saying no

A pull request that mixes ten thousand lines of generated documentation with fifty lines of mission-critical logic cannot be reviewed well by anyone, and the reviewer's attention is spread across all of it. The code review products I know treat the pull request as the unit and try to make the reviewer faster at it. I would do the opposite. The review tool's first job is to the change, not to the reviewer: reshape it before anyone reads it.

Reshape the change before you review it. A pull request as it arrived, and the same change split by what it can reach: the part that provably reaches no behaviour, the part that reaches behaviour through tested paths, and the part that reaches a mission-critical path or a layer with no test holding it. Review is sorting what can go fast from what must go slow. Counts are illustrative.

With the graph a change can be split by reach, and each part gets the review it deserves. The part that provably reaches no behaviour, documentation, generated code, formatting, a rename whose every caller moved with it, is merged by a deterministic check, with no model in line and no human. Most of the volume lives here. The part that reaches behaviour through paths that have tests, or inside a layer whose contract is held still, gets the checks run and a model's one-line summary of what moved against what was intended, and a person glances. The part that reaches a mission-critical path, a story customers depend on, or a layer with no test holding it gets people, the right people, with the right context: the architect for the component layer, the product owner for the story, the author for the method stream.

Review, in this shape, is not saying no. It is sorting what can go fast from what must go slow, and giving the slow part to people who have the context to decide well. The faster the first class moves, the more attention the third class gets. I wrote a version of this in the SecDevOps Risk Workflow book a decade ago, from the security side: "When you do a code review, you tend to visualize a slice of a model of the application. Your focus is fixed entirely on the problem at hand", so "you need systems that can flag when something is a problem, or needs to be reviewed." The graph is that system, and the slice it hands the reviewer is the method stream. And the thing that makes the whole arrangement safe is the thing the March article on the Villager phase recorded about a real review: a bug in which a model had invented an API call "would not have been caught by tests alone, because the tests would have been written against the same hallucinated API". It was caught by a person reading the code. The graph's job is to make sure that person is reading the fifty lines and not the ten thousand.

The brief on the serialised pull request from August 2026 has the other half of this. "A serialised diff is not a missing pull request. It is a different and, for agent collaboration, better one." The review surface does not have to be a hosted page with an approve button. It can be a vault holding the change and the graph of what it reaches, read by a person on another machine with a read key, which is how the patches in that Villager phase actually moved, sixteen of them, as encrypted vaults.

Replicate reality first

Everything above is reading. The reader's questions about depth, cost and what to compare against need the other half, which is what makes the graph trustworthy, and the first article's answer stands: not accuracy, reality. A graph nobody can correct is a drawing.

The order of operations I would build the company on is to replicate reality before having an opinion about it. Derive from the code what a parser can derive, deterministically, with no model in line, so that every node has a file and a line and every number is checkable. This is the rule I set out in July 2025 for every model-driven solution, "use LLMs to figure out the solution, then use code to execute the solution thereafter", and in a graph the rules that matter are plain to execute. The 2025 piece on semantic knowledge graphs for source code analysis gave the example: "we check that every endpoint function node has an edge to an Auth check node; the ones that don't are instantly known." Let a model propose what a parser cannot see, the story a command serves, the intent behind a module, the edge a dynamic dispatch hides, and mark every proposal as proposed. Then confront the graph with the three things that can correct it: users, who confirm whether a story holds; experts, who each read the layer they know; and tests, mapped from imports and confirmed by execution. Every correction is a commit to the graph files. Only then is the conversation about the code a conversation about facts, and the memo put it as "replicate reality first, then have a fact-based conversation".

Synthetic users are part of that, and earlier than most teams put them. The architecture review and the simulation come before the code in the order above, and the simulation is a run of synthetic users through the projected graph: characters, not scripts, as a brief from February 2026 insisted, "with realistic behaviour patterns, including getting lost, clicking wrong things, and taking happy paths." This site ran that against itself in September and recorded forty-three steps, fifteen questions the site did not answer and ten confusions, with the rule that "a run with no confusion anywhere is treated as a run done badly." Run against a projected graph before code exists, the same synthetic users find the stories that contradict each other and the flows that go nowhere, which is the cheapest possible time to find them.

Time is the other dimension, and the graph has to carry it. Every derived layer, every proposal, every correction and every run is a version, so the review can go back to the graph as it was when a decision was taken and ask what the change reached then. That is the twin's property: a journal plus a replay. The one-shot approach to working with models that the Send team uses is built on it, "curate reality, send it all at once, get the answer, then use the answer to improve the reality for the next prompt", and the bundles it sends are "versioned. They're replayable. I can load a previous bundle, change one element, and re-execute." A code review company would do that for the graph: load the repository's reality as it was, change one node, replay the review.

Constraints keep it in control, the five whys make it cheaper

The reader's practical worry was cost: that onboarding a repository would spike it, that the depth had to be tunable, and that a graph of everything is a graph nobody can afford. The worry is right and the answer is in two parts.

Two mechanisms. Left, a budget per altitude and per change, decided in advance: deterministic checks always run and cost compute, the delta re-derives only what the change touched, a model summarising has a small budget and a hard stop, a model proposing has a larger one and every output marked proposed, onboarding is budgeted per slice, and work handed from agent to agent carries its budget with it. Right, the five whys as the loop: a bug at one layer is a question about the layer below, and the fix is a rule there, so the next change like it is caught for the price of a query. Budgets are illustrative.

The first part is constraints, decided before anything runs. Agents keep going; that is what they do, and one agent handing work to another looks efficient while it spends. The way to stay on the right side of diminishing returns is a budget per altitude and per change, declared in advance, which is the same move the Send team's workflow brief made in August 2026: "a per-step ceiling declared before execution" that "makes a workflow's maximum cost knowable before it runs rather than discovered afterwards." Deterministic checks over the derived layers cost compute and no tokens, and always run. The delta re-derives the files the change touched and what depends on them, never the repository. A model summarising what moved against what was intended gets a small budget and a hard stop. A model proposing edges and stories gets a larger budget and every output marked proposed. Work handed from agent to agent carries its budget with it and never widens it, and the cost of a change becomes one more thing the graph can show about it.

The budgets live in the same place as the rest of the agents' rules. Every agent doing this work, the one deriving a layer, the one proposing stories, the one summarising a change, runs under an Agent Behaviour Policy written before its first run, which is how the agents behind this site are run already. The policy says which repositories the agent may read, which actions it may take, where it may write, what it may spend and where it is to focus, and it is the control on the agents' own blast radius: a reviewing agent that can read one repository and write only to a vault of findings can leak nothing from the next customer and can merge nothing on its own. Review of the code and control of the reviewers are the same shape, which is as it should be, and the footprint an agent leaves is read against its policy the way a change is read against its intent.

Onboarding is where the reader expected the spike, and the way to avoid it is to be opportunistic. Do not analyse the repository. Start at one point, the change in front of you, the command a customer is complaining about, the module somebody is about to touch, and continue from there. Derive the files that point touches and what they reach, and stop. The next change extends the graph a little further, and the graph grows along the paths people actually walk, which are the paths worth having. The first slice is dear and every later one is cheaper, because the layers it needs are increasingly already there. The depth is tunable by the same means: how far up and down from the change the derivation runs is a number, per altitude, and the villager column of the map above says where it should be high and the town planner column says where it should stop.

The second part is what makes the whole thing cheaper every time, and it is the five whys. Every finding at one layer is a question about the layer below. Why did the story stop holding? A command's behaviour changed. Why did the command change? A method on its path was edited to satisfy a test elsewhere. Why was the reach not seen? The call went through a base class and the review read the diff, not the graph. Why did no check catch it? There was no rule that a change reaching a customer-facing command needs a human. So the fix is below the bug: add the rule, as a graph query, and the next change like it is caught for the price of a query, not a review. I wrote about the five whys as a translator between domains in June 2026, and the point there holds here: the chain "is just as powerful aimed downward, into causes", capturing "the second, third, and fourth stories". The March article on who reads the code said the same thing about process: when an issue reaches the town planner that the villager phase should have caught, "that is an incident. Why did this reach the Town Planner? What was missing in the Villager's review process? Fix the process, not just the code."

This is the feedback loop, and it is the opposite of whack-a-mole. Whack-a-mole fixes the bug at the layer where it showed. The loop fixes the layer below, and the automated checks grow one rule at a time, each rule the residue of a real finding, until the review of a commodity is almost entirely deterministic and the model is left with the part that needs judgement. It also turns review into graph theory, which is where I would want it: reach, cut sets, invariants and cycles are properties the graph can compute, and the more of the review that is expressed over the graph, the more of it is deterministic.

And it is why one line can matter more than fifteen thousand. Once the layers exist, a review is about the delta, one commit against the graph. One line that changes a base-class method every command passes through has a reach that fills the ladder. Fifteen thousand lines of generated documentation have none. Without the graph those two changes look like what they are in the diff, small and large. With it they look like what they are in the system, and the review spends its attention where the reach is.

What to compare the blast radius to, and how long the report should be

The reader's last questions were the ones a product manager would ask. What is the blast radius compared against? How do you keep the report short without losing the intent? How do you use the documentation engineers already write? How do you measure small increments on a big feature?

Compare the blast radius to the intent, not to the old code. The old code is the baseline the diff already gives you. The intent is the projected graph, the commit message, the pull request description, the story the change claims to serve, and the documentation that already exists, which is the projected graph written in prose and should be parsed into it rather than ignored. The vault's delta view already puts the commit message beside the graph "as evidence to be checked against the graph rather than taken on trust", and that is the pattern: the author's claim about reach on one side, the computed reach on the other, and the review is the difference.

The report is one line per layer that moved, saying whether the move matches the intent, and nothing for the layers that did not. A tidy-a-method refactor reports two lines: methods moved, syntax moved, and the class shapes held as claimed. A fix reports the layers it climbed to the story that now holds, and whether a test was added. A change that reached a layer its description did not mention reports that layer in a different colour. The intent is kept not by writing more but by having the projected graph to compare against, so that "matches the intent" is a computed statement rather than a reviewer's paraphrase.

And small increments on a big feature are measured as the count of projected nodes that now have a derived match. A feature is a subgraph of the projected graph. Each commit either adds a derived node under one of its projected nodes or does not. That count moves before any customer sees anything, which is what lets a team see progress on a six-week feature in the second week and see that it has stalled in the fourth.

The company

If somebody built a company on this, the business model is the part I am least in doubt about, because the layers settle it. The layers are the customer's. The loop has to run where the code is, inside the customer's repositories and pipelines, reading their stories and their deploys. A company that held the graph in its own cloud and locked customers into its layers would be selling the one thing it has no right to own, and the moment the customer's layers diverged from the vendor's, which they would in the first month, the product would be fighting the customer.

So the only model I would consider is open source, in the way the Send team's strategy briefs have put it this year. "Everything the company does, code, logic, functionality, and schemas, is open source, with no proprietary core and no bait model where the best features are locked away, and the only line is at the customer boundary." The commercial hook is the customised deployment: the layers configured for this company's culture, languages and stack; the loop run and tuned for them; the rules that accumulate from their findings; and the record, versioned, that shows what every change reached and who decided what. The brief from June 2026 says why that is a business rather than a service: "one model is that we host customisations and mass-customised deployments, because it is cheaper for the customer to pay us than to run it themselves." And the July brief says what open source does to the moat: it "removes technology as the moat, and relocates the lock-in somewhere more durable and more honest". The sibling site open-source.sgit.ai says where that is: "the lock-in is in the quality and the services and the new versions and the maintainability and the certification", and puts the offer to a customer in one line, "You can take it. It will cost you that. Or you can pay me, and have it next week." For a code review company the thing they can take is the grammar and the tools, and the thing they pay for is the layers configured, the loop running and the rules accumulated for them.

What the company would not sell is the parser, which is a commodity, or the model, which is somebody else's, or the grammar, which has to be shared for the graphs of different companies to be comparable at all. What it would sell is the two moves on the map, done well and repeatedly: taking a guess that became a product and giving it the review it never had, and taking a product that became a commodity and turning its review into checks that run without anyone. Both moves produce rules, and the rules are the asset, and they belong to the customer, and the customer pays to have them made.

What the first article got wrong, and what this one leaves open

Three corrections to the first article, so they are in one place. The sentence about a refactor leaving the class shapes still was one refactor at one altitude stated as the rule; the rule is the held layer. The seven altitudes were one repository's layers presented as if they were the layers; the fractal part is the grammar. And the text of the worked example climbed from methods to stories without walking the commands and modules in between, which the delta had and the figure showed.

Two things the vault page said and the article did not, now reconciled. The vault page describes the fix shape as "bodies moved and shapes still and a new assertion" and calls a refactor "its mirror image"; under the held-layer rule both are changes that hold the class shapes, and the difference is that the fix adds an assertion and moves a story, while the refactor adds none and moves none. The page's wording stays, with this paragraph as the gloss.

What is open. The projected graph has not been built for a real project from a brief forward; this site's own briefs and reality files are the nearest thing and they are reconciled by hand. Rules and examples as the top layer, matched to tests by execution, is on the vault's list and has moved up it. The deploy layer has not been derived for any repository here. And the thing the first article ended on is still true: the review itself, the thing that sits on a change and shows it at every altitude, does not exist yet. What exists is a vault with the layers, a reader's seven questions, and this answer to them. The day after this was published, the requirement for it was written down as a build brief, Code review graphs in the repository: a folder a repository carries rather than a vault, full coverage from the first commit for new projects and delta first for existing ones, a page from story to source line and back, and the tool reviewed by a second copy of itself.

The article as one page, by another model

As with the two articles before it, Dinis fed the published text to ChatGPT and asked for an infographic, and it is below after a check against the article. The layers, the before and after, the grammar and layers distinction, the projected and derived graphs with the three outcomes of the join, the held layer, the deploy, the three stages, the sorting by reach, the budgets, the five whys and the company are all as the article has them, and nothing was invented. Two lines in quotation marks, "same reader, every altitude" and "standardise the grammar, not the customer's worldview", are the other model's condensations of sentences here rather than sentences from here; they are fair ones.

The article as a one-page state, generated with ChatGPT from the published text on 5 October 2026 and credited by it to the source. The style and the slogans in the margins are the other model's, kept as they came, because who made a condensation is part of its provenance.

Threads woven here

Sources

Threads

Graphs & knowledgeStartups & strategy This article as a graph →

Builds on

Continued by

All articles · All graphs

← All articles