for agents/llms.txtv0.7.35 · 10 Oct 2026

Home / Articles / If somebody built a company on code review / Versions / v1.2.0

If somebody built a company on code review: what changed in v1.2.0

From v1.1.0 (2026-10-05, 4d6a7b929) to v1.2.0 (2026-10-05, 9e4d1cb45), paragraph by paragraph.

13 paragraphs added, 0 removed, 4 changed in place, 112 unchanged. About 1,508 words added and 96 removed. Insertions are marked like this, deletions like this; unchanged runs are folded to one line; figures appear as their file names.

← v1.1.0 · all versions · v1.3.0 →

# If somebody built a company on code review: thehow grammarI would do it, and why it is fractal,only thenow layers are yours, and the review is the join between what was meant and what was builtpossible

Summary: TheA reader of the code review article drewreplied a careful reply from a reader,with seven questionsgood long,questions, and a voice memo fromof memine answeringanswered them.them Thiswith isa thechange writtenof answer, framed the way the memo framed it:frame: if somebody were building a company on code review, this is how I would do it. The questions were good ones. Could the graph be built from the user stories before the code exists, and the two ends married? Where were the commands and modules layers in the example? Are hand-written stories too high level, and should they be rules and examples instead? Is "a refactor leaves the class shapes still" wrong when an architecture migration keeps only the client interface? Where do deploys and infrastructure go? How do you stop the cost spiking when a repository is onboarded? What is the blast radius compared to, and how long should the report be? The answers turn on onea distinction the first article did not make clearly enough. What is fractal in a fractal semantic graph is the grammar:grammar; nodes, edges, an ontology, a taxonomy and a way to the next altitude. Whichwhich layers exist, what they are called and where they stopexist is decided by each company, its culture, its languages and its stack, and a product that standardises them away loses the thing it was meant to review. Two things make the rest possible only now. One technology can read every layer, from strategy to bytecode, so the graphs can be built at every altitude and built close to reality. And that moves code review from an art of opinion and power to a science of facts, provided the models are used to build, prune and maintain the graphs and then taken out of the line. From there: a projected graph built top down from stories,stories rulesbefore the code exists and examples, a derived graph built bottom up from the syntaxcode, tree, andwith the review as the join; a refactor as relative to the layer youheld holdstill, still;correcting the first article; the deploy as a layer; who reads the code at each stage of evolution, after Wardley; reshaping a change ratherby thanreach; reviewing itbudgets as itthe arrived;objective agood budget per altitudeenough and the five whys as the looploop; thatbehaviour makespolicies for the nextagents reviewdoing cheaper;the work; and open source as the only business model that fits, because the layers are the customer's and the loop has to run where the code is.fits.

3 unchanged paragraphs, under In short

• One technology can read every layer. People moved between strategy and source code by instinct, and it never scaled, because nothing could process every altitude in one pass. A language model can, from a board objective to a syntax tree, in one path or delegated to sub-agents. That is what makes a graph at every layer buildable, and buildable close to reality, and it is why the companies that spent fortunes trying this before could not.

• From art to science. Code review has mostly been an art: opinion, taste, power, and agendas, because there was no fact base to go back to. The graph and the map give one: we are not doing this because the project is at explorer stage; we must do this because it is now a product. The models get review there only if they are used to build, prune and maintain the graphs against reality and then taken out of the line, so that what runs on every change is graphs, deterministic checks and formulas. That is not a lowest common denominator. The better the structure, the faster you can go.

8 unchanged paragraphs

• Constraints keep it in control; the five whys make it cheaper every time. A budget per altitude and per change, decided in advance.advance, which is also the objective way to say good enough. The agents doing the work run under a behaviour policy written before the first run. Onboard a repository by starting at one point and continuing, never by analysing everything. And every finding at one layer is a question about the layer below, which is where the fix goes, as a rule or a graph query, so that the next change like it is caught for the price of a query.

2 unchanged paragraphs

One technology, every layer

Most of what follows is something good engineers have always done by instinct. A senior reviewer reads a diff and sees the story it breaks; an architect looks at a pull request and sees a boundary crossed; a product owner reads a commit message and knows the feature it does not finish. The move between altitudes, from what the business wants to what the code does, has been made in people's heads for as long as there has been code review, sometimes deliberately, mostly by instinct. What has never scaled is doing it explicitly, because nothing could process every layer. The tools that read syntax trees could not read a strategy document, and the people who read strategy documents did not read syntax trees, and the layers in between were held together by meetings.

That is the thing that changed, and it is worth stopping on how strange it is. One technology can now read a board objective, a user story, an architecture diagram, a class, a method body, a syntax tree and, if you want, the bytecode, in the same path or delegated to sub-agents that are the same technology. It is not a different tool at each altitude with a format conversion between them. It is the same reader all the way down, and it can name a node and the verb to the next one at any altitude for a fraction of a cent. That is what makes a graph at every layer buildable. And because the same reader can also be shown the code, the tests and the users, it is what makes the graph buildable in a good way, correct or close to reality, rather than a drawing somebody made in a workshop.

I say this near the top because a great many people will agree with everything below in principle and then say it is ridiculous, that companies have spent millions trying to connect requirements to code and failed, and they will be right about the history. The attempts failed because they needed a different tool and a different team at each layer, and the joins between them were hand-maintained and died. The approach here has one grammar and one reader, it grows along the paths people actually walk rather than by mapping everything, and it is driven by what the change can reach and what the business values, which is a different proposition from the ones that failed.

From art to science

Code review, like security review, has mostly been an art. It was opinionated, and the opinions were not always about the code. A good deal of it was power: who got to say no, whose taste set the style, whose agenda the sprint served, and sometimes a reviewer's quiet wish that the team were working on something else. Some of those judgements were right. There are changes that should not be made because they make no sense in context, and a reviewer who says so is doing their job. But the reason so much of it was politics is that there was no fact base to go back to. You could not say we are doing this because of that, or more usefully, we are not doing this because of that, and point at something.

The map gives the first fact. We are not adding tests to this because it is an explorer project, and that is fine. We have to add them now because it has customers, and it is a villager project. We are stopping at the boundary of this component because it is a commodity and the behaviour at the boundary is what matters. Each of those is a sentence about where the code is, not about who is in the room, and each can be checked against what the code is doing and who depends on it.

The graph gives the rest. A change reaches these commands and these stories, or it does not. The class shapes held, or they did not. A rule holds everywhere, or here are the nodes where it does not. Those are facts, and the point of this whole article is that the models can get review to them, but only if they are used in a particular way: to build the graphs, to review the graphs, to prune and maintain them against reality, to run the feedback loops that correct them, and then to get out of the line. The destination is graphs, deterministic checks, formulas and flows that run on every change without a model in them, which is what a science of code review would look like. A model in the line on every change is the art again with a different reviewer.

People sometimes hear that as a race to the lowest common denominator, as if science meant fewer judgements and worse code. I have found the opposite, every time. The better the structure, the better the principles, the better the components and the more commoditised the parts, the faster a team can go, because it can react, because it can see, and because it can stop. Budgets are the clearest case. Being able to say to a team that the budget for this is spent was one of the most useful controls I ever had, because teams are very good at keeping going. Engineering has a bad name for over-engineering, and some of it is deserved, but a lot of what gets called over-engineering is just work that should have stopped at the point of diminishing returns and had no objective way to know it had reached it. Code review is the same: technically you can keep reviewing forever. A budget per altitude is an objective good enough, and the five whys are how the good enough moves, one rule at a time, in the right direction.

And it makes code review accountable, which is what it should always have been. Code review when it works is one of the most enjoyable things in software. People like doing it, it adds value you can see, and it gives you the confidence to ship, which is the confidence not to slow down. The test of it is also plain: a review that works is not contradicted when the change hits production. That contradiction, or its absence, is the feedback loop, and it is the thing that separates what can be automated from what cannot. What a check caught, or should have caught, becomes a rule. What only a person saw is the judgement that stays with people, and the review product's job is to deliver that judgement the fifty lines it needs and not the ten thousand it does not.

57 unchanged paragraphs, under The grammar is fractal, the layers are yours, Build the graph from the top, before the code exists, The layers the example had and the text did not walk…

The budgets live in the same place as the rest of the agents' rules. Every agent doing this work, the one deriving a layer, the one proposing stories, the one summarising a change, runs under an Agent Behaviour Policy written before its first run, which is how the agents behind this site are run already. The policy says which repositories the agent may read, which actions it may take, where it may write, what it may spend and where it is to focus, and it is the control on the agents' own blast radius: a reviewing agent that can read one repository and write only to a vault of findings can leak nothing from the next customer and can merge nothing on its own. Review of the code and control of the reviewers are the same shape, which is as it should be, and the footprint an agent leaves is read against its policy the way a change is read against its intent.

21 unchanged paragraphs, under What to compare the blast radius to, and how long the report should be, The company, What the first article got wrong, and what this one leaves open…

• Footprint and blast radius, the same measure applied to agents rather than code.code, and Six agents, one inbox, the behaviour policies the reviewing agents would run under.

21 unchanged paragraphs, under Sources

← v1.1.0 · all versions · v1.3.0 →