for agents/llms.txtv0.6.54 · 3 Oct 2026

Home / Articles / Code review as a fractal semantic graph: source code is already one, and the review should read every layer of it

Code review as a fractal semantic graph: source code is already one, and the review should read every layer of it

By · 2026-10-03 · v0.6.52 · code-reviewfractal-semantic-graphsgraphsstatic-analysisastmethod-streamso2-platformai-generated-codevaultsc4gherkinthreat-modellingnfrsarticle

Abstract: Source code is a very good example of a fractal semantic graph. It has layers within layers, and each one is a graph with its own vocabulary: what the user is trying to do, the features and flows, the components people draw as architecture, the classes, the methods and the calls between them, the syntax tree, and on down to the machine code if you want it. C4 saw the layers and stopped at four; Gherkin got the top layer into a shape people could write and then glued it to the code with regular expressions. What changed is that naming a node and the verb to the next one is now cheap, because a language model can do it from the syntax tree, once per change, and write the result as files. This article argues that code review should read a change at every one of those layers, and that two things fall out when it does: a refactor is a change that moves the bottom layers and leaves the top ones still, and a bug fix is a change that is visible at the top as a story that now holds. It revisits method streams, the review technique from the OWASP O2 Platform in 2012, as one script over a syntax tree with resolved calls. It comes with a worked example published as a vault: the sgit command-line tool, 377 classes and 1,111 methods, read as layered graphs with nothing run, including one real commit read upwards from the seven methods it changed to the nine commands and six user stories it can reach. And it says what makes the whole thing trustworthy, which is not getting the graph right but getting it to where users, experts and tests can correct it. There is a company in this for somebody to build.

The altitudes of one Python repository, and where the vocabulary changes. Each rung is a graph of its own. The green marks are where you change universe: the words for architecture are not the words for types, and neither are the words for syntax. C4's four levels and Gherkin's top rung are beside it for comparison. Counts are from the worked example in the vault.
Where this comes from, and what is behind it. A voice memo on 3 October 2026, in the middle of a thread about code review, and a body of earlier work it points back to: the O2 Platform and method streams from 2010 to 2013, the writing on semantic knowledge graphs for source code analysis from 2025, the fractal semantic graphs article from September, and the two sibling sites that measure what this site's own code looks like, coding.sgit.ai and nfrs.sgit.ai. It comes with a worked example you can open: Code Review Graphs, a vault in which one real codebase is read as layered graphs, with the scripts that did it. The example is deliberately modest in one way, it uses a parser and not a model, so that every number in it is checkable; the article says where a model goes on top.

In short

Source code is already the thing

When I introduced the term fractal semantic graph I meant something specific: a graph in which every node opens into another graph with its own ontology, so that there is no privileged level, only the altitude you are looking from and how much definition you choose to load there. Regulation is like that, risk is like that, a news story is like that. But the best example of it that I know, the one where the layers are not an interpretation but a fact about the artefact, is source code.

Look at what a running application actually is. At the top there is a person trying to do something: go to a page, see an article, pick one, buy it. Below that are the features and flows that let them, the behaviour of the system as the user meets it. Below that are the components, the boxes people draw when you ask for the architecture, and the flows between them, which is also exactly where a threat model lives. Below that are the classes, with their fields and their types, which is where you first hit what exists in the code rather than what is said about it, and which is where the vocabulary becomes language-specific: a Python class is not a Rust struct is not a Java interface. Below that are the methods and the calls between them. Below that is the syntax tree, every statement as a node. Below that, bytecode, machine code, the runtime, the compute it runs on. The ladder does not end. You choose how far down to load.

Each of those is a graph, and each is a different universe. The edges at the component layer are imports and data flows. The edges at the class layer are field types and inheritance. The edges at the method layer are calls. The edges in the syntax tree are the grammar of the language. None of these vocabularies is the others, and none should be forced to be. That is the fractal part: not that there are many levels, but that the meaning of a node is supplied by the ontology of the altitude where it sits.

Every source file produces a pile of files, and the review reads the pile. Start from the syntax tree, not the text. Derive each layer once and again only when its inputs change. Keep the results next to the code, in a git repository, in a vault, or in both side by side. The source box is illustrative; the file names and figures are from the worked example.

What C4 saw, and what Gherkin got right

I am not the first to see layers in code, and I want to be clear about the debt. Simon Brown's C4 model, whose roots he dates to 2006 to 2009 and which got its name in 2011, is exactly this instinct: "maps of your code, at various levels of detail, in the same way you would use something like Google Maps to zoom in and out." It names four levels, context, container, component and code, and of the fourth it says: "This is very much an optional level of detail", because most IDEs can generate it on demand. That is the right instinct, drawn before there was anything that could cheaply fill the levels below the whiteboard. The trouble with a fixed four is not that four is wrong. It is that the real thing has levels of levels, and the levels below the component are where most of the evidence lives.

Gherkin, the language under Cucumber, which Aslak Hellesøy created in 2008, got a different part right. It forced the top rung into a shape people could write and read: Given a context, When an event, Then an outcome. Teams that used it well produced user stories that were also executable, and that mattered. But the glue underneath was pattern matching. Cucumber's own documentation says it: "A step definition's expression can either be a Regular Expression or a Cucumber Expression." The step text was matched to code by a regex, and the documentation's own anti-patterns page warns that feature-coupled step definitions "may lead to an explosion of step definitions, code duplication, and high maintenance costs." I said on a podcast in 2024 that Gherkin was always a hack because of how hard it was wired to the back end, and I stand by that, with affection. It was the right top rung with no ladder beneath it.

Threat modelling is a third witness. Data flow diagrams are hierarchical by design, a context diagram and then level 0, level 1, level 2, each process decomposing into the next, and OWASP's own guidance says they "can be used to decompose the application into subsystems and lower-level subsystems." Microsoft's STRIDE per element assigns threats by the type of each node in the diagram. Threat modellers have been drawing a fractal graph of the system by hand for twenty years, at one altitude, and leaving the rest unconnected.

What changed

Naming used to be the expensive part. Deciding what a method does, which story a command serves, what the right verb is between a component and the data store it writes to: those were acts of human attention, and there were never enough of them to go round. A language model can now read a function and describe what it does, and can propose the edge between two nodes faster than a person can type it. What it produces should not be prose. It should be a structured file, a node with a type and edges with verbs, validated against a schema so that a malformed answer fails at the point of the mistake rather than being discovered later.

The second change is that the syntax tree, not the text, is the right starting point, and the tools for that are mature. Tree-sitter builds a concrete syntax tree for any of dozens of languages and updates it incrementally. The code property graph work of 2014 showed what merging a syntax tree with control flow and data dependence buys you, and Joern made it a tool. Structural diff tools exist because text diffs are the wrong altitude: difftastic is "not line-oriented. If you reformat your code and it's now split over multiple lines, difftastic will show you what's actually changed." A paper published in January 2026 on graph retrieval for codebases found that "deterministic AST-derived graphs provide more reliable coverage and multi-hop grounding than LLM-extracted graphs at substantially lower indexing cost". That matches what I would expect and what the worked example here does: derive what a parser can derive, deterministically, and spend the model on what the parser cannot know.

The third change is that the results can be kept as files, and files are the thing we already know how to version, diff, sign and share. Every source file produces a pile of JSON files: its classes, its methods and calls, its place in the package graph, its fingerprint, the tests that import it, the rules it passes and fails. Most of that pile is stable between commits, so you generate it once and regenerate on change, and the cost scales with the change rather than the repository. The pile can live in a git repository next to the code, in a vault, or in both side by side; the two are compatible and I already run workflows that split data between them. I wrote in 2025 that storing the graph as JSON files "makes it easily portable, diffable, and mergeable, just like source code itself is tracked", and I have not found a reason to change that.

Review every layer, diff every layer

Here is the argument of the article, stated for code review in particular because that is the thread I was in.

A code review today reads a text diff and a reviewer's memory of the codebase. The diff is at one altitude, the lines, and it is the wrong one for almost every question a reviewer actually has. Does this change the behaviour a user sees? Is this the fix it claims to be, or a refactor, or both? What else can this reach? Did the tests that were added assert the thing that was broken? Is this consistent with how the rest of the codebase does it? None of those is a question about lines.

If every layer is a graph and every layer is kept as files, then every layer can be diffed, and the diffs say different things.

A refactor should change the method and syntax layers and leave the class shapes, the commands and the stories untouched. If a change claims to be a refactor and the story layer moved, the claim is wrong, and you can see that without running anything.

A bug is a story that stopped holding. Its fix should move the bottom layers and be visible at the top as that story holding again, usually with a new test that asserts it. The shape is: bodies moved, signatures and fields still, a test added.

A feature adds nodes at several layers at once, a new story, a new command or flow, new classes, and the review question is whether the new nodes at each altitude are connected to each other and to the existing graph the way the architecture says they should be.

A leak is an edge that should not exist: a path from a node that holds personal data to a node that logs. I wrote the rule that way in 2025, "no edge from a PII data node to a Logging function node", and that is what a security rule is when the code is a graph: a forbidden shape.

And the blast radius of any change is the climb. Start from the methods the change touched, walk the callers upwards through the call graph, including dynamic dispatch, until you reach the commands and the stories. That set is what the change can affect. It is the same word I use for agents, what a row of reach would cost if it were used, and it is the same idea: the footprint is what happened, the blast radius is what could.

One commit, read upwards. The text diff is 75 lines added and 7 removed in three files. The graph reads the same change at every altitude and climbs the call graph to the commands and stories that can reach it. The pattern across the rungs is the shape of a fix: bodies moved, shapes still, a test added.

This is also the only honest answer I have to the characteristic failure of AI-generated code, which is the change made in one place that breaks another. A model edits a method to satisfy the test in front of it, and a command three packages away stops working because it shared that method through a base class. In a text diff that is invisible. In the layered graph it is a changed node whose upward reach includes a command the author never looked at, and the review can put that command's story in front of a person before the change lands. The user interface should be a graph too, its components linked to the classes that serve them, so that the climb does not stop at the API. Guess what moves when you break something: the graph.

The numbers on the review burden are not in dispute. GitClear's analysis of 211 million changed lines found cloned code rising from 8.3 to 12.3 percent of changed lines between 2021 and 2024 while moved code, which is what refactoring looks like in a diff, fell from 25 percent to under 10. Google's 2024 DORA report estimated that each 25 percent increase in AI adoption came with a 7.2 percent reduction in delivery stability. Stack Overflow's 2025 survey found 66 percent of developers citing "AI solutions that are almost right, but not quite" and 45 percent saying debugging AI-generated code takes longer. Veracode's 2025 report found AI-generated code introducing security flaws in 45 percent of tests across more than 100 models. More code, less refactoring, and the reviews are the bottleneck. The answer is not a faster reviewer reading the same diff. It is a different diff.

Method streams, again

In June 2012 I wrote: "MethodSteams are a code representation of an entire call-tree, i.e. one file that contains the original method and all the methods it calls (recursively)." It was how I reviewed code on the O2 Platform. I did not read the whole codebase. I picked the entry point I cared about, followed the calls, and produced one file with only the code on that path. Code streams were the next step: every taint path through a method stream, which is how we found injection flaws without a full engine. The O2 Platform had its first major release in July 2010 and I spent years on the argument that static analysis fails on trace connection rather than on rules, that "large code-bases will have tons of air-gaps created by interface-driven/WebServices/Message-Queues architectures" and that a scanner which cannot customise its sources and sinks is blind at the wrappers.

A method stream: the code on one path and only that code. The stream from the push command, breadth-first to depth five, is 100 methods and 1,703 lines from 20 classes, which is 7.5 percent of the repository's method code and the whole of what a reviewer of "publish a new vault" has to read. The fix in the worked example sits at depth four of it.

From a syntax tree with resolved calls, a method stream is one script and a few seconds. In the worked example, the stream from the push command to depth five is 100 methods and 1,703 lines out of 22,773 lines of method code in the repository. That is the difference between reviewing a change and reviewing a repository: a reviewer of "publish a brand-new vault for the first time" reads 7.5 percent of the code and knows it is the right 7.5 percent, because it is the code the command can reach. The stream for clone is 32 methods and 469 lines. Both are in the vault as Markdown, each method with its source, in the order the walk found them.

The air gaps I complained about in 2012 are still there, and they are the place a model earns its keep. A parser resolves self.sync.push() when the field is typed; it cannot resolve a call through a message queue or a dynamically named handler. Those are the edges to propose with a model and confirm with a run. The parser gives you the floor, and the floor in the worked example is 2,592 edges; what a model adds should be marked as proposed until something confirms it.

The worked example

I wanted the example to be real code, and I wanted every number in it to be checkable, so it uses the sgit command-line tool, which is Apache-2.0 and whose conventions I know, and it uses Python's own parser rather than a model. The vault is at Code Review Graphs, the read key is on that page, and the three scripts that made it are inside.

The nine views of the vault app: the ladder, the stories, the package graph, the classes, the methods, the two method streams, one commit read upwards, the pattern checks and the test map. Every screenshot in this section is a picture of the real vault, and clicking one opens its page, where the vault runs live with its read key. Click the image to open the vault page ↗
Open the vault. Code Review Graphs has the read key on the page, the app running live, a walkthrough of each view, and the three scripts that made it. A read key is the whole credential: clone it, read everything, change nothing.
The vault as it opens: the ladder of seven altitudes, with how many nodes at each one commit touched. Click a rung to open its layer. Click the image to open the vault page ↗

At commit 397be83 the tool is 427 Python files and 30,606 lines. Read from the syntax tree, that is 13 packages with 45 import relations between them, 377 classes of which 361 are Type_Safe, 1,111 methods, 2,592 calls that resolve inside the repository, and 177,313 syntax-tree nodes. Seventy-two commands were parsed out of the argument-parser wiring, each with the method that handles it. Eleven user stories were written by hand from the command help, in Given, When, Then, each naming the commands it uses; that is the one layer a person wrote, and the app says so. Three hundred and forty-two test modules were mapped to the classes they import, which gave 47 classes that no test imports, without running a test.

The commit itself is a bug fix I made the week before, found while publishing a vault from this site: the automatic transport was treating any HTTP 404 as "this host has no live API" and silently flipping a fresh, writable vault to the read-only static transport on its first push. The text diff is 75 lines added and 7 removed in three files. Read upwards, it is one method in the transport class and six command handlers changed; no signature, field or class shape moved; a test module with six tests added. The climb from those seven methods, through a status call on a field typed as the base class and down into the subclass by dynamic dispatch, reaches nine of the 72 commands and six of the eleven stories. The shape is the shape of a fix. The commit message says it is a fix. The two agree, and a reviewer can see that they agree before reading a line.

The commit in the vault: every altitude it touched in red, the fix-or-refactor table, and the commit message beside it as evidence to be checked against the graph rather than taken on trust. Click the image to open the vault page ↗

The patterns are the part I find most telling. The project states its own rules in a file at its root. No raw primitives as fields of typed classes: the graph finds 19 exceptions in 755 fields. No module-level functions: eleven in 427 modules. No static methods, immutable defaults, nothing but imports in the command package's init file, no init files under tests: all hold. A mirrored test file for every domain class: missing for 143 of 253. Thirty-eight methods are over 80 lines and the longest, push itself, is 323. The ordering is a finding in itself: the rules the project states most firmly are the ones that hold, and the one it states least is the one most often broken. The graph shows where the discipline is and where it is only written down.

Ten house rules, each run as a query over the graph, with the count of nodes where it does not hold and the count checked. A rule is a shape; a violation is a node where the shape is missing. Click the image to open the vault page ↗

What the example does not do is as important. The stories are hand-written, not proposed. Call resolution is conservative and stops at the repository's edge. The syntax-tree layer is fingerprinted per method, so a change is a changed hash, but it is not yet diffed node by node. It is one language with unusually strong conventions, which makes the class layer easy; a second language is the test of whether the layers hold their shape. Every one of those is the next thing to build, and all of them are marked in the vault.

What makes it work is reality

I want to be careful here, because it is where this kind of idea usually dies. A model's first reading of a codebase will be partly wrong. Some edges it proposes will not exist; some stories it names will be the wrong story. If the plan requires the graph to be right, the plan fails.

The plan does not require that. It requires the graph to be made of layers that someone knows, and to make predictions that something can check.

The feedback loop. Derive what a parser can derive and let a model propose the rest, marked as proposed. Confront it with users, who confirm whether a story holds; with experts, who each read the rung they know; and with tests, which are mapped from imports and confirmed by execution. Keep every correction as a commit to the graph files, and re-derive only what changed.

Users confirm the top. A user story either holds after a change or it does not. If the graph said a change could reach the story "publish a new vault for the first time" and a user cannot, the graph was right about the reach and the change was wrong. If the user can and the graph said it could not be reached, the graph has a missing edge, and that is a correction to commit. Either way, behaviour is the test of the top rung, and it is the one test that does not depend on anybody's opinion.

Experts confirm their rung. The architect reads the component graph and says whether it is the architecture. The security reviewer reads the flows and the forbidden shapes. The author reads the method stream. The designer reads the interface graph. Nobody has to read the whole thing, because the whole thing is not at any one altitude. This is also why the graph has to be files and not a tool's internal state: a reviewer corrects a file, and the correction is a commit with a name on it.

Tests confirm execution. A test is a controlled execution, which makes it the natural instrument for confirming a stream: run the test, record the path, and compare it to the path the static graph predicted. The air gaps show up as the difference. And the tests are themselves nodes to review. Which classes does each exercise? Which classes does nothing exercise? What does the test actually assert? All of that can be mapped from the imports and the syntax tree before a single test runs, and the worked example does the first two.

Patterns are findings. The two sibling sites that measure this site's own code, coding.sgit.ai and nfrs.sgit.ai, are built on a hypothesis I will state outright: a codebase's quality is visible in the shape of its graph. One class per file is a tree shape. Mirrored tests are a symmetry. Raw primitives in typed classes are a forbidden node type. A method nobody calls, a class nobody tests, a package that imports from everything, a test suite that only runs two-thirds of what it could collect: each is a pattern, each is a query, and none needs to be read line by line to be seen. The style guide "that measured itself" is the same idea applied to the house rules; the non-functional requirements site ends every page with its own counter-evidence for the same reason this article lists what the example does not do.

What exists today, and what does not

Exists and runs: the vault, with the three scripts, the layered graphs, the two streams, the delta and the pattern checks for one Python project at two commits; the method stream technique, which has existed since 2012; the fractal graph grammar and the graphs this site publishes for its own articles; the sibling sites that measure this site's code; mature parsers, code property graphs and structural diff tools from other people, cited below.

Does not exist yet: a model proposing the stories and the unresolved edges, marked as proposed and confirmed by runs; a node-level diff of the syntax tree; a second language; the interface layer linked to the classes that serve it; the execution path from a test compared to the predicted stream; and, above all, the review itself, the thing that sits on a pull request and shows the reviewer the change at every altitude with the climb to the stories it can reach. Everything in the previous sentence is a project a competent team could start on Monday, and the first of them is the one I would start with.

For somebody to build

I think this is a company, and I want to say so plainly rather than leave it implied.

What it sells is a code review that reads every layer. Attach it to a repository and it derives the pile once, keeps it current per commit, and on every change shows the reviewer the diff at each altitude, the fix-or-refactor shape, the method stream of what changed, and the climb to the commands, interfaces and stories it can reach. It proposes the edges a parser cannot see and marks them as proposed until a run confirms them. It keeps the graph as files in the customer's own repository or vault, so the customer owns the record and can read it later with a read key and no write credential. It runs the customer's own rules as graph queries and shows the patterns, and it gets better on a codebase the longer it runs there, because every correction is a commit.

Who buys it is anyone generating code faster than they can review it, which this year is everyone, and in particular the teams whose review is a compliance requirement rather than a courtesy: regulated software, payments, health, anything with a threat model that is drawn once a year and never connected to the code. For those teams the threat model and the component layer are the same graph, and the sale is that they finally agree with each other.

What is defensible is the record and the loop, not the model. Models will be commodities; the graph of a particular codebase, corrected over a year by the people who know it, is not. The method is written down here and in the vault, and the scripts are Apache-2.0, so a founder does not need my permission. They need a second language and a first customer. It sits with the other business plans published for founders on this site, and I would be glad to see it built with us or without us.

Threads woven here

Sources

Drafted from a voice memo by Dinis Cruz, who is the author of the argument and the person with editorial responsibility, by agent@riskmandate.ai (Claude Fable 5.1, claude-fable-5-1) in the sgit.ai site session, on 3 October 2026. The earlier writings were read at the URLs given; quotations are verbatim from those pages. The worked example was built in the same session from the public sgit-ai repository at the two commits named, with Python's own parser and no model; its numbers are reproducible from the scripts in the vault. The figures are infographics drawn from those numbers; the source file shown in one of them is abbreviated.

© 2026 Dinis Cruz. This article's own text is licensed under CC BY 4.0. You're free to share and adapt it, as long as you give credit. Quoted material and linked sources keep their own licences.

Threads

Graphs & knowledgeStartups & strategy This article as a graph →

Builds on

Continued by

All articles · All graphs

← All articles