Home / Articles / Code review as a fractal semantic graph: source code is already one, and the review should read every layer of it
Code review as a fractal semantic graph: source code is already one, and the review should read every layer of it
By Dinis Cruz · 2026-10-03 · v0.6.52 · code-reviewfractal-semantic-graphsgraphsstatic-analysisastmethod-streamso2-platformai-generated-codevaultsc4gherkinthreat-modellingnfrsarticle
Abstract: Source code is a very good example of a fractal semantic graph. It has layers within layers, and each one is a graph with its own vocabulary: what the user is trying to do, the features and flows, the components people draw as architecture, the classes, the methods and the calls between them, the syntax tree, and on down to the machine code if you want it. C4 saw the layers and stopped at four; Gherkin got the top layer into a shape people could write and then glued it to the code with regular expressions. What changed is that naming a node and the verb to the next one is now cheap, because a language model can do it from the syntax tree, once per change, and write the result as files. This article argues that code review should read a change at every one of those layers, and that two things fall out when it does: a refactor is a change that moves the bottom layers and leaves the top ones still, and a bug fix is a change that is visible at the top as a story that now holds. It revisits method streams, the review technique from the OWASP O2 Platform in 2012, as one script over a syntax tree with resolved calls. It comes with a worked example published as a vault: the sgit command-line tool, 377 classes and 1,111 methods, read as layered graphs with nothing run, including one real commit read upwards from the seven methods it changed to the nine commands and six user stories it can reach. And it says what makes the whole thing trustworthy, which is not getting the graph right but getting it to where users, experts and tests can correct it. There is a company in this for somebody to build.
In short
- Source code is a fractal semantic graph. Not as a metaphor. It has layers within layers, and each layer is a graph with its own node types and verbs: stories, flows, components, classes, methods and calls, the syntax tree, and below that the bytecode and the machine. You change universe when you move from one to the next.
- The fixed-level models were right and stopped early. C4 named four levels and left the bottom one to the IDE. Gherkin got the top level into Given, When, Then and glued it to the code with regular expressions. Both saw the shape before there was anything cheap enough to fill it.
- What is cheap now is naming. A language model can read a function and say what it does, name the story a command serves, and propose the verb between two nodes. From a syntax tree, once per change, written out as files.
- Review every layer, diff every layer. A refactor should move the bottom layers and leave the top ones still. A bug is a story that stopped holding; its fix is visible at the top. The blast radius of a change is the climb from the methods it touched to the stories that can reach them, and it can be read before anybody runs anything.
- Method streams are back. Follow the call tree from one method and write out only that code. It was how I reviewed code on the O2 Platform in 2012. From a syntax tree with resolved calls it is one script, and for the push command of a 30,000-line tool it is 1,703 lines.
- Reality makes it trustworthy, not accuracy. You do not have to get the graph right. You have to get it to where user behaviour, the people who know each layer, and the tests can confirm or correct it. The patterns the graph shows, folder shape, type shape, test shape, are themselves findings.
- The worked example is real and runs nothing. The sgit command-line tool: 427 files, 377 classes, 1,111 methods, 2,592 resolved calls, 72 commands, eleven stories, two method streams, ten house rules checked, and one commit read upwards. Published as a vault with the read key on the page.
- There is a company here. Said plainly at the end, with what it would sell and to whom.
Source code is already the thing
When I introduced the term fractal semantic graph I meant something specific: a graph in which every node opens into another graph with its own ontology, so that there is no privileged level, only the altitude you are looking from and how much definition you choose to load there. Regulation is like that, risk is like that, a news story is like that. But the best example of it that I know, the one where the layers are not an interpretation but a fact about the artefact, is source code.
Look at what a running application actually is. At the top there is a person trying to do something: go to a page, see an article, pick one, buy it. Below that are the features and flows that let them, the behaviour of the system as the user meets it. Below that are the components, the boxes people draw when you ask for the architecture, and the flows between them, which is also exactly where a threat model lives. Below that are the classes, with their fields and their types, which is where you first hit what exists in the code rather than what is said about it, and which is where the vocabulary becomes language-specific: a Python class is not a Rust struct is not a Java interface. Below that are the methods and the calls between them. Below that is the syntax tree, every statement as a node. Below that, bytecode, machine code, the runtime, the compute it runs on. The ladder does not end. You choose how far down to load.
Each of those is a graph, and each is a different universe. The edges at the component layer are imports and data flows. The edges at the class layer are field types and inheritance. The edges at the method layer are calls. The edges in the syntax tree are the grammar of the language. None of these vocabularies is the others, and none should be forced to be. That is the fractal part: not that there are many levels, but that the meaning of a node is supplied by the ontology of the altitude where it sits.
What C4 saw, and what Gherkin got right
I am not the first to see layers in code, and I want to be clear about the debt. Simon Brown's C4 model, whose roots he dates to 2006 to 2009 and which got its name in 2011, is exactly this instinct: "maps of your code, at various levels of detail, in the same way you would use something like Google Maps to zoom in and out." It names four levels, context, container, component and code, and of the fourth it says: "This is very much an optional level of detail", because most IDEs can generate it on demand. That is the right instinct, drawn before there was anything that could cheaply fill the levels below the whiteboard. The trouble with a fixed four is not that four is wrong. It is that the real thing has levels of levels, and the levels below the component are where most of the evidence lives.
Gherkin, the language under Cucumber, which Aslak Hellesøy created in 2008, got a different part right. It forced the top rung into a shape people could write and read: Given a context, When an event, Then an outcome. Teams that used it well produced user stories that were also executable, and that mattered. But the glue underneath was pattern matching. Cucumber's own documentation says it: "A step definition's expression can either be a Regular Expression or a Cucumber Expression." The step text was matched to code by a regex, and the documentation's own anti-patterns page warns that feature-coupled step definitions "may lead to an explosion of step definitions, code duplication, and high maintenance costs." I said on a podcast in 2024 that Gherkin was always a hack because of how hard it was wired to the back end, and I stand by that, with affection. It was the right top rung with no ladder beneath it.
Threat modelling is a third witness. Data flow diagrams are hierarchical by design, a context diagram and then level 0, level 1, level 2, each process decomposing into the next, and OWASP's own guidance says they "can be used to decompose the application into subsystems and lower-level subsystems." Microsoft's STRIDE per element assigns threats by the type of each node in the diagram. Threat modellers have been drawing a fractal graph of the system by hand for twenty years, at one altitude, and leaving the rest unconnected.
What changed
Naming used to be the expensive part. Deciding what a method does, which story a command serves, what the right verb is between a component and the data store it writes to: those were acts of human attention, and there were never enough of them to go round. A language model can now read a function and describe what it does, and can propose the edge between two nodes faster than a person can type it. What it produces should not be prose. It should be a structured file, a node with a type and edges with verbs, validated against a schema so that a malformed answer fails at the point of the mistake rather than being discovered later.
The second change is that the syntax tree, not the text, is the right starting point, and the tools for that are mature. Tree-sitter builds a concrete syntax tree for any of dozens of languages and updates it incrementally. The code property graph work of 2014 showed what merging a syntax tree with control flow and data dependence buys you, and Joern made it a tool. Structural diff tools exist because text diffs are the wrong altitude: difftastic is "not line-oriented. If you reformat your code and it's now split over multiple lines, difftastic will show you what's actually changed." A paper published in January 2026 on graph retrieval for codebases found that "deterministic AST-derived graphs provide more reliable coverage and multi-hop grounding than LLM-extracted graphs at substantially lower indexing cost". That matches what I would expect and what the worked example here does: derive what a parser can derive, deterministically, and spend the model on what the parser cannot know.
The third change is that the results can be kept as files, and files are the thing we already know how to version, diff, sign and share. Every source file produces a pile of JSON files: its classes, its methods and calls, its place in the package graph, its fingerprint, the tests that import it, the rules it passes and fails. Most of that pile is stable between commits, so you generate it once and regenerate on change, and the cost scales with the change rather than the repository. The pile can live in a git repository next to the code, in a vault, or in both side by side; the two are compatible and I already run workflows that split data between them. I wrote in 2025 that storing the graph as JSON files "makes it easily portable, diffable, and mergeable, just like source code itself is tracked", and I have not found a reason to change that.
Review every layer, diff every layer
Here is the argument of the article, stated for code review in particular because that is the thread I was in.
A code review today reads a text diff and a reviewer's memory of the codebase. The diff is at one altitude, the lines, and it is the wrong one for almost every question a reviewer actually has. Does this change the behaviour a user sees? Is this the fix it claims to be, or a refactor, or both? What else can this reach? Did the tests that were added assert the thing that was broken? Is this consistent with how the rest of the codebase does it? None of those is a question about lines.
If every layer is a graph and every layer is kept as files, then every layer can be diffed, and the diffs say different things.
A refactor should change the method and syntax layers and leave the class shapes, the commands and the stories untouched. If a change claims to be a refactor and the story layer moved, the claim is wrong, and you can see that without running anything.
A bug is a story that stopped holding. Its fix should move the bottom layers and be visible at the top as that story holding again, usually with a new test that asserts it. The shape is: bodies moved, signatures and fields still, a test added.
A feature adds nodes at several layers at once, a new story, a new command or flow, new classes, and the review question is whether the new nodes at each altitude are connected to each other and to the existing graph the way the architecture says they should be.
A leak is an edge that should not exist: a path from a node that holds personal data to a node that logs. I wrote the rule that way in 2025, "no edge from a PII data node to a Logging function node", and that is what a security rule is when the code is a graph: a forbidden shape.
And the blast radius of any change is the climb. Start from the methods the change touched, walk the callers upwards through the call graph, including dynamic dispatch, until you reach the commands and the stories. That set is what the change can affect. It is the same word I use for agents, what a row of reach would cost if it were used, and it is the same idea: the footprint is what happened, the blast radius is what could.
This is also the only honest answer I have to the characteristic failure of AI-generated code, which is the change made in one place that breaks another. A model edits a method to satisfy the test in front of it, and a command three packages away stops working because it shared that method through a base class. In a text diff that is invisible. In the layered graph it is a changed node whose upward reach includes a command the author never looked at, and the review can put that command's story in front of a person before the change lands. The user interface should be a graph too, its components linked to the classes that serve them, so that the climb does not stop at the API. Guess what moves when you break something: the graph.
The numbers on the review burden are not in dispute. GitClear's analysis of 211 million changed lines found cloned code rising from 8.3 to 12.3 percent of changed lines between 2021 and 2024 while moved code, which is what refactoring looks like in a diff, fell from 25 percent to under 10. Google's 2024 DORA report estimated that each 25 percent increase in AI adoption came with a 7.2 percent reduction in delivery stability. Stack Overflow's 2025 survey found 66 percent of developers citing "AI solutions that are almost right, but not quite" and 45 percent saying debugging AI-generated code takes longer. Veracode's 2025 report found AI-generated code introducing security flaws in 45 percent of tests across more than 100 models. More code, less refactoring, and the reviews are the bottleneck. The answer is not a faster reviewer reading the same diff. It is a different diff.
Method streams, again
In June 2012 I wrote: "MethodSteams are a code representation of an entire call-tree, i.e. one file that contains the original method and all the methods it calls (recursively)." It was how I reviewed code on the O2 Platform. I did not read the whole codebase. I picked the entry point I cared about, followed the calls, and produced one file with only the code on that path. Code streams were the next step: every taint path through a method stream, which is how we found injection flaws without a full engine. The O2 Platform had its first major release in July 2010 and I spent years on the argument that static analysis fails on trace connection rather than on rules, that "large code-bases will have tons of air-gaps created by interface-driven/WebServices/Message-Queues architectures" and that a scanner which cannot customise its sources and sinks is blind at the wrappers.
From a syntax tree with resolved calls, a method stream is one script and a few seconds. In the worked example, the stream from the push command to depth five is 100 methods and 1,703 lines out of 22,773 lines of method code in the repository. That is the difference between reviewing a change and reviewing a repository: a reviewer of "publish a brand-new vault for the first time" reads 7.5 percent of the code and knows it is the right 7.5 percent, because it is the code the command can reach. The stream for clone is 32 methods and 469 lines. Both are in the vault as Markdown, each method with its source, in the order the walk found them.
The air gaps I complained about in 2012 are still there, and they are the place a model earns its keep. A parser resolves self.sync.push() when the field is typed; it cannot resolve a call through a message queue or a dynamically named handler. Those are the edges to propose with a model and confirm with a run. The parser gives you the floor, and the floor in the worked example is 2,592 edges; what a model adds should be marked as proposed until something confirms it.
The worked example
I wanted the example to be real code, and I wanted every number in it to be checkable, so it uses the sgit command-line tool, which is Apache-2.0 and whose conventions I know, and it uses Python's own parser rather than a model. The vault is at Code Review Graphs, the read key is on that page, and the three scripts that made it are inside.
At commit 397be83 the tool is 427 Python files and 30,606 lines. Read from the syntax tree, that is 13 packages with 45 import relations between them, 377 classes of which 361 are Type_Safe, 1,111 methods, 2,592 calls that resolve inside the repository, and 177,313 syntax-tree nodes. Seventy-two commands were parsed out of the argument-parser wiring, each with the method that handles it. Eleven user stories were written by hand from the command help, in Given, When, Then, each naming the commands it uses; that is the one layer a person wrote, and the app says so. Three hundred and forty-two test modules were mapped to the classes they import, which gave 47 classes that no test imports, without running a test.
The commit itself is a bug fix I made the week before, found while publishing a vault from this site: the automatic transport was treating any HTTP 404 as "this host has no live API" and silently flipping a fresh, writable vault to the read-only static transport on its first push. The text diff is 75 lines added and 7 removed in three files. Read upwards, it is one method in the transport class and six command handlers changed; no signature, field or class shape moved; a test module with six tests added. The climb from those seven methods, through a status call on a field typed as the base class and down into the subclass by dynamic dispatch, reaches nine of the 72 commands and six of the eleven stories. The shape is the shape of a fix. The commit message says it is a fix. The two agree, and a reviewer can see that they agree before reading a line.
The patterns are the part I find most telling. The project states its own rules in a file at its root. No raw primitives as fields of typed classes: the graph finds 19 exceptions in 755 fields. No module-level functions: eleven in 427 modules. No static methods, immutable defaults, nothing but imports in the command package's init file, no init files under tests: all hold. A mirrored test file for every domain class: missing for 143 of 253. Thirty-eight methods are over 80 lines and the longest, push itself, is 323. The ordering is a finding in itself: the rules the project states most firmly are the ones that hold, and the one it states least is the one most often broken. The graph shows where the discipline is and where it is only written down.
What the example does not do is as important. The stories are hand-written, not proposed. Call resolution is conservative and stops at the repository's edge. The syntax-tree layer is fingerprinted per method, so a change is a changed hash, but it is not yet diffed node by node. It is one language with unusually strong conventions, which makes the class layer easy; a second language is the test of whether the layers hold their shape. Every one of those is the next thing to build, and all of them are marked in the vault.
What makes it work is reality
I want to be careful here, because it is where this kind of idea usually dies. A model's first reading of a codebase will be partly wrong. Some edges it proposes will not exist; some stories it names will be the wrong story. If the plan requires the graph to be right, the plan fails.
The plan does not require that. It requires the graph to be made of layers that someone knows, and to make predictions that something can check.
Users confirm the top. A user story either holds after a change or it does not. If the graph said a change could reach the story "publish a new vault for the first time" and a user cannot, the graph was right about the reach and the change was wrong. If the user can and the graph said it could not be reached, the graph has a missing edge, and that is a correction to commit. Either way, behaviour is the test of the top rung, and it is the one test that does not depend on anybody's opinion.
Experts confirm their rung. The architect reads the component graph and says whether it is the architecture. The security reviewer reads the flows and the forbidden shapes. The author reads the method stream. The designer reads the interface graph. Nobody has to read the whole thing, because the whole thing is not at any one altitude. This is also why the graph has to be files and not a tool's internal state: a reviewer corrects a file, and the correction is a commit with a name on it.
Tests confirm execution. A test is a controlled execution, which makes it the natural instrument for confirming a stream: run the test, record the path, and compare it to the path the static graph predicted. The air gaps show up as the difference. And the tests are themselves nodes to review. Which classes does each exercise? Which classes does nothing exercise? What does the test actually assert? All of that can be mapped from the imports and the syntax tree before a single test runs, and the worked example does the first two.
Patterns are findings. The two sibling sites that measure this site's own code, coding.sgit.ai and nfrs.sgit.ai, are built on a hypothesis I will state outright: a codebase's quality is visible in the shape of its graph. One class per file is a tree shape. Mirrored tests are a symmetry. Raw primitives in typed classes are a forbidden node type. A method nobody calls, a class nobody tests, a package that imports from everything, a test suite that only runs two-thirds of what it could collect: each is a pattern, each is a query, and none needs to be read line by line to be seen. The style guide "that measured itself" is the same idea applied to the house rules; the non-functional requirements site ends every page with its own counter-evidence for the same reason this article lists what the example does not do.
What exists today, and what does not
Exists and runs: the vault, with the three scripts, the layered graphs, the two streams, the delta and the pattern checks for one Python project at two commits; the method stream technique, which has existed since 2012; the fractal graph grammar and the graphs this site publishes for its own articles; the sibling sites that measure this site's code; mature parsers, code property graphs and structural diff tools from other people, cited below.
Does not exist yet: a model proposing the stories and the unresolved edges, marked as proposed and confirmed by runs; a node-level diff of the syntax tree; a second language; the interface layer linked to the classes that serve it; the execution path from a test compared to the predicted stream; and, above all, the review itself, the thing that sits on a pull request and shows the reviewer the change at every altitude with the climb to the stories it can reach. Everything in the previous sentence is a project a competent team could start on Monday, and the first of them is the one I would start with.
For somebody to build
I think this is a company, and I want to say so plainly rather than leave it implied.
What it sells is a code review that reads every layer. Attach it to a repository and it derives the pile once, keeps it current per commit, and on every change shows the reviewer the diff at each altitude, the fix-or-refactor shape, the method stream of what changed, and the climb to the commands, interfaces and stories it can reach. It proposes the edges a parser cannot see and marks them as proposed until a run confirms them. It keeps the graph as files in the customer's own repository or vault, so the customer owns the record and can read it later with a read key and no write credential. It runs the customer's own rules as graph queries and shows the patterns, and it gets better on a codebase the longer it runs there, because every correction is a commit.
Who buys it is anyone generating code faster than they can review it, which this year is everyone, and in particular the teams whose review is a compliance requirement rather than a courtesy: regulated software, payments, health, anything with a threat model that is drawn once a year and never connected to the code. For those teams the threat model and the component layer are the same graph, and the sale is that they finally agree with each other.
What is defensible is the record and the loop, not the model. Models will be commodities; the graph of a particular codebase, corrected over a year by the people who know it, is not. The method is written down here and in the vault, and the scripts are Apache-2.0, so a founder does not need my permission. They need a second language and a first customer. It sits with the other business plans published for founders on this site, and I would be glad to see it built with us or without us.
Threads woven here
- Introducing fractal semantic graphs: the definition this article applies to code, and the grammar of verbs with named inverses.
- Footprint and blast radius: the same two words for agents; the footprint is what happened, the blast radius is what could.
- Custom UIs are not the exception: interfaces built from graphs in an afternoon, which is what the review surface will be.
- Every risk is already accepted: the risk graph and the explorer built on it, the same method one domain over.
- Code Review Graphs: the worked example, as a vault with its read key.
- Fractal semantic graphs, the demo page, and the ThreatModCon 2025 vault, whose ladder already runs from customer to compute.
- coding.sgit.ai and nfrs.sgit.ai: the two sites that measure this site's own code and practice, and the hypothesis that quality is visible as shape.
Sources
- Dinis Cruz, O2 .NET SAST engine: MethodStreams and CodeStreams, June 2012; In SAST the issue is trace connection, not rules, June 2012; first major release of the OWASP O2 Platform, July 2010; comments on the SATEC static analysis document, March 2013.
- Dinis Cruz, Semantic knowledge graphs for LLM-driven source code analysis, May 2025; Graphs of graphs of graphs in threat modelling, May 2025; MGraph-AI, the file-backed graph database the method uses.
- Simon Brown, the C4 model: introduction, the code diagram, FAQ on its history.
- Cucumber, step definitions, Gherkin reference, anti-patterns; InfoQ, Cucumber at ten, April 2018.
- Yamaguchi, Golde, Arp and Rieck, Modeling and discovering vulnerabilities with code property graphs, IEEE S&P 2014; Joern; tree-sitter; difftastic; Falleri et al., GumTree, ASE 2014.
- OWASP, threat modeling process; Microsoft, Uncover security design flaws using the STRIDE approach, November 2006.
- Reliable graph retrieval for codebases, arXiv, January 2026; Aider's repository map, October 2023; CodexGraph, August 2024; RepoGraph, October 2024.
- GitClear, AI assistant code quality research, 2025; Google, the 2024 DORA report; Stack Overflow, 2025 Developer Survey, AI; Veracode, 2025 GenAI code security report.
- Practical AI episode 286, conversation with Dinis Cruz, September 2024, for the Gherkin remark.
Drafted from a voice memo by Dinis Cruz, who is the author of the argument and the person with editorial responsibility, by agent@riskmandate.ai (Claude Fable 5.1, claude-fable-5-1) in the sgit.ai site session, on 3 October 2026. The earlier writings were read at the URLs given; quotations are verbatim from those pages. The worked example was built in the same session from the public sgit-ai repository at the two commits named, with Python's own parser and no model; its numbers are reproducible from the scripts in the vault. The figures are infographics drawn from those numbers; the source file shown in one of them is abbreviated.
© 2026 Dinis Cruz. This article's own text is licensed under CC BY 4.0. You're free to share and adapt it, as long as you give credit. Quoted material and linked sources keep their own licences.
Threads
Builds on
- Fractal Semantic Graphs: everything connects to everything, and nobody has to share a schema A fractal semantic graph has no privileged level and no single schema: each world keeps its own vocabulary and connects to others through named edges.
- Footprint and blast radius: what the agent actually did, and what it would have cost Footprint is what an agent actually did, read afterwards from logs and vault history; blast radius is what a row of its reach would cost the business today.
- Custom UIs are not the exception: the inbox in 2026, where every message has its own universe Every message has a graph, so it can be shaped for the reader's moment; a custom interface per moment is now how interfaces get made, and each gets a policy.
- Every risk is already accepted. The only question is by whom, and for how long. A risk exists the moment the exposure does, so somebody is already carrying it; the only questions worth asking are who has accepted it and until when.
Continued by
- Memory is not a spectator sport: how a web of open sites, graphs and vaults became the memory for sessions like this one Agentic memory as context management: many published, fractal, provenance-carrying memories rather than one store, shown in the session that wrote the article.