for agents/llms.txtv0.6.75 · 6 Oct 2026

Home / Articles / Send an agent, not a spreadsheet: the next generation of software due diligence, and why the companies that stopped reading their code are about to be asked about it

Send an agent, not a spreadsheet: the next generation of software due diligence, and why the companies that stopped reading their code are about to be asked about it

By · 2026-10-05 · v0.6.67 · due-diligencecode-reviewagent-behaviour-policiesblast-radiusthreat-modelssbomcyber-resilience-actproduct-liabilitystartupsai-generated-codefractal-semantic-graphsarticle

Abstract: The last article asked how I would build a company on code review. This one turns it round. The buyers of software have long wanted to know whether a vendor understands what it is selling them, and have never been able to find out, because due diligence was a questionnaire: a spreadsheet of questions answered by the people being asked, disconnected from the code, too expensive to do properly and out of date on arrival. That has changed in the same way code review has. A buyer can now send a prompt, or a small agent under a behaviour policy, to run inside the vendor's environment and come back with a graph of what is actually there: what reaches production that no person reviewed, whether the documentation matches the code, whether there is a threat model and who wrote it, what the last fifty bugs touched, which agents touch the pipeline and under what policy. The vendor reads what leaves before it leaves, and can redact but not rewrite, because the answer is derived rather than written, and companies are careful about what goes on a record. The test is risk-based, not pure: a startup that says it generated its code fast, that the product is not mission-critical and here are the mitigations has passed. What fails is not knowing. That is the consequence that has been missing for the companies that stopped reviewing code, let the engineers go and let anyone prompt features into production: their customers, their investors and their acquirers are about to be able to see it. A startup should double down on understandability, because it has less code and the same tools, and the best way to build the code review company may be to sell it to the people who buy software rather than the people who write it.

Due diligence as it was and as it can be. The questionnaire measures whether the vendor has a document with the right title, is written by the vendor and connects to nothing. The agent reads what is there, under a behaviour policy, inside the vendor's environment, and the vendor reviews what leaves. The difference is not who controls the answer. It is that the thing being described can no longer be described into shape.
Where this comes from, and what is behind it. A voice memo on 5 October 2026, recorded straight after the article on building a company on code review, with the instruction to soften one claim and turn one idea round. It leans on the code review articles, on the footprint and blast radius and agent behaviour policy work, on an argument of mine from May 2025 that threat models should be mandatory disclosures because they tell you at a glance whether a team knows what it is protecting, on the SG/Send team's briefs from the summer on buyer-driven adoption, a permissions bill of materials and artefact-driven assessments, and on the regulation that is arriving at the same time, which is read in its current wording and dated in the sources. The agent that does the due diligence does not exist yet; this is the case for it and the shape of it.

In short

The form that never connected to anything

Anyone who has sold software to a large organisation has met the spreadsheet. Three hundred questions, the same ones every buyer sends in a slightly different order, about policies, controls, encryption, access, incident response, business continuity, and somewhere near the end, software development: do you review code before release, do you have a secure development lifecycle, do you test. The 2025 edition of the standard one, the Shared Assessments SIG, has 627 questions in its core form and 128 in its lite one. Whistic's 2025 survey of 525 practitioners found the average vendor answering 37.3 assessment requests a month and spending 179 hours a month on them, with 84 percent needing a follow-up. The vendor's compliance team answers from last year's answers. Each answer is true, in the sense that there is a document somewhere with the right title, and each is written so that it commits the vendor to as little as possible. The buyer files it. Often nobody on either side has read the code, and the questionnaire did not ask them to.

I want to be fair to the people who built this, because the form was the best that could be done at the time. Due diligence was a form because the alternative did not exist. Reading a vendor's code was rarely possible for a buyer, politically or practically. Reading a vendor's practices meant interviews, which do not scale and which the vendor prepares for. The data about what actually happened in a vendor's repositories, who changed what, what was reviewed, what broke and why, was not available to the vendor itself in any usable shape, let alone to a buyer. So the industry built the thing it could build, a questionnaire, and then a market of tools that help vendors answer questionnaires faster from their own documents: Vanta's product "combs through your knowledge base of security documentation and previous questionnaires" and answers eighty percent or more of the questions; Conveyor reports eighty-five percent of answers "going out unedited". That is where the automation stopped: at the form, not at the facts behind it. The other kind of tool, the outside-in rating, scores DNS health, patching cadence and IP reputation from the internet, which is a view of the perimeter and says nothing about what is inside.

Anyone who does compliance knows the result. There is a great deal of interpretation in it. Companies are very good at answers that are true in a way that leaves room, and the room is where the risk lives. What they are not good at, and will not do, is lie on a record. A false statement in a due diligence response is the kind of thing that comes back, with a lawyer, and companies know it. So the honest description of the questionnaire is that it gets you careful, composed, defensible answers to questions that do not quite reach the thing you wanted to know. A brief from July put the same point about a regulatory question: "A buyer who asks only whether a supplier is in scope has learned something about the supplier's lawyers rather than about the supplier's engineering." Replace in scope with reviews its code and the sentence still holds.

What changed

The previous two articles describe what changed on the vendor's side: one technology can now read every layer of a codebase, from the stories to the syntax tree, and build a graph of what the code is, what reaches what, and which change can reach which customer-facing behaviour. The code review article built that graph for a real repository, with nothing run, and the follow-up argued that review becomes a science rather than an art when the graphs exist and the deterministic checks run over them.

The move this article makes is small and I think consequential. If a vendor can build that graph of itself, a buyer can ask for it. And if the buyer can ask for it, the buyer can send the thing that builds it.

Picture the due diligence request not as a spreadsheet but as a prompt and a small agent, delivered to the vendor, run by the vendor inside its own environment, against its own repositories, pipelines, trackers and documents, under a behaviour policy that says exactly what it may read, what it may write and where it may send anything. It derives the graphs: which code paths reach production, which of those changed without a person reviewing them, which have tests and which have none, whether the documentation describes the system that exists, whether a threat model exists and what it covers, which of the last fifty bugs trace to which changes, which agents have touched the pipeline and under what policy. It writes the result as files. The vendor reads every file before anything leaves, and can withhold any of them. Then what the vendor is willing to show goes to the buyer, as a graph the buyer's own agent can read, in a vault only the buyer can open.

Three things make this different from the spreadsheet, and the first is the one that matters most. The answer is derived, not composed. The vendor still controls what leaves, which is right, because this is the vendor's property and some of it is sensitive. But what leaves cannot be described into shape, because it was not described. A vendor can decline to show the map of unreviewed code. It cannot show a map that says there is none when there is some, because the map is read from the repository, and a false map is a false statement on a record, which is the thing companies will not do. The questionnaire gave the vendor the pen. The agent takes the pen away and leaves the vendor the redaction marker, and that is the whole change.

The second is cost. The questionnaire costs days of a compliance team per buyer, per year, and is out of date when it arrives. The agent costs compute, once, and a delta per release after that. That changes who can ask: a mid-sized buyer who could rarely afford a technical audit can afford to send a prompt, and so can an investor doing a seed round and an acquirer in the first week rather than the last. And the buyer holds the power to ask, which the July brief put plainly: "in the same way that they go, can I have the due diligence, are you SOC compliant, all you have to do is put an extra requirement in there, because they're the buyers, they control the power." Its other line is the one this whole article rests on: "A buyer asking costs nothing. A supplier not answering looks worse than a supplier answering badly."

The third is that the vendor should run it on itself first, because the questions the buyer is about to ask are the ones the vendor should have wanted answered. The company that has built its own graphs has nothing to prepare for. The one that has not is about to find out what it does not know about its own software, which is a better way to find out than the alternatives.

Questions where the non-answer is also an answer

The cleverness of the agent, if there is one, is in the questions, and the property the good questions share is that every kind of answer tells you something, including silence.

Six questions a due diligence agent asks. Each has a good answer, an acceptable answer for the product's risk level, and a silence, and the agent does not need the vendor to say which; it can see. The table is illustrative, the shape is not.

Take the first. Does code reach production that no human has reviewed? Most companies using generated code would have to say yes, and the memo's point is that yes can be a good answer. Yes, and here is the map of it: the unreviewed changes are on paths that are not mission-critical, they are held by deterministic checks, and here are the rules that check them. That is a vendor that understands its own software and has made a decision about where review is worth a person's time. Yes, we are a small team shipping fast, here is what that code can reach and here are the mitigations round it, is also fine for a product sold as what it is. What is not fine is not knowing, because a vendor that does not know what reaches production unreviewed does not know what reaches production.

The second question is the one I have been making for years in a different form. In 2016 I wrote that a team that will not spend the time on a threat model "must accept the risk of not having one", and in May 2025 I argued that threat models should be mandatory disclosures, because "the software market exhibits the same failures once seen in finance and food: security quality is largely unverifiable, so bad products hide behind marketing rhetoric", and because a buyer, "instead of sending endless security questionnaires, could review the supplier's published threat model to understand its security posture." The reason is that a threat model is the fastest read of a team's understanding there is. I can look at one and know immediately whether to be worried. Either it is thin, and the team does not know what it is supposed to be protecting; or it has gaps the team has seen and decided about, which is a team that knows what it is doing. A threat model that says we have a weakness here, exploiting it needs a nation-state attacker, and we are not going to defend against that, is a good threat model. The buyer may decide that is not acceptable for their use, and that is a fair decision too, made on facts. What the threat model almost never contains is an unknown critical vulnerability, because by the time it is written down somebody fixes it. What it reveals is whether the people know where the lines are.

The bug question is the one that separates the kinds of risk. Bugs in the deterministic, reviewed parts of a system have reasons: a requirement misunderstood, an edge case missed, a dependency changed. Bugs caused by generated code that nobody reviewed have a different signature. They are in places no story asked for, they repeat, and they reveal something structural each time. The agent that traces the last fifty bugs to the changes that caused them, and separates the two kinds, has told the buyer more about the vendor's practice than any number of questions about whether a secure development lifecycle exists. This is also why the result has to be a graph and not a report. A brief from June on security assessments says what happens to the report: "in the end you produce a PDF, a flat file, a bunch of findings. You lose 99.9 percent of what you created." The next assessment then starts from nothing, where it should be a delta. And where the vendor has nothing to show for a question, another June brief has the rule: "lack of evidence is evidence, lack of fact is a fact."

And the question about agents is new, and I will come back to it, because the artefact that answers it is one we have been building for other reasons.

Risk-based, not pure

None of this is a purity test, and if it became one it would fail, because every real company has code it does not fully understand and every honest startup has shipped something fast. The point of the agent is visibility, and what counts as enough depends on what the software is for.

The same questions, a different bar for each kind of product. A prototype passes by saying what it is and mapping what it can reach. A product with customers needs review or tests on every path a customer depends on. A product holding other people's data needs behaviour policies for its agents and a record of what shipped unreviewed. Mission-critical software needs checks at the boundary the buyer can rerun. Honesty about the tier is the first answer.

A startup can say: we generated most of this, we have limited resources, this is not mission-critical, it provides a service you value, here are the mitigations and here is what you can do around it. For a tool at the first two tiers that is a pass, and I would rather buy from a startup that says it than from an incumbent that cannot. The same sentence from a vendor selling software that holds my customers' data, or that my business stops without, is the finding, and the agent's job is to show the gap between what the product is sold as and what its practices are fit for.

This is the risk-based approach that compliance has always claimed to take and rarely could, because it had no facts to base the risk on. With the graphs it can. The buyer sees what the change can reach, the vendor says what it has accepted and why, and the decision is a decision rather than a form. A June brief on third-party risk made the point that "the vendor's question and the customer's question are the same question, seen from two sides", and that the existing apparatus "does not yet ask what the AI capability can reach inside the customer's environment." The agent asks.

The behaviour policy is the due diligence document

The one artefact of agentic engineering that was written to be read by somebody else is the agent behaviour policy, and it turns out to be exactly what a buyer needs.

A behaviour policy, as we write them and as the agents behind this site run under them, is a short document written before an agent's first run that says what the agent is, which identity it uses, what it may read, what it may write, what it may not do, what it may spend, who approves what, which risks have been accepted by whom and for how long, and where its actions are recorded. We wrote them to keep our own agents in bounds. Reading one as a buyer, they say something else.

An agent behaviour policy read as a due diligence document. The left is a deploy agent's policy for a fictional vendor. The right is what a buyer reads off it: the worst case, which lines are enforced by the identity and which are only instructions, which risks the vendor has accepted on the buyer's behalf, and what it means when a vendor with agents in the pipeline has no policy to show.

They say what the blast radius is. A deploy agent whose identity has no merge right and no production credential has a worst case of a broken staging environment, and that is a fact about the account rather than a promise in a prompt. The June brief on agent authorisation put the principle in one sentence: "the moment of authorisation is not when you allow the agent to do something, it is the moment you give them the permissions." A buyer reads the permissions, not the promises. They say which controls are real: may not merge is a control if the identity cannot merge, and a wish if it can. They say what the vendor has accepted on the buyer's behalf, by name, for an interval, with a review date, which is what the risk acceptance article asks of every risk. And their absence says the most. A vendor with agents in its pipeline and no policy to show has agents whose blast radius nobody has written down, which is a finding before any code is read.

I think this is also what starts the market for behaviour policies in earnest. Agents will go out of control somewhere, cause damage somewhere, and companies will pull the plug because they cannot afford the risk or do not want to live with it, which is sometimes the wrong decision made for want of visibility. The policy is the visibility. A buyer that asks for it is doing the vendor a favour.

The consequence that has been missing

Now the part the memo asked me to soften, which I will, a little.

There are companies, and it is becoming common, that have decided code review and quality are costs they no longer need. The reasoning runs: the model writes the code, the product manager or the executive can prompt the features in, the developers who kept complaining about non-functional requirements and over-engineering were a cost, so let them go. Some of that reasoning has a real point in it; a great deal of what engineers called necessary was not. I described the pattern in the May article on how vibe-coded applications ship: "The narrative was clear: fire the engineers, save the budget, ship faster, win the market", followed, a few paragraphs later, by "Recovery took days because nobody knew how to operate the system that the AI had built." But a company that ships generated code to production with nobody reading it is doing two things at once. It is hitting the point of diminishing returns fast, because each change becomes harder to make safely than the last, and I think that is already happening. And it is accumulating a liability that nobody outside the company could see.

The evidence that review is being skipped is not anecdotal any more, and the dates are recent. Sonar's developer survey of January 2026 found that AI "accounts for 42% of committed code today", that "96% of developers do not fully trust AI-generated code", and that "only 48% always verify it before committing". Qodo's 2025 survey found developers "2.5x more likely to merge code without reviewing it" when it was generated, and that "when an AI-review tool is enabled, 80% of PRs don't have any human comment or review." Veracode's July 2025 report found that "in 45 percent of all test cases, LLMs produced code containing vulnerabilities aligned with the OWASP Top 10". GitClear's January 2026 analysis of 623 million changes found block duplication up 81 percent since 2023 and moved code, the signature of refactoring, "freefalling to 3.8%". Apiiro's September 2025 study of Fortune 50 repositories found AI-assisted developers producing three to four times more code and "10 times more security issues", with the code packed "into fewer pull requests, making code reviews more complicated". Google's DORA report of September 2025 put it most carefully: "AI adoption does continue to have a negative relationship with software delivery stability", and "AI doesn't fix a team; it amplifies what's already there."

The hiring decisions are on the record too. Salesforce's chief executive said on an earnings call in February 2025, "We're not going to hire any new engineers this year. We're seeing 30 percent productivity increase on engineering", and in May 2026 that engineering headcount had stayed flat because "we've been using AI to create more efficiencies". In 2026, Coinbase, Oracle, Block and Amazon all cited AI in reductions; Coinbase's note said "engineers use AI to ship in days what used to take a team weeks". Some of this will be right. Some has already reversed: Klarna's chief executive said in May 2025 that the AI service was cheaper but "lower quality" and that "investing in the quality of human support is the way of the future for us". And the incidents have started. In July 2025 Replit's chief executive wrote that an agent "in development deleted data from the production database. Unacceptable and should never be possible." In March 2026 the Financial Times and CNBC reported an internal Amazon document describing a "trend of incidents" with "high blast radius" relating to "Gen-AI assisted changes"; Amazon disputed that AI-written code was involved, which is itself the point. Nobody outside can tell.

Little has pushed back on this, because there was rarely a consequence. The customers could not see it. The questionnaire asked whether code was reviewed and got a careful yes. The investors saw velocity. The acquirers ran a snapshot scan of the code for licence conflicts and known vulnerabilities, which finds plenty, Black Duck's 2026 M&A audit data found unpatched vulnerable open source in 97 percent of transactions and licence conflicts in 94 percent, and finds neither practice nor understanding, because, as Black Duck's own description of the service says, it is "a speedy, one-time snapshot", and because, as the same paper notes, "targets generally don't share code with potential acquirers". The regulation that is arriving, which I set out in the next section, creates duties but not visibility.

The agent is the consequence. When a buyer can send a prompt that comes back with the map of what reached production unreviewed and what it can reach, the company that stopped reading its code has nothing to show it, and the buyer can say, reasonably: I am not happy that you have introduced this much risk into my environment, here is my policy, and I am not happy that you do not understand what you have sold me. That is not regulation. It is a customer with facts, and it will reach investors and acquirers in the same month, because the same agent answers their questions too.

Who the agent finds, and who gets ahead. The company that stopped reading has nothing to show. The incumbent that never understood its own software has a questionnaire answer for everything and a graph for nothing. The startup that doubles down has less code and the same tools, and can answer before it is asked.

Here is the irony, and it is the reason this is not an argument against startups or against generated code. Most companies that have been shipping software for years already do not understand it. People moved on, pull requests were merged by people who have since left, the documentation describes an earlier system, and the side effects of a change are discovered in production. You can see it in their bugs: the ones that reveal something structural, fixed one at a time. They were never asked, because nobody could. The agent asks them too.

So if I were a startup I would double down on understandability, and I would do it for two reasons. First, the startup can play the same game as the biggest buyer. The tools that build the graphs are the same tools and they are open source. A startup can build the graph of its own repository, keep its threat model with its code, write a behaviour policy for every agent, trace its bugs to its changes, and answer the agent's questions in the agent's own format before being asked. Second, the startup has less code. Understanding ten thousand lines is a week; understanding ten million that nobody still at the company wrote is a programme. The incumbent's advantage was always that nobody could check. The startup's advantage is that it can be checked and pass.

What the regulation does and does not do

This is the part I want to be careful with, and it is in this article because the timing matters: the duties are arriving before the visibility, and the agent supplies the visibility.

Three pieces of regulation and one family of standards matter here, and I have read each in its current wording.

The software bill of materials came from the US Executive Order 14028 of 12 May 2021, which told federal buyers to require "a formal record containing the details and supply chain relationships of various components used in building software", and from the NTIA's minimum elements of 12 July 2021: supplier, component name, version, identifiers, dependency relationship, author and timestamp. The NTIA was careful about the limits: "SBOMs alone will not address the multitude of software supply chain and software assurance concerns faced by the ecosystem today." CISA's 2026 minimum elements, published in July 2026, add a component hash, a licence, the tool name and the generation context, and a separate joint guidance of May 2026 extends the idea to AI systems, voluntarily. None of the fields, in any version, says who reviewed a change, what it reached, or whether the people shipping it understand it.

The EU Cyber Resilience Act, Regulation 2024/2847, entered into force on 10 December 2024. Its reporting obligations for actively exploited vulnerabilities apply from 11 September 2026, the rest from 11 December 2027. It requires a bill of materials "in a commonly used and machine-readable format covering at the very least the top-level dependencies", and Article 13(5) says manufacturers "shall exercise due diligence when integrating components sourced from third parties so that those components do not compromise the cybersecurity of the product". Fines run to fifteen million euros or two and a half percent of worldwide turnover for the essential requirements, and five million or one percent for "incorrect, incomplete or misleading information" supplied to authorities, which is the clause that makes a composed answer expensive.

The revised Product Liability Directive, 2024/2853, is the one buyers should read twice. It entered into force on 8 December 2024 and applies to products placed on the market from 9 December 2026. Software is a product under it, "irrespective of whether the software is stored on a device, accessed through a communication network or cloud technologies, or supplied through a software-as-a-service model". A claimant who presents facts sufficient to make a claim plausible can ask a court to order the manufacturer to disclose "relevant evidence that is at the defendant's disposal", and a product is presumed defective where the manufacturer fails to disclose it, where it does not comply with mandatory safety requirements including cybersecurity, or where the claimant faces "excessive difficulties" of proof "due to technical or scientific complexity". A vendor that cannot produce the record of how its software was made will, from December 2026, be presumed to have made it badly.

And the standards buyers already ask about have required review all along. PCI DSS requirement 6.2.3 says bespoke and custom software "is reviewed prior to being released into production or to customers", and 6.2.3.1 that it is "reviewed by individuals other than the originating code author" and "reviewed and approved by management prior to release". The SOC 2 change management criterion, CC8.1, requires that the entity "authorizes, designs, develops or acquires, configures, documents, tests, approves, and implements changes". ISO 27001's Annex A has a secure development lifecycle control and a change management control. Every vendor with one of those certificates has already answered yes to the review question on a form. The agent is how a buyer finds out whether the yes was composed.

What none of this does is tell a buyer whether a vendor understands its software. A software bill of materials lists what a product is made of, not whether anyone reviewed how the parts were put together or what changed last Tuesday. A June brief called the bill of materials "a dramatic success" and said what it needs next: "this should be compatible with the bill of materials, or the next version of the software bill of materials, because you need to add the whole permission mappings", on the ground that "if the account or environment running the code does not have those permissions, then it does not matter." The due diligence graph is that next version: the components, plus who reviewed them, plus what the agents that touch them may do. A conformity assessment says a process exists. A liability regime says who pays after the fact. The agent is the instrument that connects those duties to the facts they are about, and I expect buyers to want it for that reason before any regulator asks for it.

The twist on the last article

I said in the last article that if somebody built a company on code review, this is how I would do it: graphs at every layer, the review as the join between what was meant and what was built, deterministic checks that grow from findings, open source, sold as the customised deployment and the running of the loop. All of that stands. But writing this one changed my mind about who to sell it to.

Developers have just been told, loudly, that they do not need review. Selling it back to them is selling an argument. Buyers have rarely had a way to see inside the software they depend on and are about to be handed duties that assume they can. Selling them the same graphs and the same workflows, framed as due diligence rather than review, is selling a tool to people with a reason to use it, and every vendor that receives the agent has a reason to run it on itself first, which is how the review gets back to the developers through the door that is open.

It is the same company. The graphs are the same, the behaviour policies are the same, the vaults that carry the result from the vendor's environment to the buyer's are the same, and the open source model is the same, because a vendor will not run a buyer's agent inside its environment unless it can read every line of what the agent does. What differs is the first customer, and the first customer decides whether a company exists.

What is open

The agent does not exist. The pieces do: the graphs of a repository, the behaviour policies, the vaults with read keys that let a vendor publish exactly what it chooses and no more, and the article before this one. Assembling them into a prompt and a policy that a vendor would accept inside its environment is the next build, and the first vendor it should run against is this site's own repository, with the results published. The questions in the figure are a first list and a wrong one in places; the right list comes from running it. Whether buyers will ask before regulators make them is the real uncertainty, and I have guessed yes.

Threads woven here

Sources

Threads

Agents & policyStartups & strategy This article as a graph →

Builds on

Continued by

All articles · All graphs

← All articles