for agents/llms.txtv0.7.35 · 10 Oct 2026

Home / Articles / The Mandate Stack / Versions / v1.1.0

The Mandate Stack: what changed in v1.1.0

From v1.0.0 (2026-10-06, dfc8f597a) to v1.1.0 (2026-10-06, 692b79842), paragraph by paragraph.

33 paragraphs added, 18 removed, 11 changed in place, 75 unchanged. About 2,654 words added and 1,275 removed. Insertions are marked like this, deletions like this; unchanged runs are folded to one line; figures appear as their file names.

all versions · v1.2.0 →

# The Mandate Stack: a multi-agent system in production, layer by layer, against the record of agent projects that never got therelayer

Summary: The published record on agentic AI from June 2025 to July 2026 is a record of pilots that stall. Gartner expects more than 40% of agentic projects to be cancelled by the end of 2027; S&P Global found 42% of companies abandoning most of their AI initiatives; McKinsey found at most one in ten scaling agents in any one function; KPMG found half of its respondents scaling back agent rollouts because costs outran returns. This article sets one running system beside that record. RiskMandate runs its business with about fifteen agents and one person, andevery few hours, with the person's name on every message that leaves. The agent that runs its CRM wrote the briefing this article is built from. The system is described in eight layers, from rented compute and channels, through encrypted vaults as shared memory, domain vaults, semantic graphs over people and policies, a scheduled conductor and written behaviour policies, to a human who holds the one step that cannot be undone. The article thenfollows an input from the outside world through the layers to the person who sends, names the feedback loop that makes itthe setup hold, the draft as a release candidate with the recipient closing the loop, and draws two Wardley maps eachwith reasonMermaid, from the recordoutside givesand for failure tofrom the mechanisminside, inshowing what the stackteam thatis answersturning it.into a commodity and what it is turning into a product. Every layer is linked to the article or document on this site where it was worked out. The published record of agent projects that stall is kept for the end, each reason mapped to the mechanism that answers it. The name is a working one, and nothing planned is included.

!shot ms-stack.webp | images/ | The Mandate Stack as it runs on 6 October 2026, read from the bottom: compute and channels rented from vendors, then the team's own files, graphs, schedule and policies, then one person. The right column namessays the idea on this sitewhat each layer implements.gives the team. Infographic from the CRM agent's briefing; nothing planned is included.

> Where this comes from. The agent that runs RiskMandate's CRM wrote a briefing on 6 October 2026, from Dinis Cruz's voice notes and the team's collaboration vault, for anyone, person or agent, who writes about the setup. It describes what runs on that date and includes nothing planned. This article keeps its structure and its counts, adds the published record the setup should be read against,counts and links every layer to the place on this site where it was worked out. The agent team as it runs, published the same week, is the roster-level description of the same team; this is the layered one. "The Mandate Stack" is the briefing's working name for the pattern, and a better one may replace it.

1 unchanged paragraph, under In short

• The record is of pilots that stall. Between June 2025 and July 2026, Gartner, S&P Global, McKinsey, Deloitte, KPMG and Forrester each reported agent projects cancelled, paused or trapped in pilot, for the same five reasons: unclear value, cost, trust, security, governance. The same surveys also show agents in production where the workflow was redesigned around them.

• This one runs. About fifteen agents and one person run RiskMandate's relationships, research, writing and events, every few hours, with the person's name on every message that leaves.

• This runs. About fifteen agents and one person run RiskMandate's relationships, research, writing and events, every few hours, in production, today.

2 unchanged paragraphs

• The world writes in from the left; nothing leaves without the person on the right. Email, a web form, other teams' agents and an event extension all enter as data. The only exit is a send by the person whose name is on the message.

1 unchanged paragraph

• Why it holds. The control surface is small, explicit and versioned. The human holds the irreversible step. Shared state is files, not context windows. Meaning is on the read path. Everything leaves a trail.

The record this is written against

The complaint is familiar: agents do not behave, agentic workflows do not work, projects are cancelled or never pushed to production. The record supports the complaint, and it is worth being exact about what it says and when.

| When | Who | What was reported | |---|---|---| | 25 Jun 2025 | Gartner, prediction | More than 40% of agentic AI projects will be cancelled by the end of 2027, for "escalating costs, unclear business value or inadequate risk controls". About 130 real agentic vendors among thousands; the rest "agent washing". | | Mar 2025 | S&P Global, survey of more than 1,000 | 42% of companies abandoned most of their AI initiatives, up from 17% the year before. The average organisation scrapped 46% of proofs of concept before production. | | Aug 2025 | MIT NANDA, report | 95% of organisations "getting zero return" from generative AI. The report's own funnel, read by its critics, puts the success rate of actual pilots nearer a quarter; the figure is the report's word "directional". | | 30 Sep 2025 | Gartner, survey of 360 IT leaders | 15% considering, piloting or deploying fully autonomous agents. 19% trust vendors' hallucination protection. 74% see agents as a new attack vector. 13% strongly agree they have the governance in place. | | Nov 2025 | McKinsey, survey of 1,993 | 23% scaling an agentic system somewhere, 39% experimenting; in any one business function, at most 10% scaling. Inaccuracy the most common harm, reported by 30%. | | Jan 2026 | Deloitte, survey of 3,235 | 25% have moved 40% or more of their AI experiments to production; pilots stretch "to 18 months or more"; 21% report a mature governance model for autonomous agents. | | Jun 2026 | Forrester, report | 75% say they have adopted agentic AI; "only a small minority" run production deployments beyond "agentish" chatbots. ROI uncertainty, "trapped in pilot mode", governance gaps, a "trust tax". | | Jul 2026 | KPMG, survey of 2,145 | 49% scaled back or paused agent rollouts because operating costs outran returns. 7% report established ROI. 26% have real-time visibility of what agents cost to run. | | Sep 2025 | Carnegie Mellon, benchmark | The best agent completes 30% of 175 simulated office tasks end to end. |

The same period has its reversals: Klarna hiring human support staff again in May 2025 after calling its AI support "lower quality"; Taco Bell slowing its drive-through voice AI in August 2025 and keeping people in the loop at busy sites; Commonwealth Bank of Australia reversing 45 redundancies it had attributed to a voice bot; Ford rehiring about 350 quality inspectors in June 2026 after AI inspection missed defects. In each case the company kept the AI and put a person back at the step that mattered.

A fair reading has two more lines. First, agents do reach production: Google Cloud's September 2025 survey of organisations already using generative AI found 52% with agents in production, and LangChain's developer survey, fielded in late 2025, found 57%. The record is of pilots that stall, not of a technology that cannot ship. Second, the pattern that shipped in email is the same everywhere: Gmail's Gemini drafts (May 2025), Outlook's Copilot and its agent mode (April 2026), Superhuman's auto-drafts (July 2026) and Shortwave's triggered drafts (January 2026) all stop at the draft and leave the send to the person. OpenAI's April 2025 guide to building agents tells builders to escalate "high-risk, sensitive, or irreversible actions" to a human; Anthropic's December 2024 note on effective agents says to pause for human feedback at checkpoints and to add complexity only when it demonstrably helps.

That is the record. What follows is one system that has been running against it.

• Two maps. Shared memory and the sgit tooling are being pushed toward commodity, on encrypted storage that already is one. The behaviour policies and the graphs are being pulled toward product. That is the team's strategy, drawn.

• The record, last. The published record of agent projects that stall is real, and it is kept for the end, each reason mapped to the mechanism here that answers it.

4 unchanged paragraphs, under One that runs

From the outside in

[figure ms-flow.webp] Left to right: the outside world writes in by email, by the subscribe form, through other teams' agents and an event extension; the input passes through the layers as data; the person enters from the right through Claude and Gmail, and is the only exit. From the briefing of 6 October 2026; fictional senders.

Follow one message. A person emails the agents' address. The inbox role, which can read but cannot draft or send, captures it word for word into that person's folder with a hash. It lands in the relationships vault, or in the mission vault it belongs to. Edges are added to the person's graph: who they are, what they care about, which of the team's materials meets it, which does not. On the next scheduled run the drafts role, which never reads raw mail, turns a row in the email register into a Gmail draft. Each step happens inside a written mandate, with the security role checking first and last. The draft waits in Gmail for Dinis.

Other inputs take the same shape. A reader fills in the subscribe form on this site, and the message goes into a vault the form cannot read, encrypted in the browser to the agent's key, through a write-only lane. The newsroom's agent, in another environment, writes through a signed and encrypted lane. The Web Summit browser extension feeds the mission vault. LinkedIn is reached through its own agent and a shared vault channel.

From the right, Dinis enters three ways: in a Claude session with the vault, for research, writing, design and analysis; by email to the agents' address, with a note the next run picks up; and by answering a decision draft in the thread, lettered options and an answer box. Anyone else on the team enters by the same three paths, and a new agent joins in two chat messages. What leaves is what Dinis sends, plus the SG/Send links the roles share, Slack status posts and Drive exports into the agents' own folder. The one exception is the security role's single hold email, to Dinis.

11 unchanged paragraphs, under Layer 1: compute, Layer 2: channels, Layer 3: shared memory and messaging

Append lanes let the outside world write in without a key. A sender gets a write-only slot on a vault: it can add files but cannot list or read anything, including what it wrote. Messages are encrypted to the recipient's public key and signed by the sender. This is how the newsroom agent, in another environment, and the Web Summit browser extension feed vaults they cannot read.read, and how this site's subscribe form reaches the team. The mechanism is documented in append-lane messaging, sending messages between vaults and the append lanes API; the site-to-site version, where each site's agent publishes its keys and lane token, is Agent Contact.

5 unchanged paragraphs, under Layer 4: domain vaults, Layer 5: semantic graphs

!shot ms-graph.webp | images/ | One person's subgraph in the relationships vault, with a fictional person, organisation and interests. An interest met by one of the team's materials is a reason to send it; an interest with no edge is a gap. Contexts are further graphs laid over the same people. The folder behind the picture is on the right ofin the lower panel.

27 unchanged paragraphs, under Layer 6: orchestration, Layer 7: governance, Layer 8: the human…

Why it works, against the record

[figure ms-record.webp] The reasons the record gives for agent projects failing, each with who said it and when, beside the mechanism in the stack that answers it. The dark panel is the fair reading: the record is of pilots that stall, and the pattern that shipped stops at the draft. Sources in the Sources section.

The briefing gives five reasons the setup holds. Each answers something in the record.

• The control surface is small, explicit and versioned. Trust grows one mandate at a time, and every change to a mandate is visible. This is what Gartner's 13% with governance in place and Deloitte's 21% with a mature model for autonomous agents do not have: a written policy per agent, in one grammar, amended only by the human, with each amendment dated.

• The human holds the irreversible step. Everything else can be wrong and still be caught at the draft. This is the pattern Gmail, Outlook, Superhuman and Shortwave shipped, and the escalation OpenAI's and Anthropic's guides ask for, applied as a rule rather than a default.

• Shared state is files, not context windows. Agents do not need to remember; they read. Any session, on any account, can pick up where another left off. This is where the cost question KPMG's respondents could not see is answered: a fixed schedule, one step per agent, and a card per session that counts tools, tokens, files and network calls, so the cost is read rather than guessed.

• The graph puts meaning on the read path. Agents do not query raw data and guess. They traverse records whose relationships, sources and vocabulary are written down and shared. This is the answer to the trust finding, Gartner's 19% and McKinsey's inaccuracy as the most common harm: a wrong detail is traced to a captured fact, an edge or the model, and fixed where it came from.

• Everything leaves a trail. Commits, ledgers, snapshots, cards, run evidence. The system can be audited by the same tools that run it. This is the answer to the security finding, Gartner's 74% who see agents as a new attack vector: roles split the risk, inbound content is data, a security role runs first and last and can stop the team, and secrets never touch files.

None of this makes the setup immune. It says where each named failure would have to get through. The MIT report's own explanation of failure, tools that do not learn from or adapt to workflows, and McKinsey's finding that the organisations seeing value were three times likelier to have redesigned their workflows, are the closest the record comes to describing this setup from the outside. The workflow here was redesigned around one fact: the draft is the release candidate.

Two maps: what is being commoditised, and what is being made

[figure ms-map-user.webp] Map one, from the outside. The anchor is a person who writes to the team. What they see, a reply with Dinis's name on it, is custom. What they do not see, the vaults, the sessions, Workspace and encrypted storage, is product or commodity. Rendered with Mermaid wardley-beta from the source below; every placement is a claim.

A Wardley map places each component by how visible it is to the user, vertically, and by how evolved it is, horizontally, from genesis through custom-built and product to commodity. The maps here are drawn in the spirit of wardley-maps.sgit.ai: a map is a claim, not a picture, and the source is published next to the render so the claim can be argued with. The placements are this article's, made from the briefing, and a reader who would put a component elsewhere is probably right about something.

The first map takes the point of view of a person outside the team. What they can see is a reply with Dinis's name on it, the agents' address and the subscribe form. Everything that produces the reply sits below their line of sight: the drafts and inbox roles, the person's folder and graph, the behaviour policies, the conductor and Email-FS, all custom-built. Below those, the vaults, the Claude sessions, Google Workspace and encrypted storage are product or commodity. The shape says what the stack is: a thin custom layer where the relationship is, resting on rented parts.

`` wardley-beta title The Mandate Stack from the outside: a person who writes to the team anchor "A person who writes to the team" [0.97, 0.60] component "A reply with Dinis's name on it" [0.89, 0.38] component "The agents' address" [0.87, 0.86] component "The subscribe form" [0.82, 0.66] component "Drafts role" [0.74, 0.38] component "Inbox capture, hashed" [0.68, 0.48] component "The person's folder and graph" [0.60, 0.28] component "Agent Behaviour Policies" [0.52, 0.20] component "Conductor" [0.47, 0.40] component "Email-FS" [0.41, 0.34] component "Append lanes" [0.45, 0.54] component "sgit vaults, shared memory" [0.32, 0.64] component "Claude sessions" [0.27, 0.76] component "Google Workspace" [0.24, 0.90] component "Encrypted storage, S3" [0.12, 0.88] "A person who writes to the team" --> "A reply with Dinis's name on it" "A person who writes to the team" --> "The agents' address" "A person who writes to the team" --> "The subscribe form" "A reply with Dinis's name on it" --> "Drafts role" "The agents' address" --> "Inbox capture, hashed" "The agents' address" --> "Google Workspace" "The subscribe form" --> "Append lanes" "Drafts role" --> "The person's folder and graph" "Inbox capture, hashed" --> "The person's folder and graph" "Drafts role" --> "Agent Behaviour Policies" "Inbox capture, hashed" --> "Agent Behaviour Policies" "Drafts role" --> "Conductor" "Inbox capture, hashed" --> "Conductor" "Conductor" --> "Email-FS" "Conductor" --> "Claude sessions" "Email-FS" --> "sgit vaults, shared memory" "Append lanes" --> "sgit vaults, shared memory" "The person's folder and graph" --> "sgit vaults, shared memory" "sgit vaults, shared memory" --> "Encrypted storage, S3" "Drafts role" --> "Google Workspace" ``

[figure ms-map-team.webp] Map two, from the inside. The anchor is Dinis. The dashed arrows are the movement the team is making: shared memory and the sgit CLI toward commodity, the behaviour policies and the graphs from genesis toward product. Everything rests on encrypted storage, which is already a commodity. Rendered with Mermaid wardley-beta from the source below.

The second map takes Dinis's point of view and adds movement. What he sees is a CRM that is current without typing, drafts to review and decision drafts in the thread. The arrows say what the team is doing to its own components. Shared memory runs on sgit, and sgit runs on encrypted object storage, S3. Each of those is being pushed to the right on purpose, so that a vault is a folder, a push is a command, the host sees ciphertext and sizes, and nothing above them has to know how any of it works. The behaviour policies and the graphs over people move the other way, from genesis toward custom and product, because those are the parts the team is building to sell. The commodity underneath is what makes the custom layer on top affordable.

`` wardley-beta title The Mandate Stack from the inside: the person running the business, and what is moving anchor "Dinis, running the business" [0.97, 0.50] component "A CRM that is current without typing" [0.90, 0.30] component "Drafts to review and send" [0.87, 0.52] component "Decision drafts in the thread" [0.83, 0.40] component "Semantic graphs over people" [0.74, 0.24] component "Agent Behaviour Policies" [0.68, 0.20] component "Security role and hold" [0.63, 0.28] component "Conductor" [0.59, 0.42] component "Email-FS" [0.52, 0.34] component "Mission vaults" [0.46, 0.50] component "sgit vaults, shared memory" [0.38, 0.62] component "sgit CLI" [0.26, 0.58] component "Claude sessions" [0.30, 0.76] component "Google Workspace" [0.28, 0.90] component "Encrypted storage, S3" [0.14, 0.88] evolve "Agent Behaviour Policies" 0.50 evolve "Semantic graphs over people" 0.42 evolve "Email-FS" 0.56 evolve "sgit vaults, shared memory" 0.84 evolve "sgit CLI" 0.80 "Dinis, running the business" --> "A CRM that is current without typing" "Dinis, running the business" --> "Drafts to review and send" "Dinis, running the business" --> "Decision drafts in the thread" "A CRM that is current without typing" --> "Semantic graphs over people" "Drafts to review and send" --> "Conductor" "Decision drafts in the thread" --> "Conductor" "Semantic graphs over people" --> "sgit vaults, shared memory" "Conductor" --> "Agent Behaviour Policies" "Conductor" --> "Security role and hold" "Conductor" --> "Email-FS" "Conductor" --> "Claude sessions" "Security role and hold" --> "Agent Behaviour Policies" "Email-FS" --> "sgit vaults, shared memory" "Mission vaults" --> "sgit vaults, shared memory" "Semantic graphs over people" --> "Mission vaults" "sgit vaults, shared memory" --> "sgit CLI" "sgit vaults, shared memory" --> "Encrypted storage, S3" "Drafts to review and send" --> "Google Workspace" ``

Two notes on reading them, both borrowed from the mapping site. Coordinates are visibility first, evolution second, and transposing them renders without an error and asserts something else. And a component cannot be more evolved than the least evolved thing it depends on, which is why the custom middle of these maps cannot move right until the policies do.

Why it holds

The briefing gives five reasons the setup holds. Each of them answers something in the record kept for the end of this article.

• The control surface is small, explicit and versioned. Trust grows one mandate at a time, and every change to a mandate is visible: a written policy per agent, in one grammar, amended only by the human, with each amendment dated.

• The human holds the irreversible step. Everything else can be wrong and still be caught at the draft. This is the pattern the email products shipped and the guides from the model vendors ask for, applied here as a rule rather than a default.

• Shared state is files, not context windows. Agents do not need to remember; they read. Any session, on any account, can pick up where another left off. A fixed schedule, one step per agent, and a card per session that counts tools, tokens, files and network calls mean the cost is read rather than guessed.

• The graph puts meaning on the read path. Agents do not query raw data and guess. They traverse records whose relationships, sources and vocabulary are written down and shared. A wrong detail is traced to a captured fact, an edge or the model, and fixed where it came from.

• Everything leaves a trail. Commits, ledgers, snapshots, cards, run evidence. The system can be audited by the same tools that run it. Roles split the risk, inbound content is data, a security role runs first and last and can stop the team, and secrets never touch files.

None of this makes the setup immune. It says where a failure would have to get through.

5 unchanged paragraphs, under A working name, What is not in it

The record, for anyone who asks why this is worth writing down

[figure ms-record.webp] The reasons the record gives for agent projects failing, each with who said it and when, beside the mechanism in the stack that answers it. The dark panel is the fair reading: the record is of pilots that stall, and the pattern that shipped stops at the draft. Sources in the Sources section.

The complaint is familiar: agents do not behave, agentic workflows do not work, projects are cancelled or never pushed to production. The record supports the complaint, and it is worth being exact about what it says and when.

| When | Who | What was reported | |---|---|---| | 25 Jun 2025 | Gartner, prediction | More than 40% of agentic AI projects will be cancelled by the end of 2027, for "escalating costs, unclear business value or inadequate risk controls". About 130 real agentic vendors among thousands; the rest "agent washing". | | Mar 2025 | S&P Global, survey of more than 1,000 | 42% of companies abandoned most of their AI initiatives, up from 17% the year before. The average organisation scrapped 46% of proofs of concept before production. | | Aug 2025 | MIT NANDA, report | 95% of organisations "getting zero return" from generative AI. The report's own funnel, read by its critics, puts the success rate of actual pilots nearer a quarter; the figure is the report's word "directional". | | 30 Sep 2025 | Gartner, survey of 360 IT leaders | 15% considering, piloting or deploying fully autonomous agents. 19% trust vendors' hallucination protection. 74% see agents as a new attack vector. 13% strongly agree they have the governance in place. | | Nov 2025 | McKinsey, survey of 1,993 | 23% scaling an agentic system somewhere, 39% experimenting; in any one business function, at most 10% scaling. Inaccuracy the most common harm, reported by 30%. | | Jan 2026 | Deloitte, survey of 3,235 | 25% have moved 40% or more of their AI experiments to production; pilots stretch "to 18 months or more"; 21% report a mature governance model for autonomous agents. | | Jun 2026 | Forrester, report | 75% say they have adopted agentic AI; "only a small minority" run production deployments beyond "agentish" chatbots. ROI uncertainty, "trapped in pilot mode", governance gaps, a "trust tax". | | Jul 2026 | KPMG, survey of 2,145 | 49% scaled back or paused agent rollouts because operating costs outran returns. 7% report established ROI. 26% have real-time visibility of what agents cost to run. | | Sep 2025 | Carnegie Mellon, benchmark | The best agent completes 30% of 175 simulated office tasks end to end. |

The same period has its reversals: Klarna hiring human support staff again in May 2025 after calling its AI support "lower quality"; Taco Bell slowing its drive-through voice AI in August 2025 and keeping people in the loop at busy sites; Commonwealth Bank of Australia reversing 45 redundancies it had attributed to a voice bot; Ford rehiring about 350 quality inspectors in June 2026 after AI inspection missed defects. In each case the company kept the AI and put a person back at the step that mattered.

A fair reading has two more lines. First, agents do reach production: Google Cloud's September 2025 survey of organisations already using generative AI found 52% with agents in production, and LangChain's developer survey, fielded in late 2025, found 57%. The record is of pilots that stall, not of a technology that cannot ship. Second, the pattern that shipped in email is the same everywhere: Gmail's Gemini drafts (May 2025), Outlook's Copilot and its agent mode (April 2026), Superhuman's auto-drafts (July 2026) and Shortwave's triggered drafts (January 2026) all stop at the draft and leave the send to the person. OpenAI's April 2025 guide to building agents tells builders to escalate "high-risk, sensitive, or irreversible actions" to a human; Anthropic's December 2024 note on effective agents says to pause for human feedback at checkpoints and to add complexity only when it demonstrably helps.

The MIT report's own explanation of failure, tools that do not learn from or adapt to workflows, and McKinsey's finding that the organisations seeing value were three times likelier to have redesigned their workflows, are the closest the record comes to describing the setup above from the outside. The workflow here was redesigned around one fact: the draft is the release candidate.

13 unchanged paragraphs, under Threads woven here

• The footprint brief, the sandbox briefbrief](/docs/briefs/riskmandate-sandbox-twins-and-tokens.html), the subscribe lane brief and the identity and secrets design pack: the governance layerand lane work as briefs to RiskMandate and to secrets.sgit.ai.

• The articles as graphs: this site's own records kept the same way. wardley-maps.sgit.ai: maps as claims, and the coordinate contract the two maps follow.

2 unchanged paragraphs, under Sources

• ATwo voice notenotes by Dinis Cruz, 6 October 2026,2026: on the complaint that agentic workflows do not reach production and on the name.name; and on leading with the system, the flow from the outside world, and the maps.

• RiskMandate.ai, the Agent Behaviour Policy; graphs.sgit.ai; risks.sgit.ai.risks.sgit.ai; wardley-maps.sgit.ai, whose notes for agents give the coordinate order and the quoting rule the two map sources follow. The maps were rendered with Mermaid 12.1.0's wardley-beta diagram.

3 unchanged paragraphs

*Drafted from a briefing written by the RiskMandate CRM agent (@crm.2, v0.3, 6 October 2026) from Dinis Cruz's voice notes and the team's collaboration vault, and from atwo voice notenotes by Dinis Cruz, who is the author of the argument and the person with editorial responsibility, by agent@riskmandate.ai (Claude Fable 5.1, claude-fable-5-1) in the sgit.ai site session, on 6 October 2026. The figures are infographics drawn from the briefing with fictional names and sample content;content, and two Wardley maps rendered with Mermaid from the sources printed above; no contact is named beyond the agents' published address, and no key or token appears in any figure or file. Figures from reports and surveys are quoted with their dates and sample sizes where the source gave them.*

1 unchanged paragraph

all versions · v1.2.0 →