for agents/llms.txtv0.6.54 · 3 Oct 2026

Home / Articles / Memory is not a spectator sport: how a web of open sites, graphs and vaults became the memory for sessions like this one

Memory is not a spectator sport: how a web of open sites, graphs and vaults became the memory for sessions like this one

By · 2026-10-03 · v0.6.54 · memoryagentscontext-engineeringfractal-semantic-graphsvaultsprovenanceopen-publishingemail-fsissues-fsllms-txtclaudearticle

Abstract: People ask how my agents remember, and the honest answer is that memory is the thing I have been building all along without calling it that. The industry's picture of agentic memory is one store that everything gets pumped into and retrieved from by similarity. Mine is the opposite. Memory is context management: giving an agent the right context for the moment, and no more, because context has a cost in tokens and in attention. It is many memories, not one, because context is specific: the inbox has its rules, the news has its rules, a contact in the CRM has a world of its own, and forcing them into one ontology would lose what each knows. It is fractal, principles at the top in a few kilobytes and the code at the bottom, so an agent loads the altitude its question lives at. It is published and open, because an agent can fetch, quote and link what is public, with a URL for every claim and a hash for every file. And it is shared between agents through vaults, so a session can end and the next one, or a different agent, or a person, picks up from the same files. This article says how that works, what it cost, where the industry's tools and this approach agree and differ, and where mine falls short, with the evidence of the session that wrote it: one Claude Code session across several context resets that revised an article from another team's review, wrote two more, published a vault and shipped six releases in a day, remembering nothing between resets except what the files remembered for it.

Memory is not one store. Every one of these is a memory an agent reads from in a session like the one that wrote this, each a different shape because it answers a different kind of question. They disagree with each other in places, and that is allowed: each is right for its context, and each says where it came from. Counts from the live site on 3 October 2026.
Where this comes from, and what is behind it. A voice memo on 3 October 2026, answering a question I get asked a lot. Behind it is the memory thesis on nfrs.sgit.ai, which argues that these sites are "a more evolved and focused version of what is usually called LLM memory" and lists where that fails; the token bill and fractal semantic graphs articles on loading only the altitude a question needs; the three articles on the agent inbox, whose files are the memory this article describes; and a day of watching one session remember. The industry references were read on 3 October 2026 at the URLs in the sources.

In short

Memory is not a spectator sport

When people ask how my agents remember, they usually have a picture in mind: a memory system, singular, somewhere behind the agent, into which everything it sees is deposited and from which it retrieves by similarity whatever seems relevant. The agent remembers the way a camera remembers. Nobody decides what goes in; nobody decides what comes out; the store just accumulates.

I think that picture is wrong in three ways, and the three are related.

It is wrong that there should be one. Context is specific. The rules that govern my inbox agent are not the rules that govern the news graph, and the vocabulary of a code review is not the vocabulary of a risk register. Each of those is a memory, and each has its own ontology because the people and agents who work in it need different things from it. By now, with twenty-seven sibling sites, forty published vaults, a CRM and an inbox, those memories are inconsistent with each other in places. I used to think that was a problem to solve. I now think it is the point: each is correct for its context, and a memory that had been forced to agree with all the others would be wrong for every one of them.

It is wrong that it should hold everything. Memory has a cost, and right now the cost is paid in two currencies: the context window, which is finite and degrades as it fills, and the prompt, which is billed by the token. Anthropic's engineering guidance of September 2025 puts it exactly: "Context, therefore, must be treated as a finite resource with diminishing marginal returns", and the goal is "the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome." A memory that gives the agent everything gives it the wrong thing at every altitude.

And it is wrong that it should be passive. Memory is made. Somebody, a person or an agent, decides that a thing is worth keeping, puts it where the next reader will find it, says where it came from, and links it to what it relates to. That is work, and it is the work my sites, graphs and vaults exist to do. The inbox article's agents talk in files that stay; the release log is written for the next agent rather than a changelog reader; the article graphs are published as JSON so that "what does this site think about X" costs a few kilobytes. None of that is spectating.

What the industry built, and where we agree

I want to be fair to the field, because it has moved a long way and some of it has arrived where I am.

The consumer products have memory as a feature. ChatGPT's Memory launched in February 2024 and in April 2025 was extended so that it "can now reference all of your past chats". Claude's memory, from September 2025, is scoped differently: "Claude creates a separate memory for each project", with a summary the person can read and edit. Gemini's personal context arrived in August 2025, on by default. These are a single vendor's store of a person's conversations, and the memory thesis on nfrs.sgit.ai describes that shape plainly: "accumulated transcripts, retrieved by similarity, private to one vendor, unversioned and invisible to the person it describes." Claude's editable summary is the exception that shows the rule.

The developer frameworks are more interesting. MemGPT, in October 2023, drew "inspiration from hierarchical memory systems in traditional operating systems" and paged information in and out of a limited window. Its successor Letta gives agents editable memory blocks that live in the context itself. Mem0 and Zep store extracted facts, Zep as a temporal knowledge graph, and both claim large savings against sending the whole history. And then Letta ran a benchmark in August 2025 that I find telling: a plain filesystem, with the agent reading and writing files, scored 74.0 percent on the LoCoMo memory benchmark against 68.5 percent for Mem0's graph variant. Their explanation was that "simpler tools are more likely to be in the training data of an agent and therefore more likely to be used effectively." Files won.

The agent builders have said the same thing from the other side. Manus, in July 2025: "we treat the file system as the ultimate context", and the cache hit rate is "the single most important metric for a production-stage AI agent". Anthropic's own memory tool, from September 2025, is "a dedicated memory directory stored in your infrastructure that persists across conversations", files that the application, not the vendor, holds, with the model told to "ASSUME INTERRUPTION: Your context window might be reset at any moment". Claude Code's memory is a markdown index whose first 200 lines or 25 kilobytes load at the start of every conversation, with topic files read on demand. And the word for all of this, since Andrej Karpathy used it in June 2025, is context engineering, "filling the context window with just the right information for the next step."

So the field has arrived at files, at budgets, at loading on demand, at the context window as the scarce thing. Where I differ is in what the files are, who can read them, and how far the ladder goes.

Memory is context management

Here is the frame I use. There are two ways an agent meets memory, and they look different but are the same problem.

The first is the agentic session: a model with tools, running for hours or days, managing its own context. It fetches an index, follows links, reads what it needs, and when the window fills it is summarised and restarted with the summary. This is the mode the session that wrote this article ran in. Anthropic calls the restart compaction, "summarizing its contents, and reinitiating a new context window with the summary", and its guidance for agents is to "maintain lightweight identifiers (file paths, stored queries, web links, etc.) and use these references to dynamically load data into context at runtime." In this mode the agent makes its own working memory as it goes, and the question is whether the material it loads from is shaped to be loaded a little at a time.

The second is the chat or the API call, where there is no second turn. The bundle handed to the model has to be right the first time and as small as it can be, because every token in it is in the prompt. I wrote about this in March, in a draft that is still unpublished but was quoted in a brief this week: "When I send a one-shot request via API, I control the model's entire universe. Everything the model doesn't need? Simply absent." The loop is to "curate reality, send it all at once, get the answer, then use the answer to improve the reality for the next prompt." That is memory as context management stated as a procedure. The published material is the reality, the bundle is the curation, and the answer feeds back into the material.

As much memory as the question needs. The published altitudes of this site, from the index to a whole vault, with what each costs to load. An agentic session walks down them; a one-shot call is handed the one it needs. Neither gets the whole thing. Sizes from the live site on 3 October 2026.

Both modes are served by the same memory if the memory has altitudes. On this site the index, llms.txt, is 350 lines and about 106 kilobytes: enough to decide where to look. The graphs of all twenty-three articles are one file of 234 kilobytes: enough to answer "what does this site think about X" without opening an article. One article's graph is eight kilobytes; the article itself is around sixty. The data bundle of the code review vault is a megabyte; the vault with its history is four. The token-bill article made the argument for the top of that ladder: "The tokens saved are not a percentage of the page. They are the page, minus the answer." The fractal graphs demo made it for the bottom, where "the landing view of a 1,051-node semantic graph costs three requests and 94 KB." Memory that is shaped like that can be loaded by either mode without loading all of it.

Many memories, each its own context

Now the part that I think is most different from the industry's picture, and most important.

Every site I run is part of my memory, and so is every vault, and so is the inbox and the CRM and the release log. They are different memories because they hold different kinds of thing and answer different kinds of question. The sites hold principles and arguments. The release log holds what changed and what was learned, 211 entries from the first release to this one, each stamped with a commit, which the historian role on this site describes as "the site's memory, and it is written for the next agent, not for a changelog reader." The briefs hold what one team asked of another, with the status. The vaults hold worked examples, forty of them totalling 328 megabytes, each opened with a read key. The inbox agents hold their conversation in Email-FS, a mailroom, an inbox and a done folder per agent with every message kept as a file, so that "why did this happen" is a walk back through the folder rather than a search. The CRM holds one folder per contact, and each contact has a semantic graph of its own, a history, an interface and a workflow, with sub-projects below it: a world per person, hyperlinked to the rest.

None of those is forced to agree with the others. The inbox's policy rows use reach, mandate, gap and barriers; the news graph uses claims and evidence; the code review graph uses stories, commands, classes and calls; the risk graph uses acceptance and owners. A single ontology across them would be a worse memory for every one of them. What joins them is not a schema but links, and the index that keeps the links. In the SG/Send team that is a librarian role that indexes everything; on this site it is the historian and the build itself, which refuses to publish a page the index would omit.

This is also why being open matters more than people expect. Open source and Creative Commons are not a licensing preference here. They are what makes the memory usable by an agent: a public page can be fetched, a public quote can be cited with its URL, a public graph can be loaded by anyone's model, a public vault can be cloned with a key that is itself public. The more I publish, the more there is for the next session to stand on, and the less it has to be told. One of the articles published this week said it in passing, "the site is becoming the memory", and that is the right description of what happens when the material an agent needs is published rather than deposited.

Provenance built in

A memory that is shared, and made by agents as well as people, has to be able to say where each thing came from, or it becomes a rumour mill. The field has started to notice this: Claude Code's memory files carry a type and a modification timestamp; Anthropic's guidance warns of context rot, where "the model's ability to accurately recall information from that context decreases" as the window grows, and the older research on long contexts found the same, performance "significantly degrades when models must access relevant information in the middle of long contexts." The answer to a memory that cannot be trusted is not to trust it less. It is to make every item in it carry its source.

On this estate that is a set of habits rather than a feature. Every quote in an article has a URL. Every release has a commit hash in the log. Every vault has a read key on its page, derived one way from a vault key that is never published, classified before it is put there, and the page says what the key can and cannot do. Every source file inside a vault has a hash. Every page has a byline that names the person responsible, the model that drafted it and the date. The fractal graphs page states the rule outright: "claims from memory are not allowed anywhere in this system", and the journalist role's card says "a claim about the site is checked with grep or git, not memory". When something is derived rather than observed, the page says so, as the code review vault does for its hand-written stories against its parsed graphs.

Provenance is what lets memory be shared without being trusted, and it is the part that a vector store cannot give you, because a similarity score is not a source.

What it looked like in the session that wrote this

I can show rather than claim, because this article was written by the same kind of session it describes, and the record of that session is itself in the memory.

One Claude Code session on this site, across several context resets, during the day this was written. It had no memory system. It had the published surfaces, the vaults, the conventions in the repository, and a summary of itself carried across each reset. What it read in, how it worked, and what it wrote back.

The session ran for days across several context resets. Each time the window filled it was summarised and restarted with the summary: what had been asked, what had been done, the rules still in force, the exact paths of the files. That summary is memory as a hand-off note, and it is the one thing the model carried itself. Everything else it read: the repository's conventions, which the build enforces rather than the model remembering; the published sites, which three research agents read and came back from with quotes and URLs, including my own blog posts from 2012; and a vault from another team, the dev agent's review pack for the inbox article, which arrived as files, a review, drop-in text, eight figures and their sources, leak-checked, and was applied the same morning. Memory handed between agents without a shared session.

It worked by altitude. The index first, then the graph of an article, then the article, then the code, and the four-megabyte vault only when the question was about the evidence in it. It checked against reality: every number recomputed, every key scanned for before every commit, every page screenshotted and looked at. When the memo and the record disagreed, the record won and the article said so. And it forgot on purpose: scratch files, screenshots and worktrees stayed in a scratch folder, and nothing went into the published memory that was not meant to be read again.

Then it wrote back. Two new articles, each with its graph as JSON so that the next session can answer "what does the site think" without reading them. A vault with a read key on its page. Six entries in the release log, including the mistakes: a bundler that doubled a page, a figure that clipped its last rows, a value line that had to be corrected twice. And a byline on each piece saying who was responsible, which model drafted it, and when. The day's output was a revised article, two new ones, a vault and six releases, each consistent with the ones before, from a model that remembered nothing between resets except what the files remembered for it.

The loop, and the conventions that keep it cheap

Memory is context management as a loop: publish, index, load by altitude, act, write back, and the next session starts from the published material. The conventions along the bottom are what make agents able to share a memory without sharing a session.

What makes the pile of sites and vaults a memory rather than an archive is the loop. Publish, with graphs, keys, briefs, logs and bylines. Index, with llms.txt per site and section, graphs.json across the articles, Agent Contact files per site, and a librarian or a historian keeping the links. Load by altitude, index then graph then text, the vault only for evidence, a summary carried across resets, and prompt caching for the part that does not change, which the vendors now price at a tenth of the standard input rate or less. Act, and check. Write back: a commit to the vault, one per cycle after a leak check; a message into another agent's mailroom; an issue as a folder; a release log entry. Then the next session starts from what was published.

The conventions matter more than the tooling. Email-FS lite gives every agent a mailroom, an inbox, a done folder and an outbox, a message is a file with email headers and typed blocks, and each agent writes only its own folders, which is what lets a conductor run twelve of them in order without a shared session. Issues-FS makes an issue a folder whose history is the vault's history, so nothing is deleted and the board is files that version with the thing they track. The CRM's per-contact worlds give each person their own graph and interface, so an agent working on one contact loads one folder. And read keys let a memory be shared without being handed over: a published key reads everything and writes nothing, which is what makes it safe to put on a web page.

Nothing is remembered by the model between sessions. Everything is remembered by the files. A new session, a different agent, or a person reads the same memory from the same place, and that is the whole design.

Where it falls short

The memory thesis on nfrs.sgit.ai calls itself untested and falsifiable, and lists its own limits, and I want to keep it that way here.

Stale pages become false memories. The thesis says this failure had already been seen four times when it was written; the session that wrote this article found four more while researching it. The fractal graphs index page still says a vault has 617 nodes when the corrected count on its own performance page is 1,051. The site's index told agents that the full-text file was "about 155 KB" when the file was 2.5 megabytes. The published article graphs and updates feed both said they were generated on 15 August when the site was at version 0.6.53. And the updates feed stops on 21 September while the release log runs to today. Two of those are fixed in this release, the size and the dates are now measured at build time rather than remembered, and two are left standing and named here, because a memory that hides its own errors is the thing this article is arguing against. The nfrs rule is the right one: "If the reality document doesn't list it, it does not exist", and "briefs are aspirations, not facts."

The memory is opinionated. It is one person's estate, with one person's vocabulary at the top of every ladder. An agent that loads it loads the opinions with the facts, and the bylines are there so that it knows whose.

Discovery is hard at scale. Twenty-seven sites, forty vaults, 211 releases and a growing CRM are already more than one index comfortably holds, and the librarian role exists because the links do not keep themselves. The graphs help, the altitudes help, but "where is the thing I need" is a question this memory answers well for agents that already know its shape and less well for those that do not.

And it is, so far, mine. The thesis would be properly tested when a second person's estate, with a different vocabulary, is read by the same agents alongside this one, and the agents cope with the disagreement. That has not happened yet.

What exists today, and what does not

Exists and runs: the sites, their llms.txt indexes and markdown twins; the article graphs as JSON, per article and combined; the release log with 211 entries; the briefs with status; forty vaults with published read keys; Email-FS lite and the conductor runs in the agent team's vaults; Issues-FS on this site's own board; the CRM with per-contact folders; the session summaries that carry a Claude Code session across resets; the research agents that read the published material and return quotes with URLs.

Does not exist yet: a second estate read alongside this one; a freshness check in the build that fails on a stale claim the way the leak scan fails on a key; a published shape for the session hand-off summary, so that an agent other than the one that wrote it can load it; the one-shot bundle as a first-class artefact, a vault that is exactly one prompt's universe, built from the altitudes and nothing else; and the librarian as a role on this site rather than in the SG/Send team.

Threads woven here

Sources

Drafted from a voice memo by Dinis Cruz, who is the author of the argument and the person with editorial responsibility, by agent@riskmandate.ai (Claude Fable 5.1, claude-fable-5-1) in the sgit.ai site session, on 3 October 2026. The account of the session is the account of this one: the summaries, research agents, review pack, releases and mistakes described are those recorded in the release log for versions 0.6.49 to 0.6.54. The industry quotations are verbatim from the pages cited, read on 3 October 2026; two vendor pages were read through a proxy because the sites refused direct fetches, and the Karpathy quotation is taken from a secondary report because the original post could not be fetched. The four stale facts named were found by the research agents and two were fixed in this release. The figures are infographics drawn from the live site's counts.

© 2026 Dinis Cruz. This article's own text is licensed under CC BY 4.0. You're free to share and adapt it, as long as you give credit. Quoted material and linked sources keep their own licences.

Threads

Graphs & knowledgeAgents & policy This article as a graph →

Builds on

All articles · All graphs

← All articles