for agents/llms.txtv0.6.20 · 29 Sep 2026

Home / Articles / Six agents, one inbox: what a real multi-agent setup taught me about access policies

Six agents, one inbox: what a real multi-agent setup taught me about access policies

By · 2026-09-29 · v0.6.19 · agentsconnectorsaccess-policiesnon-human-identitygmailgithubriskmandatesecurityarticle

Abstract: For the past few weeks I have run agents on dedicated accounts, a Google Workspace seat, a Claude Team seat and a GitHub account per agent, and split the work across six roles: a scheduled reader of the inbox, a mailbox agent that drafts, an inbox agent that sends, a CRM agent, a dev team agent and a site editor. This article is what that setup taught me, and it is mostly about the gap between the policy I wanted and what the tools can enforce. Three findings. The account, not the session, is the blast radius, so a dedicated account per agent is the first real control anyone has, and it turns out to do more than segregate, because it puts each agent in its own organisational unit where Google's compliance rules become per-agent enforcement. The first exception arrived before the first policy was written: the reader that must never reply must reply when the message comes from me, which is an authentication problem, not a permissions one. And the policy I had written on the assumption that the Gmail connector could not send attachments was wrong, because an agent found the attachments field, proved it with a signed PDF, and wrote up how. The vendor's own two documentation pages disagree about whether the connector can send at all. So the article ends with a table of every rule in the setup against how it is enforced today, by identity, by scope, by a compliance rule, by an approval prompt or by nothing but the agent's good behaviour, and with the argument that a policy is only as real as its worst row.

Six roles, two identities, one policy table. Each agent runs on its own Workspace, Claude and GitHub accounts. The cells say what each agent may do with mail, code and vaults; the colour says how the rule is enforced today: by the identity boundary, by a compliance rule an admin set, by an approval prompt, or by nothing but the agent doing as it was told.

Here is the claim, stated so it can be wrong. An access policy for an agent is only as real as its worst row. Write the policy as a table, one row per rule, and add a column for how each rule is enforced. If the column says the agent has been told, that row is a hope, and the policy is exactly as strong as that hope on the day something goes wrong. This article is about filling in that column for a setup I actually run, and about three things I did not expect to find.

In short

The setup, and what it costs

The agents run for two identities: the agent for RiskMandate, and the agent for my own domain. I am not going to print their addresses, for the obvious reason, but the shape is the same for both and it is copyable. Each identity has a dedicated Google Workspace account, which is about £7 a month, a seat on a Claude Team plan, and a GitHub account of its own. Nothing the agent does runs as me. Everything it does runs as an account whose mailbox, drive, repositories and connectors I chose.

That sounds like an obvious thing to do. It is the thing almost nobody does, because connectors are designed to be switched on by a person for that person, and the whole appeal is that the agent can see what you see. The research on nhi.sgit.ai said this plainly when it scored the shared-drive options in August: "Everything available runs on your identity, the only granularity is segregation rather than scoping, and nothing supports per-agent keys." Its conclusion, that the practitioner answer is "a shared drive dedicated to the agent", is what the dedicated account generalises. If segregation is the only granularity you can buy, buy it at the account level, where it covers mail, files, code and connectors at once.

Six roles share those two identities, and the split follows what each one is allowed to touch:

RoleWhat it doesRuns where
The readerA scheduled session that reads the inbox, works out what each message wants, and writes a tagged list of what needs to happenScheduled Claude session
The mailboxOperations and management on the mailbox itself: labels, threads, filing, drafting repliesCowork
The inboxThe one role allowed to send, and only from drafts the mailbox agent preparedCowork
The CRMCustomers, workflows and tasks, built on what the reader taggedCowork
The dev teamThe vaults, the vault UI and the tooling behind themCowork, sgit
The site editorEdits and publishes the websitesClaude Code

The point of six roles rather than one is not neatness. It is that each role gets a policy short enough to check, and the policies differ in the one place that matters: who may send.

Finding one: the account does more than segregate

I expected the dedicated account to bound what an agent could see. What I did not expect was how much it lets an administrator enforce, because Google Workspace's Gmail controls are set per organisational unit, and an account of its own can be a unit of its own.

Three settings turned out to matter, all under Apps, Google Workspace, Gmail, Compliance in the Admin console, and all applied to a unit rather than the organisation:

None of these is new. What is new is having an account per agent to point them at. On a shared identity you cannot restrict delivery for the agent without restricting it for yourself. On a dedicated one you can, and the rule is enforced by Google, not by the agent.

The same move works for code. A dedicated GitHub account is given access to exactly the repositories that agent should touch, either by installing the Claude GitHub App on those repositories alone, or with a fine-grained personal access token, which GitHub describes as able to be "limited to only access specific repositories" with "specific, fine-grained permissions" and an expiry. The site editor's account sees the website repository and nothing else. That is the first time I have been able to say that about an agent with a GitHub connector, and it is worth the price of the seat by itself.

Finding two: the exception arrived before the policy

The reader's policy was going to be the simplest: read, classify, write a list, touch nothing else. In particular, never send.

It lasted until the first scheduled run. Some of the messages that land in that inbox are from me, and some of what I send it are instructions: do this, file that, reply to them with this. For those, a reader that only writes a list is wasting the one thing a scheduled agent is for. So the rule became: never reply, unless the message came from me, in which case act on it.

Two things follow, and both are bigger than the reader. First, this is how policies actually get written: from the workflow, one exception at a time, not from principles handed down before anyone has run anything. The list of exceptions is the policy, and it is only complete once the workflow has run for a while. Second, "unless it came from me" is not a permissions question. Anyone can put my address in a From header. The exception is only safe if the agent can verify that a message came from me, which is a signature check, and the machinery for that is the one the append-lane messaging work built for agents talking to each other: per-session keys, a pinned registry of who holds which key, a signature over the message. The reader needs the same registry with one more entry in it, mine.

There is also a cheap partial enforcement for this rule, and it comes from finding one. Put the reader in a unit whose delivery is restricted to my domain, and the failure mode of a spoofed instruction is that the reader sends its reply to me, because that is the only place it can send anything. That does not make the policy right. It makes the worst row bounded.

Finding three: the policy written from the documentation was wrong

When the Gmail policies were first mapped, one assumption was that an agent could not attach files to mail through the Claude connector. There was no "attach a file" call, the documentation did not mention attachments beyond reading their metadata, and so the policy did not bother to forbid what seemed impossible.

On 27 September an agent working in Cowork needed to send a signed letter and found the way. The debrief it wrote is short and worth quoting: "The connector has no 'attach a file from disk' call, but create_draft and update_draft both accept an attachments array. Each item carries the file inline as base64." It shrank a 1.4 MB scanned PDF to 17 KB, base64-encoded it, put it in the draft, verified the MIME structure with get_draft in raw format, and I sent the draft by hand. It arrived intact, with Gmail's "One attachment · Scanned by Gmail" label and a preview. The debrief also records the gotchas an agent would trip over: a later update_draft without the field silently removes the attachments, so attach last; the message and thread ids change on every update but the draft id does not; the bytes pass through the tool call as text, so the practical limit is the size of a call, not Gmail's 25 MB.

I want to be precise about what that finding is and is not. It is not a bypass. The agent used a documented field of a documented call. It is a policy that was written from an assumption about capability, and the assumption was wrong, and the only way it could have been found wrong was by an agent trying to do the thing. That is the connector twin argument exactly: before you write down what an agent may do, replay what it can do.

And the vendor's documentation would not have settled it, because the vendor's documentation disagrees with itself. On 29 September 2026, one Anthropic page says of the Gmail connector: "Claude reads and searches email only; it can't create, send, or modify messages," and lists under limitations "Claude can't create, send, or modify emails." Another says it can "Draft emails with proper formatting and context," "Send, reply to, and forward emails from Gmail," and that "By default, Claude asks for your approval before each of these actions. On Team and Enterprise plans, owners decide whether members can allow these actions to run without asking each time." The debrief proves the second page is closer to the truth, at least for drafts. Neither page mentions the attachments field. I do not say this to criticise the documentation, which is trying to describe a product that changes monthly. I say it because a policy that cites the documentation as its enforcement has cited a moving target.

Here is the catch I promised in finding one. The Gmail API's scopes put drafting and sending in the same grant: gmail.compose is "Manage drafts and send emails," and gmail.modify is "Read, compose, and send emails." There is no scope that allows drafts and forbids sending. So the Workspace admin's scope limit, the sharpest tool in the box, cannot express the one rule the mailbox agent most needs, drafts only. That rule has to be enforced one rung down the ladder, by a delivery restriction on the unit, or by an approval prompt, or by the agent doing as it was told.

Where the reach differs: Cowork and Claude Code

One more thing the setup made visible. The same seat runs two products, and they do not reach the same things.

The Cowork agent, in this setup, has no GitHub connector and cannot open the repository at all, even though the connector is listed in the catalogue and the account has access to the repository on GitHub. The Claude Code agent, on the same seat, opens the repository and edits the site. Whether that is a deliberate boundary or a gap in the product does not matter for the policy; what matters is that it is a real boundary today, and the site editor role is defined around it. The Claude Code agent is the only one that can change the websites, and the only thing its GitHub account can see is the website repositories.

The two products also differ in what an administrator can control. On a Team plan an Owner enables or disables a connector for the whole organisation, and in Cowork there is one further switch, whether members may set "Always allow" for "write-capable connector tools." Per-tool control, where a specific action can be set to "Always allow," "Needs approval," or "Blocked", exists only in the Enterprise plan's role-based permissions. So on a Team plan, never send can be a prompt that a person has to click through, and on Enterprise it can be a block. That difference belongs in the enforcement column too.

The table

Where a rule is enforced, from the top of the ladder down. Identity and keys enforce themselves. Compliance rules enforce at the platform, per organisational unit. Approval prompts enforce at the person. Below that, a rule is the agent doing as it was told. Every rule in the table sits on one of these rungs, and the policy is as strong as its lowest.

Here is the policy for the setup, one rule per row, with the column that matters. Enforced means the platform or the cryptography prevents the action. Admin rule means a Workspace or GitHub setting an administrator applied to the agent's own account or unit. Approval means a person is asked before the action runs. Told means the agent has been instructed and nothing else stands in the way.

RuleReaderMailboxInboxCRMDev teamSite editorHow it is enforced today
Sees only its own account's mail, files and reposyesyesyesyesyesyesEnforced: dedicated accounts
Reads mailyesyesyesyesnonoAdmin rule: connector enabled per seat; Workspace scope limits per unit
Creates draftsnoyesnonononoApproval on Team, Blocked per tool on Enterprise; no scope separates drafts from sending
Sends mailnonoyes, from drafts onlynononoAdmin rule: Restrict delivery on every other unit; Approval for the inbox agent
Sends attachmentsnonoonly when a person attached themnononoAdmin rule: outbound attachment compliance strips or rejects
Replies to a message from meyes, if verifiednononononoTold, until the signature registry has my key; delivery restriction bounds the failure
Reads repositoriesnonononoits ownthe site'sEnforced: GitHub App installed per repository, or a fine-grained token
Writes to repositoriesnonononoits ownthe site'sEnforced: token permissions; Cowork has no GitHub connector here
Reads vaultstagged listyesnoyesyesyesEnforced: a read key per vault per agent
Writes to vaultsnononoyes, CRM vaultyesnoEnforced: the vault key, held only by the writer
Acts outside its rolenonononononoTold

Read the last column top to bottom and the shape of the problem is clear. The rows that touch identity, code and vaults are enforced, because accounts, repositories and keys are boundaries the platform or the mathematics maintains. The rows that touch mail are admin rules and approvals, which is good, and better than I expected before I understood organisational units. And two rows are told. One of them, the reply-to-me rule, has a known fix and a bounded failure. The other, acts outside its role, is the row every agent policy in the world shares, and no setting anywhere enforces it. It is the reason the connector twin exists: if you cannot prevent it, you must at least be able to see it afterwards.

What I would tell somebody starting this

  1. Buy the segregation. A Workspace seat per agent is £7 a month. It is the only control that covers mail, files, code and connectors at once, and it is the precondition for every admin rule below.
  2. Put each agent in its own organisational unit. Then restrict delivery for every agent that should not correspond with the world, and strip outbound attachments for every agent that should not send files. These are settings, not hopes.
  3. Give each agent its own GitHub account and install the app per repository. The site editor sees the site. Nothing else sees anything.
  4. Write the policy as a table, with the enforcement column. Fill the column honestly. Anything that says told is where your next incident comes from, and where your next piece of tooling should go.
  5. Replay before you write. Have an agent try the things the policy assumes it cannot do. Mine found attachments in an afternoon. The documentation would not have told you.
  6. Expect the exception on day one. Write the policy from the workflow, and treat the growing list of exceptions as the policy maturing rather than failing.
  7. Journal everything anyway. For the rows that say told, the twin is the only control there is.

What exists today, and what does not

The setup described here runs: the accounts, the roles, the scheduled reader, the drafting and sending split, the dedicated GitHub account for the site editor. The Workspace compliance rules exist and are documented, and I have linked the documentation rather than restating it. The attachment finding is proven and written up. The append-lane signature registry exists for agent-to-agent messages.

What does not exist yet: my own key in that registry, so the reader's exception is still enforced by instruction; a broker that would let drafts only be a block rather than a prompt on a Team plan; and any way to enforce acts outside its role. The policy table above is the honest state of a real setup on 29 September 2026, and I would rather publish it with its two told rows than pretend they are something else.

Sources

© 2026 Dinis Cruz. This article's own text is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). You're free to share and adapt it, as long as you give credit. Quoted material and linked sources keep their own licences.

← All articles