for agents/llms.txtv0.7.35 · 10 Oct 2026

Home / Articles / The investigation GitHub owes its customers / Versions / v1.1.0

The investigation GitHub owes its customers: what changed in v1.1.0

From v1.0.0 (2026-10-05, ae21f7588) to v1.1.0 (2026-10-05, dc11503b4), paragraph by paragraph.

0 paragraphs added, 0 removed, 15 changed in place, 66 unchanged. About 78 words added and 1 removed. Insertions are marked like this, deletions like this; unchanged runs are folded to one line; figures appear as their file names.

all versions · v1.1.1 →

1 unchanged paragraph

Summary: At 19:11 UTC on 5 October 2026 GitHub's status page said it was investigating degraded performance for Actions. For the next two hours, every customercustomers in every region who ran on GitHub's hosted runners could not reliablyrely starton a workflow,workflow starting, which for most of the world'ssoftware softwarethat deploys through Actions is the same as not being able to deploy. This site's own release sat in the queue, was cancelled by the incident, and went live two hours late on a retry. The status page said "delays", then "degraded availability", and will say "resolved". NobodyUnless GitHub chooses to publish it, the outside GitHubworld will not learn which component failed, what it could reach, why the scope of one fault was everyone,everyone on hosted runners, or whether the same roll of the dice had come up before. This article argues that incidents at platforms this critical should be treated the way aviation treats them: investigated by somebody independent whose only job is prevention, reported whether or not the consequence was severe, published with the evidence, and followed through to the second story, why the system allowed it, and the third, why nobodythe fix was not paid to change it.for. The usual objection has always been that the evidence is confidential, enormous and expensive to gather. It is not any more. Encrypted vaults with one-way read keys, signed records, per-party access and agents that read a graph make the aviation docket affordable for a ninety-minute fault. The companies that depend on GitHub cannot see how close to the wind it flies, and that, not the outage, is the risk nobody has signed for.

3 unchanged paragraphs, under In short

• What happened. From 19:11 UTC, GitHub-hosted runners stopped being reliably assigned to Actions jobs foracross every customer in everythe runner configuration;configurations GitHub named; by 20:47 the status page called it degraded availability; by 21:09 repository lists and billing pages were failing for some customers too. For mostthe organisations,many Actionsorganisations is thewhose only path to production,production soruns through Actions, and a 2025 survey found most organisations run a single CI/CD tool, this was a global outage of the ability to deploy.

3 unchanged paragraphs

• The three stories. The first is what happened. The second is why the system allowed it. The third is why nobodythe fix was not paid to change it.for. Software incident write-ups, when they exist, stop at the first and gesture at the second. The third is where the money is, and it is never told.

1 unchanged paragraph

• The objection, and why it has expired. Independent investigation of software incidents was alwayslong said to be impossibleimpractical because the evidence is confidential, enormous and expensive to gather. Encrypted vaults with one-way read keys, signed and versioned records, per-party access and agents that read the graph make it affordable for a ninety-minute fault.

3 unchanged paragraphs, under What happened, as far as anyone outside can tell

That is the whole of what the world is told, and the words are chosen with care. Delays. Some workflows. Degraded. Each is true and each undersells it, because "some workflows may take longer to start across runner configurations" describes, from the inside, a condition in which no customer anywhereon hosted runners could rely on a workflow starting, and a workflow that does not start is a release that does not ship.

I know because I watched one. This site's previous release, the one that published the evidence vault behind the article before this, was pushed at 19:23. Its validation job ran in eight seconds on the first runner it found. Its tag job then waited for a runner from 19:55, was cancelled at 20:10, and the deploy was skipped; the run was marked failed. It would have stayed failed, because nothing in our workflow retries a cancelled deploy, until a fresh commit at 21:26 started a new run, which, in a window the incident happened to leave open, ran in sixty-five seconds. The release went live two hours late and I found out by waiting, which is how every one of GitHub's customers found out tonight.

1 unchanged paragraph, under Why "everybody" is the point

Actions and push-and-merge are the two surfaces GitHub has with the widest blast radius, because they are the ones through which every customer changes its own software. GitHub's own figures put that at more than 180 million developers, 630 million repositories and 71 million Actions jobs a day; a 2025 survey of 805 organisations found 59 percent running a single CI/CD tool, and GitHub Actions the most used. When they fail, every customer that depends on them loses the same ability at the same time. For a company that needed to ship a security fix, a time-sensitive update or an incident response of its own in those two hours, this was not a delay. It was an outage of its ability to respond, and for some of them, somewhere, tonight, that was catastrophic. We will not hear about those either.

A fault that reaches one customer, or one region, is an incident. A fault that reaches every hosted-runner customer in every region at once is a statement about how the system is built: that on the one component which gates every deploy there is nolittle isolation between customers or regions, so that the scope of one fault is close to everyone. That is not bad luck. It is a design outcome, and it may well be a rational one, chosen against costs nobodythat outsideare cannot see.visible from outside. The shape a service at this scale should have is regional faults, for some customers, some of the time. When the shape is global, it is a signal about isolation, change control and resilience that senior management should be able to see before an incident shows it to them, and that customers should be able to see before they build their only path to production through it.

It cuts the other way too. If your only path to production runs through Actions, GitHub is inside your blast radius. Can you ship a fix when your pipeline vendor cannot? MostFew companies have never been asked, which is the due diligence point of the previous article turned on the buyer.

11 unchanged paragraphs, under What aviation does instead, Near misses, and counting the rolls of the dice, The first, second and third story

The frame I use for incidents comes from Three Mile Island, and I set it out in an essay in February 2025. The first story is what happened: the operator did this, the valve stuck, the file was malformed. It is nearly always available, because it is the story the organisation tells. The second story is why the system allowed it: the design, the constraints, the normalised deviations that made the first story possible. "Second stories shift the narrative from individual blame to the systemic conditions that make failures more likely." And there is a third, which the five whys reach if you keep going: why nobodythe change was not paid to change the systemfor when the second story was already known.

3 unchanged paragraphs

Tonight's first story is not yet published. The second story is visible from the outside: on the one component that gates every deploy, the scope of a fault was everyone. The third story is the one nobodythat cannot be told from outside can tell and the one that matters, because it is about trade-offs. Somewhere inside GitHub there is a reason the runner assignment path is not isolated by region or by customer, and it is probably a reasonable reason: cost, complexity, a migration not yet finished, a system too fragile to touch. Those are exactly the reasons that in aviation are called the finding, not the excuse. A system too complicated to understand and too fragile to change is not a reason to leave it alone. It is the single point of failure, named.

3 unchanged paragraphs, under The what-if ladder

What if GitHub were down for a day? Every organisation whose only path to production is a workflow is frozen; the ones with a second path are not;not, nobodyand hassome countedhave a second and a third; I know of no public count of which is which. A week? Releases, dependency updates and incident response across the whole industry stall together, which means vendors cannot patch while their customers are being attacked, and I do not know of a regulator that has modelled what a week without GitHub costs a country. Corrupted, or not restorable? Git is distributed, so the code survives; the issues, the reviews, the pipelines and the record of who decided what are not, and few organisations hold a copy they could prove is the same. Withdrawn or weaponised, access denied to an organisation, a sector or a country by decision rather than by fault? GitHub is mission-critical to governments that do not control it, which is a sovereignty question that somebody, somewhere, has implicitly accepted on their behalf. I wrote about the grid version of this after the Iberian blackout in April 2025: "Europe cannot afford single points of failure; a more federated architecture would enable graceful islanding, where unaffected areas detach and survive." The software version is less visible and no less true.

2 unchanged paragraphs, under Why the market does not fix this

The reason nothing pushes back is that the market economics do not work, and they do not work because customers cannot see. Maybe it is fine for GitHub to have this level of resilience and this kind of failure every so often; that is a legitimate decision for a business to make. But its customers have to be able to see how close to the wind it flies and what the true level of resilience is, and today they cannot. We only know about the incidents that reach the status page; the ones that stayed inside are not reported because nothingno rule requires them to be. Uptime, for a platform with a barrier to exit this high, is mostly a marketing and damage-limitation function, and inside the company the strategic calculation is implicit: customers will not move unless it gets much worse, so the investment goes elsewhere, and the growth that would be the moment to harden the infrastructure is spent on the next feature.

8 unchanged paragraphs, under The objection, and why it has expired, What the regulation does and does not do

What none of this does is produce the second story. The reporting duties produce notifications: that an incident happened, its rough scope, and in a month a final report to a regulator, which is not published. The critical third party regimes produce oversight of a handful of hyperscalers for the benefit of financial firms, not of the platform the rest of the economy deploys through. GitHub's own practice is a monthly availability report on its blog with one paragraph per incident, a cause and a remediation, which is more than most platforms give and much less than a docket; its chief technology officer wrote in April 2026, after a merge-queue bug reverted changes in more than two thousand pull requests, that the company's priority was "availability first, then capacity, then new features", and independent trackers counted 257 incidents in the twelve months to that month, a fifth of them in Actions. The scale that depends on it is public in its own numbers: more than 180 million developers, 630 million repositories, 71 million Actions jobs a day, 143 US federal civilian agencies, the UK government's own code. ThereI isknow of no body, anywhere,body whose job is to read the record of what happened tonight and say so in public.

5 unchanged paragraphs, under What I would ask for, What is open

The cause of tonight's incident, which GitHub may publish in its monthly availability report and may not. The count of how many organisations have a second path to production, which nobodyI has.have not found; some certainly do, and they are the ones this evening did not touch. A model of what a week without GitHub costs, which no regulator I know of has published. The incident vault described above, which exists in pieces on this site and has not been assembled for a real incident. And the question of whether the industry will build the aviation habit before or after the step on the ladder that makes it unavoidable. This site's own release got through on a retry tonight. Many will not have.

22 unchanged paragraphs, under Threads woven here, Sources

all versions · v1.1.1 →