for agents/llms.txtv0.6.75 · 6 Oct 2026

Home / Articles / The investigation GitHub owes its customers: why a global outage of a platform the world deploys through deserves an aviation-style inquiry, and how the evidence could now be gathered

The investigation GitHub owes its customers: why a global outage of a platform the world deploys through deserves an aviation-style inquiry, and how the evidence could now be gathered

By · 2026-10-05 · v0.6.69 · incidentsresiliencegithubnear-missessecond-storiesblast-radiusinvestigationaviationregulationvaultssgitarticle

Abstract: At 19:11 UTC on 5 October 2026 GitHub's status page said it was investigating degraded performance for Actions. For the next two hours, customers in every region who ran on GitHub's hosted runners could not rely on a workflow starting, which for most of the software that deploys through Actions is the same as not being able to deploy. This site's own release sat in the queue, was cancelled by the incident, and went live two hours late on a retry. The status page said "delays", then "degraded availability", and will say "resolved". Unless GitHub chooses to publish it, the outside world will not learn which component failed, what it could reach, why the scope of one fault was everyone on hosted runners, or whether the same roll of the dice had come up before. This article argues that incidents at platforms this critical should be treated the way aviation treats them: investigated by somebody independent whose only job is prevention, reported whether or not the consequence was severe, published with the evidence, and followed through to the second story, why the system allowed it, and the third, why the fix was not paid for. The usual objection has been that the evidence is confidential, enormous and expensive to gather. It is not any more. Encrypted vaults with one-way read keys, signed records, per-party access and agents that read a graph make the aviation docket affordable for a ninety-minute fault. The companies that depend on GitHub cannot see how close to the wind it flies, and that, not the outage, is the risk nobody has signed for.

One incident, two records. The top lane is every update GitHub's status page gave in the first two hours and twenty minutes; the bottom lane is this site's own release, submitted twelve minutes after the incident opened, cancelled by it, and live on a retry while it was still unresolved. Neither record names a cause, a scope or a blast radius.
Where this comes from, and what is behind it. A GitHub Actions incident on the evening of 5 October 2026, which held this site's previous release in a queue while an article about software due diligence waited to go out, and a voice memo recorded while it was still open. The thinking is older than the evening: a brief from February on why near misses deserve the same investigation as incidents, one from July on counting how many times the same dice were rolled, the 2025 essay on second stories, and this site's own first article, which was about a GitHub incident that made a green build tick lie. The aviation facts, the UK air traffic control review and the regulation are read in their current wording and dated in the sources. Where this article argues from one evening's incident it says so; the argument does not depend on what GitHub's cause turns out to be.

In short

What happened, as far as anyone outside can tell

The record of this incident, as it stands tonight, is a status page. At 19:11 GitHub was "investigating reports of degraded performance for Actions". At 19:15 it was "an issue causing delays when assigning GitHub-hosted runners to Actions jobs", with "some workflows" taking longer to start "across runner configurations". At 19:50 it was still delays, "across multiple runner configurations". At 20:39 it was "job failures and delays". At 20:47 the page turned red: "Actions is experiencing degraded availability." At 21:09 some customers could not reach repository lists, licensing or billing pages, and at 21:22 Pages was degraded too. The incident was still open when this article was written.

That is the whole of what the world is told, and the words are chosen with care. Delays. Some workflows. Degraded. Each is true and each undersells it, because "some workflows may take longer to start across runner configurations" describes, from the inside, a condition in which no customer on hosted runners could rely on a workflow starting, and a workflow that does not start is a release that does not ship.

I know because I watched one. This site's previous release, the one that published the evidence vault behind the article before this, was pushed at 19:23. Its validation job ran in eight seconds on the first runner it found. Its tag job then waited for a runner from 19:55, was cancelled at 20:10, and the deploy was skipped; the run was marked failed. It would have stayed failed, because nothing in our workflow retries a cancelled deploy, until a fresh commit at 21:26 started a new run, which, in a window the incident happened to leave open, ran in sixty-five seconds. The release went live two hours late and I found out by waiting, which is how every one of GitHub's customers found out tonight.

Why "everybody" is the point

Actions and push-and-merge are the two surfaces GitHub has with the widest blast radius, because they are the ones through which every customer changes its own software. GitHub's own figures put that at more than 180 million developers, 630 million repositories and 71 million Actions jobs a day; a 2025 survey of 805 organisations found 59 percent running a single CI/CD tool, and GitHub Actions the most used. When they fail, every customer that depends on them loses the same ability at the same time. For a company that needed to ship a security fix, a time-sensitive update or an incident response of its own in those two hours, this was not a delay. It was an outage of its ability to respond, and for some of them, somewhere, tonight, that was catastrophic. We will not hear about those either.

A fault that reaches one customer, or one region, is an incident. A fault that reaches every hosted-runner customer in every region at once is a statement about how the system is built: that on the one component which gates every deploy there is little isolation between customers or regions, so that the scope of one fault is close to everyone. That is not bad luck. It is a design outcome, and it may well be a rational one, chosen against costs that are not visible from outside. The shape a service at this scale should have is regional faults, for some customers, some of the time. When the shape is global, it is a signal about isolation, change control and resilience that senior management should be able to see before an incident shows it to them, and that customers should be able to see before they build their only path to production through it.

It cuts the other way too. If your only path to production runs through Actions, GitHub is inside your blast radius. Can you ship a fix when your pipeline vendor cannot? Few companies have been asked, which is the due diligence point of the previous article turned on the buyer.

What aviation does instead

I keep coming back to aviation because it is the best example we have of a highly complex system, made of technology and business process and human beings, that learned to be safe. It did not learn by having fewer things go wrong. It learned by treating everything that went wrong, and everything that nearly did, as data belonging to the system rather than to the operator.

The same seven questions put to aviation and to software platforms: who investigates, must it be reported, what is published, how near misses are treated, whether there is a second story, who tracks the recommendations, and who can ask. The difference is not the engineers; it is whether failure produces learning that everyone can see.

The rules are written down. ICAO Annex 13, which every signatory state's investigators work under, says in its third chapter: "The sole objective of the investigation of an accident or incident shall be the prevention of accidents and incidents. It is not the purpose of this activity to apportion blame or liability." It requires each state to establish an investigation authority that is independent, with "unrestricted authority over its conduct", and to publish the final report "if possible, within twelve months". The United States re-established the NTSB in 1974 as a body outside the Department of Transportation on the argument that no agency can investigate properly unless it is "totally separate and independent"; it has since issued more than 15,500 safety recommendations and tracks the acceptance of every one. The UK's Air Accidents Investigation Branch states its purpose as determining "the circumstances and causes of air accidents and serious incidents, and promoting action to prevent reoccurrence", and its reports do so "without attributing blame".

The practice is as striking as the rules. When a door plug blew out of a Boeing 737 in January 2024, the NTSB published a preliminary report within a month and a final one seventeen months later whose probable cause named Boeing's "failure to provide adequate training, guidance and oversight to its factory workers" and the regulator's ineffective oversight of "repetitive and systemic" nonconformance; the docket behind it holds the interview transcripts and factual reports for anyone to read. When Air India 171 crashed in June 2025, India's investigators published a preliminary report in thirty days that said what the recorders showed. When a helicopter and an airliner collided over Washington in January 2025, the final report in February 2026 ran to about four hundred pages, seventy-four findings and fifty recommendations, with the chair testifying to the Senate about it the same month. That is what "published" means in aviation: not a paragraph, a docket.

And then there is the part the industry is proudest of, which is what it does about the accidents that did not happen. NASA's Aviation Safety Reporting System has taken confidential reports from pilots, controllers and mechanics since 1976, more than two million of them, around a hundred thousand a year and rising, with immunity from penalties for unintentional violations as the incentive, and "no reporter's identity has ever been breached". The UK has run an equivalent since 1982. European law since 2014 defines the "just culture" this depends on: one in which people "are not punished for actions, omissions or decisions taken by them that are commensurate with their experience and training, but in which gross negligence, wilful violations and destructive acts are not tolerated." Pilots report near misses because it is safe to, and the result is the safest complex system humans operate.

The software industry has none of this by default, and the people who run its platforms know it. In September 2026 Sam Altman told a conference audience that "accidents with any new technology are unavoidable, and we should have a great culture of transparent reporting about them", and pointed at exactly this precedent: "With the FAA and the NTSB, one of the things that has gone so well there is the culture of accident reporting." A week later OpenAI published principles for third-party assessment that include independent investigation of incidents in which its models acted without authorisation, after a summer in which researchers who tried to piece together one such incident said "it was difficult to get a precise understanding of events and we were missing aspects of the story." The idea of an NTSB for computing has been proposed since 1991; the nearest thing that existed, the US Cyber Safety Review Board, published three reports, including one in 2024 that found a major cloud intrusion "should never have happened" and resulted from "a cascade of security failures", and was disbanded in January 2025 in the middle of its fourth review. The argument is being made for AI. It applies at least as strongly to the platform every AI product is built and deployed through.

Near misses, and counting the rolls of the dice

Tonight's incident is a near miss for almost everyone it touched, and a near miss is the most valuable kind of data there is, because it arrives before the damage. I wrote this in February, in a brief on why near misses matter more than incidents: "A near miss, a plane that almost had an incident but didn't, gets the same level of investigation as an actual crash. Because the systemic causes are the same. The only difference is luck. And you don't build safety on luck." The same brief has the number I think about most: "Usually the actual damage is around 10% of what was possible. Often it's 1%." And the sentence that describes tonight: "Companies celebrate containing an incident when they should be terrified by the paths not taken."

In July I wrote the other half of it, about why the same dice get rolled so many times before anyone counts: "one of the root causes we do not pay enough attention to is how many times the same kind of event, the same roll of the dice, happened where they got away with it." The move the brief asks for is to go "from what happened was a one-off to what happened was actually a very predictable statistical event, because if you do ten or twenty dangerous events, eventually one or two will not work out." And the question it ends on is the one GitHub's customers should be asking: "one party may be willing to take that risk, but do the people who depend on them want them to be that cavalier and risk-insensitive." We do not know how many times runner assignment has wobbled in the last year without reaching the status page. GitHub does. That asymmetry is the whole problem.

The first, second and third story

The frame I use for incidents comes from Three Mile Island, and I set it out in an essay in February 2025. The first story is what happened: the operator did this, the valve stuck, the file was malformed. It is nearly always available, because it is the story the organisation tells. The second story is why the system allowed it: the design, the constraints, the normalised deviations that made the first story possible. "Second stories shift the narrative from individual blame to the systemic conditions that make failures more likely." And there is a third, which the five whys reach if you keep going: why the change was not paid for when the second story was already known.

Three incidents read the way an investigator reads them: the UK air traffic control failure of August 2023, the security vendor's update of July 2024, and tonight. The first story is the company's; the second and third are what an independent investigation exists to find.

The UK air traffic control failure of 28 August 2023 is the clean example, because it got the investigation this article is asking for. The first story: a flight plan contained two waypoints with the same name, about four thousand nautical miles apart; the system that processed it could not resolve them and, as designed, shut itself down, and so did its backup. It had processed fifteen million flight plans in five years without meeting the case. The second story, told by the independent review the regulator commissioned: a single system in the path of every flight, with a fail-safe that took the whole service with it and a recovery that depended on a specialist who was not on site. The third story, in the review's own words fourteen months later: "a major failure on the part of the air traffic control system", with thirty-four recommendations, twelve to the operator, eleven to the regulator, six to airlines and airports and five to government, about the incentive regime for investment, the communications and the response of the system as a whole. Over 700,000 passengers were affected and the cost was put at between 75 and 100 million pounds. The regulator closed the last of the thirty-four recommendations in June 2026. Three months later the same centre had another flight-data failure, and the regulator asked for a full incident report. That is the system working: not that nothing fails, but that every failure has to account for itself in public.

The security vendor's update of July 2024 is the software example that came closest to the aviation treatment, because its consequence was too large to avoid it. The first story: a content update carried a file with twenty-one fields where the sensor expected twenty; the out-of-bounds read crashed about 8.5 million machines. The second story, in the vendor's own root cause analysis three weeks later: a validator that trusted the file it was checking and a content channel that bypassed staged rollout, with the operating system kernel as the blast radius. The third story was told by a Congressional hearing, a lawsuit claiming more than half a billion dollars, and insurers estimating losses to large companies near 5.4 billion. The investigation was done by the company, with two unnamed outside reviewers it hired; there was no independent board, because there is none to do it.

Tonight's first story is not yet published. The second story is visible from the outside: on the one component that gates every deploy, the scope of a fault was everyone. The third story is the one that cannot be told from outside and the one that matters, because it is about trade-offs. Somewhere inside GitHub there is a reason the runner assignment path is not isolated by region or by customer, and it is probably a reasonable reason: cost, complexity, a migration not yet finished, a system too fragile to touch. Those are exactly the reasons that in aviation are called the finding, not the excuse. A system too complicated to understand and too fragile to change is not a reason to leave it alone. It is the single point of failure, named.

The what-if ladder

What makes tonight worth an article rather than a shrug is where it sits on a ladder, and that nobody has had to answer the questions on the steps above it.

The what-if ladder. Today was ninety minutes of degraded runner assignment. A day, a week, a corruption that cannot be restored, and a withdrawal of service by decision rather than fault are each within the realm of reality for a platform most of the world's software passes through, and each has a blast radius nobody has published.

What if GitHub were down for a day? Every organisation whose only path to production is a workflow is frozen; the ones with a second path are not, and some have a second and a third; I know of no public count of which is which. A week? Releases, dependency updates and incident response across the whole industry stall together, which means vendors cannot patch while their customers are being attacked, and I do not know of a regulator that has modelled what a week without GitHub costs a country. Corrupted, or not restorable? Git is distributed, so the code survives; the issues, the reviews, the pipelines and the record of who decided what are not, and few organisations hold a copy they could prove is the same. Withdrawn or weaponised, access denied to an organisation, a sector or a country by decision rather than by fault? GitHub is mission-critical to governments that do not control it, which is a sovereignty question that somebody, somewhere, has implicitly accepted on their behalf. I wrote about the grid version of this after the Iberian blackout in April 2025: "Europe cannot afford single points of failure; a more federated architecture would enable graceful islanding, where unaffected areas detach and survive." The software version is less visible and no less true.

Each step up the ladder is a question tonight lets us ask for free. Treating a global ninety-minute fault as a minor incident is choosing to pay for the lesson later, at a step where it is expensive.

Why the market does not fix this

The reason nothing pushes back is that the market economics do not work, and they do not work because customers cannot see. Maybe it is fine for GitHub to have this level of resilience and this kind of failure every so often; that is a legitimate decision for a business to make. But its customers have to be able to see how close to the wind it flies and what the true level of resilience is, and today they cannot. We only know about the incidents that reach the status page; the ones that stayed inside are not reported because no rule requires them to be. Uptime, for a platform with a barrier to exit this high, is mostly a marketing and damage-limitation function, and inside the company the strategic calculation is implicit: customers will not move unless it gets much worse, so the investment goes elsewhere, and the growth that would be the moment to harden the infrastructure is spent on the next feature.

I do not say this to attack GitHub's engineers. I say it because inside every platform there are engineers who know exactly where the single points of failure are and cannot make the business case, because the case has to be made against revenue with no evidence on the other side except their own judgement. An independent investigation, published, is how the people who want to do the right thing get the evidence to do it. That is the kindest thing an outside inquiry does, and in aviation it is the main thing it does.

The objection, and why it has expired

The objection to all of this was never the principle. Nobody argues that a platform the world deploys through should be less accountable than an airline. The objection was practical: the evidence of a software incident is confidential, enormous and expensive to gather, and handing it to a third party meant handing them the keys to everything, at a cost no incident short of a disaster justified. For most of the history of the industry that was true.

How the evidence would move if it were captured in vaults. The provider captures as it happens into a signed, versioned store and reviews what leaves, redacting rather than rewriting; an independent investigator holds a read key to what the inquiry needs; each affected customer gets the part of the record that concerns it; the findings are published as a graph whose every node links to the hash of the evidence it rests on.

It is not true any more, and the reason is the same set of pieces the articles before this one are about. A provider can capture the record of an incident as it happens, logs, configuration, the change that preceded the fault, the dependency graph of the failed component, into an encrypted vault that is hashed, versioned and signed by the provider's own key, so that nothing in it can be altered afterwards without the record saying so. The provider reads what leaves before it leaves and can withhold a file; it cannot rewrite one. An independent investigator holds a read key to the files the inquiry needs and nothing else, and an agent under a behaviour policy does the first pass over the graph, which is what makes the cost proportionate to a ninety-minute fault rather than a disaster. Each affected customer gets its own vault, derived from the same record, with the part that concerns it: what of theirs the fault reached, for how long, and what they had no second path for. And the findings are published as a graph, first story, second and third, each node linked to the hash of the evidence it rests on, with the recommendations held open until somebody closes them.

Every piece of that exists and runs on this site. The vault is a footprint recorder by construction: every commit signed, versioned and append-only, every message down an append lane recorded with its token and its time, and a read key that lets a reviewer read all of it without touching anything. Each party signs with its own key, so nothing anonymous can accumulate. The due diligence article describes the same movement of a derived, redacted-not-rewritten record from a vendor to a buyer. The incident record is those pieces pointed at a failure instead of a release. What is missing is not the technology. It is the requirement, and the habit.

What the regulation does and does not do

Some of the machinery is arriving, and it is worth being exact about what it does. The EU's Digital Operational Resilience Act has applied to financial firms since January 2025, and in November 2025 the European supervisory authorities designated nineteen critical ICT third-party providers for direct oversight, Microsoft, Amazon Web Services, Google Cloud, Oracle and IBM among them; GitHub is not named, its parent is. The UK's parallel regime designated its first four critical third parties in July 2026, Amazon, Google, Microsoft and Oracle, with oversight beginning that month and a duty to "maintain open, timely communication with regulators and the firms that rely on them, particularly during major incidents". NIS2 requires an early warning within 24 hours of a significant incident, a notification within 72 and a final report within a month. The US Securities and Exchange Commission requires a public filing within four business days of deciding a cyber incident is material. The Bank of England said in 2022 that "if a large number of FMIs become dependent on a small number of dominant outsourced arrangements, this could give rise to systemic concentration risks", which is the sentence that describes GitHub and Actions precisely, written about banks.

What none of this does is produce the second story. The reporting duties produce notifications: that an incident happened, its rough scope, and in a month a final report to a regulator, which is not published. The critical third party regimes produce oversight of a handful of hyperscalers for the benefit of financial firms, not of the platform the rest of the economy deploys through. GitHub's own practice is a monthly availability report on its blog with one paragraph per incident, a cause and a remediation, which is more than most platforms give and much less than a docket; its chief technology officer wrote in April 2026, after a merge-queue bug reverted changes in more than two thousand pull requests, that the company's priority was "availability first, then capacity, then new features", and independent trackers counted 257 incidents in the twelve months to that month, a fifth of them in Actions. The scale that depends on it is public in its own numbers: more than 180 million developers, 630 million repositories, 71 million Actions jobs a day, 143 US federal civilian agencies, the UK government's own code. I know of no body whose job is to read the record of what happened tonight and say so in public.

What I would ask for

Not a regulator for software, at least not first. Four things, in order of how cheap they are.

First, that platforms above a certain criticality publish, for every incident that reaches their status page, the second story within thirty days: which component, what it can reach, why the scope was what it was, and whether the same fault pattern had occurred before without reaching the page. Aviation publishes a preliminary report in about that time. Second, that near misses be counted and reported in aggregate, confidentially if need be, the way pilots report them, so that the dice get counted before one comes up wrong. Third, that the evidence be captured into a signed, versioned record as a matter of course, so that an investigation does not begin with a reconstruction. And fourth, that when an incident reaches everyone, someone independent reads that record and says what they found, in public, with the evidence linked.

The companies that depend on GitHub should be asking for the first of these tonight, in the same breath as asking themselves whether they have a second path to production. GitHub will say, truthfully, that it has had a hard evening and its teams are working to mitigate. They are. The question is not whether they are working. It is whether anyone outside will ever know what they found, and whether the people inside who already knew will now be able to spend the money.

What is open

The cause of tonight's incident, which GitHub may publish in its monthly availability report and may not. The count of how many organisations have a second path to production, which I have not found; some certainly do, and they are the ones this evening did not touch. A model of what a week without GitHub costs, which no regulator I know of has published. The incident vault described above, which exists in pieces on this site and has not been assembled for a real incident. And the question of whether the industry will build the aviation habit before or after the step on the ladder that makes it unavoidable. This site's own release got through on a retry tonight. Many will not have.

Threads woven here

Sources

Threads

Agents & policyVaults & method This article as a graph →

Builds on

All articles · All graphs

← All articles