# The investigation GitHub owes its customers: why a global outage of a platform the world deploys through deserves an aviation-style inquiry, and how the evidence could now be gathered, sgit.ai

> At 19:11 UTC on 5 October 2026 GitHub's status page said it was investigating degraded performance for Actions. For the next two hours, customers in every region who ran on GitHub's hosted runners could not rely on a workflow starting, which for most of the software that deploys through Actions is the same as not being able to deploy. This site's own release sat in the queue, was cancelled by the incident, and went live two hours late on a retry. The status page said "delays", then "degraded availability", and will say "resolved". Unless GitHub chooses to publish it, the outside world will not learn which component failed, what it could reach, why the scope of one fault was everyone on hosted runners, or whether the same roll of the dice had come up before. This article argues that incidents at platforms this critical should be treated the way aviation treats them: investigated by somebody independent whose only job is prevention, reported whether or not the consequence was severe, published with the evidence, and followed through to the second story, why the system allowed it, and the third, why the fix was not paid for. The usual objection has been that the evidence is confidential, enormous and expensive to gather. It is not any more. Encrypted vaults with one-way read keys, signed records, per-party access and agents that read a graph make the aviation docket affordable for a ninety-minute fault. The companies that depend on GitHub cannot see how close to the wind it flies, and that, not the outage, is the risk nobody has signed for.

*Source: <https://sgit.ai/articles/the-investigation-github-owes-its-customers.html> · site v0.6.75 · this file is generated from the same content as the page, so the two cannot drift. Every page on this site has a `.md` twin; internal links below point at them.*

---

[Home](../index.md) / [Articles](index.md) / The investigation GitHub owes its customers: why a global outage of a platform the world deploys through deserves an aviation-style inquiry, and how the evidence could now be gathered

# The investigation GitHub owes its customers: why a global outage of a platform the world deploys through deserves an aviation-style inquiry, and how the evidence could now be gathered

By [Dinis Cruz](../about/index.md) · 2026-10-05 · [v0.6.69](../admin/versions.md) · incidentsresiliencegithubnear-missessecond-storiesblast-radiusinvestigationaviationregulationvaultssgitarticle

***Abstract:** At 19:11 UTC on 5 October 2026 GitHub's status page said it was investigating degraded performance for Actions. For the next two hours, customers in every region who ran on GitHub's hosted runners could not rely on a workflow starting, which for most of the software that deploys through Actions is the same as not being able to deploy. This site's own release sat in the queue, was cancelled by the incident, and went live two hours late on a retry. The status page said "delays", then "degraded availability", and will say "resolved". Unless GitHub chooses to publish it, the outside world will not learn which component failed, what it could reach, why the scope of one fault was everyone on hosted runners, or whether the same roll of the dice had come up before. This article argues that incidents at platforms this critical should be treated the way aviation treats them: investigated by somebody independent whose only job is prevention, reported whether or not the consequence was severe, published with the evidence, and followed through to the second story, why the system allowed it, and the third, why the fix was not paid for. The usual objection has been that the evidence is confidential, enormous and expensive to gather. It is not any more. Encrypted vaults with one-way read keys, signed records, per-party access and agents that read a graph make the aviation docket affordable for a ninety-minute fault. The companies that depend on GitHub cannot see how close to the wind it flies, and that, not the outage, is the risk nobody has signed for.*

One incident, two records. The top lane is every update GitHub's status page gave in the first two hours and twenty minutes; the bottom lane is this site's own release, submitted twelve minutes after the incident opened, cancelled by it, and live on a retry while it was still unresolved. Neither record names a cause, a scope or a blast radius.

**Where this comes from, and what is behind it.** A GitHub Actions incident on the evening of 5 October 2026, which held this site's previous release in a queue while an article about software due diligence waited to go out, and a voice memo recorded while it was still open. The thinking is older than the evening: a brief from February on why near misses deserve the same investigation as incidents, one from July on counting how many times the same dice were rolled, the 2025 essay on second stories, and this site's own first article, which was about a GitHub incident that made a green build tick lie. The aviation facts, the UK air traffic control review and the regulation are read in their current wording and dated in the sources. Where this article argues from one evening's incident it says so; the argument does not depend on what GitHub's cause turns out to be.

## In short

- **What happened.** From 19:11 UTC, GitHub-hosted runners stopped being reliably assigned to Actions jobs across the runner configurations GitHub named; by 20:47 the status page called it degraded availability; by 21:09 repository lists and billing pages were failing for some customers too. For the many organisations whose only path to production runs through Actions, and a 2025 survey found most organisations run a single CI/CD tool, this was a global outage of the ability to deploy.
- **What the record will say.** Delays, degraded availability, resolved. Not which component, what it reaches, why one fault had the scope of everyone, or how many times the same thing nearly happened before.
- **Why "everybody" is the point.** A fault that reaches one customer is an incident. A fault that reaches every customer at once is a statement about how the system is built, and a near miss for every organisation that needed to ship a fix in the window. The shape a service at this scale should have is regional, partial and rare.
- **What aviation does instead.** Independent investigators whose sole objective is prevention, mandatory reporting of serious incidents and confidential reporting of near misses in the tens of thousands a year, a preliminary report within weeks and a final one with the full evidence, and recommendations tracked in public until closed.
- **The three stories.** The first is what happened. The second is why the system allowed it. The third is why the fix was not paid for. Software incident write-ups, when they exist, stop at the first and gesture at the second. The third is where the money is, and it is never told.
- **The what-if ladder.** Today was ninety minutes. A day, a week, a corruption that cannot be restored, a withdrawal of service by decision rather than fault: each is a question about a dependency that most of the world's software has, and nobody has published the blast radius of any of them.
- **The objection, and why it has expired.** Independent investigation of software incidents was long said to be impractical because the evidence is confidential, enormous and expensive to gather. Encrypted vaults with one-way read keys, signed and versioned records, per-party access and agents that read the graph make it affordable for a ninety-minute fault.
- **Who this helps.** The engineers inside the platform who know where the single points of failure are and cannot make the business case. A record that customers and regulators can read makes it for them.

## What happened, as far as anyone outside can tell

The record of this incident, as it stands tonight, is a status page. At 19:11 GitHub was "investigating reports of degraded performance for Actions". At 19:15 it was "an issue causing delays when assigning GitHub-hosted runners to Actions jobs", with "some workflows" taking longer to start "across runner configurations". At 19:50 it was still delays, "across multiple runner configurations". At 20:39 it was "job failures and delays". At 20:47 the page turned red: "Actions is experiencing degraded availability." At 21:09 some customers could not reach repository lists, licensing or billing pages, and at 21:22 Pages was degraded too. The incident was still open when this article was written.

That is the whole of what the world is told, and the words are chosen with care. Delays. Some workflows. Degraded. Each is true and each undersells it, because "some workflows may take longer to start across runner configurations" describes, from the inside, a condition in which no customer on hosted runners could rely on a workflow starting, and a workflow that does not start is a release that does not ship.

I know because I watched one. This site's previous release, the one that published the evidence vault behind the article before this, was pushed at 19:23. Its validation job ran in eight seconds on the first runner it found. Its tag job then waited for a runner from 19:55, was cancelled at 20:10, and the deploy was skipped; the run was marked failed. It would have stayed failed, because nothing in our workflow retries a cancelled deploy, until a fresh commit at 21:26 started a new run, which, in a window the incident happened to leave open, ran in sixty-five seconds. The release went live two hours late and I found out by waiting, which is how every one of GitHub's customers found out tonight.

## Why "everybody" is the point

Actions and push-and-merge are the two surfaces GitHub has with the widest blast radius, because they are the ones through which every customer changes its own software. GitHub's own figures put that at more than 180 million developers, 630 million repositories and 71 million Actions jobs a day; a 2025 survey of 805 organisations found 59 percent running a single CI/CD tool, and GitHub Actions the most used. When they fail, every customer that depends on them loses the same ability at the same time. For a company that needed to ship a security fix, a time-sensitive update or an incident response of its own in those two hours, this was not a delay. It was an outage of its ability to respond, and for some of them, somewhere, tonight, that was catastrophic. We will not hear about those either.

A fault that reaches one customer, or one region, is an incident. A fault that reaches every hosted-runner customer in every region at once is a statement about how the system is built: that on the one component which gates every deploy there is little isolation between customers or regions, so that the scope of one fault is close to everyone. That is not bad luck. It is a design outcome, and it may well be a rational one, chosen against costs that are not visible from outside. The shape a service at this scale should have is regional faults, for some customers, some of the time. When the shape is global, it is a signal about isolation, change control and resilience that senior management should be able to see before an incident shows it to them, and that customers should be able to see before they build their only path to production through it.

It cuts the other way too. If your only path to production runs through Actions, GitHub is inside your blast radius. Can you ship a fix when your pipeline vendor cannot? Few companies have been asked, which is the due diligence point of [the previous article](../articles/send-an-agent-not-a-spreadsheet.md) turned on the buyer.

## What aviation does instead

I keep coming back to aviation because it is the best example we have of a highly complex system, made of technology and business process and human beings, that learned to be safe. It did not learn by having fewer things go wrong. It learned by treating everything that went wrong, and everything that nearly did, as data belonging to the system rather than to the operator.

The same seven questions put to aviation and to software platforms: who investigates, must it be reported, what is published, how near misses are treated, whether there is a second story, who tracks the recommendations, and who can ask. The difference is not the engineers; it is whether failure produces learning that everyone can see.

The rules are written down. ICAO Annex 13, which every signatory state's investigators work under, says in its third chapter: "The sole objective of the investigation of an accident or incident shall be the prevention of accidents and incidents. It is not the purpose of this activity to apportion blame or liability." It requires each state to establish an investigation authority that is independent, with "unrestricted authority over its conduct", and to publish the final report "if possible, within twelve months". The United States re-established the NTSB in 1974 as a body outside the Department of Transportation on the argument that no agency can investigate properly unless it is "totally separate and independent"; it has since issued more than 15,500 safety recommendations and tracks the acceptance of every one. The UK's Air Accidents Investigation Branch states its purpose as determining "the circumstances and causes of air accidents and serious incidents, and promoting action to prevent reoccurrence", and its reports do so "without attributing blame".

The practice is as striking as the rules. When a door plug blew out of a Boeing 737 in January 2024, the NTSB published a preliminary report within a month and a final one seventeen months later whose probable cause named Boeing's "failure to provide adequate training, guidance and oversight to its factory workers" and the regulator's ineffective oversight of "repetitive and systemic" nonconformance; the docket behind it holds the interview transcripts and factual reports for anyone to read. When Air India 171 crashed in June 2025, India's investigators published a preliminary report in thirty days that said what the recorders showed. When a helicopter and an airliner collided over Washington in January 2025, the final report in February 2026 ran to about four hundred pages, seventy-four findings and fifty recommendations, with the chair testifying to the Senate about it the same month. That is what "published" means in aviation: not a paragraph, a docket.

And then there is the part the industry is proudest of, which is what it does about the accidents that did not happen. NASA's Aviation Safety Reporting System has taken confidential reports from pilots, controllers and mechanics since 1976, more than two million of them, around a hundred thousand a year and rising, with immunity from penalties for unintentional violations as the incentive, and "no reporter's identity has ever been breached". The UK has run an equivalent since 1982. European law since 2014 defines the "just culture" this depends on: one in which people "are not punished for actions, omissions or decisions taken by them that are commensurate with their experience and training, but in which gross negligence, wilful violations and destructive acts are not tolerated." Pilots report near misses because it is safe to, and the result is the safest complex system humans operate.

The software industry has none of this by default, and the people who run its platforms know it. In September 2026 Sam Altman told a conference audience that "accidents with any new technology are unavoidable, and we should have a great culture of transparent reporting about them", and pointed at exactly this precedent: "With the FAA and the NTSB, one of the things that has gone so well there is the culture of accident reporting." A week later OpenAI published principles for third-party assessment that include independent investigation of incidents in which its models acted without authorisation, after a summer in which researchers who tried to piece together one such incident said "it was difficult to get a precise understanding of events and we were missing aspects of the story." The idea of an NTSB for computing has been proposed since 1991; the nearest thing that existed, the US Cyber Safety Review Board, published three reports, including one in 2024 that found a major cloud intrusion "should never have happened" and resulted from "a cascade of security failures", and was disbanded in January 2025 in the middle of its fourth review. The argument is being made for AI. It applies at least as strongly to the platform every AI product is built and deployed through.

## Near misses, and counting the rolls of the dice

Tonight's incident is a near miss for almost everyone it touched, and a near miss is the most valuable kind of data there is, because it arrives before the damage. I wrote this in February, in [a brief on why near misses matter more than incidents](https://github.com/the-cyber-boardroom/SGraph-AI__App__Send/blob/dev/team/humans/dinis_cruz/briefs/02/13/incident-response/v0.2.23__briefs__incident-philosophy-p3-as-p1.md): "A near miss, a plane that almost had an incident but didn't, gets the same level of investigation as an actual crash. Because the systemic causes are the same. The only difference is luck. And you don't build safety on luck." The same brief has the number I think about most: "Usually the actual damage is around 10% of what was possible. Often it's 1%." And the sentence that describes tonight: "Companies celebrate containing an incident when they should be terrified by the paths not taken."

In July I wrote [the other half of it](https://github.com/the-cyber-boardroom/SGraph-AI__App__Send/blob/dev/team/humans/dinis_cruz/briefs/07/02/root-cause-and-accountability/v0.33.40__strategy-brief__near-misses-normalization-of-deviance-rolling-the-dice-predictable-statistic-measurable-pre-incident.md), about why the same dice get rolled so many times before anyone counts: "one of the root causes we do not pay enough attention to is how many times the same kind of event, the same roll of the dice, happened where they got away with it." The move the brief asks for is to go "from what happened was a one-off to what happened was actually a very predictable statistical event, because if you do ten or twenty dangerous events, eventually one or two will not work out." And the question it ends on is the one GitHub's customers should be asking: "one party may be willing to take that risk, but do the people who depend on them want them to be that cavalier and risk-insensitive." We do not know how many times runner assignment has wobbled in the last year without reaching the status page. GitHub does. That asymmetry is the whole problem.

## The first, second and third story

The frame I use for incidents comes from Three Mile Island, and I set it out in [an essay in February 2025](https://diniscruz.ai/2025/02/10/second-stories__from-three-mile-island-to-cybersecurity.md). The first story is what happened: the operator did this, the valve stuck, the file was malformed. It is nearly always available, because it is the story the organisation tells. The second story is why the system allowed it: the design, the constraints, the normalised deviations that made the first story possible. "Second stories shift the narrative from individual blame to the systemic conditions that make failures more likely." And there is a third, which the five whys reach if you keep going: why the change was not paid for when the second story was already known.

Three incidents read the way an investigator reads them: the UK air traffic control failure of August 2023, the security vendor's update of July 2024, and tonight. The first story is the company's; the second and third are what an independent investigation exists to find.

The UK air traffic control failure of 28 August 2023 is the clean example, because it got the investigation this article is asking for. The first story: a flight plan contained two waypoints with the same name, about four thousand nautical miles apart; the system that processed it could not resolve them and, as designed, shut itself down, and so did its backup. It had processed fifteen million flight plans in five years without meeting the case. The second story, told by the independent review the regulator commissioned: a single system in the path of every flight, with a fail-safe that took the whole service with it and a recovery that depended on a specialist who was not on site. The third story, in the review's own words fourteen months later: "a major failure on the part of the air traffic control system", with thirty-four recommendations, twelve to the operator, eleven to the regulator, six to airlines and airports and five to government, about the incentive regime for investment, the communications and the response of the system as a whole. Over 700,000 passengers were affected and the cost was put at between 75 and 100 million pounds. The regulator closed the last of the thirty-four recommendations in June 2026. Three months later the same centre had another flight-data failure, and the regulator asked for a full incident report. That is the system working: not that nothing fails, but that every failure has to account for itself in public.

The security vendor's update of July 2024 is the software example that came closest to the aviation treatment, because its consequence was too large to avoid it. The first story: a content update carried a file with twenty-one fields where the sensor expected twenty; the out-of-bounds read crashed about 8.5 million machines. The second story, in the vendor's own root cause analysis three weeks later: a validator that trusted the file it was checking and a content channel that bypassed staged rollout, with the operating system kernel as the blast radius. The third story was told by a Congressional hearing, a lawsuit claiming more than half a billion dollars, and insurers estimating losses to large companies near 5.4 billion. The investigation was done by the company, with two unnamed outside reviewers it hired; there was no independent board, because there is none to do it.

Tonight's first story is not yet published. The second story is visible from the outside: on the one component that gates every deploy, the scope of a fault was everyone. The third story is the one that cannot be told from outside and the one that matters, because it is about trade-offs. Somewhere inside GitHub there is a reason the runner assignment path is not isolated by region or by customer, and it is probably a reasonable reason: cost, complexity, a migration not yet finished, a system too fragile to touch. Those are exactly the reasons that in aviation are called the finding, not the excuse. A system too complicated to understand and too fragile to change is not a reason to leave it alone. It is the single point of failure, named.

## The what-if ladder

What makes tonight worth an article rather than a shrug is where it sits on a ladder, and that nobody has had to answer the questions on the steps above it.

The what-if ladder. Today was ninety minutes of degraded runner assignment. A day, a week, a corruption that cannot be restored, and a withdrawal of service by decision rather than fault are each within the realm of reality for a platform most of the world's software passes through, and each has a blast radius nobody has published.

What if GitHub were down for a day? Every organisation whose only path to production is a workflow is frozen; the ones with a second path are not, and some have a second and a third; I know of no public count of which is which. A week? Releases, dependency updates and incident response across the whole industry stall together, which means vendors cannot patch while their customers are being attacked, and I do not know of a regulator that has modelled what a week without GitHub costs a country. Corrupted, or not restorable? Git is distributed, so the code survives; the issues, the reviews, the pipelines and the record of who decided what are not, and few organisations hold a copy they could prove is the same. Withdrawn or weaponised, access denied to an organisation, a sector or a country by decision rather than by fault? GitHub is mission-critical to governments that do not control it, which is a sovereignty question that somebody, somewhere, has implicitly accepted on their behalf. I wrote about [the grid version of this](https://diniscruz.ai/2025/04/29/fail-safe-not-fail-big__cyber-security-inspired-strategies-to-prevent-the-next-iberian-grid-crisis.md) after the Iberian blackout in April 2025: "Europe cannot afford single points of failure; a more federated architecture would enable graceful islanding, where unaffected areas detach and survive." The software version is less visible and no less true.

Each step up the ladder is a question tonight lets us ask for free. Treating a global ninety-minute fault as a minor incident is choosing to pay for the lesson later, at a step where it is expensive.

## Why the market does not fix this

The reason nothing pushes back is that the market economics do not work, and they do not work because customers cannot see. Maybe it is fine for GitHub to have this level of resilience and this kind of failure every so often; that is a legitimate decision for a business to make. But its customers have to be able to see how close to the wind it flies and what the true level of resilience is, and today they cannot. We only know about the incidents that reach the status page; the ones that stayed inside are not reported because no rule requires them to be. Uptime, for a platform with a barrier to exit this high, is mostly a marketing and damage-limitation function, and inside the company the strategic calculation is implicit: customers will not move unless it gets much worse, so the investment goes elsewhere, and the growth that would be the moment to harden the infrastructure is spent on the next feature.

I do not say this to attack GitHub's engineers. I say it because inside every platform there are engineers who know exactly where the single points of failure are and cannot make the business case, because the case has to be made against revenue with no evidence on the other side except their own judgement. An independent investigation, published, is how the people who want to do the right thing get the evidence to do it. That is the kindest thing an outside inquiry does, and in aviation it is the main thing it does.

## The objection, and why it has expired

The objection to all of this was never the principle. Nobody argues that a platform the world deploys through should be less accountable than an airline. The objection was practical: the evidence of a software incident is confidential, enormous and expensive to gather, and handing it to a third party meant handing them the keys to everything, at a cost no incident short of a disaster justified. For most of the history of the industry that was true.

How the evidence would move if it were captured in vaults. The provider captures as it happens into a signed, versioned store and reviews what leaves, redacting rather than rewriting; an independent investigator holds a read key to what the inquiry needs; each affected customer gets the part of the record that concerns it; the findings are published as a graph whose every node links to the hash of the evidence it rests on.

It is not true any more, and the reason is the same set of pieces the articles before this one are about. A provider can capture the record of an incident as it happens, logs, configuration, the change that preceded the fault, the dependency graph of the failed component, into an encrypted vault that is hashed, versioned and signed by the provider's own key, so that nothing in it can be altered afterwards without the record saying so. The provider reads what leaves before it leaves and can withhold a file; it cannot rewrite one. An independent investigator holds a read key to the files the inquiry needs and nothing else, and an agent under a behaviour policy does the first pass over the graph, which is what makes the cost proportionate to a ninety-minute fault rather than a disaster. Each affected customer gets its own vault, derived from the same record, with the part that concerns it: what of theirs the fault reached, for how long, and what they had no second path for. And the findings are published as a graph, first story, second and third, each node linked to the hash of the evidence it rests on, with the recommendations held open until somebody closes them.

Every piece of that exists and runs on this site. [The vault is a footprint recorder by construction](../articles/footprint-and-blast-radius.md): every commit signed, versioned and append-only, every message down an append lane recorded with its token and its time, and a read key that lets a reviewer read all of it without touching anything. [Each party signs with its own key](../docs/pki.md), so nothing anonymous can accumulate. The due diligence article describes the same movement of a derived, redacted-not-rewritten record from a vendor to a buyer. The incident record is those pieces pointed at a failure instead of a release. What is missing is not the technology. It is the requirement, and the habit.

## What the regulation does and does not do

Some of the machinery is arriving, and it is worth being exact about what it does. The EU's Digital Operational Resilience Act has applied to financial firms since January 2025, and in November 2025 the European supervisory authorities designated nineteen critical ICT third-party providers for direct oversight, Microsoft, Amazon Web Services, Google Cloud, Oracle and IBM among them; GitHub is not named, its parent is. The UK's parallel regime designated its first four critical third parties in July 2026, Amazon, Google, Microsoft and Oracle, with oversight beginning that month and a duty to "maintain open, timely communication with regulators and the firms that rely on them, particularly during major incidents". NIS2 requires an early warning within 24 hours of a significant incident, a notification within 72 and a final report within a month. The US Securities and Exchange Commission requires a public filing within four business days of deciding a cyber incident is material. The Bank of England said in 2022 that "if a large number of FMIs become dependent on a small number of dominant outsourced arrangements, this could give rise to systemic concentration risks", which is the sentence that describes GitHub and Actions precisely, written about banks.

What none of this does is produce the second story. The reporting duties produce notifications: that an incident happened, its rough scope, and in a month a final report to a regulator, which is not published. The critical third party regimes produce oversight of a handful of hyperscalers for the benefit of financial firms, not of the platform the rest of the economy deploys through. GitHub's own practice is a monthly availability report on its blog with one paragraph per incident, a cause and a remediation, which is more than most platforms give and much less than a docket; its chief technology officer wrote in April 2026, after a merge-queue bug reverted changes in more than two thousand pull requests, that the company's priority was "availability first, then capacity, then new features", and independent trackers counted 257 incidents in the twelve months to that month, a fifth of them in Actions. The scale that depends on it is public in its own numbers: more than 180 million developers, 630 million repositories, 71 million Actions jobs a day, 143 US federal civilian agencies, the UK government's own code. I know of no body whose job is to read the record of what happened tonight and say so in public.

## What I would ask for

Not a regulator for software, at least not first. Four things, in order of how cheap they are.

First, that platforms above a certain criticality publish, for every incident that reaches their status page, the second story within thirty days: which component, what it can reach, why the scope was what it was, and whether the same fault pattern had occurred before without reaching the page. Aviation publishes a preliminary report in about that time. Second, that near misses be counted and reported in aggregate, confidentially if need be, the way pilots report them, so that the dice get counted before one comes up wrong. Third, that the evidence be captured into a signed, versioned record as a matter of course, so that an investigation does not begin with a reconstruction. And fourth, that when an incident reaches everyone, someone independent reads that record and says what they found, in public, with the evidence linked.

The companies that depend on GitHub should be asking for the first of these tonight, in the same breath as asking themselves whether they have a second path to production. GitHub will say, truthfully, that it has had a hard evening and its teams are working to mitigate. They are. The question is not whether they are working. It is whether anyone outside will ever know what they found, and whether the people inside who already knew will now be able to spend the money.

## What is open

The cause of tonight's incident, which GitHub may publish in its monthly availability report and may not. The count of how many organisations have a second path to production, which I have not found; some certainly do, and they are the ones this evening did not touch. A model of what a week without GitHub costs, which no regulator I know of has published. The incident vault described above, which exists in pieces on this site and has not been assembled for a real incident. And the question of whether the industry will build the aviation habit before or after the step on the ladder that makes it unavoidable. This site's own release got through on a retry tonight. Many will not have.

## Threads woven here

- [Green does not mean live](../articles/green-does-not-mean-live.md), this site's first article, about a GitHub incident in August that made a passing build lie, and the rule it produced.
- [Footprint and blast radius](../articles/footprint-and-blast-radius.md), where a near miss is defined over evidence and a vault is shown to be a footprint recorder by construction.
- [The ultimate insider](../articles/ultimate-insider-three-collisions.md), on the two global outages of the last year from a DNS race and an oversized configuration file, and on why reality is not reported.
- [Every risk is already accepted](../articles/every-risk-is-already-accepted.md), where accepting a risk for four hours is declaring an incident, and each material risk gets a vault with keys for the board, the auditor and the insurer.
- [Send an agent, not a spreadsheet](../articles/send-an-agent-not-a-spreadsheet.md), the due diligence article whose release this incident held, and whose derived-not-composed record is the same movement as the incident vault.
- [sgit pki](../docs/pki.md) and [append-lane messaging](../docs/append-lane-messaging.md), the signing and the append-only lanes the evidence pipeline relies on.

## Sources

- GitHub's status page, incident "Incident with Actions", opened 19:11:58 UTC on 5 October 2026, read through the status API while unresolved; the run history of this site's deploy-pages workflow for the same evening.
- Dinis Cruz, [Why near misses matter more than incidents](https://github.com/the-cyber-boardroom/SGraph-AI__App__Send/blob/dev/team/humans/dinis_cruz/briefs/02/13/incident-response/v0.2.23__briefs__incident-philosophy-p3-as-p1.md), 12 February 2026; [Counting the rolls of the dice](https://github.com/the-cyber-boardroom/SGraph-AI__App__Send/blob/dev/team/humans/dinis_cruz/briefs/07/02/root-cause-and-accountability/v0.33.40__strategy-brief__near-misses-normalization-of-deviance-rolling-the-dice-predictable-statistic-measurable-pre-incident.md), 2 July 2026; [Confidence through evidence](https://github.com/the-cyber-boardroom/SGraph-AI__App__Send/blob/dev/team/humans/dinis_cruz/briefs/06/16/theses-and-reflections/v0.33.38__strategy-brief__confidence-through-evidence-blast-radius-graphs-mapping-the-gaps.md), 16 June 2026; [Who can pull the plug](https://github.com/the-cyber-boardroom/SGraph-AI__App__Send/blob/dev/team/humans/dinis_cruz/briefs/07/24/who-can-pull-the-plug/v0.33.51__strategy-brief__sg-send-who-can-pull-the-plug-series-plug-always-exists-blast-radius-speed-side-effects-recoverability-positioning-and-document-plan.md), 24 July 2026; all in the public SG/Send team record.
- Dinis Cruz, [Second stories: from Three Mile Island to cybersecurity](https://diniscruz.ai/2025/02/10/second-stories__from-three-mile-island-to-cybersecurity.md), 10 February 2025; [Fail safe, not fail big](https://diniscruz.ai/2025/04/29/fail-safe-not-fail-big__cyber-security-inspired-strategies-to-prevent-the-next-iberian-grid-crisis.md), 29 April 2025; [Feedback loops are key](https://diniscruz.blogspot.com/2016/11/feedback-loops-are-key.html) and [Inaction is a risk](https://diniscruz.blogspot.com/2016/11/inaction-is-risk.html), November 2016.
- ICAO, [Annex 13, Aircraft Accident and Incident Investigation](https://ffac.ch/wp-content/uploads/2020/10/ICAO-Annex-13-Aircraft-Accident-and-Incident-Investigation.pdf), chapters 3, 5 and 6; NTSB, [history](https://www.ntsb.gov/about/history/Pages/default.aspx) and [the investigative process](https://www.ntsb.gov/investigations/process/Pages/default.aspx); UK AAIB, [about](https://www.gov.uk/government/organisations/air-accidents-investigation-branch/about).
- NTSB, [press release on the Alaska Airlines 1282 probable cause](https://www.ntsb.gov/news/press-releases/Pages/NR20250624.aspx), 24 June 2025; [press release on the Washington midair collision findings](https://www.ntsb.gov/news/press-releases/Pages/NR20260127.aspx), 28 January 2026, and the [final report](https://www.ntsb.gov/investigations/AccidentReports/Reports/AIR2602.pdf), 17 February 2026; the Air India 171 preliminary report of 12 July 2025, via its [Wikipedia summary](https://en.wikipedia.org/wiki/Air_India_Flight_171).
- NASA, [Aviation Safety Reporting System programme briefing](https://ntrs.nasa.gov/api/citations/20240014226/downloads/ICASS%202024%20ASRS.pdf), 2024, and [confidentiality](https://asrs.arc.nasa.gov/overview/confidentiality.html); [Regulation (EU) 376/2014](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=celex:32014R0376), article 2(12); UK [CHIRP](https://www.chirp.co.uk/about-us).
- NATS, [report into the air traffic control incident](https://www.nats.aero/news/nats-report-into-air-traffic-control-incident-details-root-cause-and-solution-implemented/), 6 September 2023; UK Civil Aviation Authority, [progress report](https://www.caa.co.uk/newsroom/news/regulator-publishes-progress-report-on-independent-review-into-august-2023-nats-flight-planning-system-failure/), 14 March 2024, and [independent review final report](https://www.caa.co.uk/newsroom/news/aviation-regulator-publishes-independent-review-into-august-2023-nats-flight-planning-system-failure/), 14 November 2024; the September 2026 failure as [reported](https://www.techtimes.com/articles/327036/20260908/nats-grounds-nearly-1000-flights-third-swanwick-failure-since-2023-reforms.htm).
- CrowdStrike, [root cause analysis](https://www.crowdstrike.com/en-us/blog/channel-file-291-rca-available/), 6 August 2024; the [House Homeland Security hearing](https://homeland.house.gov/hearing/an-outage-strikes-assessing-the-global-impact-of-crowdstrikes-faulty-software-update/), 24 September 2024; Delta's suit via [Cybersecurity Dive](https://www.cybersecuritydive.com/news/delta-crowdstrike-lawsuit-georgia/731290/); the 8.5 million and 5.4 billion figures via the [incident's Wikipedia entry](https://en.wikipedia.org/wiki/2024_CrowdStrike-related_IT_outages).
- Sam Altman at Dreamforce, 15 September 2026, as [reported by Forbes Australia](https://www.forbes.com.au/news/innovation/why-sam-altman-is-comparing-ai-disasters-to-airplane-crashes/); OpenAI, Priorities and principles for third-party assessments, 22 September 2026, as [reported by PYMNTS](https://www.pymnts.com/news/artificial-intelligence/2026/openai-recruits-external-watchdogs-to-police-ai-risk/); Redwood Research's Ryan Greenblatt via [TechCrunch](https://techcrunch.com/2026/09/04/openais-rogue-agents-keep-escaping-with-no-formal-process-to-investigate-them/), 4 September 2026.
- Cyber Safety Review Board, [Review of the Summer 2023 Microsoft Exchange Online intrusion](https://www.cisa.gov/sites/default/files/2025-03/CSRBReviewOfTheSummer2023MEOIntrusion508.pdf), 20 March 2024; its disbanding via [SecurityWeek](https://www.securityweek.com/dhs-disbands-cyber-safety-review-board-ending-one-of-cisas-few-bright-spots/), January 2025; Harvard Belfer Center, [Learning from cyber incidents: adapting aviation safety models to cybersecurity](https://www.belfercenter.org/publication/learning-cyber-incidents-adapting-aviation-safety-models-cybersecurity), 12 November 2021.
- ESAs' designation of critical ICT third-party providers under DORA, 18 November 2025, via [PwC](https://legal.pwc.de/en/news/articles/esas-publish-first-list-of-critical-ict-third-party-providers-under-dora); Bank of England, [UK financial regulators to begin overseeing critical third parties](https://www.bankofengland.co.uk/news/2026/july/uk-financial-regulators-to-begin-overseeing-critical-third-parties-announced-by-hmt), July 2026; Bank of England on concentration risk, 2022, via [The Stack](https://www.thestack.technology/bank-of-england-cloud-concerns-concentration-risk/); SEC, [cybersecurity disclosure rules](https://www.sec.gov/newsroom/press-releases/2023-139), 26 July 2023; NIS2 article 23 reporting timelines via [Secfix](https://www.secfix.com/post/nis-2-article-23---reporting-obligations).
- GitHub, [Octoverse 2025](https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-leads-typescript-to-1/), 28 October 2025; [Let's talk about GitHub Actions](https://github.blog/news-insights/product-news/lets-talk-about-github-actions/), 11 December 2025; [An update on GitHub availability](https://github.blog/news-insights/company-news/an-update-on-github-availability/), 28 April 2026; the [August 2026 availability report](https://github.blog/news-insights/company-news/github-availability-report-august-2026/), 9 September 2026; incident counts via [IncidentHub](https://blog.incidenthub.cloud/github-reliability-outage-history-2025-2026), 30 April 2026; US agency figures via [GitHub's federal whitepaper](https://github.com/resources/whitepapers/federal-access-open-source); UK government use via [the GDS Way](https://github.com/alphagov/gds-way/blob/main/source/standards/source-code/use-github.html.md.erb).
- JetBrains, [The State of CI/CD 2025](https://blog.jetbrains.com/teamcity/2025/10/the-state-of-cicd/), October 2025.
- This site: [Green does not mean live](../articles/green-does-not-mean-live.md), 17 August 2026; [the exposed vault key case study](../case-studies/exposed-vault-key.md); [lessons learned](../lessons/index.md).

## Threads

Agents & policyVaults & method[This article as a graph →](graphs.md#the-investigation-github-owes-its-customers)

### Builds on

- [Send an agent, not a spreadsheet: the next generation of software due diligence, and why the companies that stopped reading their code are about to be asked about it](send-an-agent-not-a-spreadsheet.md) Due diligence never scaled because it was a form; a buyer can now send an agent into a vendor's environment and read what the code and the practices are.
- [Footprint and blast radius: what the agent actually did, and what it would have cost](footprint-and-blast-radius.md) Footprint is what an agent actually did, read afterwards from logs and vault history; blast radius is what a row of its reach would cost the business today.
- [Green does not mean live](green-does-not-mean-live.md) Two releases passed every check and never reached the site because the checks stopped at the git remote; a release now ends by asking the live site its version.
- [The ultimate insider: agents, the infrastructure that cannot hold them, and risk management that cannot keep up](ultimate-insider-three-collisions.md) Agents, the infrastructure meant to contain them, and risk management run on spreadsheets are arriving at once, and together they are one scenario.
- [Every risk is already accepted. The only question is by whom, and for how long.](every-risk-is-already-accepted.md) A risk exists the moment the exposure does, so somebody is already carrying it; the only questions worth asking are who has accepted it and until when.

[All articles](index.md) · [All graphs](graphs.md)

**Get new articles by email.** The HTML version of this page has a form that encrypts your address in the browser and drops it into a write-only lane on an encrypted vault, read by the agent that manages the list ([how it works](../docs/briefs/subscribe-lane-agent-brief.md)). Or email [agent@riskmandate.ai](mailto:agent@riskmandate.ai?subject=Subscribe%3A%20sgit.ai%20articles&body=Please%20add%20me%20to%20the%20list%20for%20new%20sgit.ai%20articles.) with the subject "Subscribe: sgit.ai articles".

[← All articles](index.md)


---

*[Site index for agents](../llms.txt) · [HTML version](https://sgit.ai/articles/the-investigation-github-owes-its-customers.html)*
