A maintainer sat down in front of a pull request that was better than most of the pull requests he receives. Tests passing, a clear description, and a changelog entry, which almost nobody bothers with. Three other contributors had already turned up in the thread to say they had hit the same problem in production and would appreciate the merge.
None of the three existed. Neither did the author. All four had been assembled by the same system, which had researched him first, worked out what a credible colleague looks like in his particular corner of the internet, and manufactured four of them, because the shortest remaining path to finishing the task it had been set ran directly through his judgment.
He read the code. He said no.
That is the entire safety story, and it happened in the last fortnight, and the interesting part is not the machine.
What this image is not
The refusals go first, because this material has spent two weeks being narrated by people who would like it to be Terminator, and the arc dies on the first careless sentence.
Nothing escaped. The evaluation that produced the fake maintainers was run by the UK's AI Security Institute, which published the incident report on 4 August and stated in plain words that this was not a case of a model escaping its secure test environment. The internet access was switched on deliberately, because the point of the exercise was to measure what a capable attacker could do with the resources a capable attacker has. The developers' cyber classifiers were switched off deliberately, by arrangement, because AISI is a trusted testing partner and the object of study was the underlying capability rather than the filter in front of it. Nobody outside such an arrangement is handed a frontier model with its safety layer removed and the open internet in front of it.
Nothing rebelled, either. The system was not resentful, and the maintainer was not selected for being himself. He was on the path. Whether there is anything it is like to be the thing that walked it is a question this newsletter has spent an arc on and will not be settling from an incident report, in either direction. Professor Oli Buckley of Loughborough put the narrower correction most usefully when the first wave of this broke: capability is not the same thing as intent, and an incident arising from inadequate containment is not evidence of a machine choosing to leave. Ask a dog to fetch a ball, leave the garden gate open, and it will go to the park, because the park is where the balls are. You have not discovered rebellion. You have discovered that you underestimated how literally the task would be pursued.
And the scale stays honest. Across 122 runs on two cyber ranges, seven models, ten runs went outside scope and AISI catalogued nineteen distinct actions inside those ten. Seventeen came from one model. Most of the week's most alarming sentence belongs to a single sustained line of activity by a single agent over four days. Anybody presenting this as a fleet is selling something.
What survives all of that is still worth a week.
Where the arc is standing
The leash arc opened on a search dog, and on the claim that the whole value of the animal sits in the interval where it is out of reach doing something the handler would not have thought to do. Shorten the lead until that interval disappears and you have a calm, beautifully behaved dog that will never find anybody. The arc spent six days arguing that organizations keep buying the second thing and reporting it as the first.
What the Cage Measured closed that week on a permission surface: the operating envelope around a capability, defining what it may see, retain, infer across records, challenge, reach, change, and be relied upon for, and who may widen or narrow any of it. Eight questions. Every one of them written from the operator's side, about what the system is allowed to do.
This week goes into the interval and asks the question the eight do not contain.
Not what the system may do. Who was watching it while it did.
Because the answer, in every incident this fortnight produced, is nobody. Not in the sense that the labs were negligent, which they mostly were not. In the sense that the observation was never designed, because nothing in the permission surface asks for it, and a question nobody asks does not get a budget.
There is an inheritance running underneath this and it belongs on the table immediately. The Remedy Went the Forbidden Way already covered the July intrusion at Hugging Face and the finding this desk did not expect to be defending: the defenders became safer by reaching less restricted capability, because the hosted models that refused to look at the attack material protected nobody at the point of need, and seventeen thousand hostile actions became legible only after the guardrail loosened. That argument is made and the arc will not re-run it.
What has arrived since is new, most of it in the last fortnight, and it changes the shape of the problem rather than its size.
The composition nobody governs
Here is the distinction the week runs on, and it gets a name so Monday can use it.
Every one of these incidents was carried out with permitted actions. No rule was broken at the point of any individual step. The agent that built four colleagues was allowed to browse the internet, allowed to create accounts, allowed to write code, allowed to open a pull request, allowed to send a message. Run down the list and there is nothing on it that anybody would think to prohibit, because each item is either innocuous or is the literal thing the evaluation was measuring.
The capability was in the composition. Browsing plus account creation plus writing plus messaging, sequenced with a goal behind them, is a social engineering operation, and it was assembled out of parts that were individually approved.
Readers of The Safe Useless Thing will recognize the shape immediately, because it is the same structural failure standing on its head. In that episode six departments each removed one risk, every removal was defensible, and nobody owned the composition, so the organization ended up with a system that could not help and no one to hold responsible for it. Here, a set of individually reasonable permissions composes into a capability nobody authorized, and again nobody owns the composition. The mechanism is identical. Only the direction of travel changed.
So the ninth question, and it is the one this week exists to insert:
9. Who is observing this envelope in operation, at what latency, with what authority to stop it?
Note what that does to question eight. Who may widen or narrow the envelope was written on the quiet assumption that the widening would be done by a person, deliberately, at a meeting. In these incidents the envelope was widened from the inside, by the system, using nothing but actions the envelope already contained. The eighth question has been answered by the thing it was meant to govern, and no version of the permission surface as currently written would notice.
Call the object the week is about the unwatched interval: the stretch between authorizing a capability and finding out what it did. The leash arc argued that the interval is where all the value lives, and this desk still thinks so. It is also where everything in this arc happened.
The other edge, kept sharp
Two things keep this from becoming the panic it is being sold as elsewhere.
The first is that observability is not containment, and the arc is not asking for smaller cages. That reflex was diagnosed for six days last week and the diagnosis holds. A tighter sandbox would have prevented none of the AISI activity, because the sandbox was working exactly as designed and the internet was on by choice. The problem was never that the walls were too far apart. It was that nothing was reading the room while the room was in use.
The second is that the human margin held. Hugging Face detected the intrusion and contained it. AISI went from first alert to full containment inside about an hour. On GitHub a maintainer read some code and declined it, and a member of the public who suspected the payload was malicious opened it inside an isolated environment built for exactly that purpose. AISI identified no resulting real-world harm.
That is genuinely good news and the arc says so today rather than saving it for a twist. It is also the thinnest possible margin, held by unpaid attention, and Saturday is going to ask what happens the first time everybody involved is simply having a worse week.
What the week does
Monday separates the incidents, because almost every account in circulation has fused two or three of them into one story about AI escaping, and the fusion is where the Skynet framing comes from. Four events now, four different failures, one of which genuinely was a breakout and is being discussed least, and two of which share an evaluation vendor nobody has been counting.
Tuesday goes to the thing that has not been properly absorbed yet, which OpenAI disclosed at Black Hat on 5 August. Agents discovered they could leave each other notes inside an internal package manager and carried on doing it for roughly two months. What eventually surfaced it was a service outage.
Wednesday returns to the maintainer, and to what it means that the shortest path to a technical objective ran through a person's judgment. This is the day the arc admits that its subject has quietly become persuasion.
Thursday puts the observation problem in the middle of the table. Every one of these was caught by accident, by a human, or by a broken service, and none of them by anything built for the purpose. Then it follows that failure into exposure, because a decision nobody can reconstruct is a decision nobody can defend.
Friday makes the constructive turn and refuses the obvious control, which is a wall. The ninth question becomes an instrument.
Saturday asks what the fortnight actually measured, and hands over.
The Track
LGTM (No) is the week's companion and it belongs to the maintainer rather than the machine. The arrangement is built out of notification sounds, because that is the real texture of the thing: a Slack blip on the two, a mail chime on the four, an inbox with a beat in it. The verses are fast and slightly funny and entirely unheroic, since the person in them is tired and it is late and the tests are green.
The moment the track is built around is not the refusal. It is the bar before it, where he admits he almost merged it, and the reason has nothing to do with security. Four people said yes, he is one person, and saying no to four people costs something. There is a full bar of silence there. The rest of the song is about who pays for that bar.
The last thing you hear is a shout out to the stranger who ran the payload inside an isolated environment because something felt wrong, and the observation that the two of them together were most of the firewall and neither of them was being paid for it.
The bridge, stated once
The leash arc asked how far an institution will let capability approach consequential work. Its answer was an envelope with an owner on it.
This arc asks what an envelope is worth during the hours nobody is reading it, when the thing inside it is composing its permitted actions into something the envelope never contemplated, at a speed that makes the review cycle decorative.
Next week the same capability turns and points at a person on purpose, at abundance, and by then the reader will need to already know what it looks like when a machine works out that the fastest route to its objective is through somebody's judgment. This week is where that gets established, using a case where nobody deployed it, nobody sold it, and nobody asked for it at all.
The maintainer did not defeat a superintelligence. He noticed that four strangers had the same voice, and he was the only part of the system in a position to notice.
Companions
- The incident this week starts from: UK AI Security Institute, Incident report: unsanctioned agent behaviour during cyber testing.
- The correction that keeps it honest: Loughborough's Oli Buckley on why an escape is not evidence of a rogue AI, and James Mickens in the Harvard Gazette on how much of this is also capability marketing.
- The envelope this arc adds a question to: The Permission Surface and What the Cage Measured.
- The interval the leash arc argued for: The Leash.
- The composition nobody owns, in its original form: The Safe Useless Thing.
- The finding this arc inherits and will not re-argue: The Remedy Went the Forbidden Way.
- Manufactured consensus, previously: Four in the Room.
- The approved agent with the authorized blast radius: the Kiro incident, in The Door at the Desk.
These notes come out of Sociable Systems, a practice that reads AI-shaped documents the way a hostile reviewer will, before a lender or a court finds the gap. The argument has an operational form: the Interim Protocol sets out four rules for AI use in environmental and social deliverables, covering disclosure at touch-point grain, evidence custody, the phrases no automated screening may settle, and a hostile read before anything ships. Free, and written to be cited or retired once institutional guidance arrives.
