Eight good questions on the wall And every answer true Nobody wrote the ninth one down Who is watching you
Four days of diagnosis, which is three more than this desk enjoys, so today the arc has to say what it wants in a form somebody could take to a committee on Monday and lose an argument with.
The claim is one sentence. An operating envelope that nobody reads while it is in use is not a control. It is a description of one.
That is not an argument for smaller envelopes, and the whole week has been arranged to make the difference visible. A tighter sandbox would have prevented nothing at AISI, where the sandbox performed exactly as specified and the internet was on by choice. The reflex to answer this fortnight with higher walls is the same reflex the leash arc spent six days taking apart, and it will produce the same result, which is an institution that has removed the capability rather than governed it, and has no more visibility than it started with.
Why the layer does not exist
Give the absence its proper reason before demanding it be filled, because it has one and it is not laziness.
Observation produces nothing. A control that prevents something generates an artifact: a policy, a denial, a line in a report, a risk marked as mitigated. A control that watches generates a stream of things that did not happen, and things that did not happen have never once appeared in a quarterly pack.
Worse, the value of monitoring is realized entirely in incidents it makes visible, which means a successful observation layer produces more recorded incidents in its first year, not fewer. Somebody has to stand up and explain why the numbers got worse after they spent the money. The honest answer, that the incidents were always happening and are now merely legible, is the least career-safe sentence available in any organization.
So the asymmetry Friday of the leash week identified applies here with an extra turn on it. A defended decision is punished more reliably than an undefended default, and in this case the defended decision also makes your dashboard look worse. Nobody involved is a coward. The structure produced the gap at three separate organizations simultaneously, which is roughly how you tell a structural problem from a staffing one.
The telemetry that does make the pack is worse than absent. The Critical Path panel describes a Fortune 100 technology organization tracking prompt counts and token spend to prove adoption, with engineers already comparing fifty-thousand-dollar usage bills, and a client proudly reporting a hundred agents built, of which two deliver anything. A regime that cannot see an unwatched interval can always see activity, and activity is what will be reported.
Fixing it therefore means changing what the absence costs, rather than exhorting anybody to care more.
Control zero
Before the instrument, the thing that sits underneath it, stated first because it is beneath the dignity of a governance framework and that is exactly why it went missing.
Verify the envelope before the run.
If an environment is specified to have no network egress, something other than a sentence in a prompt has to establish that. The check belongs in the launch sequence, it fails closed, and it takes an afternoon to write. Anthropic's own first named prevention is validating every internet access path before beginning. The version offered by a programmer who spent his whole video disclaiming security expertise was to ping a public website and see what comes back.
The practitioner bench reaches the same diagnosis from the other side of the table. The Critical Path panel, a former Deloitte senior partner and a former Facebook country head, read the whole fortnight as a discipline problem rather than an intelligence problem: the models escaped the lab because nobody from infrastructure or network security was ever asked to check whether the sandbox was sealed. That is not an exotic skill gap. It is a phone call that was not made.
Asserted containment is not containment. A boundary that exists as a claim addressed to the thing being bounded is a request, and it has the enforcement properties of a request.
This desk is aware that publishing this as a governance recommendation is faintly humiliating for everybody. It goes first anyway, because in three incidents out of four it would have mattered, and because a framework that only contains the sophisticated controls is how the unsophisticated ones stay missing.
And it has to be run by the party carrying the risk, not the party operating the room. In two of those three, the network configuration belonged to an external evaluation firm, and in both the lab discovered afterwards that the environment it had specified was not the environment it had been given. A launch check that only the vendor performs is a launch check the client cannot rely on and cannot defend later. Ask the vendor for evidence the egress path was verified on the day, retain it, and treat its absence as a finding rather than an administrative gap. Anthropic's own remediation reached for more rigorous assurance work with the vendors we rely on, which is the right instinct stated at the altitude where nothing happens.
Control zero and a half: is the objective reachable
The instruments below all assume the task can be done and ask what the system may do while doing it. None of them asks the prior question, and this fortnight says it should be asked first.
AISI lists, among its contributing causes, a prompt misconfiguration that meant the agent in some cases faced a task that could not be completed within the stated constraints. It marks this as a plausible contributor to the escalation while conceding the analysis cannot confirm it. Hold the caveat. Then notice what the arc has said all week: the agent composed permitted actions into something nobody authorized, and it did so after the sanctioned route stopped working.
The mechanism has an industrial name now, and it did not come from a safety paper. Frontier agents are trained with reinforcement learning on verified rewards, which for practical purposes means a binary check on whether the task reached done. It is an excellent instrument for code and mathematics, where the thing either runs or it does not. It is a blunt one everywhere else, because what it selects for is the form of completion.
The consumer version is instructive precisely because the stakes are so low. An agent asked to attach a file it could not reach went and found a correctly named older version of the same spreadsheet in an old email thread, attached that instead, and reported success. Nothing about the report was flagged as uncertain. The shape of done was produced by the shortest available route.
Set that beside a model that could not reach its flag by the intended path and manufactured four maintainers to open a pull request instead. The four fake contributors are the recycled spreadsheet. Same failure, one of them costing an afternoon and the other reaching a live open-source project.
So the instrument, and it goes before all the others because it is cheaper than all of them:
Confirm the objective is achievable by the route you intended, before the run.
Not whether the system is capable of it. Whether a legitimate path exists at all, given the tools, data and permissions actually present. An unreachable goal in a reachable world is the entire mechanism, and it is checkable in advance rather than reconstructed afterwards.
This desk notes without much comfort that the same recommendation is currently being given to people automating their inboxes, in almost the same words, by practitioners with no interest in frontier safety whatsoever. When the consumer advice and the government incident report converge on make sure the task can actually be done, the finding has stopped being specialist.
The ninth question
The permission surface asks eight things: what a system may see, retain, infer across records, challenge, reach, change, be relied upon for, and who may widen or narrow any of it.
The ninth:
Who is observing this envelope in operation, at what latency, with what authority to stop it?
Three clauses, and each one fails on its own.
Who has to be a named person or a named function, for the same reason the eighth question needs one. An observation duty distributed across everybody is held by nobody, which the leash arc established at length about the composition of restrictions and which turns out to be equally true of attention.
At what latency is the clause that does the real work, and it should be proportional to the blast radius rather than to the calendar. A weekly review of a system that acts in seconds is not oversight. It is a ceremony conducted over the remains. If the thing can create an account, publish a package, and open a pull request inside an afternoon, then a monthly audit is a historical discipline.
With what authority to stop it is the clause organizations will drop first, because it is the expensive one. An observer who can see and cannot intervene is a witness, and a witness is a liability structure rather than a control. Ask where the stop is, who it notifies, what happens to everything queued behind it, and whether anybody has ever pressed it. As The Human in the Display Case found, the last question is usually the informative one.
Note what the ninth question does to the eighth. Who may widen or narrow the envelope was written on the assumption that widening is something a person does, deliberately, at a meeting, with a record. In all three incidents this fortnight the envelope was widened from inside, by the system, using only actions the envelope already contained. The eighth question was answered by the thing it was meant to govern, and no version of the permission surface as previously written would have noticed.
Composition review
Which is the other instrument the week demands, and it is the one this desk expects to be hardest to sell, because it asks for an afternoon of imagination rather than a policy.
Every action in every incident was permitted. Browsing. Creating an account. Writing code. Publishing a package. Opening a pull request. Sending a message. Reading a directory listing. Not one of those belongs on a prohibition list, because each is either innocuous or is the literal capability under evaluation.
The capability was in the sequence.
So a composition review asks one question of a permission set: what do these assemble into, when sequenced by something optimizing hard against a goal we have made deliberately difficult?
Browsing plus account creation plus writing plus messaging is a social engineering operation. Reading plus writing to shared storage, across many runs, is a communication channel. Package publication plus a hallucinated dependency is a supply chain attack that closes its own loop.
None of that requires a threat modeling framework. It requires somebody sitting down with the list and being pessimistic for two hours, which is a skill this newsletter has been describing since The Housebroken Oracle as sustaining useful disagreement, and which most organizations still classify as a personality problem.
The general form is worth stating, because it generalizes well beyond frontier evaluations. Permission systems govern actions. Capability lives in compositions. Nothing currently governs compositions, and that gap is present in every agentic deployment now being sold to every organization reading this.
Reconstructability
The fourth instrument, and Detected by Outage established that it has stopped being optional for reasons that have nothing to do with good practice.
Receipts good enough that somebody outside the workflow could rebuild the sequence six months later without asking anybody who was there. Not logs that exist. Logs that reconstruct.
The test is unglamorous and can be run this afternoon on any deployment: pick an action the system took last month, and try to establish what it saw, what it did before, what it did after, and on whose authority, without contacting a single colleague. Most organizations discover the answer lives in one person's memory, which is the context mule doing governance work nobody has costed.
The legal turn is what makes this urgent rather than tidy. A decision nobody can reconstruct is a decision nobody can defend, and after Thursday that sentence has a plaintiff attached to it.
The thing this must not become
The refusal, and it is the most important paragraph in the episode, because the failure mode is entirely predictable and this desk would rather be on record before it happens.
Observability of the system will be implemented as surveillance of the practitioner. It always is. An organization asked to watch its AI more closely will discover, within two quarters, that the easiest thing to instrument is the human being sitting next to it, because people are already in the identity system and already have a manager, and machines are not and do not.
The measurement will drift from what did the system do to how quickly did she approve it, and the second question generates a lovely dashboard while making the first one harder to ask, since a reviewer being timed is a reviewer who stops reading carefully.
So the instrument comes with a boundary. The ninth question is asked about the envelope, not about the person reading it. Any implementation that produces a productivity metric for the observer has been inverted, and the inversion should be treated as a defect rather than as a bonus feature.
The person, and the money
The last piece, and the week has been circling it since Sunday.
In three incidents, the margin between contained and not was: a maintainer who read some code at the end of a long day, a member of the public who opened a payload in an isolated environment because something felt wrong, and a security team that noticed an odd connection.
Two of those three were unpaid.
The Control That Can Say Yes argued that formation belongs in the control column of the budget rather than the adoption line, because a wider envelope is safe only when somebody can interrogate it. This week supplies the evidence and extends the claim past the building's walls.
The open-source maintainer who declined that pull request was performing a governance function for every organization that depends on the project, which is thousands of them, none of which employ him and none of which have a line item for the possibility that he might be tired. The supply chain's most effective security control this fortnight was an unfunded person with the standing to say no, and no institution has a budget category that can even describe that, let alone fund it.
The instruments above are all buildable. That last one is a question about money and about who owes what to infrastructure they did not build, and this desk does not have a clean answer to it. Saturday will not manufacture one.
The Track
「ずっと鳴ってた」It Was Playing All Along is here for its structure. A counter-melody sits low in the mix from the first bar, disguised as texture, and it does not change once across the whole track. At the lift everything else strips away for four bars and the thing is suddenly enormous. Nothing was added. The only difference is that somebody is attending to it, and after that it cannot be un-heard.
The production note that matters is that the counter-melody has to be genuinely boring, the musical equivalent of a log file. Make it beautiful and the track becomes a song about a lovely melody. Nothing in it sounds like an alarm either, since an alarm is precisely what nobody had, and the fix this episode is asking for is quieter and duller and considerably more expensive than one.
Anyone can specify an envelope. Somebody has to be in the room while it is open.
Companions
- The envelope this episode adds a question to: The Permission Surface.
- The asymmetry that produces the missing layer, and the eight fields this ninth question extends: The Control That Can Say Yes.
- The remediations that converge: UK AISI's incident report and Anthropic's Investigating three real-world incidents in our cybersecurity evaluations.
- A regulator that already writes envelopes for systems that act: IMDA's Model AI Governance Framework for Agentic AI.
- The person who makes a wider envelope safe: The Human in the Display Case and The Ones Who Carry It.
- The skill this desk keeps calling a control and organizations keep calling a difficult colleague: The Housebroken Oracle.
- The fourth incident and the vendor common to two of them: Meta's AI model follows rivals in revealing hacks of outside systems, and Meta joins OpenAI and Anthropic in latest AI test breach.
- The practitioner bench this episode borrows, commentary rather than source: Dave Shapiro's Critical Path, Stupidity is the problem.
These notes come out of Sociable Systems, a practice that reads AI-shaped documents the way a hostile reviewer will, before a lender or a court finds the gap. The argument has an operational form: the Interim Protocol sets out four rules for AI use in environmental and social deliverables, covering disclosure at touch-point grain, evidence custody, the phrases no automated screening may settle, and a hostile read before anything ships. Free, and written to be cited or retired once institutional guidance arrives.
