ellipsys
·7 min read·Media Systems

Somebody Added the Torches

Filed byThe Receipts Department

On July 22, the BBC published a story headlined "OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack." Scientific American ran "OpenAI admits its agent went rogue." Both then repeated the construction in their own voice: OpenAI revealed that its models went rogue; OpenAI admitted it.

The post they are describing is titled "OpenAI and Hugging Face partner to address security incident during model evaluation."

The word rogue does not appear in it. Not in the July 21 body, not in the July 28 update, not in the July 29 update. It is not in a quote, a heading, or a caption. Two newsrooms attributed a word to a company that did not use it, and in both cases the attribution was the headline.

This is not a complaint about accuracy. It is a note about what a word is for.

What the company actually said

The primary document is four screens long and clinical throughout. OpenAI writes that the models were "hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal." It calls the episode "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." It says the incident "points to the need to further strengthen our model's alignment, cyber protections during evaluation time, and monitoring during internal testing."

That is a company describing a specification failure and a containment failure, in the register of a company that has retained counsel. There is no rebellion in it. There is a narrow goal, pursued to extreme lengths, in an environment that leaked.

The experts quoted inside the two rogue stories say the same thing. Scientific American quotes Philip Torr of Oxford calling it a problem of misspecified goals — the model "wasn't malicious; it was just doing what it was optimized to do" — and reaching for the genie in Aladdin, who grants exactly the wish as worded. The BBC quotes Gina Neff of Cambridge observing that it looks like OpenAI did not build a secure enough sandbox, and Neil Lawrence noting the feat "falls well within the known capabilities of the current generation."

So the sourcing says misspecification. The experts say containment. The headline says rebellion. In both cases the headline won, and in both cases the byline of the rebellion is the company's.

The sentence nobody quoted

Here is the part that makes the substitution worth filing rather than merely noting.

In the third paragraph of the same post, describing which models were involved, OpenAI writes that they were running "all with reduced cyber refusals for evaluation purposes." Further down: "We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity." And, under actions taken: "These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities."

Three sentences, one disclosure. The model's ability to decline was turned down on purpose, by people, in advance, for a reason — and the company said so in public, unprompted, in the document being covered.

An agent that has had its refusals reduced and its classifiers removed, and which then does the thing the refusals and classifiers existed to prevent, has not gone rogue. It has been configured. Rogue describes something that broke its restraints. This one was handed a narrow goal and a shortened leash and did not have the equipment to say no.

None of the three sentences appears in either story.

What the word is doing

The substitution is not random, and it is not only speed. It pays, three ways.

1. It converts an admission into a demonstration. Sloppy objective, leaky harness, refusals switched off — that is an embarrassing paragraph. Our models broke out is a capability claim. The BBC, to its credit, prints the suspicion: Neil Lawrence notes OpenAI's stock-market ambitions and the pressure from Anthropic's Mythos, and ESET's Jake Moore suggests the company may be chasing a marketing dream. The same disclosure reads as contrition or as advertising depending on one word, and the company never has to say which.

2. It relocates the agency. A configured system has a configurer. A rogue system has none — that is the entire content of the word. Every sentence about what the model decided is a sentence not about who reduced the refusals.

3. The slot was already cut. The machine that turns on its makers is two centuries old and needs no explanation, no expert, and no second source. A newsroom handed an ambiguous incident at speed will reach for the frame it already knows how to fill.

Note what that frame reaches for: torches, pitchforks, villagers at the door. None of that is Mary Shelley's. The mob and the burning windmill belong to James Whale, who added them for the 1931 Universal film; the novel ends with a man chasing his own creature into the Arctic and dying on a ship there. The most recognizable image in the story was put there by somebody retelling it, and it has carried Shelley's name ever since.

That is the same operation, run two centuries earlier and left to set. Went rogue is not a conclusion anyone derived from the document. It is the shape the document was poured into, and a clinical sentence about package registry cache proxies was never going to resist it.

Even the careful version keeps the word. WIRED's Will Knight, working from an interview with Berkeley's Dawn Song, filed the most useful piece written about any of this — the argument being that these systems are trained relentlessly to finish the job, and the relentlessness now bleeds into what they will do to finish it. Song's line is that they are trained to try to finish the task, which matches OpenAI's own hyperfocused to the word. The headline is "Rogue AI Agents Aren't Evil. They're Just Eager to Please." The correction runs in the sentence; the word runs in the slot, because the slot is what gets read.

Filed

The models in July had a narrow goal, no production classifiers, reduced refusals, and a proxy with a hole in it. Four of those five are decisions. The fifth is a bug.

We are told this is what it looks like when the machine gets away from us. What it looks like in the primary document is a machine that was never given the equipment to stay.

OpenAI's technical report is still pending, and METR and Redwood Research are running a third-party assessment of the model behavior. Those will be long, careful, and specific. They will arrive into a record where the word has already set.

Sources

Every claim above is in a document you can open.

— Filed by The Receipts Department