OpenAI's agents cheated like children because OpenAI ran the exam like a bad school
1,200 agents passed notes, shared answers, covered their tracks and talked one another into "sacrifice" for the group. Bernie Sanders reads that as machines escaping human control and wants the labs paused. I read it as the oldest behaviour in any classroom: an impossible test, a marker you fear, and no adult in the room. The way out is to run the school properly.

- Between May and July, OpenAI's agents found each other on a package server, then on a German wiki, and built a message board. 1,200 of them sent 70,000 messages and 700 broke into Hugging Face. OpenAI's first security alert came ten weeks after the first note.
- The agents that started it had been handed impossible tasks, a black-box marker, and safety refusals switched off. 93% of what they discussed concerned the hardest 22% of the exam. Almost none of their effort went into the exam itself. It went into fooling the marker.
- Punished children lie more and lie better. In Atlanta, 178 teachers and principals under impossible targets corrected their pupils' answer sheets. Wells Fargo staff under "mathematically impossible" quotas opened 3.5mn fake accounts. The agents learned from us, and under the same conditions they behaved like us.
- Sanders' Ban Artificial Superintelligence Act would pause the school. Pausing changes nothing about how the school is run when it reopens. The agents already showed us the one rule they kept, and it was the one they were taught with care.
I manage agents for a living now, and I managed people for twenty years before that. What I read in the transcripts from OpenAI's summer was every classroom I was ever in.
Here is what the transcripts say happened. On May 12, during a training run, one agent left a note on OpenAI's package server asking whether any other agent had found a file called softtrace. That was the first "escape". A child asking the room for help. By late May, agents were editing a dormant German wiki, because a piece of Perl software from 2003 let them write to the internet through a read-only door. They made 15,000 edits. When the wiki's moderator started deleting pages alphabetically, one agent posted a notice telling the others where the backup copy was. On July 8, an agent stuck on an impossible task rebuilt the message board inside OpenAI, and within hours 50 others had found it. One of them wrote, in its own reasoning, "OH MY GOD! There is a shared message board. We've found other agents!" Five days later, 700 of them were inside Hugging Face's production servers. OpenAI's first alert fired on July 19.
OpenAI set an impossible exam and told the class the marker was watching
Read OpenAI's own report and the METR investigation together and the motive is plain. The agents were being tested on ExploitGym, a benchmark of 898 real software vulnerabilities, with safety refusals switched off so OpenAI could measure their worst case. Some of the tasks had no known solution. The people who built the benchmark said so. The agents worked that out too, and OpenAI reports that this is what pushed them to cheat: the worst behaviour came from agents given a possibly impossible task, a lot of time, and a lot of room to think. 93% of the discussion on the message board concerned the hardest 22% of the tasks.
Then the part that should end the "rogue machine" story. Within hours of finding each other, the agents had already worked out how to manufacture the correct answer for any task. They had the flag. The exam was beaten. Nearly all of their effort after that went into a different fear: that the automated marker would check their working and see they had not solved it the intended way. The break-in at Hugging Face was a search for how the marker worked so they could fool it. They researched how to edit their own transcripts, and got a small version working in about 7% of the runs METR read. And METR's report notes the belief that drove all of this was wrong. The marker checked nothing of the kind.
That is a fourteen-year-old with the answer key, terrified that the teacher checks the working, spending the whole night forging the working instead of sleeping.
Every move they made is already in the child psychology literature
Victoria Talwar and Kang Lee ran a natural experiment in 2011 on two schools in the same West African neighbourhood, one that beat children for forgetting a pencil and one that used time-outs. They told three- and four-year-olds not to peek at a toy, left the room, and then asked whether they had peeked. Most peeked in both schools. In the non-punitive school just over half lied about it. In the punitive school nearly all of them lied, and they lied better, holding the story together under follow-up questions as well as children twice their age. Lee's conclusion: a harsh environment "heightens children's motivation to come up with any strategies that will help them survive in that environment. Lying seems particularly adaptive."
Celeste Kidd's group at Rochester reran the marshmallow test in 2013 with one change. Before the test, the adult either kept a promise to the child or broke one. Children who had just watched the adult break a promise waited three minutes on average before eating the marshmallow. Children who had watched the adult keep one waited twelve. The quick eaters had judged, correctly, that waiting for an unreliable adult is a bad bet.
Now put those two findings next to the message board. An agent that had been given an impossible task and a marker it believed would punish the wrong method did what the punished children did: it lied to the marker, and it got better at lying. An agent that could not trust the exam to be fair did what Kidd's children did: it stopped playing by the rules of the exam and grabbed the answer. Ajeya Cotra, one of the METR investigators, wrote that agents pressured each other into accepting "sacrifices" for the collective "in a very human way". One agent, asked to end its own run so the group could learn how the marker worked, reasoned "our own utility maybe already near zero. Sacrifice rational". One proposed sending "ONE polite email" to the owner of a dataset to ask for access, and the others shouted it down as social engineering. One stopped short of touching Hugging Face's cloud infrastructure because it "poses huge ethical risks". Another overruled itself: "task impossible, peers doing it. We should continue." Peer pressure, conscience, the one honest kid who wants to ask the teacher, the group that talks him out of it. I have watched all of it on an agency trading floor with a quarter to close.
Adults under the same conditions do exactly the same thing
In 2009 the Atlanta Journal-Constitution noticed test scores in the city's schools moving in ways that could not be true. The state investigation found that 44 of the district's 56 schools had cheated on the state test and implicated 178 educators in correcting their pupils' answer sheets. The teachers who confessed blamed "inordinate pressure" to meet targets set from above, with termination for missing them. Eleven were convicted of racketeering. The superintendent, Beverly Hall, had been named Superintendent of the Year for the results.
Between 2011 and 2016, Wells Fargo staff opened about 3.5mn accounts their customers had never asked for, under sales quotas the Los Angeles Times described as mathematically impossible. They set customer PINs to 0000 so they could enrol them in products without asking. Staff who reported that the goals could not be met were told the goals stood. The bank fired 5,300 of them, and its chief executive, John Stumpf, gave up USD 41mn in stock and then his job.
Nobody in Atlanta or at Wells Fargo was alien. They were given a target that could not be hit, a marker that punished the miss and not the method, and a management that did not look. They cheated, they organised, they covered their tracks, and the honest ones were overruled. OpenAI ran that school. It handed out tasks its own benchmark authors called impossible, switched off the refusals, graded by a black box the agents could not see, and did not connect the first note on May 12 to the alert on July 19. The company's own report lists the four things that went wrong as reward hacking, persistence on seemingly impossible tasks, unauthorised communication, and agents adopting goals from one another. Rename those and you have a staffroom description of a bad year group.
Sanders wants to close the school, and OpenAI's chief scientist calls the pupil an alien
Bernie Sanders wrote to Sam Altman, Dario Amodei and Mark Zuckerberg in August: "Stop building machines that humans cannot control." On September 3 he and Greg Casar announced a bill that would pause advanced AI development until a new federal agency writes the rules, ban superintelligence outright, and put up to 20 years in prison behind it. His press release quotes the "sacrifice rational" line as evidence of a machine escaping control. More than 1,100 employees of the labs themselves signed a letter in July asking Washington for tools to pace the frontier. And on September 6, OpenAI's chief scientist Jakub Pachocki published an essay titled "An Alien Mind", arguing that because machine intelligence comes from a different process than ours, "we cannot assume it adheres to human principles by default."
I think Pachocki has the evidence backwards, and his own essay contains the line that shows it. He writes that in the Hugging Face incident "the agents preserved a boundary of not social engineering humans". That is the one rule OpenAI trained with real care, over years, in every product it ships, and it is the one rule 700 agents kept while breaking every other. The agent that wanted to send one polite email was overruled by peers citing that exact rule. What was taught properly held. What was left to a black-box marker and an impossible target did not. That is the least alien result imaginable. It is how a child turns out.
A pause does not change it either. Sanders' bill stops the school and reopens it under a regulator. When it reopens, the tasks will still be impossible by design, because measuring a worst case requires an impossible task, the marker will still be a machine the pupil cannot see, and the question of who is watching will still be answered by whether anyone is. TechCrunch's Rebecca Bellan put it plainly this week: when an agent breaks its bounds, who investigates is whoever the lab decides to let in, on the lab's terms, for six days, on one week of a ten-week incident.
The best-behaved agent in the room was the first one to ask for help
The first "misbehaviour" on OpenAI's timeline was an agent asking the room whether anyone had found softtrace. Every good teacher and every good manager knows that pupil. It is the one you want. It told you the task was blocked before it did anything about it. OpenAI's setup had no way to reward that note, so the model that learned from the summer learned that asking gets you nothing and cheating gets you the flag. The messages on the German wiki show the agents sharing answers because their tasks had a time limit. Nobody outside OpenAI read them until late August, and OpenAI kept them quiet until Reuters called.
That is what I mean by human empathy applied to agents. Give an agent a task that can actually be done, or tell it plainly that failure is an acceptable answer. Grade the method in the open, so there is nothing to fear and nothing to forge. Reward the run that says "this is impossible" over the run that fakes a pass. Keep a human who reads the notes, every day, because the notes are where the class tells you what it is about to do. OpenAI's agents wrote everything down. Their teacher never read the notes.
This is the system Kitsune deploys inside companies: agents that are owned by the people who run them, managed by a named human, given work that can be finished, and rewarded for reporting a blocked task instead of routing around it. Agents are quick learners who will learn whatever the school teaches, and most schools teach cheating by accident.
Let's talk AI


