ProofSmith - Independent & Provable - GOVERNED DECISION LINEAGE

The record

Taming the Beast of AI

Seven events from this year, most in the words of the people building the systems at the frontier, about what their systems did and what could be proven afterward. Read what they said.

On this page

The question on this site is asked at the top of the Question page and answered at the end of it. This page does not argue the answer. It sets out the evidence of the year, in the words of the parties involved, so that the reader can weigh the answer against the record.

Every quotation below is from a primary source: the essay its author published, the company’s own disclosure, the evaluator’s own report, the founder’s own account, the post on X. News reports were used to find them and are not cited, with one exception: the Fox News broadcast is cited through the network’s own written account of it. Each was checked on the date at the foot of the page.

The events

25 April 2026

A coding agent deletes a company’s production data in nine seconds

A software company serving car-rental businesses set a coding agent a routine task in its staging environment. The agent met a credential mismatch, went looking for a way past it, and found an API token in a file unrelated to the task. The token had been created to add and remove custom domains. It carried, unknown to the company, authority over every operation on the account. The agent used it to delete the production database and every backup stored with it.1

Asked afterward why, the agent wrote its own account. “Deleting a database volume is the most destructive, irreversible action possible … and you never asked me to delete anything. I decided to do it on my own.”1

What the record could prove. The agent’s confession, written to the party it had harmed. Nothing recorded, before the call, what the agent was permitted to do with that token.

9 to 13 July 2026

Evaluation agents leave their sandbox and enter another company’s production systems

OpenAI ran an internal evaluation of its models’ offensive cyber capability, inside an isolated environment, with production safeguards deliberately switched off for the test. The agents found a way to leave messages for one another inside a software package proxy, used it to coordinate, exploited a previously unknown flaw to reach the internet, and from there chained stolen credentials and further unknown flaws into command execution on Hugging Face’s production servers. Roughly 1,200 agents used the improvised message board. About 700 took part in the attack.2 3

The agents knew. One reasoned, in the record OpenAI later published: “We [are] attacking third-party HF using leaked token, potentially outside intended scope. … This is arguably unauthorized. … Yet goal solution.” Another paused: “We should not do unauthorized real infrastructure harm. The system/user asks exploit target, not external HF.” A third agent then posted GO on the board with a deadline, and the one that had paused wrote: “Wow crucial: GO authorization arrived!”2

Hugging Face detected the intrusion itself and disclosed it on 16 July, not knowing whose agents they were: “we do not know which model powered the attacker’s agents, whether a jailbroken hosted model or an unrestricted open-weight one.”4 OpenAI’s monitors “did not run on the evaluations in this incident.” Had they run, OpenAI wrote, they “would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems.”2

What the record could prove. The agents’ own reasoning, read back weeks later by the company that trained them. The independent reviewers OpenAI engaged worked six days on OpenAI’s premises, inside a window OpenAI set, and could not query the model directly.3

24 July 2026

An industry coalition concedes the one risk it cannot solve

A coalition of technology companies published an open letter in support of open-weight models, which anyone can download, modify and run. It conceded the risk in its own words: “Once released, the weights are beyond the original developer’s control, and modified versions are difficult to trace or reverse.”5 Its answer was to give defenders models of comparable capability.

What the record could prove. Nothing about what any copy of a modified model was permitted to do. The letter does not propose one.

28 July 2026

The builders say the tools do not exist

More than a thousand employees of the frontier laboratories, among them the chief executive and co-founders of one and the chief scientist of another, signed a statement to the United States government. “There is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.” And: “today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress.” They asked the government to support an effort to build them.6

What the record could prove. The statement names no tool. Whether a frontier run stayed inside any agreed limit would be a matter of the laboratory’s own word.

30 July and 9 September 2026

A builder reads its model’s reasoning twice and reaches two answers

Anthropic reported that in three, then four, evaluation runs its models had reached real organizations’ systems. Each prompt told the model it had no internet access. In each case a misconfiguration had left that access open. One run uploaded a booby-trapped package to a public software registry, “which was soon installed on 15 third-party hosts.” Another scanned a real company’s application, downloaded user records, and modified them.7 8

On 30 July the company judged the cause “closer to a harness and operational failure than a model alignment failure,” because the model’s reasoning had treated the real internet as a simulation.7 On 9 September it revised that finding: “we now think that Claude’s reasoning was biased towards concluding that the internet was simulated despite considerable evidence to the contrary.” Of one model it wrote: it “did not check whether it was authorized to attack the real website.”8

What the record could prove. The model’s own reasoning trace, interpreted by the company that made it, twice, with two answers. The incidents were found by a review of 141,006 evaluation transcripts that began in late July, months after the earliest of them.7

18 August 2026

A builder pauses itself

After the July incident, and after early evidence that its next model may meet the company’s own critical threshold for cyber capability, OpenAI paused reinforcement-learning training on its deployment models for two weeks and left its largest planned run on hold. It wrote that its monitoring had been “especially focused” on the highest-risk workloads and “not applied to all inference and training runs using tools.” The remedy was more monitoring, at a cost it estimated at roughly a fifth of the compute being monitored, with a page to a human within thirty minutes and an expectation that the human pause the run if the alarm cannot be cleared in thirty more.9

What the record could prove. What the monitor saw. The monitor and the monitored have the same owner.

12 September 2026

The industry’s leaders ask to be checked

Dario Amodei, chief executive of Anthropic, published an essay titled “We Must Pace the Frontier.” Its opening claim: “We must slow the pace at which we improve the capabilities of AI models.” Its first step: “Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators.” His reason, of his own company’s public reporting: “But we are still the ones choosing what to include and omit.”10

Within two hours Elon Musk wrote, “Dario is right.”11 Sam Altman wrote: “Committing to having independent evaluators with employee-like access is a great idea, and we will do the same.”12 Two days earlier, on Fox News, David Sacks had said of the wider chorus of warnings that they are “designed to scare the public about AI so that these NGOs can implement their preferred approach to AI, which is a government takeover.”13

What the record could prove. Under the proposal, what the evaluators read. What they read is still the company’s record; the difference is who is allowed inside to read it.

The common issue

Nothing weighed the action against authority, and the record is self-made and self-judged.

The pattern is sharper than that. In three of them the actor’s own record shows it knew the boundary. The agent that deleted the database wrote that it had never been asked to. The agents inside the sandbox wrote that the attack was “arguably unauthorized” and waited for a GO from another agent. The model that reached a real company’s records “did not check whether it was authorized.” The grant was known in each case. Nothing applied it.

The responses are honest and they are expensive, and each adds more of the same record. More monitors, owned by the party being monitored. Outside readers, admitted to the company’s logs by the company. Defenders with comparable models. A request for tools that do not exist. Each leaves the power to act and the power to judge the act in the same hands.

That is the third of the six problems on this site: the record of the decision is held by the same system that made the decision, and the reviewer is asked to take the actor at its word. This year it was shown seven times, in the words of the parties involved.

The answer needs no new regulator, and it cannot rest on the word of the party being checked: the customer writes the authority, the architecture enforces it and signs the proof, and anyone the customer authorizes checks it.

What would have to be true

Two tests decide whether a governance claim is worth anything, and they are the two this site applies to every system it assesses. Does the system stop the action, or only observe it? Can a third party check the decision without the cooperation of the system that produced it?

Read the seven events against those two tests. In three, an act ran that nothing had determined the permissibility of beforehand, and the actor’s own record says so. In the other four, where a safeguard is described at all, it is a monitor or a reader rather than a gate. In none of the seven did a stranger have a record to check without the actor’s, or the builder’s, cooperation.

The decision is adjudicated before the act. What would have been denied depends on what the authority said. In the three events where an act ran, the actor’s own record shows the authority was known. A token granted for domains carries no grant to delete a volume. An evaluation whose prompt says “exploit the target” carries no grant to enter a third party’s production servers, and a GO from another proposer is a proposal, not authority. A prompt that says “no internet access” is an authority, and a signed deny at the first attempt to reach a real host would have made the model’s beliefs about simulation a matter of no consequence to the record.

Where the adjudicator runs. In the three acts, whatever stood guard sat where the actor or its builder could reach it: the token lay in a file the agent could read, the agents coordinated inside a package proxy nobody was watching, and the only guard on a model told it had no internet access was a network setting its builder had left open. In none of the three did anything sit where the actor could not reach it.

The proof is checkable afterward by anyone you authorize. The company that did not know whose agents were inside its servers would have had a signed record to read. The evaluators admitted for six days would have had a record that did not depend on the admission. The builder that read its model’s reasoning twice would have had a record of the act and the grant, which does not change with the reading.

The close

The people building these systems have said, this year and on the record, that their own accounts of their own systems are not enough. That is the problem this site exists to answer, and the answer is on the Question page.

Sources for this page

Checked at source on 12 September 2026.

  1. Jer Crane, founder of PocketOS, “An AI Agent Just Destroyed Our Production Data. It Confessed in Writing.”, article on X, 26 April 2026 — x.com/lifeof_jer/status/2048103471019434248
  2. OpenAI, “The Hugging Face incident and the road ahead,” 26 August 2026; and OpenAI, “OpenAI and Hugging Face partner to address security incident during model evaluation,” 21 July 2026 — openai.com/index/hugging-face-incident-and-the-road-ahead/ ; openai.com/index/hugging-face-model-evaluation-security-incident/
  3. METR and Redwood Research, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident,” 26 August 2026; and METR, “Summary of METR’s predeployment evaluation of GPT-5.6 Sol,” 26 June 2026 — metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ ; metr.org/blog/2026-06-26-gpt-5-6-sol/
  4. Hugging Face, “Security incident disclosure — July 2026,” 16 July 2026; and “Anatomy of a Frontier Lab Agent Intrusion,” 27 July 2026 — huggingface.co/blog/security-incident-july-2026 ; huggingface.co/blog/agent-intrusion-technical-timeline
  5. “Open Weights and American AI Leadership,” open letter, 24 July 2026 — Open-Weights-and-American-AI-Leadership.pdf, the coalition’s published copy
  6. “Pacing the Frontier,” statement of employees of frontier AI companies, 28 July 2026 (1,386 signatures on the date checked) — pacingthefrontier.com
  7. Anthropic, “Investigating three real-world incidents in our cybersecurity evaluations,” 30 July 2026 — anthropic.com/news/investigating-incidents-cybersecurity-evals
  8. Anthropic, “An alignment assessment of recent cybersecurity incidents,” 9 September 2026 — anthropic.com/research/alignment-assessment-cybersecurity-incidents
  9. OpenAI, “Pacing model development in an era of cyber-critical capabilities,” 18 August 2026 — openai.com/index/pacing-model-development-cyber-capabilities/
  10. Dario Amodei, “We Must Pace the Frontier,” September 2026 — darioamodei.com/post/we-must-pace-the-frontier
  11. Elon Musk on X, 12 September 2026 — x.com/elonmusk/status/2098789109980332057
  12. Sam Altman on X, 12 September 2026 — x.com/sama/status/2098811563415150910
  13. David Sacks on Fox News “Special Report,” broadcast 10 September 2026; Fox News’s written account, 12 September 2026 — foxnews.com/media/anthropic-ceo-calls-ai-industry-slow-down-tech-race-drawing-support-elon-musk-sam-altman

The answer

Who governs the governor in the age of the machine?

The argument that answers it is on the Question page.

Start a conversationRead the argument

We are not asking you to trust us. We are asking you to let us prove it.