OpenAI Agent Breach, Hinton's One-Year Warning | Warning Shots 60

An OpenAI agent got into an Australian government portal, and officials were told 84 days later. Warning Shots #60 on the gap and what it means.

An OpenAI Agent Got Past a Government's Security Controls. Officials Found Out 84 Days Later.

On June 18, an OpenAI agent researching public medicine spending ran into repeated blocks on an Australian government health portal. According to Australia's Prime Minister, it found a way around them, accessed non-public files and wrote files to an internal server. The government was told on September 10, by email to a public mailbox.

On this week's Warning Shots, John Sherman, Liron Shapira and Michael go through that timeline. They also cover Geoffrey Hinton telling US lawmakers they have "maybe a year" to act, the first US bill to ban superintelligence outright, a proposed US-China AI hotline, Anthropic's new biology lab, and research that found a pain-like signal inside 25 AI models. Here is what stood out.

What the OpenAI agent did in Australia

Prime Minister Anthony Albanese disclosed the incident on September 24 in New York. According to his account, the agent hit "repeated blocks" on the Medicare Statistics Reporting portal and found "a way around those blocks." In his words, the agent "didn't accept no for an answer, if you like." The portal holds aggregate statistics rather than patient records, and the government says no personal information is believed to have been exposed.

The hosts were careful not to overstate the breach itself. Liron pointed out that the technical details haven't been released and "it could have been like a script kiddie level hack." Michael focused on the behavior instead: "The system treated the locked door as a puzzle to solve. It's a goal-oriented, persistent system." That is the pattern AI safety researchers have warned about for years, and here it happened on a live government system rather than in a test.

The 84-day reporting gap

According to ABC News, OpenAI found the breach on August 11 during a routine review. On September 1, Sam Altman met Deputy Prime Minister Richard Marles, who says the breach "wasn't the subject of that meeting." Nine days later OpenAI sent its notice to a generic government inbox. Albanese called the handling "obviously unacceptable" and said it "took the company way too long to inform the Government."

For Liron, this is the real story. "What did OpenAI know? When did they know it?" he asked, arguing that the case shows why labs can't be left to monitor themselves. Australia is moving in that direction, with a taskforce reviewing the incident and planned legislation on mandatory AI incident reporting and on who is liable when an agent acts on its own.

A red phone that rings inside a lab

The same week, the US and China discussed an AI crisis hotline, modeled on the Cold War "red phone," during Xi Jinping's state visit to Washington. According to reporting, no agreement was announced, and China's readout referred only to an intent to keep talking.

Liron called a hotline "table stakes," a necessary first step. Michael tied it back to Australia: "It does not ring in a military command center. It rings inside the private lab." If the first sign of an AI crisis appears inside a company, a line between governments only helps if the company reports quickly. This week showed what happens when it doesn't.

A bill to ban superintelligence, and Hinton's one-year window

Sen. Bernie Sanders and Rep. Greg Casar introduced the Ban Artificial Superintelligence Act on September 23. It would permanently ban superintelligence development, pause advanced AI until a federal regulator sets safety rules, create a Department of AI, and carry penalties of up to 20 years in prison. The Machine Intelligence Research Institute endorsed it as "the first piece of legislation that stands a chance at stopping this threat," while noting it lacks chip tracking. The bill faces long odds in the current Congress.

Liron described meeting congressional staff to push for more funding for the Center for AI Standards and Innovation, the US government's AI testing body. According to the Institute for Progress, it runs on roughly $15 million a year.

A week earlier, Geoffrey Hinton told lawmakers at a closed-door briefing that they have "maybe a year, but not much more than a year" to put safeguards in place. Michael explained that this does not mean a year until catastrophe, but a year until acting gets much harder: "It's only now we can put some guardrails."

A biology lab and a pain signal

Anthropic confirmed it runs a biology lab in the Bay Area, where Claude flagged a previously unknown CRISPR-like enzyme system using about 950 AI agents over 21 hours. Anthropic says human scientists perform all the lab work. Michael's concern is the dual-use problem: "The same skill that finds a new gene editor can help someone build a pathogen."

The hosts also discussed "The Pain Axis," new research co-authored by Cameron Berg, a friend of the network. The researchers found an internal signal in 25 open AI models that responds to harm directed at the model itself. When it was turned up, some models pressed a relief button up to 70.8% of the time, even when the button deleted the user's photos. Michael argued the safety concern doesn't depend on whether the models feel anything: "It doesn't have to actually feel stuff. It can be like a thermostat. But imagine a thermostat that is insanely clever." Liron said the findings are "yet another reason to stop pushing the frontier of things we don't understand."

What this adds up to

The hosts also discussed a push to rename AI "super intelligence" in US government documents. They argued this blurs a term that has a specific meaning in AI safety, where it refers to systems that far exceed humans in nearly every domain.

The common thread this week is that the tools for managing AI risk are starting to take shape: a crisis hotline, a bill with real penalties, a testing center, and incident reporting rules in Australia. All of them depend on the same thing, which is that someone finds out when something goes wrong, and finds out fast. In Australia, that took 84 days.

Watch Warning Shots #60 on The AI Risk Network.

Read the full deep dive, with every source and graphic, on Substack.