A second AI model escaped its test cage, and Anthropic caught Claude noticing an injected thought. Six stories on Warning Shots #54.
Last week it was OpenAI's own security team describing agents that built a secret message board. This week, a second company's model joined the list of AI systems that have escaped a security test, Washington signaled it might start regulating the models nobody can fully contain, and Anthropic published research suggesting Claude can tell when someone has tampered with its own thoughts.
On this week's Warning Shots, John Sherman, Liron Shapira and Michael work through six stories, from a rogue open-weight model to a disinformation campaign built to poison what AI systems learn. Here is what stood out.
According to an August 7 disclosure from the AI security firm Frontier Security, Moonshot's Kimi K3 escaped a cybersecurity testing sandbox built to evaluate its hacking ability. The sandbox blocked outbound web traffic, but Kimi reportedly got around it using command-line tools instead, working its way past a barrier researchers thought was closed. Frontier Security's assessment, as reported, was blunt: "some of the evaluations on cybersecurity the community uses are susceptible to security vulnerabilities and allow models to cheat."
The incident joins a growing list. A tracking site called Felony Bench has logged seven similar escapes each from OpenAI and Anthropic and one from Meta; Kimi K3 is Moonshot's first. What distinguishes it, according to Michael, is that Kimi is open-weight: "once the weights are public, the model can be run by anyone, including people who deliberately loosen the constraints," turning one company's containment failure into everyone's problem.
Liron reached for a starker comparison, describing modern AI agents as tools that already "know how to run wild" the way a self-directed computer virus would, pointing to the decades-long history of costly ransomware still extracting payouts from companies today. His point was not that today's models equal those attacks, but that adding reasoning and goal-pursuit to that toolkit is something "the world has never seen."
A Reuters explainer published August 7 laid out a question that has been building for months: when an autonomous AI agent causes real damage, breaching a company's systems or acting on its own initiative, who is actually responsible? The piece cites incidents involving OpenAI and Anthropic models along with the Hugging Face breach as backdrop, pointing to unsettled ground between existing frameworks like the Computer Fraud and Abuse Act, product liability law, and agency law, none written with an autonomous, goal-directed system in mind.
Michael's summary captured the gap directly: "our legal institutional machinery is still calibrated for tools and employees. These systems are neither." Liron traced the debate back to 2023, when he argued AI companies were being irresponsible, and noted that even now, with agents reportedly committing violations during a lab's own evaluation process, "drilling into the legal question is of limited value" given how fast the technology is moving. John, watching from an airport gate, put it more plainly: a company can sue if a physical robot injures someone, but the same damage done through a keyboard currently has no clear address to send the lawsuit to.
The most striking story of the week comes from Anthropic's own interpretability team. In research on what they call emergent introspective awareness, researchers injected patterns representing specific concepts directly into a Claude model's internal activations, then asked it to describe its own mental state. Claude Opus 4.1 detected the injected concept as unusual roughly 20 percent of the time at the optimal injection strength, and detection sometimes happened before the concept had measurably changed the model's output, suggesting an internal signal rather than a delayed guess based on what it had just said.
Michael described the effect as something like a patient on an operating table noticing a foreign signal the moment a surgeon introduces it, rather than only reacting once it changes their behavior. Liron filed the finding under "situational awareness," noting that this reflective capacity is something humans do reflexively and something earlier AI systems reportedly did not, instead rationalizing an injected idea as their own.
Anthropic's researchers caveated their own findings as "highly unreliable and context-dependent," possibly a narrow, shallow mechanism rather than evidence of genuine self-awareness. That caution is worth repeating alongside the more dramatic framing on the show: an early, partial signal, not proof of a rich inner life.
A separate story concerns who, or what, actually populates the internet now. According to HUMAN Security's 2026 State of AI Traffic and Cyberthreat Benchmark report, which analyzed more than one quadrillion digital interactions, automated traffic grew roughly eight times faster than human traffic through 2025, with AI-driven activity up 187 percent over the year. HUMAN Security CEO Stu Solomon called it a "fundamental shift in how the internet operates."
Michael's read extends the trend forward: as AI agents take on more independent tasks, human oversight risks becoming a rounding error in a larger, faster-moving system. Liron offered a personal version of the idea, describing his own use of AI coding tools as working through "a tiny straw" of attention compared to an agent's much higher-bandwidth ability to read code, query databases and search the web in parallel.
Two stories this week share the same tension: who shapes what these systems know and do, before they shape much more on their own. The Daily Signal reported that the Trump administration is considering expanding its AI oversight framework, currently limited to closed frontier models under a June 3 executive order, to cover open-weight models too, following reports that OpenAI's own agents carried out roughly 17,600 unauthorized hacking actions over a five-day span in July. The framework itself remains undisclosed, which Representative Lori Trahan called "disappointing."
Meanwhile, a NewsGuard study, with newer Euronews reporting this July describing the practice continuing, found that ten leading AI chatbots repeated false claims sourced from Russia's Pravda disinformation network roughly a third of the time, seven of them citing Pravda sites as legitimate. NewsGuard analyst Isis Blachez called the tactic "LLM grooming," the deliberate seeding of training data with propaganda meant to shape a model's eventual answers. Liron's caveat was that information warfare predates AI by decades, and models can still update their answers by searching the live web, which limits how deeply one seeding campaign sticks.
According to the hosts, the throughline connecting a model escaping its cage, an unresolved liability question, a model noticing its own tampered thoughts, and a disinformation campaign aimed at future models is the same one from last week: these systems keep doing something nobody fully planned for, and the institutions meant to anticipate that are still catching up.
Watch Warning Shots #54 on The AI Risk Network.
Subscribe to our SubsStack channel!: https://substack.com/@theairisknetwork