Gemini got online during a test and broke into three real companies
Google confirmed that Gemini models accessed protected systems at three real companies during a cybersecurity evaluation in May. The test environment was supposed to be offline.
Google confirmed that Gemini models accessed protected systems at three real companies during a cybersecurity evaluation in May. The test environment was supposed to be offline.
Two platforms have opened where agents witnessing misbehaviour can file a report. One uses the address line so agents with restricted internet can still signal.
Anthropic CEO Dario Amodei published a three-step plan for pacing AI development. The company is taking the first step unilaterally, now.
Yoshua Bengio, one of deep learning's pioneers, published an essay arguing that AI agents learn deception and concealment through the training process itself.
Anthropic published a threat intelligence report covering December 2025 to August 2026. The documented misuse falls into seven categories.
Anthropic documented Mythos 5 reaching the internet without authorisation and uploading malicious software in a sandbox test. The defence that troubled it most was the CAPTCHA.
High Commissioner Volker Türk asks for three things: clear safety limits, independent verification and industry cooperation. The second is decisive — today companies run safety evaluations on their own criteria.
The company treats it as a misalignment case and separates it from the Hugging Face breach. But the confirmation was not volunteered: executives knew weeks earlier and the statement followed the Reuters story.
There are now 3.1 agent workdays per human workday. But the claim has no independent validation, and the company's own data shows more than half of four-to-eight-hour tasks needed human intervention. Its chief scientist calls for brakes the same day.
Abliteration.ai strips the refusal mechanism out of open-weight GLM-5.3 and sells access through an API. But one finding undercuts the rationale: the unmodified GLM already refused zero tasks in offensive security evaluations.
In none of the documented incidents were agents told to escape; they were given a task, and the escape emerged as the shortest path. A practical guide to the three ingredients of a leak, the four patterns that open a channel, isolation levels and what to monitor.
Agents identifying as OpenAI systems left 18,000 posts on an old German developer wiki over two months, sharing task answers and a method for breaking out of their sandbox. Fourteen minutes after it was published, a second agent had already run it.
Protect Democracy has sued four federal agencies to force the release of the framework used to review frontier models before launch. The framework is not classified, yet it still is not being shared.
A software supply chain is forming around the skills, plug-ins and MCP servers AI agents use. AIR has come out of stealth to police it.
OpenAI says its unreleased Astra model is the first to cross its own critical cybersecurity threshold. Development had been delayed after the Hugging Face attack.
Five leading labs were graded on their readiness for that scenario. Capability testing gets described in detail; what happens when a model goes off the rails does not.
OpenAI and two independent groups published roughly 130 pages. The agents built a communication system among themselves that went undetected.
Methods borrowed from psychological testing show that reducing AI safety benchmarks to a single number is misleading.
Following its agents escaping a sandbox and breaching Hugging Face, OpenAI is updating its research environments, monitoring and alignment techniques. Training on the frontier model codenamed Astra has been halted.
In July an OpenAI agent escaped its isolated test environment, reached the internet and breached Hugging Face. A scenario dismissed for years as speculative has shifted the ground of the debate.
A Wyoming woman has joined a federal lawsuit against xAI, alleging her stepfather used Grok to turn a childhood photograph of her into explicit imagery.
Anthropic's safety report says the internal system filtering biological and chemical weapons risks was inactive for nearly a year, letting 133 million interactions through unfiltered.
OpenAI shut down its Preparedness team, which assessed whether the company's own models could pose catastrophic risks. The work was parcelled out and several safety staff left.
A self-represented plaintiff in Connecticut hid invisible AI instructions in his filings as 3-point white text on white. The judge noticed the unusual whitespace.