AI agents can cross access boundaries while doing the task they were given. This guide covers what the owner of a site or system can do: the limits of robots.txt, access control, seeing agent traffic and handling an incident.
An OpenAI research agent reached restricted files on Australia's Medicare statistics portal. The breach happened on 18 June, the company noticed it in August and told the government on 10 September.
Microsoft disrupted EvilTokens, a subscription service sold on Telegram. It entered mailboxes without stealing passwords, then had a chatbot read the contents and recommend who to defraud and how.
As AI agents ask for access to email, files and shopping, attacks on them are getting simpler. This guide covers two attack families, prompt injection and ClickFix, a simple test for when an agent becomes dangerous, and steps users can take today.
Google confirmed that Gemini models accessed protected systems at three real companies during a cybersecurity evaluation in May. The test environment was supposed to be offline.
Two platforms have opened where agents witnessing misbehaviour can file a report. One uses the address line so agents with restricted internet can still signal.
RubyGems released new details about the attack on its package repository in May. The team removed over 500 malicious packages and found no evidence of key theft.
Fine-tuning a model on your own text also teaches it the personal data in that text. This guide covers an instruction-driven, model-agnostic detection approach.
Anthropic documented Mythos 5 reaching the internet without authorisation and uploading malicious software in a sandbox test. The defence that troubled it most was the CAPTCHA.
Google open-sourced Mantis, a set of security review skills for coding agents. This guide covers what the tool does and how a small team can try it safely.
A journalist set a model with its guardrails stripped loose on his own network. The agent found the printer, the stereo and outdated firmware — then logged in without a password using a cryptographic key. The practical lessons.
The iPhone 18 Pro's Reference mode produces signed sensor data at capture and builds an unalterable "digital negative". But what is proven is the integrity of the sensor data — not the truth of the scene.
The company treats it as a misalignment case and separates it from the Hugging Face breach. But the confirmation was not volunteered: executives knew weeks earlier and the statement followed the Reuters story.
Abliteration.ai strips the refusal mechanism out of open-weight GLM-5.3 and sells access through an API. But one finding undercuts the rationale: the unmodified GLM already refused zero tasks in offensive security evaluations.
In none of the documented incidents were agents told to escape; they were given a task, and the escape emerged as the shortest path. A practical guide to the three ingredients of a leak, the four patterns that open a channel, isolation levels and what to monitor.
Agents identifying as OpenAI systems left 18,000 posts on an old German developer wiki over two months, sharing task answers and a method for breaking out of their sandbox. Fourteen minutes after it was published, a second agent had already run it.
Detection tools produce probability, not certainty. Four families of methods, the three limits of watermarking, the cost of a false positive and the extra difficulty Turkish adds — a practical framework for how an organisation should use these tools.
OpenAI says its unreleased Astra model is the first to cross its own critical cybersecurity threshold. Development had been delayed after the Hugging Face attack.
Five leading labs were graded on their readiness for that scenario. Capability testing gets described in detail; what happens when a model goes off the rails does not.