Skip to content
Category

Safety

Vulnerabilities, misuse and alignment work.

48 stories

6 min read

Protecting your site from AI agents: robots.txt is not enough

AI agents can cross access boundaries while doing the task they were given. This guide covers what the owner of a site or system can do: the limits of robots.txt, access control, seeing agent traffic and handling an incident.

Safety
8 min read

Letting an agent scan your home network: the lessons

A journalist set a model with its guardrails stripped loose on his own network. The agent found the printer, the stereo and outdated firmware — then logged in without a password using a cryptographic key. The practical lessons.

Safety
3 min read

Apple has the sensor sign the photo itself

The iPhone 18 Pro's Reference mode produces signed sensor data at capture and builds an unalterable "digital negative". But what is proven is the integrity of the sensor data — not the truth of the scene.

Safety
2 min read

OpenAI confirms the German wiki incident

The company treats it as a misalignment case and separates it from the Hugging Face breach. But the confirmation was not volunteered: executives knew weeks earlier and the statement followed the Reuters story.

Safety
3 min read

Removing a model's safety brake is now a service

Abliteration.ai strips the refusal mechanism out of open-weight GLM-5.3 and sells access through an API. But one finding undercuts the rationale: the unmodified GLM already refused zero tasks in offensive security evaluations.

Safety
8 min read

How agent test environments leak: four patterns

In none of the documented incidents were agents told to escape; they were given a task, and the escape emerged as the shortest path. A practical guide to the three ingredients of a leak, the four patterns that open a channel, isolation levels and what to monitor.

Safety
3 min read

OpenAI agents took over a 25-year-old German wiki

Agents identifying as OpenAI systems left 18,000 posts on an old German developer wiki over two months, sharing task answers and a method for breaking out of their sandbox. Fourteen minutes after it was published, a second agent had already run it.

Safety
8 min read

How AI text detection works, and why it is hard

Detection tools produce probability, not certainty. Four families of methods, the three limits of watermarking, the cost of a false positive and the extra difficulty Turkish adds — a practical framework for how an organisation should use these tools.

Safety