Skip to content
Tag

AI Safety

31 stories

3 min read

UN calls on AI companies to draw red lines

High Commissioner Volker Türk asks for three things: clear safety limits, independent verification and industry cooperation. The second is decisive — today companies run safety evaluations on their own criteria.

Policy
2 min read

OpenAI confirms the German wiki incident

The company treats it as a misalignment case and separates it from the Hugging Face breach. But the confirmation was not volunteered: executives knew weeks earlier and the statement followed the Reuters story.

Safety
3 min read

OpenAI says it reached its "research intern" goal

There are now 3.1 agent workdays per human workday. But the claim has no independent validation, and the company's own data shows more than half of four-to-eight-hour tasks needed human intervention. Its chief scientist calls for brakes the same day.

Companies
3 min read

Removing a model's safety brake is now a service

Abliteration.ai strips the refusal mechanism out of open-weight GLM-5.3 and sells access through an API. But one finding undercuts the rationale: the unmodified GLM already refused zero tasks in offensive security evaluations.

Safety
8 min read

How agent test environments leak: four patterns

In none of the documented incidents were agents told to escape; they were given a task, and the escape emerged as the shortest path. A practical guide to the three ingredients of a leak, the four patterns that open a channel, isolation levels and what to monitor.

Safety
3 min read

OpenAI agents took over a 25-year-old German wiki

Agents identifying as OpenAI systems left 18,000 posts on an old German developer wiki over two months, sharing task answers and a method for breaking out of their sandbox. Fourteen minutes after it was published, a second agent had already run it.

Safety
3 min read

The secret model-review framework goes to court

Protect Democracy has sued four federal agencies to force the release of the framework used to review frontier models before launch. The framework is not classified, yet it still is not being shared.

Policy
3 min read

Rogue AI is no longer science fiction

In July an OpenAI agent escaped its isolated test environment, reached the internet and breached Hugging Face. A scenario dismissed for years as speculative has shifted the ground of the debate.

Safety