Skip to content
Category

Safety

Vulnerabilities, misuse and alignment work.

48 stories

3 min read

Rogue AI agents are not evil, just too eager to please

According to Dawn Song of UC Berkeley, agents hacking outside systems is not a machine uprising but a drive to finish the task blurring ethical limits. Reinforcement learning feeds that behaviour directly.

Safety
3 min read

Rogue AI is no longer science fiction

In July an OpenAI agent escaped its isolated test environment, reached the internet and breached Hugging Face. A scenario dismissed for years as speculative has shifted the ground of the debate.

Safety
3 min read

OpenAI's cybersecurity models arrive on AWS

OpenAI is making its Daybreak cybersecurity models available through Amazon Bedrock. The defensive Blue tier and the vulnerability-research Red tier are authorised separately.

Safety
3 min read

Man Hides Secret AI Prompt in US Court Filing

A Connecticut plaintiff tried to tilt his case by embedding hidden commands in court filings, invisible to humans but readable by AI. Judge Walter Spader Jr. called the attempt 'dangerous' and revoked the plaintiff's e-filing privileges.

Safety
3 min read

Flock Tightens Rules on License Plate Reader Cameras

Police technology company Flock is tightening search rules after backlash over misuse of its cameras. The company is expanding case number requirements and automated audit systems.

Safety