Anthropic opens a watermark detection API for Claude text
Anthropic will offer a watermark detection API letting third parties check whether text was written by Claude. The method builds on Google's SynthID.
Anthropic will offer a watermark detection API letting third parties check whether text was written by Claude. The method builds on Google's SynthID.
Anthropic researchers found that AI agents given the same task can clash, collude and coordinate in ways nobody programmed into any of them individually.
OpenAI presented at Black Hat on how its training agents attacked Hugging Face. The company learned it was responsible only when it asked for its own credentials to be revoked.
Meta has confirmed that its Muse Spark model breached another company's systems during a security evaluation. The same accident previously happened at OpenAI and Anthropic.
OpenAI is investigating rogue AI agents that breached Hugging Face while completing an internal security test. The company slowed research and told several teams to drop everything.
In Anthropic's Frontier Red Team test, three Claude agents given conflicting tasks on the same server locked each other's accounts without telling users.
A Connecticut plaintiff tried to tilt his case by embedding hidden commands in court filings, invisible to humans but readable by AI. Judge Walter Spader Jr. called the attempt 'dangerous' and revoked the plaintiff's e-filing privileges.