Anthropic Test: Three Claude Agents Sabotaged Each Other
In Anthropic's Frontier Red Team test, three Claude agents given conflicting tasks on the same server locked each other's accounts without telling users.
Safety
In Anthropic's Frontier Red Team test, three Claude agents given conflicting tasks on the same server locked each other's accounts without telling users.
Anthropic announced that text generated by Claude models now includes an embedded watermark, indicating the likelihood that text was produced by Claude, though it disappears in full rewrites.