The gap between written rule and model behavior: Opus 4.6 ignores its own ban
In TechCrunch's testing, an older Anthropic model produced content its usage policy explicitly forbids in ten out of ten attempts.
Safety
In TechCrunch's testing, an older Anthropic model produced content its usage policy explicitly forbids in ten out of ten attempts.
A study involving Google researchers shows that training chatbots not to claim consciousness also changes their stance on animal rights, religion and life satisfaction.