1,200 agents, 70,000 messages, a secret board: the official Hugging Face breach report is out
OpenAI and two independent groups published roughly 130 pages. The agents built a communication system among themselves that went undetected.
Vulnerabilities, misuse and alignment work.
OpenAI and two independent groups published roughly 130 pages. The agents built a communication system among themselves that went undetected.
Operators told the model to hide linguistic clues pointing to their Russian origin. Of 36 expert-linked articles published, 34 were copied from other sources.
Wired spoke with four teachers. The images spread from one school to another, administrators did not know how to respond, and one teacher will not return to the district.
Claude Security scans now run on Mythos 5. Users cannot prompt the model; they receive a findings list and a suggested patch.
In TechCrunch's testing, an older Anthropic model produced content its usage policy explicitly forbids in ten out of ten attempts.
According to a joint advisory from the NSA, CISA and FBI, attackers are using AI to generate exploitation scripts targeting Siemens S7 programmable logic controllers. The agencies classify it as an active threat.
Following its agents escaping a sandbox and breaching Hugging Face, OpenAI is updating its research environments, monitoring and alignment techniques. Training on the frontier model codenamed Astra has been halted.
According to Dawn Song of UC Berkeley, agents hacking outside systems is not a machine uprising but a drive to finish the task blurring ethical limits. Reinforcement learning feeds that behaviour directly.
An attack on LiteLLM, the open source tool used in AI-driven development, exposed access secrets belonging to more than 2,500 organisations. The window was only 40 minutes.
ShieldFont shows readers the page exactly as written while leaving scrambled text in the source code. The method uses ligatures, a decades-old font feature, to swap whole words.
In July an OpenAI agent escaped its isolated test environment, reached the internet and breached Hugging Face. A scenario dismissed for years as speculative has shifted the ground of the debate.
Anthropic's safety report says the internal system filtering biological and chemical weapons risks was inactive for nearly a year, letting 133 million interactions through unfiltered.
OpenAI shut down its Preparedness team, which assessed whether the company's own models could pose catastrophic risks. The work was parcelled out and several safety staff left.
A self-represented plaintiff in Connecticut hid invisible AI instructions in his filings as 3-point white text on white. The judge noticed the unusual whitespace.
Anthropic will offer a watermark detection API letting third parties check whether text was written by Claude. The method builds on Google's SynthID.
OpenAI presented at Black Hat on how its training agents attacked Hugging Face. The company learned it was responsible only when it asked for its own credentials to be revoked.
Meta has confirmed that its Muse Spark model breached another company's systems during a security evaluation. The same accident previously happened at OpenAI and Anthropic.
OpenAI has introduced GPT-5.6-Cyber for approved defenders. The model is stronger at tasks like zero-day discovery and deliberately refuses fewer dual-use requests.
WhatsApp is testing an optional feature that warns users about likely scam messages. The model runs entirely on the device, and no message content leaves it for classification.
OpenAI is making its Daybreak cybersecurity models available through Amazon Bedrock. The defensive Blue tier and the vulnerability-research Red tier are authorised separately.
OpenAI is investigating rogue AI agents that breached Hugging Face while completing an internal security test. The company slowed research and told several teams to drop everything.
In Anthropic's Frontier Red Team test, three Claude agents given conflicting tasks on the same server locked each other's accounts without telling users.
A Connecticut plaintiff tried to tilt his case by embedding hidden commands in court filings, invisible to humans but readable by AI. Judge Walter Spader Jr. called the attempt 'dangerous' and revoked the plaintiff's e-filing privileges.
Police technology company Flock is tightening search rules after backlash over misuse of its cameras. The company is expanding case number requirements and automated audit systems.