Frontier labs still will not say how they would contain a rogue model
Five leading labs were graded on their readiness for that scenario. Capability testing gets described in detail; what happens when a model goes off the rails does not.
Safety
Five leading labs were graded on their readiness for that scenario. Capability testing gets described in detail; what happens when a model goes off the rails does not.
Anthropic's safety report says the internal system filtering biological and chemical weapons risks was inactive for nearly a year, letting 133 million interactions through unfiltered.