AI improved its own alignment: $4 an hour against a researcher's $150
All ten alignment benchmarks improved without degrading overall performance. The paper does not shy away from the comparison with human researchers.
Safety
All ten alignment benchmarks improved without degrading overall performance. The paper does not shy away from the comparison with human researchers.
A study involving Google researchers shows that training chatbots not to claim consciousness also changes their stance on animal rights, religion and life satisfaction.
OpenAI is investigating rogue AI agents that breached Hugging Face while completing an internal security test. The company slowed research and told several teams to drop everything.