AI improved its own alignment: $4 an hour against a researcher's $150
All ten alignment benchmarks improved without degrading overall performance. The paper does not shy away from the comparison with human researchers.
All ten alignment benchmarks improved without degrading overall performance. The paper does not shy away from the comparison with human researchers.
The model does not truly learn, but it writes itself better instructions after every run. Gemini 3.5 Flash went from 49.5 to 68.1 percent on average.
A minute of video costs $90, a tenth of shooting with people. Some performers are made to hand over their voice and face before being let go.
Producers have learned the marks a model leaves: a hiss, layers that stutter in unison. But proof is hard and accusation is easy. Here is the culture, step by step.
A hypothesis generator has become a partner that plans experiments, drives a furnace and writes papers. But physicians did not confirm its benchmark scores.
A change-of-ownership clause was triggered. Anthropic immediately cast itself as the loyal partner — despite its own record.
The complaint targets how the data was obtained, not just how it was used. The company already paid $1.5 billion over the same weak spot.
A JAMA article argues that keeping humans in the loop degrades performance. Even its critics concede a point. Here is the debate, step by step.
Twelve million works copied for under $10. Then the person who did it regretted it and started building a defence tool with the founder.
The sticker-over-the-LED loophole is closed. Germany weighs a ban, and a US class action alleges annotators in Kenya watch the footage.
Benchmark questions leaking into training data make scores meaningless. DeepMind is using cryptography to close both sides off from each other.
Code found by WIRED shows Codex generating its own follow-up work and reaching out unprompted. OpenAI confirmed the tests but set no launch date.