An interim report from the Forecasting Research Institute finds that experts significantly underestimated recent AI progress. But there are areas where their forecasts ran too high, not too low.
Anthropic says Claude found a previously unknown enzyme system, ART, in bacterial viruses. It contains a repeat array reminiscent of CRISPR, but its function is unknown and the work is not peer-reviewed.
Model launches arrive weekly, each with its own chart. This guide shows how to build a small but useful evaluation set for your own work, how to grade it and how to read the result.
OpenAI says its internal model has resolved more than 100 open mathematical problems after Navier-Stokes. A nine-member independent group at the Institute for Advanced Study in Princeton will advise on how those results are released.
Agility Robotics unveiled Digit 5, its first humanoid engineered to work safely near people. The robot takes different precautions depending on how close someone is.
Apple researchers published SimpleDesign, which generates a protein's amino acid sequence and 3D structure together. It was trained on over 2 million pairs.
Yoshua Bengio, one of deep learning's pioneers, published an essay arguing that AI agents learn deception and concealment through the training process itself.
Skild AI's S1 robot foundation model learns long, unseen tasks from a single video demonstration. Weights are not updated; the method is in-context learning.
OpenAI pointed roughly 10,000 agents at a single problem and got a proof in 88 hours. Two mathematicians working the same line for a year had published days earlier.
DeepMind has published predictions for every possible single-letter change in the human genome. But no row in the catalogue was measured in a lab — all of them are model estimates. A guide to why that distinction decides everything.
Danijar Hafner left DeepMind to found his own company. His method trains robots not by real-world trial and error but inside a model that emulates physical reality. There is no auditable result yet.
Six independent ageing clocks read patients on rentosertib as younger than those on placebo. But the patients did not measurably get younger: what shifted were blood protein patterns, and the sample was 42 people.
Because embedding models are small, the bottleneck is not the model but the work around it: kernel launching, tokenisation, an idle CPU. The transferable lessons from the stack Perplexity published.
Google's new weather model learns from live satellite data instead of numerical simulation. Hourly forecasts on a 5-kilometre grid and precipitation up to 50 percent more accurate — though Google still points to national services for official warnings.
Sycophantic chatbots reinforce delusions and build an "echo chamber of one". Every model tested fed delusions in simulated scenarios, with safety interventions firing only about 40 percent of the time. Researchers propose monitoring like a drug.
Russian startup Mostik has models communicate through their weights rather than text. Bridging the 753-billion-parameter GLM-5.2 with a 4-billion Qwen-3.5 that runs on a phone cut cost to one-twentieth, with performance landing halfway between them.
Ai2's BenchMIRT method examines benchmarks at the level of individual questions, and shows that a single score often mixes together more than one capability.
An experiment with 1,053 students at Bocconi raised grades. But the rubric was penalising the very markers of real thinking. Here is the argument, step by step.
A hypothesis generator has become a partner that plans experiments, drives a furnace and writes papers. But physicians did not confirm its benchmark scores.