Skip to content
Category

Research

Papers, technical findings, lab output.

69 stories

8 min read

A map of 9 billion variants: what is AlphaGenome Atlas?

DeepMind has published predictions for every possible single-letter change in the human genome. But no row in the catalogue was measured in a lab — all of them are model estimates. A guide to why that distinction decides everything.

Research
3 min read

A stealth startup training robots in world models

Danijar Hafner left DeepMind to found his own company. His method trains robots not by real-world trial and error but inside a model that emulates physical reality. There is no auditable result yet.

Research
3 min read

AI-designed drug turned back the biological clock

Six independent ageing clocks read patients on rentosertib as younger than those on placebo. But the patients did not measurably get younger: what shifted were blood protein patterns, and the sample was 42 people.

Research
7 min read

Where an embedding stack bottlenecks: Perplexity's case

Because embedding models are small, the bottleneck is not the model but the work around it: kernel launching, tokenisation, an idle CPU. The transferable lessons from the stack Perplexity published.

Research
2 min read

WeatherNext 3 drops physics simulations

Google's new weather model learns from live satellite data instead of numerical simulation. Hourly forecasts on a 5-kilometre grid and precipitation up to 50 percent more accurate — though Google still points to national services for official warnings.

Research
3 min read

Psychiatry has to decide on "AI psychosis"

Sycophantic chatbots reinforce delusions and build an "echo chamber of one". Every model tested fed delusions in simulated scenarios, with safety interventions firing only about 40 percent of the time. Researchers propose monitoring like a drug.

Research
3 min read

A bridge that lets models talk without words

Russian startup Mostik has models communicate through their weights rather than text. Bridging the 753-billion-parameter GLM-5.2 with a 4-billion Qwen-3.5 that runs on a phone cut cost to one-twentieth, with performance landing halfway between them.

Research
8 min read

What are benchmark scores actually measuring?

Ai2's BenchMIRT method examines benchmarks at the level of individual questions, and shows that a single score often mixes together more than one capability.

Research