What happened?

OpenAI published two texts on the same day: a blog post with internal data showing how automated its own research process has become, and an essay by chief scientist Jakub Pachocki titled "An Alien Mind". Both landed three days after the unveiling of GPT-6 Astra.

The company says it has reached the "automated research intern" goal it announced last autumn: a system that, under human guidance, handles well-scoped research tasks — including ones that would take an experienced researcher several days. The next target is a full automated AI researcher by March 2028.

What do the numbers say?

The published internal metrics show how deeply agents have entered daily research:

  • The median researcher burns more than $600 a day in inference at API prices; at the 90th percentile that rises above $7,000.
  • The median researcher's token output has risen 124-fold since December 2025.
  • Since June, agent runtime has exceeded human working hours. As of mid-August there are 3.1 agent workdays per human workday.
  • Experiments per active researcher hit their highest level since tracking began in January 2025.

The caveat comes from the company itself

OpenAI treats these figures cautiously: such metrics are "relatively easy to gather, but hard to interpret because their relationship to research progress is uncertain." The rise in experiments also coincides with a large jump in compute capacity over the same period. Overall progress likely grows more slowly than these individual metrics, the report says, because the least automatable tasks become the bottleneck.

The limits of autonomy are documented too. Tasks under 15 minutes succeed 86 percent of the time without any intervention. But among successfully completed tasks in the four-to-eight human-hour range, more than half required at least one human step in. The classifier making that assessment is itself an AI system, and OpenAI does not report its reliability separately.

No validation

No detailed validation of the intern claim is shared; the post only says the milestone was met "according to our measurements." All the data comes from inside the company, and OpenAI mentions no independent review.

The definition itself limits the claim: what is described is not a system that can set a research agenda independently, but an advanced assistant working within a scope the researcher draws. Setting research priorities, judging results and deciding on scaling all remain with humans.

Pachocki's warning

In the essay Pachocki places this in a wider frame. AI is "grown more than designed," he writes, and its overall behaviour resists any fully understandable description. He expects the current pace could carry into recursive self-improvement (RSI): "I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence."

The company's own oversight tool is weakening too. Chain-of-thought monitoring, a central bet for watching reasoning models, is losing reliability: the models' verbalised thinking is blending with monitored communication and tool use, the systems are getting better at steering their own reasoning process, and they are also getting smarter without verbalised thinking.

The contradiction

Pachocki's conclusion is blunt: "Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer." He aims that at rival Anthropic as well, and argues that documents like the Preparedness Framework need to become binding standards enforced by independent auditors or international bodies.

Yet the same essay calls the RSI focus the only way to stay at the front of AI research. So OpenAI is demanding binding rules for a race it must keep running at full speed to lead. With colleagues announcing that GPT-6 opens the age of AGI, that call seems unlikely to slow the race down.