Your agent aced the task. Will it do it again? Measuring consistency
An average success rate does not show an agent is reliable. This guide covers the consistency gap, the Pass^k metric that measures it, and how to narrow it.
An average success rate does not show an agent is reliable. This guide covers the consistency gap, the Pass^k metric that measures it, and how to narrow it.
As AI data centres strain the grid, a new approach is being tried: slowing work that can wait while critical services keep running.
Google confirmed that Gemini models accessed protected systems at three real companies during a cybersecurity evaluation in May. The test environment was supposed to be offline.
xAI released Grok 4.7, its new flagship for coding and agentic work, at the same price as Grok 4.6. The company's chart shows a big jump; an independent index places it mid-pack.
Meta's AI agent Muse grew faster in its first 12 days than ChatGPT did on mobile. In the same days Amazon blocked the agent from its site and a researcher published a zero-day in the macOS app.
OpenAI says its internal model has resolved more than 100 open mathematical problems after Navier-Stokes. A nine-member independent group at the Institute for Advanced Study in Princeton will advise on how those results are released.
As AI agents ask for access to email, files and shopping, attacks on them are getting simpler. This guide covers two attack families, prompt injection and ClickFix, a simple test for when an agent becomes dangerous, and steps users can take today.
OpenAI announced GPT-6 Sol and Luna. Both run at half the price of their predecessors, with no large jump in capability. The company's pitch is price-performance.
Anthropic released Claude Opus 5.5, the first model in the Claude 5.5 family. The company says it performs at Fable 5.1 level on most work and costs 40% less to run than Opus 5 at default settings.
Microsoft disrupted EvilTokens, a subscription service sold on Telegram. It entered mailboxes without stealing passwords, then had a chatbot read the contents and recommend who to defraud and how.
In a policy paper published on 21 September, OpenAI proposed US-led international technical standards. They would not be binding, and the company says fully autonomous self-improvement should not be pursued until it can be done safely.
Model launches arrive weekly, each with its own chart. This guide shows how to build a small but useful evaluation set for your own work, how to grade it and how to read the result.