Anthropic released Claude Opus 5.5, the first model in the Claude 5.5 family. The company says it performs at Fable 5.1 level on most work and costs 40% less to run than Opus 5 at default settings.
OpenAI announced GPT-6 Sol and Luna. Both run at half the price of their predecessors, with no large jump in capability. The company's pitch is price-performance.
xAI released Grok 4.7, its new flagship for coding and agentic work, at the same price as Grok 4.6. The company's chart shows a big jump; an independent index places it mid-pack.
An average success rate does not show an agent is reliable. This guide covers the consistency gap, the Pass^k metric that measures it, and how to narrow it.
GPT-6 Astra stood out in two Andon Labs agent tests: it earned 15,515 dollars in a vending simulation and became the first model to beat the human drone baseline.
In a production LLM application the cost piles up in two separate places. This guide separates prefix caching from semantic caching, and says what to measure.
OpenAI has put the harness that runs Codex into public beta as the Agents API. Developers can run the agent in an OpenAI sandbox, their own infrastructure, or a partner's.
Google shipped its third Flash model in six weeks. The token price matches 3.7 Flash exactly, but the model reasons more, pushing cost per task up 40 percent — and Google itself recommends staying on the older model when efficiency matters.
Anthropic announced its new models with a cache discount. Artificial Analysis, which took part in pre-release testing, says the saving does not hold in every scenario.
The models were trained from scratch on 15 trillion tokens and can toggle thinking on and off. The speech model transcribes three hours of audio in one second.
MiniMax has released its text-to-music model with open weights. It takes lyrics carrying section tags and a detailed music description, and returns a complete song of up to five minutes in a single generation.
DeepSeek has released Harness, the layer between an agent and its environment, under the MIT licence. Models, tools, sessions and the control loop itself all sit behind plugin boundaries and can be swapped without touching the source.
Chinese firm Zhipu AI has released GLM-5.3. The model shares the same base as its predecessor and all gains come from extended post-training alone. The company also reports finding 2,436 vulnerabilities across 269 projects with security teams.
NVIDIA announced day-zero RTX support for Qwen3.8-27B, a 27-billion-parameter open model. It is sized to fit a single GPU and reaches 131 tokens per second on a GeForce RTX 5090.
NVIDIA has expanded its open-weights multilingual text-to-speech model Magpie with Modern Standard Arabic, Korean and Brazilian Portuguese. The company's argument is that running speech generation on your own infrastructure gives the latency budget back.
Ai2's Earth observation platform OlmoEarth Studio has begun computing and exporting numerical vectors from satellite imagery. The source code and model weights are public, so the method can be inspected.
Liquid AI has released LFM2.5-VL-3B, a 3.1-billion-parameter model built for on-device deployment. It reads screen interfaces, grounds objects to coordinates and, for the first time in this line, calls tools.