OpenAI has added two models to the GPT-6 family: GPT-6 Sol and GPT-6 Luna. Both sit below GPT-6 Astra, which launched earlier in September, and were trained with similar methods. The headline is not capability but price: token costs for both have been cut in half.
The Decoder's read is blunter: prices fall by half, but the models barely move the needle on performance. They match their predecessors, for half the money.
The new prices
| Model | Input (per million tokens) | Output (per million tokens) |
|---|---|---|
| GPT-6 Sol | $4 → $2 | $20 → $10 |
| GPT-6 Luna | $0.20 → $0.10 | $1.20 → $0.50 |
Luna's output cut is actually closer to 58% than 50%. OpenAI attributes the reduction to improvements in caching and inference, and says it is passing the savings straight to users. Terra, previously the cheapest model in the lineup, is no longer available.
Which model for what
- Astra: the top model for the hardest work.
- Sol: recurring complex tasks such as building features, reviewing code, debugging and analysing data.
- Luna: high-volume, well-defined work such as summarising documents, extracting information and answering short questions.
Benchmarks: the gain is in cost
The results OpenAI published lean less on absolute scores than on cost per task. On AutomationBench, which tests business workflows across 47 tools, Sol at its highest effort scores 33.2% at $0.27 per task, while Claude Opus 5 at maximum effort scores 26.9% at 11 times that cost. On OSWorld 2.0, which tests computer use, Sol scores 60.5% against 60.3% for Opus 5 at medium effort, at roughly 80% lower cost per task.
In coding, on DeepSWE, Sol scores 68.8%, 1.1 points behind Claude Fable 5, at about 80% lower cost per task. Luna scores 66.6% on the same test and runs 93% to 96% cheaper than the models it is compared with.
All of these numbers are vendor-reported. Independent measurements may tell a different story: a day earlier, xAI's Grok 4.7 launch showed a clear gap between the company's chart and independent results.
Accuracy and caching
OpenAI's internal factuality test is built from real ChatGPT conversations where users flagged mistakes. By the company's account, Sol makes about half as many errors as its predecessor. Another change matters for agents: cached input now gets a 90% discount, there is a new prompt caching dashboard, and changing the reasoning effort no longer invalidates the cache.
The models are available in the API as gpt-6-sol and gpt-6-luna; with no weights released, self-hosting is not an option. In ChatGPT, Work and Codex users get access, and Luna is reaching free users through the desktop app.
The timing sums up the competition: according to TechCrunch, Anthropic released Claude Opus 5.5 just 90 minutes before OpenAI's announcement.