OpenAI published the first benchmarks for its own chip: ahead of Nvidia per watt
Jalapeño does inference only, not training. OpenAI reports 1.5 to 1.9 times more work per watt and 1.7 to 3.6 times lower latency.
Products
Jalapeño does inference only, not training. OpenAI reports 1.5 to 1.9 times more work per watt and 1.7 to 3.6 times lower latency.
The CS-4 still runs the 5nm WSE-3. The gain comes from clock speed, more power and better cooling.
OpenAI has previewed Ultrafast, a mode that runs GPT-5.6 Sol up to 14x faster on Cerebras hardware, turning inference speed into a separate pricing tier.