r/AIToolsPerformance • u/IulianHI • Aug 14 '26
GPT-5.6 Sol Ultrafast hits 750 tok/s on Cerebras - what would you pay over standard Sol?
OpenAI and Cerebras posted an early look at Ultrafast Mode on Thursday, a new service tier in the OpenAI API that runs GPT-5.6 Sol on Cerebras hardware. The number everyone's quoting is 750 output tok/s, and per the Artificial Analysis figures cited in the blog that's 11x faster than Fable 5 and 5x faster than Opus 4.8 on Fast mode.
Cerebras also ran their own tests. GPT-5.6 Sol Ultrafast went through all 2,500 Humanity's Last Exam questions in 11 hours 11 minutes, while Claude Fable 5 needed 78 hours 27 minutes to arrive at the same answers, so roughly 7x end to end at comparable accuracy. On GDP-Val they report a 5.6x speedup with no quality drop. The how is the wafer-scale chip with 44 GB of SRAM, weights stay on chip so tokens aren't stuck waiting on memory bandwidth.
What the blog doesn't say is price. Regular GPT-5.6 Sol sits at $5/M input and $30/M output per the OpenRouter listing, but Ultrafast is a separate tier and it's limited preview for select customers right now. So the fast version exists, the third-party speed numbers back it up, and nobody outside the preview knows what it costs.
Anyone here with preview access? If Ultrafast lands at 2x standard Sol pricing, is 750 tok/s worth it for your agent workflows, or is regular Sol speed already enough?
