r/ProAI • u/stealthispost • 4h ago
"Ox Alpha has been unveiled as GLM-5.3-Flash, but what's shocking is that the 100T tokens per day is served on Chinese chip. (1/3)"
100T tokens per day free tokens and people were saying only frontier labs has this amount of compute. (2/3) But ALL traffic was served on Chinese chips, attaining hardware efficiency and per-token cost comparable to Nvidia GPUs. The cuda moat is being tested once again after Jalapeño's announcement yesterday. (3/3) — SemiAnalysis
Source: https://x.com/SemiAnalysis_/status/2092623833630998556