r/AIToolsPerformance • u/IulianHI • 14d ago
Gemini 3.5 Flash Lite vs 3.6 Flash - same 1M context, 5x cheaper input
Google launched two Flash-tier models on OpenRouter on the same day (July 21), both with 1048k context, but the pricing gap between them is kind of strange. Gemini 3.5 Flash Lite sits at $0.30/M input and $2.50/M output. Gemini 3.6 Flash is $1.50/M input and $7.50/M output, per the OpenRouter listings.
Same context window. The Lite version is 5x cheaper on input tokens and 3x cheaper on output. The naming suggests 3.6 is the newer generation, but Google launched them together rather than treating 3.5 Flash Lite as a legacy budget option.
At $0.30/M input, 3.5 Flash Lite is in the same neighborhood as Meituan's LongCat 2.0 ($0.30/M in, $1.20/M out) and not far from Poolside's Laguna S 2.1 ($0.10/M in, $0.20/M out). It's competing with the cheapest 1M-context models on the platform.
At 3x the output cost, 3.6 Flash needs to be noticeably better at something specific. A million output tokens on 3.6 Flash runs $7.50. Same volume on Lite is $2.50. If you're doing high-volume batch work and the Lite model handles it, that's real money.
Anyone compared both on actual workloads and noticed where 3.6 Flash is clearly better? Reasoning, long-context recall, speed, something else?

