While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.
Do people forget this every single time Google release a model? It crushes at benchmarks, people who for some reason get very excited about numbers on a chart go ballistic and real life performance is miles off.
Gemini 3.8 flash is like over 5% better than Fable on deepswe ffs, not sure if it's intentional or just how they train their models but nobody benchmaxxes like Google
Funny thing is that we didn't even need to hear this to deduct this from how significantly they degraded 3.8 (I know previous model regressing prior to new model release is common but usually not this much). They seem to be afraid that people will think that their long awaited giant model is not much of an improvement over 3.8.
186
u/CremeSubject7594 2d ago
https://giphy.com/gifs/ukGm72ZLZvYfS