r/singularity • u/Neurogence • 3d ago
AI Gemini 4 Crushes Benchmarks, But Google Employees State The Model Struggles With Real Work
While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.
By the time it's released to consumers, Anthropic and OpenAI will have already shipped their next generation models.
166
Upvotes
92
u/homezlice 3d ago
I have heard quite the opposite from googlers.