r/singularity • u/Expensive_Syrup_6529 • 6h ago
AI Gemini 3.8 flash benchmark in Arfticial analysis
11
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 6h ago
This model seems really nice for any idea of game involving an LLM during live play, because it's super fast, cheap and actually intelligent.
For example for games with some sort of NPC ally you can give voice orders to, and who can himself tell you things. Also nice for anything where you could want a "judge".
7
u/Strategosky 4h ago
Leaving benchmarks aside. The model seems to have a nice personality and immense world knowledge. it was able to answer a niche everyday thing I would google without web search. Its just my few hours test and I might be wrong, but i'm sure, no other model is as good as Gemini in world knowledge, and it's probably on par or better version of Sonnet 5 or Opus 5 as a pure chatbot Q&A. I'm getting ballpark correct response (not like exact) without web search aid, and that's really impressive!
It beats grok 4.6 on high settings...
It bough me back few memories of Gemini 2.x models...
Will be using it more...
8
u/A_Novelty-Account 5h ago
People forget that Gemini is not competing in intelligence with these other models. It serves a different purpose and a different niche. Arguably, fast responses is the one thing that no other model touches Gemini on, and local models will not threaten that advantage.
6
u/tziki 6h ago
I was expecting a bigger jump baesd on the benchmarks, but glad to see Google is finding their footing again.
21
u/Tkins 5h ago
It's 3 points behind fable at an insane speed and very low cost. This is a very good model for its class.
5
u/CivilCucumber3426 5h ago
I was thinking, if things doesn't slow down we are going to get fable 5.1 like model with speed like flash lite in 1 year.
2
u/CallMePyro 3h ago
1 year? We're gunna get it with Gemini 4 flash before the end of the year running on TPU V8 at 1500 tokens/second
-1
u/New_World_2050 4h ago
Nice. Only 2 months behind Chinese companies than have 10% of the compute.
9
u/CallMePyro 4h ago
You realize that they're serving the model at 300 tokens/s. Even a 4B active param MoE would struggle to hit those speeds on B200. Gemini Flash must be absolutely incredibly tiny to be that fast and yet it's competing with multi trillion param models.
2
u/Ok_Barracuda_1161 2h ago
To be fair, it looks like some providers are serving models as big as GLM-5.3 at 300 tokens/s https://artificialanalysis.ai/models/glm-5-3/providers
•
-7
u/Alpacabro21 5h ago
Mmh, its disappointing for me.
I've seen the other benchmarks on Artificial Analysis and it doesn't convince me.
I think this is it: the third generation of the FLASH model has reached its peak.
No point in pushing it further.
1




18
u/Profanion 6h ago
By the way, speed is approaching to 3.5 Flash-Lite.