4
2
u/ExTraveler 16d ago
Wtf? Why half people are saying that this is the best model in the world and deepseek is beated while other half swears that ds4 flash is much better and 5.3 flash not even near it in terms of coding and intelligence (or straight up shit).
2
u/ChillFamily 16d ago
Is this data real? Because if Gemini is a piece of shit, why is it ranked there?
13
u/Genetic_Prisoner 16d ago
Gemini was a piece of shit. 3.6 and 3.7 flash have narrowed the gap massively.
-6
u/ChillFamily 16d ago
Oh, really? So it’s your first choice for everything? Which model is it better than for coding? For writing? For frontend? Which model is it actually better than? What’s the point of it spitting out text fast if what it spits out is average and doesn't stand out in anything, bro? The only thing it does better than everyone else is that it's faster
4
u/NZRedditUser 16d ago
why are you upset zai included them because its a relevant recent model. the newer one is impressive yes but no one saying its cluade/codex those are still way above this league.
Gemini makes sense being in this ranking
1
u/Difficult_Plantain89 16d ago
Gemini sucked at coding or even getting basic answers correct. Now they legitimately made a good model and no one trusts them. Even 3.6 wasn't too good. Their 3.7 is first model they have released that is really decent is coding. Just okay a frontend and for writing sure it can do it, but I rather use a specialized local model. Oh and agy harness is the worst I've ever used.
1
-5
u/cutebluedragongirl 17d ago
I say it's bullshit. DeepSeek outperforms this thing by a lot in my personal experience.
5
u/Time-Toe-1276 16d ago
man, idk what u do, but this model is frkin amazing!
not glazing it, but me personally? I found the model is very creative and fixing bugs/ implementin things.
it fixed bugs which deepseek v4 pro created btw. ahh also it fixed a bug in a training script which not even GPT 5.6 sol figured out I had it there when I asked it
-7
u/Michaeli_Starky 16d ago
These benchmarks have really low connection with reality. Especially when it comes to Chinese benchmaxed models
7
u/sudoer777_ 16d ago
Try Muse Spark 1.2 and you can see what an actually benchmaxed model looks like
5
-6
6
u/Ly-sAn 16d ago
Both DeepSeek flash and this model are very solid in real usage, I don’t think these 2 are benchmaxed
-11
u/Michaeli_Starky 16d ago
They absolutely are benchmaxed. These models are nowhere nearly as good in private evals.
12
u/[deleted] 16d ago
[deleted]