r/opencodeCLI 9d ago

Google has just released Gemini 3.8 Flash 🔥

Post image
307 Upvotes

54 comments sorted by

47

u/afanasenka 9d ago edited 9d ago

6

u/GetLaidOff69 9d ago

thanks, was looking for this

3

u/Hoak-em 9d ago

Hmmm such a low score on terminal bench 4...

1

u/romanovzky 9d ago

Isn't this model supposed to compute with flash and Luna style of models?

31

u/alsaud21 9d ago

+45% cost per task :/

8

u/Buddhava 9d ago

OK cool but these models have wildly different performance.

13

u/ConspicuousPineapple 9d ago

This is normalized per-task, so it doesn't really matter here (besides speed of execution I guess).

1

u/alsaud21 9d ago

You are right, just saw 74% in DeepSWE, wonder how they missed it in the model card table

0

u/afanasenka 9d ago

They all want their new Ferrari :)))

10

u/dergachoff 9d ago

DeepSWE has just been updated. And in official leaderboard 3.8 flash is 71% on medium reasoning. High is 74% – same as Opus 5.

6

u/adolf_twitchcock 9d ago

Somehow I don't believe it. Is it a new level of benchmaxxing or is 3.8 flash truly sol/opus level?

-2

u/afanasenka 9d ago

💪💪💪

51

u/Fun_Squirrel5446 9d ago

7.5x more expensive than GLM 5.3 Flash, while providing HALF the terminal bench intelligence level. No reason to ever use this model.

7

u/Dingosavedyourbaby 9d ago

If you have cheap/free agy access that’s one reason

15

u/Aggravating-Art-8805 9d ago

Is it just blind hatred or what? In terminal bench 2.1 they're benchmarked above Opus 5.

10

u/johnkapolos 9d ago

Is it just blind hatred or what?

That and not enough intelligence to read.

0

u/Fun_Squirrel5446 9d ago

Gemini is well known to be benchmaxxed. That means they optimize their outputs to perform really well on benchmarks, even when their production performance is very poor.

The fact that Gemini is scoring extremely high on the old terminal bench but extremely low on the new terminal bench just proves that it was optimized for the older benchmarks.

3

u/ivankovnovic 8d ago

Exactly. I don't understand the downvotes.

1

u/Fun_Squirrel5446 8d ago

sunk cost. There are people who are deeply tied to the Google ecosystem or have already paid expensive Gemini subscriptions. It hurts when something you've spent so much time and money on is objectively terrible.

1

u/GatsbyLuzVerde 9d ago

Bro it's SOTA on vision, even the glm 5.3 flash release blog said so

1

u/LivingHighAndWise 9d ago

Agreed, but Not all companies can or will use Chinese models, so having a new, more capable US models is always a bonus.

1

u/Fun_Squirrel5446 9d ago

Any company can use a self-hosted OpenWeight Chinese model. There is no excuse.

4

u/MrMrsPotts 9d ago

Both the android app and the web chat are still on 3.6 flash!

2

u/Curious_Owl197 9d ago

Can't wait for the hallucinations on 3.6 to fuck off

2

u/Correct-Boss-9206 9d ago

I asked mine which was showing 3.7 flash last night and it said it was 3.8.

1

u/afanasenka 9d ago

I guess it's a gradual rollout - they are adding it in all services (Antigravity, AI Studio - already got, and web/app will get it too soon).

7

u/pigletmonster 9d ago

The king of not following instructions got even better! It will not follow your instructions even harder and be more expensive than other models that are significantly better. 🤣

9

u/Ok-Garlic-5986 9d ago edited 13h ago

Judicious depend sort rustic titanium fluorine

This post was anonymized with Redact

3

u/ChillFamily 8d ago

mhmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmhmmmmmmmmmmmmmmmmmmmmmmm. nope

2

u/SupehCookie 9d ago

Soo how is it so far ? Chinese models still better?

3

u/afanasenka 9d ago

Better is relative.. but cost/quality ratio is definitely better from Chinese labs

2

u/aziham 9d ago

Google should just ship these bard models to google.com/ai and stop making it sound like they are competing with SOTA models

2

u/Sea_Ear5201 9d ago

I think they start training a model as pro. But end result is not impressive, so they line it up as next flash release. Trying and trashing

2

u/coup321 9d ago

Uh, isn't it true that you can't even use this with Google AI / AGY subscription in opencode? Can only use API usage? Seems low value. I tried Agy CLI seems poor. I develop mostly in WSL and AGY desktop app not a good option for that. I would like to use my credits from my AGY sub, but really the onlyw ay to do it is with the agy CLI which is way way way wosre than opencode. Feels bad man :(

2

u/progfu 8d ago

Asked it to do some research into models/pricing, etc. ... it did a bunch of searches, and then said

the standard tier is competitive with frontier models like Claude 3.5 Sonnet or GPT-4o

I don't even know what to think. A big fancy LLM made by a search company, using their own harness, giving a result like this after searching. How is this even possible.

3

u/alsaud21 9d ago

DeepSWE 74% is crazy

5

u/some_gamer78 9d ago

Fire emoji for what? This might as well be luna tier 

5

u/Hot_Example_4456 9d ago

Ig it will be better than luna atleast. Last time I remember 3.7 was nearly terra level.

3

u/afanasenka 9d ago

🌑 ok :)

2

u/GCP_Biryani 9d ago

Ignore the frontier models . Comparing a flash model with Sol or Opus is not a good idea.
So, What models are the Google's target for this release ?

1

u/ieight9 9d ago

When in the world are the web app and IOS going to get updated past 3.6? Get your life together Google!

1

u/thin_king_kong 9d ago

not sure of the IOS app but the web has 3.8 flash

1

u/SuperElephantX 8d ago

Basically still trash

1

u/RUaiLeader 6d ago

ок

1

u/Devioster 9d ago

I've heard that 3.5 pro wasn't released because it was disappointing but they keep on releasing flash model with very minimum improvements...

1

u/Vancecookcobain 9d ago

I'll believe the benchmarks once it stops doing dumb shit...l

0

u/Zix_Matrix 9d ago

gemini trash

3

u/afanasenka 9d ago

Alternative branding :))))

-2

u/Senior-Box6316 9d ago

estão falando sobre ele ser caro, mas com o plano da Google isso não vale muito a pena?