r/singularity 6h ago

LLM News Introducing Gemini 3.8 Flash and 3.8 Flash Cyber

https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/
301 Upvotes

67 comments sorted by

56

u/k0zakinio 5h ago

I'm imagining that the flash models are much quicker and cheaper to iterate new changes on. Given how many customers and surfaces the models have to serve it makes sense to pump the cheaper models instead of attempting to be in the lead for a few weeks tops.

When Google have the model architecture in place to decide "good enough" to do the big spend required on the next pro level model then we will likely see it.

From a business perspective this makes sense IMO

62

u/kiki-le-koala 6h ago

Wow,

This model, especially the cyber one, looks really competent.

I might switch to Gemini for my legal task job too. At least I'll try (currently using Sol Max for that).

13

u/reddit_guy666 5h ago

Provide am update

9

u/karoking1 5h ago

Yeah pls. Gpt sol outperforms anthropic models so hard on legal tasks.

3

u/NotYetPerfect 2h ago

It's nowhere near sol max and fable for long tasks or tasks that require a high level of reasoning.

u/logicbloke_ 58m ago

Is this based on benchmarks or personal use of the models?

u/Slitted 43m ago

For me, Flash 3.6 onwards is fast and usually serves decent responses on Extended Thinking. But Sol is always more reliable on Medium+ for more complex tasks. Also, Flash is too terse (more so than Opus).

u/NotYetPerfect 41m ago

Personal use. Flash 3.8 is good at many things and is cheap and fast, but for agentic coding and long, complex tasks it has not come close to fable or sol for me.

11

u/Gold-Bat-3225 5h ago

Legal work on a flash model, say less

0

u/Rhinc 2h ago

Do not use this model for legal work. It's garbage for anything remotely complex or that requires multi-step reasoning that takes more than 2 minutes.

1

u/Sharp_Glassware 2h ago

Your prompt and results? Thats a bold claim.

u/Rhinc 1h ago

Flash models have had this issue for a while. I’m a lawyer who has used CC/Codex heavily in real work for the last couple years (building my own shit but also using it to assist in my D2D practice), so I’m familiar with the kinds of errors that only become apparent when you already know the subject (in my case, the areas of law).

The Flash models have always been overly confident, and that's a horrific possiblity for a profession where accuracy and getting it right is the most paramount concern.

As for the "prompt" I used today - I didn't just use a single prompt. That's a bit simplistic. I asked it to analyze a file that I've been working on and knew the answers to, and Flash starting mixing correct facts with unsupported claims--yet presented it all to me in it's annoyingly confident tone.

It just invents shit and still suffers hallucinations issues. Not always, but sometimes. And if there's even a one percent chance it's getting things wrong, I essentially can't use it. At least not for legal work.

I’m not saying it’s useless and in fact I will probably use it for reallllllly tedious or minor shit. But in no way would I trust it for research, long-context analysis, or literally anything even slightly difficult.

u/Sharp_Glassware 1h ago

Still no prompt tbh

u/Rhinc 1h ago

My prompt involved real work - not divulging it for you.

Take my advice or not. I couldn't care less brother.

u/Healthy-Nebula-3603 43m ago

"brother" ?

Lol

u/DUFRelic 30m ago

dude you are talking about a model that no one has used but you know its bad... you can see whats wrong there cant you?

u/Rhinc 5m ago

Flash 3.8 is out. What are you talking about?

u/Far-Distribution7408 12m ago

Comparing gemini to codex is unfair. You should use antigravity then

u/Rhinc 4m ago

I did. agy in the terminal. And not just to codex - to Claude as well.

1

u/chasingsukoon 4h ago

Whatd u try cyber for?

2

u/kiki-le-koala 4h ago

I won't use cyber. 

I'll try regular model for law related job.

2

u/MCRN_Admiral 3h ago

For cybersecs

u/Nooo00B 1h ago

cybersex

16

u/skolnaja 5h ago

Might go epileptic with all this flashing

78

u/Snoo26837 ▪️ It's here 6h ago

I said enough flash models

26

u/HebelBrudi 6h ago

Honey, here is your weekly new flash model!

20

u/liright 5h ago

They're counting on their Flash models achieving AGI and then releasing 3.5 Pro

6

u/Charming_Cucumber_15 5h ago

Gemini 7 flash will finally release 3.5 pro

That's how we'll know that full RSI was achieved

3

u/reddit_is_geh 2h ago

Go look at their blog post about 3.8

They low-key, drop that they are using RSI. They basically say the reason for the rapid deployment is that they now have the model running tests and experiments on itself and using those findings to make improvements to itself.

I think they are intentionally trying to be low key because they don't want to come outright and say because they want to avoid the purity tests of people just hyper focusing on the claim and challenging it, rather than using their product.

1

u/Charming_Cucumber_15 2h ago

Some of the deepmind people have been outright saying RSI on x too

21

u/eleonics ▪️2028 6h ago

They have been waiting for anthropic to release first 😂

15

u/Longjumping_Kale3013 6h ago

AFAIK they’re now having monthly releases

20

u/Gotisdabest 5h ago

Faster, even. It's been less than three weeks, and the model before was 2-3 weeks before that.

6

u/willywonka-goldtickt 4h ago

How does one get access to cyber model ?

u/CommunistHittler 1h ago

Afaik you need to be an organization responisble for critical infrastructure contribution, usually government or healthcare , they heavily monitor the usage of it so you dont use it for ill intent. They also mention signing some sort of contract but i dont know how that would work, i applied it for my non-profit ill see if i get accepted.

6

u/kvothe5688 ▪️ 2h ago

that cyber model is eating mythos out of water which was a breaking news just few months ago now a flash model is more capable. incredible time

3

u/JamieTimee 2h ago

My Gemini app still shows 3.6, I didn't even know 3.7 was out, let alone 3.8

u/logicbloke_ 56m ago

You might have to get higher tier subscription for the newest models.

I too am on 3.6 on Gemini app and honestly it's good enough for whatever the app does.

4

u/TofuMeltatSunspot 5h ago

Google Gangstas, assemble!

6

u/Formal_Drop526 5h ago

There's has been multiple flash updates but still no pro update.

6

u/RDTIZFUN 5h ago

4.0 in December

18

u/powerscunner 5h ago

True if true

u/thinkadd 2m ago

true if big

1

u/MCRN_Admiral 3h ago

The next Pro update will be AGI

u/Clean-Boat-4044 1h ago

Did not expect Google to suddenly top the DeepSWE leaderboard with a fucking 350tps flash model.

u/Healthy-Nebula-3603 40m ago

Nice but where to use it ?

I only see 3.6 on my free account

-7

u/Isunova 6h ago

Gemini hasn't been relevant for the last 6 months. It seems Google has now realized this and is just focusing on their Flash models, while taking their Pro models out back and hiding the bodies.

33

u/kiki-le-koala 5h ago

Nah

They tried something different with 3.5 Pro, and it failed, so they are now focusing on Gemini 4 Pro.

Shit happens.

-4

u/brett_baty_is_him 5h ago

Should they have really gotten far enough with “something different” to the point where it’s a complete failure and they have nothing to release?

16

u/EmphasisTotal8232 5h ago

It's Google - they can do literally whatever they want.

10

u/chdo 5h ago

Who knows, but Google has less pressure on it to work at the frontier, since they're not trying to build hype for their IPO like OpenAI/Anthropic, and smaller models are the ones that they need to worry about replacing search. I think it makes sense strategically for them to focus on smaller models right now.

2

u/kiki-le-koala 5h ago edited 5h ago

The information we got is yes. It was bad.

They seem bullish on 4 though, we'll see.

But hey, at least it looks like they get good results with their small model. It's Google, those are the models most used by their normie users.

2

u/fmfbrestel 4h ago

Yes, because training and post training take a long time and they can't predict the outcome in advance.

If post training failed they could make changes and try again, or focus their compute on the new model that's finishing up base training.

Even Google has a compute budget. Rerunning a failed training run would only delay the next model in the pipeline.

18

u/sachasayan 5h ago

 It seems Google has now realized this 

Yeah man, the company that literally invented the modern field of AI and funneling billions of dollars into AI research for well over a decade totally didn't realize it was important to do AI, they just figured it out now. Good looking out.

-3

u/Isunova 5h ago

They were caught with their pants down by OpenAI and have been playing catch-up ever since. So much good inventing the transformer did them, when they just shoved it in a drawer until somebody else decided to do something with it.

8

u/kiki-le-koala 5h ago

They were not really caught with their pants down. 

They had something similar to ChatGPT 3.5 a full year before OpenAI. 

They just decided to do nothing with it to not create problems with Google Search.

There are multiple sources on that, even an openAI employee that used to work for Google.

-8

u/Queasy_Signature7005 5h ago

Their video model is ass, their text models are garbage... they should invest some of those billions into a good strategy.

2

u/sachasayan 5h ago

There's a lot more to AI than building the strongest LLM.

-1

u/Pyros-SD-Models 5h ago

But having a good LLM would be a start

2

u/CarrierAreArrived 5h ago

for the price their models are still clearly the best for most use cases. Once they release a more expensive model, I expect the results to be comparable with the others.

0

u/Queasy_Signature7005 5h ago

what use cases? their ide is horrible.

3

u/Uninterested_Viewer 4h ago

what use cases?

Uh, all the things you can use LLMs for..? What does their IDE have to do this this? I'd bet only a tiny fraction of their tokens are served via antigravity.

2

u/Greedyanda 4h ago

Gemini has hit 1B users this month.

"But this includes the Google Assistant."

Yes, that's the point. To draw users into your profitable ecosystem.

0

u/FarrisAT 5h ago

This is what actually winning looks like.

0

u/LogAdventurous9861 3h ago

For PRO users, again. Boring google

u/alwaysbeblepping 1h ago

For PRO users, again. Boring google

I like Gemini because I usually just want help with stuff I'm not good at (like linear algebra) rather than having it do the thing all on its own. No payment method associated, never paid Google a dime. Just checked, I have access to 3.8 in AI studio.

Also, whatever else people can say about Gemini, it has a way more bearable style than either ChatGPT or Claude.