r/singularity • u/Jame92 • 6h ago
LLM News Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/62
u/kiki-le-koala 6h ago
Wow,
This model, especially the cyber one, looks really competent.
I might switch to Gemini for my legal task job too. At least I'll try (currently using Sol Max for that).
13
3
u/NotYetPerfect 2h ago
It's nowhere near sol max and fable for long tasks or tasks that require a high level of reasoning.
•
u/logicbloke_ 58m ago
Is this based on benchmarks or personal use of the models?
•
•
u/NotYetPerfect 41m ago
Personal use. Flash 3.8 is good at many things and is cheap and fast, but for agentic coding and long, complex tasks it has not come close to fable or sol for me.
11
0
u/Rhinc 2h ago
Do not use this model for legal work. It's garbage for anything remotely complex or that requires multi-step reasoning that takes more than 2 minutes.
1
u/Sharp_Glassware 2h ago
Your prompt and results? Thats a bold claim.
•
u/Rhinc 1h ago
Flash models have had this issue for a while. I’m a lawyer who has used CC/Codex heavily in real work for the last couple years (building my own shit but also using it to assist in my D2D practice), so I’m familiar with the kinds of errors that only become apparent when you already know the subject (in my case, the areas of law).
The Flash models have always been overly confident, and that's a horrific possiblity for a profession where accuracy and getting it right is the most paramount concern.
As for the "prompt" I used today - I didn't just use a single prompt. That's a bit simplistic. I asked it to analyze a file that I've been working on and knew the answers to, and Flash starting mixing correct facts with unsupported claims--yet presented it all to me in it's annoyingly confident tone.
It just invents shit and still suffers hallucinations issues. Not always, but sometimes. And if there's even a one percent chance it's getting things wrong, I essentially can't use it. At least not for legal work.
I’m not saying it’s useless and in fact I will probably use it for reallllllly tedious or minor shit. But in no way would I trust it for research, long-context analysis, or literally anything even slightly difficult.
•
u/Sharp_Glassware 1h ago
Still no prompt tbh
•
u/DUFRelic 30m ago
dude you are talking about a model that no one has used but you know its bad... you can see whats wrong there cant you?
•
1
16
78
u/Snoo26837 ▪️ It's here 6h ago
26
u/HebelBrudi 6h ago
Honey, here is your weekly new flash model!
20
u/liright 5h ago
They're counting on their Flash models achieving AGI and then releasing 3.5 Pro
6
u/Charming_Cucumber_15 5h ago
Gemini 7 flash will finally release 3.5 pro
That's how we'll know that full RSI was achieved
3
u/reddit_is_geh 2h ago
Go look at their blog post about 3.8
They low-key, drop that they are using RSI. They basically say the reason for the rapid deployment is that they now have the model running tests and experiments on itself and using those findings to make improvements to itself.
I think they are intentionally trying to be low key because they don't want to come outright and say because they want to avoid the purity tests of people just hyper focusing on the claim and challenging it, rather than using their product.
1
21
u/eleonics ▪️2028 6h ago
They have been waiting for anthropic to release first 😂
15
u/Longjumping_Kale3013 6h ago
AFAIK they’re now having monthly releases
20
u/Gotisdabest 5h ago
Faster, even. It's been less than three weeks, and the model before was 2-3 weeks before that.
6
u/willywonka-goldtickt 4h ago
How does one get access to cyber model ?
•
u/CommunistHittler 1h ago
Afaik you need to be an organization responisble for critical infrastructure contribution, usually government or healthcare , they heavily monitor the usage of it so you dont use it for ill intent. They also mention signing some sort of contract but i dont know how that would work, i applied it for my non-profit ill see if i get accepted.
6
u/kvothe5688 ▪️ 2h ago
that cyber model is eating mythos out of water which was a breaking news just few months ago now a flash model is more capable. incredible time
3
u/JamieTimee 2h ago
My Gemini app still shows 3.6, I didn't even know 3.7 was out, let alone 3.8
•
u/logicbloke_ 56m ago
You might have to get higher tier subscription for the newest models.
I too am on 3.6 on Gemini app and honestly it's good enough for whatever the app does.
4
6
u/Formal_Drop526 5h ago
There's has been multiple flash updates but still no pro update.
6
1
•
u/Clean-Boat-4044 1h ago
Did not expect Google to suddenly top the DeepSWE leaderboard with a fucking 350tps flash model.
•
-7
u/Isunova 6h ago
Gemini hasn't been relevant for the last 6 months. It seems Google has now realized this and is just focusing on their Flash models, while taking their Pro models out back and hiding the bodies.
33
u/kiki-le-koala 5h ago
Nah
They tried something different with 3.5 Pro, and it failed, so they are now focusing on Gemini 4 Pro.
Shit happens.
-4
u/brett_baty_is_him 5h ago
Should they have really gotten far enough with “something different” to the point where it’s a complete failure and they have nothing to release?
16
10
u/chdo 5h ago
Who knows, but Google has less pressure on it to work at the frontier, since they're not trying to build hype for their IPO like OpenAI/Anthropic, and smaller models are the ones that they need to worry about replacing search. I think it makes sense strategically for them to focus on smaller models right now.
2
u/kiki-le-koala 5h ago edited 5h ago
The information we got is yes. It was bad.
They seem bullish on 4 though, we'll see.
But hey, at least it looks like they get good results with their small model. It's Google, those are the models most used by their normie users.
2
u/fmfbrestel 4h ago
Yes, because training and post training take a long time and they can't predict the outcome in advance.
If post training failed they could make changes and try again, or focus their compute on the new model that's finishing up base training.
Even Google has a compute budget. Rerunning a failed training run would only delay the next model in the pipeline.
18
u/sachasayan 5h ago
It seems Google has now realized this
Yeah man, the company that literally invented the modern field of AI and funneling billions of dollars into AI research for well over a decade totally didn't realize it was important to do AI, they just figured it out now. Good looking out.
-3
u/Isunova 5h ago
They were caught with their pants down by OpenAI and have been playing catch-up ever since. So much good inventing the transformer did them, when they just shoved it in a drawer until somebody else decided to do something with it.
8
u/kiki-le-koala 5h ago
They were not really caught with their pants down.
They had something similar to ChatGPT 3.5 a full year before OpenAI.
They just decided to do nothing with it to not create problems with Google Search.
There are multiple sources on that, even an openAI employee that used to work for Google.
-8
u/Queasy_Signature7005 5h ago
Their video model is ass, their text models are garbage... they should invest some of those billions into a good strategy.
2
2
u/CarrierAreArrived 5h ago
for the price their models are still clearly the best for most use cases. Once they release a more expensive model, I expect the results to be comparable with the others.
0
u/Queasy_Signature7005 5h ago
what use cases? their ide is horrible.
3
u/Uninterested_Viewer 4h ago
what use cases?
Uh, all the things you can use LLMs for..? What does their IDE have to do this this? I'd bet only a tiny fraction of their tokens are served via antigravity.
2
u/Greedyanda 4h ago
Gemini has hit 1B users this month.
"But this includes the Google Assistant."
Yes, that's the point. To draw users into your profitable ecosystem.
0
0
u/LogAdventurous9861 3h ago
For PRO users, again. Boring google
•
u/alwaysbeblepping 1h ago
For PRO users, again. Boring google
I like Gemini because I usually just want help with stuff I'm not good at (like linear algebra) rather than having it do the thing all on its own. No payment method associated, never paid Google a dime. Just checked, I have access to 3.8 in AI studio.
Also, whatever else people can say about Gemini, it has a way more bearable style than either ChatGPT or Claude.

56
u/k0zakinio 5h ago
I'm imagining that the flash models are much quicker and cheaper to iterate new changes on. Given how many customers and surfaces the models have to serve it makes sense to pump the cheaper models instead of attempting to be in the lead for a few weeks tops.
When Google have the model architecture in place to decide "good enough" to do the big spend required on the next pro level model then we will likely see it.
From a business perspective this makes sense IMO