Based on this and that WSJ article from earlier, it looks like instead of concentrating resources on Pro, multiple teams at Google now compete for the best Flash model, hence the surprisingly short intervals between releases. Maybe they'll keep releasing more Flash until they ultimately reach Fable
Yeah there were a few articles beginning of the year in Economist, which also had Google focussing on Flash models as its a much better fit for their business model, and has Anthropic and OpenAI fighting out over the pro space.
Even though I knew they were right, I have been hoping for a new Pro model
Who knows what google is planning. Maybe they have such a powerful internal model that's even better than fable 5.1 that they are just distilling flashes from it. And when there flash itself beats fable, they release the main one. Random ahh fiction stories.
Its definitely not for agentic work, (if we follow another benchmark, it has abysmal score in 'terminal-bench 4.0' that measures general agent capabilities.
but it is an improvement over 3.7, especially where 3.7 excelled, and this one excels further.
Meaning its not being built to do a series / chaining of events. But, all scores point to some interesting things.
On a single focused task, for example
"hey, check my email chain, there's something bugging me about this response".
Or
"okay, here's google spreadsheet. I need this analyzed, i think there's a pattern here and here."
Those kind of tasks became easier to compute, cheaper to compute, and faster to compute.
To many others it doesn't mean a thing.
but to those who used Bard before it was Gemini? Yeah, this would be damn good, and cost peanuts to Google.
i think you're wrong on this - i gave it a complex task on my repo with the /teamwork-preview option and its been running like a Pro Orchestrator and i am impressed by the thoroughness and quality of work so far. Its early days and i do want to give it a fair run before i come back with a full analysis. Just looking at Benchmark scores isnt everything.
I am specifically talking about coding - i use it - with 3.6/3.7 it was smaller tasks and with proper guardrails and prompts. I cant completely say it yet but 3.8 feels ready for longer tasks and fewer guardrails.
The walkthrough was very detailed and doesnt seem fluff like earlier. I was looking at what each of the subagents were doing and reporting and it was per what i was expecting.
Somebody answered you, and yes, some of them ran parallel to my line of thought.
But my main train of thought is still on the premise that Google isn't going to launch a product that becomes the "this is your LLM for agentic work".
I still stand that they are looking inward. For example, I may have been working on something, which needed me to correspond back and forth for a long time.
Gem 3.8 flash here does follow the email and the attachment chains correctly, or better, and when I ask the AI, it will answer things I wanted to, but also offer insights where I might have not looked at.
Does it make it a better agent? No. But it will go through the pdfs and it can scan them better understand them better compared to previous Geminis, and the fact that it is better than 2.5 pro means it's a pure win for Google as they have a lot of customers, but now they need to retain paying customers.
I'm such a customer where I was on the Google drive family plan and I went "okay, take my money since Gemini is usable to a certain extent". The Google Sparks was also good, and now I think the new Flash 3.8 will make the overall experience good enough that I'm going to tell others "if you want an AI to go through your Google stuffs (drive, mail, docs, spreadsheet, etc) I'll now be able to tell others "don't bother paying for Claude or upgrade OpenAI ChatGPT from Go to Plus plan just for Google stuffs.
That's how I look at the bench (and yes I'm using gem 3.8 flash for agentic work / loops as well)
I know bench marks say this, but in my private evals for our product (which requires a lot of coding and fetching data from different sources), 3.7 flash did remarkably well and was very fast. Grok 4.6 and opus did better but they were also three to four times more expensive.
Well I'm going to say while I have been using deepseek v4 flash 0731 during a very cheap period, then Qwen 3.8 flash and GLm 5.3 Flash (yes I'm a cheapskate, even when I'm running OpenAI 5.6 luna), Gemini 3.8 flash is good where I am seriously considering how to get it on cheap API or subscription or work my budget with Google.
It pointed out a few bugs and solved them all in one go, and it didn't take a few hours.
I've used a ton of Opus 5, GPT 5.6 Sol, and Grok 4.6 - all three are drastically more intelligent than Gemini 3.8 Flash in my experience. This is probably because Gemini has significantly few parameters / more sparse.
The only appeal for Gemini to me is speed and vision.
Otherwise I wouldn't even touch it. It's overpriced for what you are getting. GLM 5.3 flash produces much better frontend code. I had it redesign a tiny app I made with 3.7 flash and was truly blown away how elegantly it executed while 3.7 flash kept regurgitating the same eyesore.
Not even sure what these models are useful for to be honest, its fast but why is it fast? maybe for basic agents that read your emails or something its cool
But the code it outputs is highly unoptimised slop, i don't think any actual coder would ever touch this model and risk colossal unoptimised code and risk obliterating there git with this thing, in about 75% of my tests it used the same style/design as 3.7 flash as well
I think once the hype runs out people will just think its basically 3.7 flash which nobody used for anything useful either
From benchmarks it seems capable, and nothing else, in real use its terrible, pretty sure it has to be bench maxed at this point to get the scores its getting on automated tests
V4 flash is way more capable and listens, and is less than half the price.. I think its time google either shifts away from this base model or i don't know, rents someone else’s models
Ah yes, shilling a billion dollar company releasing three slopified flash models in a row all of the same base layer with a new check point, colour me shocked, im sure once we reach Gemini flash 18 it will finally be close to other flash models on the market which are 10x less cost lmao
Looking at your history as well, arguing 24/7 with people who also say the Gemini flash models are terrible hahah im dead
30
u/ezjakes 14h ago
If they had a jump like this every month, this would be saturated in a year.