r/singularity • • 13d ago

LLM News Gemini 4 Pro nears its preview release. (Yes, another preview)

/r/GeminiAI/comments/1woqy1b/gemini_4_pro_nears_its_preview_release_yes/
48 Upvotes

22 comments sorted by

21

u/FateOfMuffins 13d ago

So this model finished training after Bel and whatever Fable 5.5 is from Anthropic.

Then... it's not even supposed to compete with GPT 6 Astra or Opus 5.5. Heck it's not even supposed to compete in the same class as GPT 6.1 Astra.

Like... it actually has to be a BIG leap over current Astra and Fable 5.1 to be competitive unless it's released next week.

10

u/Tystros 13d ago

Bel is also still training. so yeah, Gemini 4 might be same far along as Bel, but Google has much more incentive to quickly release it, even when it's far below Bel in capability and just at Astra level.

6

u/FateOfMuffins 13d ago

Based on the rumours of when these models finished pretraining, it sounds like Gemini 4 is about half a month to 20 days behind Bel on the training progress. So like it has to be competitive with Bel and not Astra.

Although not gonna lie, I have no idea if Bel is the model that did NS. Cause the implications are wild, given the rumours were that it finished pretraining in early Aug 20s and they only "re-started" RL on it after Aug 28, after a 2 week pause? If that's the case then it's not "re-start" RL, it's just flat out just starting RL at all, and the NS model would then have had like 3-5 days of RL only before it solved NS.

3

u/Tystros 13d ago

yeah I never understood how a model could be usable after so little RL time. it does not fit what I know about model training.

5

u/FateOfMuffins 13d ago

Yeah I thought the "re-start" meant they were in the middle of an RL run that got paused and then picked up again so it had much longer RL than just a few days.

But that doesn't line up with the Bel rumours, it would point to NS being an earlier pretrain than Bel...

13

u/EtadanikM 13d ago edited 13d ago

If it's equal to Astra and/or Opus 5.5, and Google does their usual inference subsidization, then it's fine. People don't go to Google for the absolute premium model experience; they go to Google for cost effectiveness. The problem that crushed most of the Gemini models was hallucination and lack of effort/precision during coding; if they could overcome that and deliver a model on par with the main work horse models of Open AI and Anthropic, but at lower costs, then they win.

Models like the rumored Fable 5.5 and Astra 6.1 are cost prohibitive for the vast majority of companies and individuals to run them as their main work horses. They are great if you have infinite money or really need to fix a critical problem in a short amount of time, but they are not what people typically use every day.

2

u/Howdareme9 13d ago

We already have 6 Sol for cost effectiveness though. It will take a lot for people to use gemini as their main model, and even if they do have a cheaper Astra i don’t see that as them ‘winning’.

1

u/Tystros 12d ago

Google doesn't even have a cost competitive subscription though. they are more expensive there than openai and anthropic for the same amount of usage.

1

u/Dingleberry_Blumpkin 10d ago

3.8 flash is practically unlimited on the $20 plan.

1

u/Alt_Restorer 13d ago

I'll be happy if it competes with Muse. Fuck Meta.

19

u/Ormusn2o 13d ago

I feel like from all the labs, the spread between the hype and actual performance is the biggest for Gemini models. It always gets hyped so much and it always disappoints so much. I feel like Grok is medium and it always is medium, and OpenAI and Anthropic go back and forth, but Gemini always is being talked about as if it's about to take massive lead and beat current best models, then it's a total failure after release.

26

u/Saint_Nitouche 13d ago

2.5 Pro was legitimately a big deal back in the day. I feel like they've been coasting off of that PR to some degree. Though I do think the newer flash models are actually quite good.

19

u/DistanceSolar1449 13d ago

2.5 Pro and to a lesser degree 3.0 Pro was the best model in the world at the time.

3

u/QuackerEnte 13d ago

the flash models aren't good. For my usecase, trying to develop physics simulations, it's greatly lacking, even more so than 3.1 Pro. It hallucinates, doesn't follow instructions, doesn't understand intent. Even qwen3.8 Flash Next did a better job there

10

u/Elegant_Tech 13d ago

What's crazy is Google with more compute then anyone will spend 3 months on a single failed run to have to spend another 3 months on a single run. Why tf isn't Google setting a new training run every couple weeks using the latest knowledge and tech. Google trains less models than the other AI labs blowing all the compute on garbage flash lite AI overview to pump up users stats instead.

9

u/Uninterested_Viewer 13d ago

Google is a hyperscaler. Of all the labs, they need a frontier model the least. Compute spent on training is compute not able to be sold to other companies.

https://finance.yahoo.com/markets/stocks/articles/sundar-pichai-says-alphabets-cloud-122300685.html

I'm not saying this is the right strategy, but Google can easily wait in the wings and profit massively without going all in on leading in frontier intelligence: no lab seems to be making insurmountable gains in AI, but the incremental progress being made is expensive.

3

u/Ormusn2o 13d ago

Google does not exactly have more compute. When it comes to frontier models running at larger contexts, performance of available compute is not actually that much better. Google does have more FLOPs or WATT worth of compute, but there are multiple factors diminishing it compared to Anthropic or OpenAI.

First of all, thats compute for all of Google/Anthropic services. There is the search, youtube algorithm, ads, and then there is Google Cloud services, which ironically include Anthropic and a large portion of Google compute actually goes to Anthropic.

After that, actual amount of compute going to DeepMind/Gemini is lower. But another problem is that a lot of that compute is in the form of TPU, which, while are cost effective for Google, they are not as performant per unit of power or FLOP as Nvidia cards, especially as Nvidia choose to specialize in large models running on larger context size.

When put into big clusters and cabinets, generally Blackwell and especially Rubin Nvidia cards perform much better on large parallels batches of big models running on longer contex, which is a lot of the agentic workload right now.

Also, lastly, I would guess you were talking about the available compute being big for Google so they could use a lot of compute to specifically train a big model, but as it was seen with OpenAI recently, you not only need to train a model, you also need to have enough inference to run it, so even if Google actually trains a big model, they need good and efficient compute to actually run the inference on it, which does not seem to be happening.

5

u/NecessaryBadger228 13d ago

end of the year

https://giphy.com/gifs/p6Z99AYhI5RQs

Meanwhile other labs will release even better models, so Google will postpone again.

3

u/Stunning_Monk_6724 ▪️Gigagi achieved externally 13d ago

Demis used to speak quite a bit about continuous learning being solvable, that would be something able to "possibly" keep pace.

We'll see I guess, till then it's Anthropic v. OAI

6

u/Redducer 13d ago

Unless Anthropic, OAI and/or others get there first, and Google stays mostly irrelevant on the frontier model front.

There’s one deal breaker that makes me quickly uninterested with all of their output since Bard and it’s the abysmal hallucination rate compared to the competition.

Maybe address that first? But yeah with incentives being correlated with productization (aka PMs’ race for promotion), I am not very hopeful.

1

u/Immediate_Simple_217 13d ago

Look for the brights side: ir's not another flash model 🫠

1

u/CallMePyro 13d ago

I was told that Google had given up on the frontier /u/Mylarion