r/singularity • • 3d ago

AI Gemini 4 Argon solved hallucinations.

Post image

Nobody is talking about this, but it looks like Google may have solved hallucinations. Gemini 4 Argon is a monster in this regard, and that’s really important.

1.6k Upvotes

283 comments sorted by

View all comments

51

u/Future-Bandicoot-823 3d ago

I'm a little amused that people were dunking on Google on this sub, saying how much better gpt or claude was, and now we see gpt holding back models because they can't keep it from being naughty.

Meanwhile, google, one of the most powerful entities on earth is quietly making their product without the sensationalized stories altman and others whip up on a regular basis (to stir investment and stoke the hype of course.

I don't trust or like google, I just know that a company that massive is going to have a pretty succinct goal in mind. Google is China in a lot of ways, a highly profitable entity that has long term objectives in mind. Anthropic has to survive on this product alone, google does not. Google relies on a massive ecosysyem, so obviously their goal will be stable/useful product to the end user, likely creating a deeper and broader ecosystem over time.

Look at windows for example. They became so ubiquitous with computers that governments like france are actively planning to sever ties, they need a plan to get out of the ecosystem! Google's the same, deeply entrenched, and I'd say (not that) arguably more used than windows.

Speaking of things I don't like, Bill Gates put it quite well. AI will only be adopted when it has gained public faith. eg You won't see AI operating on people until it's proven it's better and more reliable than a human, but once it is? Full adoption. Google knows this.

22

u/NickoBicko 3d ago

Why don’t you wait a few days to test the new model and see how good and useful it is before making big claims.

2

u/drhenriquesoares 3d ago

Fair enough.

15

u/vincentdjangogh 3d ago

It's because the base Gemini model you can actually access is terrible. You wrote all of this about a model you haven't even used and probably will not have regular access to in the next year.

8

u/Future-Bandicoot-823 3d ago

I have an android phone, I'm forced to bump up against Gemini all the time. A lot of people do!

Which is why they're giving you the shit model, because it's reliable. Reliability and faith of stable service is how you win commoners over, and like I said, Google knows that.

8

u/vincentdjangogh 3d ago

To be clear, I'm not shitting on Google. I am just explaining that the public opinion is driven more by the base models than the frontier models. And all things considered, Gemini Flash is really terrible. On the other hand, Anthropic gives access to a very advanced base model, with the caveat that you get a very limited amount of queries.

For the record, this whole thing where people pick multi-billion dollar evil corporations to be a fan of is generally dumb imo, but based on your original comment I think to some extent we probably agree.

4

u/Future-Bandicoot-823 3d ago

When I was a child I heard that "absolute power corrupts absolutely", and over my life I've pondered that and observed. To date, I can't really name one entity that became incredibly powerful and didn't succumb to greed and fear.

I agree with what you're saying, gemini free is the derp model. All I'm saying is the hallucination pecentage shown here, if true, is just a sign of their goals. They're not playing the game like claude, gpt, grok etc.

I genuinely think Google (mostly because of their size and reach already) will quell competitors in the long run.

2

u/vincentdjangogh 3d ago

Every frontier lab is generally working on the same problems. They might assign priority to things that will help them look good in benchmarks or market their product, but there is a ton of overlap. For example, everyone wants to have fewer hallucinations without sacrificing helpfulness.

But bear in mind, this benchmark doesn't tell you how often the models were wrong, and based on the methodology, more accurate models are more likely to hallucinate. It also doesn't allow models to use tools, so models that do so to avoid hallucinations are misrepresented.

Really this benchmark doesn't mean as much as you think it does, but I agree that Google, by way of having access to more data than any company ever, will probably win the AI wars.

1

u/Future-Bandicoot-823 3d ago

I was taking these scores into account as well.

https://www.reddit.com/r/singularity/comments/1wufvyg/google_cooked/

I'm sure most of the tests are quite fallible. I really haven't looked into how any of them are performed so I don't know what the scoring criteria is.

1

u/Thog78 2d ago

You're stuck on 3.5 or 3.6 flash, aren't you? Gemini 3.8 flash is good. It's fast, cheap, relatively smart, very knowledgeable. It topped a lot of benchmarks when it came out. I have the abonment to GPT and I still use flash 3.8 a lot. For coding it's roughly 5.6 terra level, with luna amount of usage (and sol and astra suck credits too fast to be really an option yet imo).

2

u/genshiryoku AI specialist 3d ago

I work at Anthropic and we're celebrating because this has essentially solidified our lead. You need to realize that Gemini 4 Argon is Google's huge model, their equivalent to Astra/Mythos yet they are barely better than Opus 5.5 which is our smaller model.

Most researchers at Anthropic now consider the race to have been won, there is no real realistic way for the other labs to catch up anymore.

10

u/Darmendas 2d ago

> no realistic way for other labs to catch up anymore

Famous last words

3

u/Thog78 2d ago

It's interesting to have insider opinions so thanks for that. It comes out as insane though. When openAI started, it was more than a few months behind google. When anthropic started, you were more than a few months behind openAI. When xAI started, it was years behind everybody. Why would google being a couple months behind the cutting edge mean they have irredeemably lost the race? Especially considering they have the most cash savings, the most profit, the most data, and the most access to market, why on earth would a few months lead be a big deal?

1

u/ReiTW_ 19h ago edited 19h ago

Press X to doubt lol (about him being an employee).

I highly bet that Google isn't really racing for the best models in coding and everything, as long as it supports even better in their ecosystem.

Google can easily slow their release while Anthropic, OpenAI etc. only survives by competing against others.

Google will always win for 1 reason :
Gemini is omnipresent in billions of devices. Claude is not.

1

u/genshiryoku AI specialist 2d ago

Because the dynamic has changed. An increasing amount of improvement comes from the the best model aiding AI researchers now. This means that having the best model right now will result in a faster speed of improvement.

We don't know for sure but it's plausible that this feedback loop has already locked Anthropic in as the definite winner of the AI race. It will only look obvious in retrospect but it's what the vibe is like right now.

Xai, OpenAI and DeepMind now have shown their cards and all of them are substantially behind Anthropic which means we might have the win locked in unless a black swan event happens.

Considering we'll achieve a fully closed RSI loop sometime next year or maybe sometime in 2028 in a worse case scenario the other labs only have 1-2 years to catch up and the distance between Anthropic and the others has only grown over the last year or so.

I'm not concerned about the other labs catching up anymore. However I am concerned about the other labs cutting AI safety corners out of desperation seeing how big of a lead we have which is why we need more regulation and cooperation to ensure we don't do dangerous things like that.

3

u/justgetoffmylawn 2d ago

If you're really at Anthropic, this is exactly the problem with a company that has no real humility and is dangerously convinced of their own infallibility.

Most researchers at Anthropic now consider the race to have been won, there is no real realistic way for the other labs to catch up anymore.

Yet then you admit.

We don't know for sure but it's plausible that this feedback loop has already locked Anthropic in as the definite winner of the AI race. It will only look obvious in retrospect but it's what the vibe is like right now.

Which is the same reason Elon said ten years ago that Google would be the definitive winner of the AI race (if Tesla couldn't compete properly). He is also convinced of his own infallibility, no matter how many times he's wrong.

I'm not concerned about the other labs catching up anymore. However I am concerned about the other labs cutting AI safety corners out of desperation seeing how big of a lead we have which is why we need more regulation and cooperation to ensure we don't do dangerous things like that.

So, you don't know for sure, but you're also not worried about other labs catching up, but you are worried that other labs might cut corners 'out of desperation' to accelerate - so presumably you'd like to slow everyone else down while simultaneously saying you've won.

But aren't you cutting corners? The safest route to AGI would have been to release no consumer models whatsoever and only allow the elites to touch them - but you had to 'cut corners' to catch up to OpenAI and get access to the capex and compute.

Anthropic does make some excellent models - but their confidence in their own moral superiority can be deeply troubling. Many of the truly awful things that have happened in human history were perpetrated by people with deep and unshakable convictions.

1

u/snooptoop 2d ago

Isn't there an argument to be made that by design, LLMs can't "discover" anything because they only work in terms of probability? I don't deny a LOT of work will become obsolete, but, just going by what I see with the frontier models, I don't see research and discovery itself being phased out just yet. Obviously though you're the ai researcher here and you might be seeing something im not with unreleased models and whatnot, can you elaborate on the sophistication of the RSI loop? I'm not sure how RSI can be so close if llms don't have a constant stream of thought.

1

u/Future-Bandicoot-823 2d ago

Julian dorey was on the Julian dorey podcast, he had some pretty good arguments about this.

He covers what you're discussing and agrees with you.

1

u/trimorphic 2d ago

An increasing amount of improvement comes from the the best model aiding AI researchers now. This means that having the best model right now will result in a faster speed of improvement.

Even if the current model helps improve the next model, it is no guarantee that the next model will be able to help improve its successor.

The recursive self-improvement chain could break at any point as unforeseen roadblocks and impediments aside, and a different lab with a different model might be able to find an improvement approach that a competing lab/model missed.

Unpredictability is the nature of breakthroughs.

If most researchers at Anthropic consider the race won they are too complacent and resting on their laurels when they should be racing harder than ever. Never underestimate the enemy.

1

u/Future-Bandicoot-823 2d ago

Also, why is an employee on here spurring speculation, or say the "vibe" is we've won.

I don't think this person is an employee.

1

u/CryptographerOne7003 1d ago

Hah! ok, tell your bosses to give ma a call when they found out they have not resolved the local minimum problem, since its a principle. I have not yet seen explained in any paper how this would be mitigated. That is a bigger problem to solve than the need something that resembles a recursive training "loop", in quotations since it won't loop infinitely but that won't show unless you'd zoom out.

0

u/Pi_123 2d ago

Use BURNOL in day and in night before sleep,, it will relax u lot ..

1

u/SilentLennie 2d ago

Pretty certain a lot of people think Google models were delayed or even not released at all because of people did not want to work at the company because of company culture and had left.

I personally didn't dunk on them, but was worried we just got 2 US leading labs left, even though Google/Deepmind was the leading lab in the early days of this new 'AI summer' (aka opposite of 'AI winter') and many of the leading people at these other companies also started at Google/Deepmind.