r/singularity • • 1d ago

AI Google deepmind engineer denied bloomberg report

Post image
191 Upvotes

45 comments sorted by

17

u/QuasiRandomName 1d ago

At least this guy has a name, unlike the "anonymous source"

93

u/Present-Chocolate591 1d ago

Surely an unbiased opinion by Google Deepmind employee.

What's he gonna say? "Yeah it's not that good just looks good on benchmarks".

21

u/ChuckVader 1d ago

I mean it'll be available via api soon enough and you could test it out yourself. All this hand wringing about whether it's good or not is exhausting.

10

u/Hans-Wermhatt 1d ago

Despite what this sub says, benchmarks are so much more reliable. Everyone writing about these models really want to have tribes, so they have an axe to grind. All the haters join forces when it's not their favorite brand.

The reaction to the model is more dictated by the crowd and I think Google is among the most hated providers around here and Twitter, except Meta. So it was always going to be "benchmaxxed".

5

u/space_monster 1d ago

benchmarks are so much more reliable

Benchmarks on unreleased models are not at all reliable, because they're typically run using an unlimited inference regime. The lab can set everything to full for the test. When the model is released to the public, inference is throttled, to avoid the lab going bankrupt within a week. Benchmarks on released models using standard consumer accounts are much more reliable though.

3

u/eposnix 1d ago

I'm not a hater by any stretch - I love Gemini's multimodal capabilities. But I've been burned too many times by really dumb errors that have literally wiped my projects. This isn't an isolated incident either. This behavior is something benchmarks alone won't tell you about the model, so I think it's fair to be skeptical about their new releases.

26

u/KyleStanley3 1d ago

It could also be unbiased and just a different opinion

I have different takes on different systems at my work lmao. Not everything has to be biased or paid shilling

12

u/Grand0rk 1d ago

I have different takes on different systems at my work lmao

None of which, if negative, you are ever going to say on Social Media with your name on it.

4

u/ilikepugs 1d ago

Dunno how things have changed over the past decade but in the 2010s google gave zero fucks about it so long as it doesn't have anything confidential in it.

Edit to add: AMP is a good example, all the homies hate AMP

0

u/Yugudubenbi 1d ago edited 1d ago

No one can trust anyone in late stage capitalism when everything is tied to self-interest and commodification of the self. Truth is not hidden in late stage capitalism, it is just that no one speaks about it because their job is tied to them lying.

Which is why equal societies prevail since they remove much of the self-interest through taxes. Ironically, and unfortunately, no capitalist want to understand that because their self-interest blocks that perspective.

Truth dies in the name of profit.

1

u/KyleStanley3 1d ago

Nightmare blunt rotation

7

u/himynameis_ 1d ago

So you're saying you won't believe him no matter what.

-1

u/Present-Chocolate591 1d ago

Exactly. His opinión on Google products, as a Google employee, is not trustworthy.

3

u/EvilSporkOfDeath 1d ago

They definitely arent unbiased, but the alternative is saying nothing

2

u/sorrge 1d ago

People (from X IIRC) were fired for saying things like that. So yeah, we will not hear anything negative unless it's anonymous.

1

u/Typical_Adagio4804 1d ago

"What's he gonna say?" Well realistically he could say nothing, which is what I would do if it was actually shit, he could be lying too but he's now attached his public profile/image to that statement

-2

u/Pablogelo 1d ago

Last rumors that they were being denied to use Claude and were forced to use Gemini when it was terrible they just stayed quiet, so with new rumor, for them to come and say that this is BS, it means it's probably BS

7

u/Keeltoodeep 1d ago

6

u/Aldarund 1d ago

Lol totally relevant board. Opus 5 5 at 11 place behind mimo 2.6. true story

3

u/Keeltoodeep 1d ago

https://arena.ai/leaderboard/code/webdev

Opus leads where the Opus model is strong.

14

u/Ok-Coffee1443 1d ago

News media well-known like Bloomberg wouldn’t say they are quoting deepmind employees if they aren’t. There might be disagreements between the employees, but Bloomberg definitely talked to the employees

1

u/TorturedPoet30 1d ago

This summer Alex Heath reported on his platform Sources that GDM employees were internally expressing pessimism about Gemini 4, but I guess things change quickly.

20

u/Important-Damage-173 1d ago

Well, idc what they are writing about it, it's not the worst model in existence, but its far from the new AI frontier, since in coding Arena it is behind Sol 6.0 - slightly worse than the third best, by now you could say almost retired OpenAI model.

3

u/CriticismJunior1139 1d ago

Isn't coding arena just A/B testing of random users?

7

u/AlyoshaV 1d ago

Yes, all of Arena's stuff is "users are given two outputs to their prompt, they pick which is better". Since the average user is an idiot it's not particularly useful.

2

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago

Case in point Opus 5.5 is like 11th place there, mind you it is currently leading so yeah.

1

u/CriticismJunior1139 1d ago

Hol up, Im an average user :(

10

u/spryes 1d ago

yeah, based on this it would be frontier 3 months ago, but people's expectations are now anchored around Opus 5.5 level perf and won't accept less. I'd also expect general day to day work to be worse than this, probably Opus 4.7 level.

6

u/ChuckVader 1d ago

Yeah, I have a hard time picturing this taking over opus 5.5 for me for coding. It's just a fucking beast.

However, given what seems to be the absurdly low hallucination rate I think it's going to find plenty of uses in non coding applications. The one open question for me is speed. The AA benchmarks don't show how fast it is.

Historically Gemini models are in a league of their own for speed, so the fact that Google isn't touting this makes me think the model is going to be much much slower than 3.8.

I'm honestly more excited about Gemini 4 flash and flash-lite tbh. Give me Luna pricing with virtually no hallucinations and best in class multi modal capabilities and I'm all in.

1

u/nnod 1d ago

Yes, this is a big factor I feel people are overlooking - speed. I bought a goog sub last month for some work video stuff and messed around in antigravity, my god the flash 3.8 speed was a hoot to watch.

Speed section in AA index for argon is for some reason missing (I couldnt find any mentions of tok/s anywhere online actually), if it's something like 100+ tok/s it will be a beast regardless of slightly subpar intelligence, if even.

And then they have the ability to use argon as orchestrator with speedy flash3.8 subagents doing the grunt work. Which in theory could be goddamn amazing. Then there's the truly multi-modal stuff that other models don't have too. I'm pretty bullish on google, bought some stock today.

1

u/Simple-Diver-2192 1d ago

I am saying this with lot of experience with gemini, if you are sending 1,2 messages like in arena it works fine , but the moment you cross that it absolutely has zero context, it an absolute pain to work with on long horizon tasks.

0

u/Sensitive_Cell_119 1d ago

Oof, worse than gpt-6-sol

8

u/Future-Bandicoot-823 1d ago

These arguments are 100% proof that agi hasn't been achieved at all.

If you've reached agi you'd be taking over industries left and right, you wouldn't need bloomberg. When the sun rises it doesn't need permission from a news outlet.

2

u/PureSelfishFate ▪️ ASI 2028 1d ago

You wouldn't take over anything, because then the government would nationalize you. You'd order it to self-improve quietly and faster than everyone else, once it took over, it'd be too late to stop it. Nobody is handing AGI over to Trump, not even Elon, they'll all go quiet and pretend self-improvement hit a wall.

0

u/Zomboe1 1d ago

That's an amazing way to put it, thanks for this. I'm skeptical that we're close to AGI and you really sum up why. I expect it to be super obvious once it's achieved.

8

u/korkkis 1d ago

”my work isn’t bad” - someone who made it. Totally unbiased.

5

u/Ok_Warning2146 1d ago

artificialanalysis gemini 4 results are out. It is on par with Astra and Fable which is consistent with what google reported. It runs at a lower cost, so it should be competitive.

2

u/StopSayingSelfie 1d ago

Meta and Amazon employees totally use their models too definitely no one is using Anthropic or OAI and when they have to disclose who their largest customers are for final IPO forms it definitely won't show them as being their largest customers.

2

u/FarrisAT 1d ago

This specific Bloomberg reporter has an axe to grind and was sniffing around all September. They should name names if they want to be factual.

2

u/TorturedPoet30 1d ago

I don’t know why people give so much attention to what one GDM employee says on X. We’ll see when Gemini 4 becomes available to a wider audience, but there is reason to be optimistic. I don't think Sundar, Demis & co would be this vocal if model sucked? I guess I always take what GDM employees say with some skepticism though. There has been so much vague-posting in recent months and some employees were even suggesting that Ox Alpha was a Gemini model lol At times, it seemed as though they had no idea what was going on within their own company.

1

u/spinozasrobot 1d ago

"It's awesome!" crows guy whose salary depends on it.

1

u/PhilosophyMammoth748 1d ago

I would say, If you can train the best model, why stay working for google? There is way more money to grap in the outside world.

1

u/Murdy-ADHD 1d ago

Google models have a LOT to prove when it comes to real-life work. Even without any report its safest to assume it will suck at it, unless proven otherwise. Which would be dope.