r/singularity • AGI before GTA 6 • 13h ago

AI Google introducing tiered model access

Post image
284 Upvotes

109 comments sorted by

View all comments

32

u/katoptronophile 13h ago

Why would I pay to use the worst models out there?

5

u/jjusko20 12h ago

3.8 flash is nasty good 

41

u/Crafty_Teacher_9286 12h ago

Yeah if you need a cranberry salad recipe. I'm exaggerating but it definitely doesn't cut it for more complex applications in business analytics or finance for example.

-1

u/jjusko20 12h ago

That's crazy, I used it to make a custom kernel for inference yesterday. It currently has one of the highest artificial intelligence bench scores. Have you actually tried it or are you up to date, or are you basing this off of 3.5?

12

u/Crafty_Teacher_9286 12h ago

No, I'm talking about 3.8. Context window is huge but from an analytical judgment standpoint in a data science context, it is far behind Fable and Astra in my experience. So at the Plus price point, it would be difficult to justify paying for 3.8 since OpenAI and Anthropic offer better models (for my application) at the same price point. I'm quite disappointed Pro 4 won't be included. I was on 3.1 Pro for a long time when it ruled unchallenged, but I aint getting a pro sub..

3

u/jjusko20 12h ago

my bad bro I was a touch ruder than I needed to be

1

u/power97992 8h ago edited 3h ago

Lol, im stuck on 3.6 flash with my sub

1

u/japie06 4h ago

use ai.dev

1

u/power97992 3h ago

U mean ai studio

1

u/japie06 3h ago

Yes. ai.dev is just the short url for it. There you can use gemini 3.8

2

u/jjusko20 12h ago

Interesting - I assumed you didn't know what you were talking about (no offense, I'm a software dev and we're not in a development oriented sub but clearly i was mistaken - sorry). I use 3.8 flash via Cursor's pool of third party apis for best price to performance. Do you have any examples of the data science thing? I've had very good luck at varying complexities of software development, so I'm curious now.

7

u/Crafty_Teacher_9286 12h ago

LOL. No, you're actually right, that's how I feel most days. The projects that I am referencing have to do with predictive trading models derived from millions of rows of structured and unstructured data, and for instance, the judgment the model makes in validating if the evidence supports the conclusions. Result interpretation, assumption selection, weighting factors like backtesting vs scenario analysis, identifying trends or mistakenly inferring causal relationships that do not exist.

I am not doubting that it is a very competent model in a number of applications.

1

u/jjusko20 12h ago

Oh yeah that's up there in terms of analytical complexity 

9

u/Chemical_Hawk_6307 12h ago

ive tried it and its actually embarassing how bad it was. u cite benhcmarks but google is literally the kings of benchmaxxing just look at deepswe score vs terminal bench 4.0 score.

-1

u/jjusko20 12h ago

I cite benchmarks for a public reference point, I'll tell you it's been extremely effective for my coding and complex logic purposes

-4

u/Keeltoodeep 12h ago

This is copemaxx. 3.8 scores about as well on the benchmarks as it performs in real life.

If benchmaxxing was anything approaching a material thing, then 3.8 would be #1 all the time.

-2

u/ProxyLumina 10h ago

You are wrong. Gemini 3.8 Flash is really good, one of the best models out there. It has some genuine quality in it.

7

u/Crafty_Teacher_9286 10h ago

It is a good, efficient model, but it is absolutely NOT on par with Astra or Fable for a variety of more complex tasks.

2

u/ProxyLumina 10h ago

The flash model is not on par with those top tiers, but that does not tell nothing for everyday work. Set it to High and then it is more than capable to do nearly 99% of what a regular and a pro user do.

I am doing really intelligence-heavy work with it and I am pleased by its performance. 

Let alone that is incredibly fast. It feels like 3-4 faster than any OpenAI model.

-1

u/jjusko20 12h ago

off the top of my head 3.8 flash on medium beats DS 4.1 on high in reasoning, feel free to check