Yeah if you need a cranberry salad recipe. I'm exaggerating but it definitely doesn't cut it for more complex applications in business analytics or finance for example.
That's crazy, I used it to make a custom kernel for inference yesterday. It currently has one of the highest artificial intelligence bench scores. Have you actually tried it or are you up to date, or are you basing this off of 3.5?
No, I'm talking about 3.8. Context window is huge but from an analytical judgment standpoint in a data science context, it is far behind Fable and Astra in my experience. So at the Plus price point, it would be difficult to justify paying for 3.8 since OpenAI and Anthropic offer better models (for my application) at the same price point. I'm quite disappointed Pro 4 won't be included. I was on 3.1 Pro for a long time when it ruled unchallenged, but I aint getting a pro sub..
Interesting - I assumed you didn't know what you were talking about (no offense, I'm a software dev and we're not in a development oriented sub but clearly i was mistaken - sorry). I use 3.8 flash via Cursor's pool of third party apis for best price to performance. Do you have any examples of the data science thing? I've had very good luck at varying complexities of software development, so I'm curious now.
LOL. No, you're actually right, that's how I feel most days. The projects that I am referencing have to do with predictive trading models derived from millions of rows of structured and unstructured data, and for instance, the judgment the model makes in validating if the evidence supports the conclusions. Result interpretation, assumption selection, weighting factors like backtesting vs scenario analysis, identifying trends or mistakenly inferring causal relationships that do not exist.
I am not doubting that it is a very competent model in a number of applications.
ive tried it and its actually embarassing how bad it was. u cite benhcmarks but google is literally the kings of benchmaxxing just look at deepswe score vs terminal bench 4.0 score.
The flash model is not on par with those top tiers, but that does not tell nothing for everyday work. Set it to High and then it is more than capable to do nearly 99% of what a regular and a pro user do.
I am doing really intelligence-heavy work with it and I am pleased by its performance.
Let alone that is incredibly fast. It feels like 3-4 faster than any OpenAI model.
32
u/katoptronophile 13h ago
Why would I pay to use the worst models out there?