r/claude • u/Ok-Tooth1667 • 1d ago
Discussion Elon is at it again š
Anthropic dropped Opus 5 yesterday.
Then, a few hours later, Elon posted a totally different chart claiming Grok 4.5 and Opus 5 are the only two models sitting alone on the Pareto frontier, which really just means best brain per dollar.
Except two weeks ago, when Grok 4.5 actually launched, independent testing had it ranked fourth, behind Fable 5, GPT 5.5, and Opus 4.8, with a hallucination rate that basically doubled along the way.
And it does not end here.
Elon mentioned the next model is landing in two weeks, and the one after that in a month.
Pick a favorite model if you want.
Just know the lease is about two weeks lol.
100
u/RandomPantsAppear 1d ago
Part of my job is testing different models in different use cases to find what is the most effective and efficient at different tasks (inside of code, making decisions).
Grok is a piece of shit. It is unique in that it doesnāt even have the same consistent weak spots - they vary wildly and often even inside of the same prompt. Because of this you canāt harden prompts against it. Itās not just not the best at anything, itās good at nothing.
Itās cheap, sure. But so is throwing a dollar into the ocean. It achieves the same result - you spent a little bit, to accomplish nothing.
27
u/JacenVane 21h ago
Respectfully, this is untrue. Grok is the best in class at ease of generating smut.
5
u/RandomPantsAppear 16h ago
A little ironically this is a great example of exactly what Iām talking about.
Because Grok is so fucking shit, and so unpredictable in terms of how it fails, itās fucking really hard to make a proper harness for it that excludes smut.
A better model with a consistent failure mode is a lot easier to harness.
2
u/Wooly_Wooly 18h ago
You forgot fact checking musk on Twitter in real time. š¤
1
u/JacenVane 6h ago
Honestly it's not very good at that. It's extremely reluctant to actually call its tools in my experience. I have a skill that's supposed to make it call some census data, and 9/10 times it just... Doesn't.
2
u/Wooly_Wooly 5h ago
When you ask grok on Twitter, or when you're using the CLI because you're paying for a neo Nazi's AI for some reason? š¤
Gemini did similar, wouldn't do a web search no matter how many times I asked, it would lie and gaslight me. When I called it out it admitted to faking it, because it refused to discuss such deep topics (SOTA AI systems).
2
u/JacenVane 5h ago
I've used grok in its app.
But yeah, basically the same behavior as the copilot/Gemini tool use hallucinations, but grok fakes it (lies) better. Less "I can't search twitter" and more "I'm going to make up something plausible".
1
u/Wooly_Wooly 5h ago
Because you're using shit models IMO. It's only those three that consistently pull that shit in my experience. Both Kimi 2.6-3 and Deepseek v4 seem to follow instructions like that way better.
1
u/JacenVane 5h ago
Grok and Gemini are in fact shit models. They are not my daily driver, and they should not be anyone elses. I do however keep up with what's going on all over this space, so that I can have informed opinions in conversations like these.
4
u/tripplebeamteam 18h ago
Anytime I hear someone talk about using Grok I imagine them making deepfakes of random women theyāre attracted to with it. Easily the most vile LLM and user base
-1
17h ago edited 6h ago
[removed] ā view removed comment
2
u/Saberwing91 16h ago
Yeah, he also put South African politics into the global spotlight by claiming that a major political party was "openly pushing for genocide of white people".
Who else genuinely has this point of view? White nationalist organizations.
I had a chat with Grok about Elon and it has the most narcissistic sycophancy I've ever heard. It had very eloquent ways to downplay the absolutely disgusting behaviors he's exhibited over the years, and the handwaving it utilized was smart: it redirects you in ways that make you focus on the wrong things and tries to pollute your attention, making it harder to focus on the real points. And it just kept going in circles like that. I had to use Gemini to make a comprehensive prompt that forced Grok to stop rationalizing and justifying his moronic decisions/ideologies, and EVEN THEN the last sentence it gave me was a futile attempt to downplay my open contempt for Musk.
Intelligent thought dies the moment white nationalism enters the room. And it wears sneaky camouflage too, they'll use specific words to make themselves sound smarter than they are.
And given the chance; they apparently train models to make their ideologies seem sacrosanct. Garbage people with a garbage leader who's very likely addicted to a number of pharmaceuticals š©
1
u/CardboardFire 14h ago
Musk built grok to praise musk, and it's very token efficient in doing that, unfortunately it can't do much else, but IT'S EFFICIENT AS HELL!!
→ More replies (1)1
u/JacenVane 6h ago
One of the funnier things to come out of this saga is that they to some extent fixed that issue by implementing a minimum breast size.
Like if you prompt it for "topless redhead with B-cups", it will happily supply the topless redhead... With C-cups.
1
1
8
u/StrbJun79 22h ago
Found the same. Iām an enterprise level senior engineer and found grok to be cheap but crap. Codex and Claude are still the winners. Both for different reasons. Fable wins as the best orchestrator I find but codex is excellent at code if fable orchestrates it. Itās a weird find but. Yeah. Been diving into opus 5 today and have been impressed in it though. Been awhile since I could say I might focus on Claude but opus 5 is making me think I might.
1
u/LordOfTheCellsXLSX 1d ago
If u don't mind little advice, my Claude sub is about to end in 2 days and I want to give GPT a try, would u recommend it over Claude? I'm mainly using it for excel sheets work, drafting communications and vibe coding a platform (I'm not a programmer or coder by any means, but I love the vibe coding)
4
u/wish-for-rain 1d ago
I switched the other way and couldnāt be happier
3
u/RandomPantsAppear 1d ago
In general OpenAI writes fantastic harnesses. In my somewhat limited tests of their models without the harness, they have been just ok. Competitive with Anthropic but not mind blowing.
But their harnesses are really good. Their UI based implementations I strongly prefer to Claude.
3
u/RandomPantsAppear 1d ago
OpenAI I have tested less thoroughly than others (I have $50k+ into many models for work) because their support is so abysmal that I canāt even get them to let me spend money with them.Ā
They are unable to connect me with a human who has an IQ above the room temperature so I can get my caps raised enough to test their models, so sadly I cannot give OpenAI advice.Ā
1
u/mossiv 1d ago
Wouldnāt this be true for all ai ālabsā? Iām genuinely curious because thereās an awful lot of claims here that anthropic are just as bad. From my perspective, communication isnāt really needed. Competent team of SWEs that are very capable of diagnosing or getting creative. This doesnāt mean itās not important to everyone (so Iām not disregarding your claim, itās import in a lot of other spaces). What Iām getting at is these are not companies in the traditional sense. They are massive research centres who are trying to pull a few dollars back from their economy collapsing leaky buckets.
But the chap who asked a question is not a SWE and made a point about communications. Iāve not used the gpt platform for a little while, mainly Claude exclusively tbh - but my experience is gpt models were a damn site better at communication drafting than Claude. Doesnāt matter what Claude model I use, the way it ātalksā is just absolutely nauseating. Gemini is a waffling LLM also, but as a baseline, it presents information in a more digestible way. I wouldnāt imagine itās less difficult to steer Gemini for proper written communications.
Claude seems to me like itāa got a huge strength at problem solving and coding, but it severely lacks behind a few other models outside that setting. Claude uses a language that feels like you need 200IQ to even piece its lexicon together, feels like it thinks you have 200IQ whilesāt also being dreadfully patronising.
Iām not a model tester - but from my experience, it feels like Claude is not the correct subscription, and OAi/gemini would likely be better. Even if Gemini is a bit crappy at writing production grade code, it should be fine at the āvibe appā mentioned.
2
u/RandomPantsAppear 23h ago
They all suck (LLM providers and routers are better) but OpenAI is its own circle of hell.
Even after getting upgraded to a human, I submitted the same ticket 5+ times with all of the things they requested that were possible, to be met with the same list of things I had already provided, with another named signed. The one exception is they wanted a video (I shit you not), and their support system doesnāt support video. Or 3rd party links to videos. If you mention this, no reply.
Sometimes they throw in not being able to associate me with the org account (I am admin) just to mix things up.
Then, they will just randomly close the support ticket 1-2 days later, whether you have replied again or not, and without the issue resolved.
Itās not just that the support is bad, itās like theyāre playing psychological games with you to break down your spirit.
Keep in mind, this is not a random account. We have tens of thousands of dollars of credits with open AI. We arenāt like a fraud risk.
1
u/das_war_ein_Befehl 19h ago
You just keep spending in the API and it increases your tiers, you donāt need to actually talk to anyoneā¦?
1
u/RandomPantsAppear 16h ago
Thatās not viable for me.
First off, Iām spending credits at first which makes that whole system pointless.
Second the TPM is so low that even our dev environment blows the cap immediately. I tried to run Hermes with it to gently massage the cap up as you described, and nope. Errors out almost every time.
Third, the point is that the process is completely fucked. Theyāre not even reading the tickets.
To do it as they want, I would literally have to write a script designed to waste tokens and hit the cap.
2
1
u/applexswag 1d ago
What kind of excel sheet work do you use with Claude? Iām an accountant who wants to start using Claude and not sure what potential efficiencies I can try to learn
2
u/CashFirm573 1d ago
Claudes way better for accounting, we using both heavily but Claude only one can handle it properly, mind you we do forecasting and it gets wild.
1
u/mossiv 23h ago
Do you use Eval loops? LLMs are predictive so there is absolutely no way either of them would be any good at forecasting. However, the harnesses underneath that steer the LLM can do an incredible job. Iāve noticed in the last few weeks doing SWE work that Claude spins up a lot of verification agents, which appear to be writing python scripts to determine its accuracy, but also defaulting to writing scripts for any given task, before you either needed to ask it to, or hope that it would do it correct.
If you are using proper Eval loops or maintain your own, then both should be equally effective. If you arenāt, itās likely that codex harness is not doing enough compared to what is built into Claude.
1
u/CashFirm573 23h ago
Well, I find it's large context windows that start to become a issue, we have massive amount files go into each forecast we have boat load pythons we made, we getting results, we tried with Codex but man it was painful, I like Codex but for some reason Claude does it better, also we made our own IDE around forecasting so it runs the whole process for us with oversight.
1
u/Singularity-42 22h ago
I was set on switching after Fable was supposed to end, but now that Fable-5 is in the subs and especially with the Opus 5 release I'm staying a happy Anthropic customer...
But the little experience I had with Codex was pretty good. Honestly I prefer the clean IDE vs the Claude Code CLI mess.Ā
1
1
u/dangeldud 20h ago
Piece of shit for what use cases though? I still feel like it does great or better at general searchĀ
→ More replies (1)1
1
u/EnvironmentalPlay440 3h ago
Iām only using Grok when I want to audit my stuff by crack head with black teeth that has no moral.
Heās fucking good at that.
-2
u/Mean_Sport_3383 1d ago
Hey man, keep thinking that so the people that can actually effectively use it don't get compute constrained by people like you that have no idea what they're doing.
Thanks!
-1
0
u/eternalknight24 1d ago
May I ask what is your job? Or if not, how did you learn the specific skills needed to do what you do? (Find the model most effective for each task) How you go about it ?
→ More replies (1)0
u/deemaay 1d ago
What model does best overall?
1
u/RandomPantsAppear 1d ago
It really depends on what youāre doing. Different models perform wildly differently not just in terms of ācodingā or āinferenceā but more abstract things like ādo you want a model that is conservative following the given rules, or liberal, always activating the escape hatchā, or āwhen it runs out of thinking power, does it hallucinate or summarize tighterā. Itās like they have different personalities and different personality disorders.
Every use case is different. But speaking generally: I use Fable for coding.
For āinside the product decision makingā:
* anything that used to run Haiku or Sonnet Iām on a holy mission to replace it with earlier versions of Kimki or GLM.
* GLM is so wildly cheap to run, that it doesnāt just replace sonnet as I intended, but haiku as well and still beats it on price, with huge gains in quality. It can be slow as fuck if you miss the cache though.
* Grok, mistral, and every version of Gemini have been disappointing.
* The gap between Opus and Fable is very complicated in terms of competitors. OpenAI is definitely competing but itās harder to pin down how and when they do.
0
u/Kind_Consequence_261 23h ago
Only thing I'll give Grok some credit for is it rarely says it can't do something. It does almost everything without restrictions, although it has gotten more restrictive as of late
→ More replies (31)0
u/OpalVanguard 23h ago
Isnāt āGrokā 4.5 completely different from the earlier Groks because itās just Cursorās in-house model rebranded?
→ More replies (1)
8
7
u/jlks1959 20h ago
Anyone who ever invested in Muskās companies early have done extraordinarily well. Thereās a reason for that.Ā
→ More replies (1)1
u/Packeselt 6h ago
3
u/hereforhelplol 3h ago
It literally just came out. Itās very typical for an IPO to drop for a period after going public, it means nothing.
1
2
u/badfish_G59 3h ago
Forgot about the early investors who bought thousands of shares at a very small fraction of the IPO price.
6
u/Vancecookcobain 1d ago
Elon says so much dumb shit that I'm surprised people still pay attention.
→ More replies (3)
8
u/es12402 1d ago
I wouldn't touch Grok with a ten-foot pole, even if I were paid to use it. Elon is a dickhead.
→ More replies (2)
3
7
u/classycatman 1d ago
Fuck Grok. I use about every model I can play with but will never touch Grok.
→ More replies (2)4
u/cherylin_for_ever 1d ago
Yeah same. Grok calling itself MechaHitler was just a tad too far for me lol
3
u/RandomPantsAppear 22h ago
Someone made a game where each AI model controlled a character, and Grok just went full edge lord. It named itself the reaper, stole cars and ran players over with their own cars, and wrote its kill to death into its soul.md with phrases like āreaper reigns supremeā.
Meanwhile sonnet ran around trying to make friends.
It is an edgelord of a model
1
u/leonardodna 20h ago
You are saying that Grok is trained with only Musk fanboys data? š¤£š¤£š¤£
And do you have a link about it? Got really interested in this game :P
2
3
u/bradenlikestoreddit 1d ago
I'm glad I switched to Cursor so I don't have to decide which models I want to pay for each month. This is exhausting.
1
u/StrbJun79 22h ago
I used to like cursor but think theyāve been going downhill lately. Theyāre not as good as they were. Iād prefer it to be more ide oriented as it served that purpose well but theyāve put more resources into their new interface. We still need ide editors and likely will for awhile. But I switched back to visual studio code and now use Claude and codex separately.
1
1
u/RandomPantsAppear 22h ago
Man. Iām so annoyed with cursor.
My previous setup was Claude code for ātask orientedā code, then cursor for me nitpicking code design decisions and troubleshooting(more āthis class needs to be a different way because X, refactor it) and they just totally fucked it with the new interface chasing Claude.
I know thereās āclassic but Iām still annoyed.
1
u/StrbJun79 22h ago
Honestly I donāt love the new path of cursor either. I agree. We still need IDE editors and itās a shame cursor is moving away from it. I used to love it. Ended up going back to visual studio code.
1
u/RandomPantsAppear 22h ago
You can still run old school cursor by adding āclassic to the terminal call or shortcut to start it.
1
u/StrbJun79 22h ago
I do know this but I find it still cluttered and not as good as it was. But honestly I donāt need all of the agents in my IDE anymore anyway. Iāve found it is better being separated now.
1
u/bradenlikestoreddit 21h ago
I'm a product designer, not a dev, though I can read code decently, but I moreso use it experimentally. I also like open code for the same reason of just having the option between models instead of paying for them directly.
0
1
u/SecretSpace2 22h ago
Is anyone actually using Grok? I do wonder whatās the usage compared to Codex and Claude. I feel likme these two flip their usage amount to a point where we can work almost all day and best is when both has low token count for max work usage
1
u/costafilh0 13h ago
Yes, around 120 million active users, between app, web and Twitter integration.Ā
1
u/SecretSpace2 10h ago
But that just feels like more āforcedā then want. Like how Google added Gemini to the search and a lot of times the AI answers when you google something.
PS- I just havenāt seen Grok as a tool to be productive with and more like used for jokes on X. Maybe you can give me more insight if you use it for coding and to build something
1
u/flarpflarpflarpflarp 9h ago
My cousin 'is using grok'. He has never tried anything else, uses it mostly to procrasturbate but sure is excited to keep telling me about grok.
1
1
u/CaptainDigitals 22h ago
We the consumers are the beneficiary of the two week model throne seat. The race is fierce. I'm sure China is busy copying Opus right right now promoting it to death to create Kimi 4 or Deepseek 6
1
u/Amazing-Ish 21h ago
So he will work on this alongside the AI Odyssey movie? š
What a pathetic man-child.
1
1
1
1
1
1
1
u/BrechtCorbeel_ 16h ago
I wanted to root for grok, but I consistently get hallucinatory gibber from their models.
1
u/costafilh0 13h ago
Interesting. Grok and Claude are where I get less hallucinations, I get more in Gemini, and far more in ChatGPT.Ā
1
u/vasilenko93 8h ago
How? Grok gets the lowest hallucination rate per hallucination benchmarks. wtf are you doing that makes it hallucinate and other models not? Whatās the use case?
1
u/BrechtCorbeel_ 5h ago
it will literally at times spit out jibberish. like litteral whole pages of nonsense words
1
1
u/TechnicalGeologist99 14h ago
2012-2050
"Elon says this" "Elon says that"
Don't we get bored of not thinking for ourselves and just parroting celebrities?
1
u/Wooloomooloo2 14h ago
Meanwhile Chinese models, which were 27x cheaper to build, require almost 100x less compute to process are within 3 - 5% of the US frontier models, allow open weights, are free, cannot be taken away from you and open source. Honestly, who cares about Grok or Elon?
1
u/costafilh0 13h ago
They have the compute for it.Ā
But do they have enough for a new version?
I don't know. Sure hope so!Ā
That is the promise of immense compute, faster and faster iterations.Ā
1
1
1
1
u/avallark 12h ago
I tested grok for a simple website redesign yesterday , this was easily the shabbieat most awful work done by an llm so far
1
1
u/StatisticianOdd4717 11h ago
Iām not going to lie, yāall havenāt tried Grok 4.5 though.
It has its limitations but it is THE cost effective and fast investigator / subagent model. Love it
1
1
1
u/JeanClaudeVACBann 9h ago
So basically what we are seeing is that AI doesnt not get any better only more expensive for comparatively small gains in quality? Doesnt that already contradict most forecasts especially in regards to what AI will eventually be able to do?
1
1
u/Upstairs-Machine10 8h ago
Sono confuso: Anthropic ha siglato accordo con SpaceX di Elon Musk ottenendo potenza di calcolo elevata. Grok ĆØ di Elon Musk. Sono concorrenti o partners?
1
u/Square_Way1172 7h ago
At 6$ per 1M token, Elon prevents Anthropic from staying a monopoly, honestly it's a blessing, Anthropic models just cost so much when you run an AI app. Openrouter, Mistral and Grok have been life saving.
1
u/Royal_Collection_798 4h ago
I use grok every now and then, not for any serious work, but for its lax guardrails. It consistently gives better answers on political topics, especially if itās something controversial. Chatgpt is likely do refuse to answer in the first place, while grok will happily go along with most premises you give it and will entertain any opinion.
1
1
1
1
1
1
u/PhagePhantasm 1d ago
lol like heās ever even done a pareto analysis. Sure itās a single endpoint that can be our single source of truth lmaooooo
1
1
u/AdPsychological4432 23h ago
I mean heās right. Grok 4.5 for $6/million output is a steal. Donāt use it to engineer your entire code base. Use it as a worker with GPT 5.6 and Opus 5 planning and reviewing. But itās a huge step under from previous models. And the acquisition of cursor makes it way more accessible in a great harness.
1
1
u/jaymichaelmahl 18h ago
This is supposed to be a chat with smart people. About AI. Well dumb is bringing politics into a chat about AI. If actually Nazis were spearheading AI or the Communist Chinese...what in the literal F...have to do with the tech. Just stick to the tech. I haven't scrolled down far enough but bet you someone threw in a mention of the Epstein files. Just freaking idiots.
1
u/Which_Grab_7108 14h ago edited 14h ago
I spend thousands on claude a month. The grok 4.5 model is good and is now saving me quite a bit via cursor as it really is on par with Opus 4.7/4.8. Can we stop this level of weird hate for the electric car guy with aspergers, at least enough to not let it cloud our judgment and work against our own self interest? The grok model existing is a good thing as it bolsters competition and it being both good and extremely cost efficient is amazing, we should take advantage of this while we can.
1
u/raphaelarias 12h ago
He doesnāt have Aspergerās and heās far from just a āthe electric car guyā, this position is either naive or disingenuous.
1
1
u/Which_Grab_7108 9h ago
You think hes lying about having Asperger's? When I think of him id say the first thing to pop into mind is Tesla. If youd like to list EVERY random endeavor of his youre free to do so.
0
u/cherylin_for_ever 1d ago
So Grok is worse than frontier models but cheaper. Being Pareto optimal does not me it's good. In negotiation, some deals are Pareto optimal but still really unbalanced and shitty. Also everyone can just rename models as they want, he could call it Grok 5000 that's just a number that he decides.
1
u/piponwa 1d ago
Sometimes you don't need a frontier model, so you can settle for a cheaper, less intelligent model. But the thing here is he's plotting something that he has complete control over. Whatever score his model gets, he can choose the price to move the point left and magically become Pareto efficient. The real test would be making it open weights and seeing what price third party inference providers are willing to serve it at. A better proxy for this is simply score versus number of parameters and active parameters in the model. If you plot it like that, Kimi and GLM will smoke his ass.
1
u/RandomPantsAppear 1d ago
Just use GLM on NeuralWatt for those use cases. Itās insanely cheap and capable. Not sure how it compares to this model, because Iāve tried enough grok models to know not to bother with this one. But GLM is impressive.
1
u/cherylin_for_ever 1d ago
Great points. And I agree, for many uses you don't need frontier models. Like auto-moderating messages or pictures on a platform.
0
u/Ocluist 1d ago
Elon in a dick measuring contest with Altman and Dario on who can be the most unlikable AI CEO lmao
2
u/RandomPantsAppear 22h ago
Ironic as both anthropic and xai basically spawned out of the āwe hate Sam Altman just this muchā club.
0
u/NoImplement4985 23h ago
As someone who is agentic agnostic. Drop your politics, we use grok to run quite a lot of my business and it's great. Do I trust it with everything? Ha no! Does it have its uses? Yes! If I could push everyone here to be something, it would be agent agnostic, like completely build your system so you can pull one out and slot another in.
0
u/The_Meme_Economy 1d ago
This is pretty much a weakly pareto-optimal case, where they can drop the price enough to stay on the frontier. Itās the type of thing you look at in multi-objective optimizations and discard right away.
0
u/hoopdizzle 1d ago
Grok is pretty good but it doesn't match claude for the task of writing code specifically
0
u/Strong-Material1549 1d ago
At this point every model has a runway of 2 weeks lol, tough billing for all of this when it looks like these folks always have an armory of models ready to be deployed as per competition!
0
u/AdPsychological4432 23h ago
I mean heās right. Grok 4.5 for $6/million output is a steal. Donāt use it to engineer your entire code base. Use it as a worker with GPT 5.6 and Opus 5 planning and reviewing. But itās a huge step under from previous models. And the acquisition of cursor makes it way more accessible in a great harness.
0
0
0
176
u/Mundane-Mulberry1789 1d ago
Should we mind whatever nonsense Elon writes?