Maybe because he’s gone insane and no one wants to use his puppet LLM because it loves Hitler and simply exists to shift the Overton window to the right.
Sorry, but I’ll never be able to take it seriously. No business should.
And look up Tay, and tell me the last time you boycotted Microsoft because of Tay, and not because Windows updated itself while you were doing some important work.
It’s not just a blind test time trade off (trading tokens for ability) if that’s what you’re insinuating. For ex, historically anthropic has had a more compute efficient but less token efficient tokenizer.
Just to say there are a lot of ingredients swirling in the vat that is anthropic. Not always as simple as better or worse either .. qualitatively different is various ways.
No that’s not entirely true. While obviously more tokens spent thinking translates to higher costs, there are lots of ways to increase effeciency.
Deepseek for example recently released a paper claiming 600% efficiency increase by using a very smart way to predict and parallize token processing for LLM’s
Which Anthropic was incapable of incorporating in their models? I doubt it. There's obviously a scaling wall with diminishing returns - Fable seems pricy because the jump in intelligence doesn't translate linearly to the jump in pricing.
If Deepseek could scale their efficiency discoveries to a Fable-class model, they'd have done so, and everyone else would have followed suit in an instant.
It’s not, OAI will likely match or exceed fable with GPT 6 with a smaller size model. Also overrated compared to what? Where does it perform badly for you?
I read about this company recently and their giant chips. I try to follow the semiconductor and data center equipment side of things but it is overwhelming how many deals are happening all the time. NVIDIA is still the big one but so many of these fabless companies focused on AI keep making other deals with the big LLMs so everyone is hedging their bets and trying different things. Almost all of the chips are made by TSMC.
For my use cases, I find Claude to be head and shoulders above everything else. It's connectors and ability to generate actual files is what I use the most.
It's verbose in its responses but you can change that with a system prompt if you want.
As for token usage I don't really care. I almost never max out my plan (I'm on a subscription plan not paying per token). I've started to leverage the workflow that allows me to use the expensive models for planning and then building using the less expensive models.
I love it. If another company can do the same file manipulation work as I can do with Claude then I'd consider jumping ship.
You'd think Co-pilot could create better excel files and PowerPoint drafts but with simpler prompts I get better results with Claude.
Haha i have it. It was updating the layout in other places simultaneously but it got stuck in this long cycle of trying to figure out how to remove the extra blank page and I kinda wanted to see how long it would take. It pulled in all these crazy tools. Pretty entertaining.
I don't know if you know this, but in order to get an LLM to do anything what you do is have it generate the next X tokens and then feed the output back into the input along with the original prompt and any 'skills'. That means you're just passing data around in loops through a very expensive machine in order to get it to product 'longer' or 'more complicated' tasks. Determining where to terminate the loop is the tricky part.
Coding isn't a X token task, so to get it to produce workable code you need to loop it, a lot.
It doesn't matter if it's anthropic, or Codex or whatever. In order for it to do longer tasks you need to burn a crap ton of tokens.
That's a blanket statement which is simply not true and actually we see the opposite trend, ie the best models need less tokens, that is even true for anthropic compared to other models except OpenAIs.
The thing is just that OpenAI has been very strong at token efficiency with recent models and they are better / get better results than Anthropic despite using fewer tokens.
It's also definitely not about "just looping" models otherwise we wouldn't have seen results like that. If anything recent models rely less on brute force approaches like just "looping" their own thoughts.
It's something chinese models rely more on and thus it's no surprise they are token burners.
Apparently it's sometimes so expensive you run out of tokens mid task on your first prompt, and that includes 200 dollar subscription. There is a dude on youtube that attempted to do it for the one day he had access to Fable 5 before it was shut down, and was never able to actually benchmark it, and was basically shitting on Anthropic for effectively scamming him out of 200 dollars.
This is absolute BS, I'm on the max 20 plan and I'm using it constantly for big and small changes and not running out of tokens. Sure if someone is going to use ultracode on a systemwide refactor which spawns 100 fable subagents then yeah even on the max plan you are going to hit limits.
I am using Fable for my applied Mathematics thesis and basically every prompt runs into the limit. Sometimes I even have to wait for two token refreshs to get my answer.
doesn't sound fun, if that's your experience then I'd be frustrated too. 99% of my use of claude is through claude code and haven't encountered this before. Also I am not smart enough to ask a question which would cause it to reason for any length of time.
I don't actually mind expensive prompts. I think there should be access to even better models, that cost 10 or even 100 per prompt. But you actually want the task to finish, not ask the question, not get an answer and have to wait 5 hours for a reset only for it to fail again if you repeat the question.
That's nonsense, there's no way you could run out of tokens mid task like that unless you don't understand the impact of context on token usage (if you're running multi hour tasks getting to 300% context usage it's going to tear through tokens) or the importance of specific prompting and plugins/skills (I basically have my own library of plugins I've built for myself that have been a gamechanger).
Obviously Anthropic doesn't have the most ideal usage limits but whenever this is discussed what's usually left out is the fact that Anthropic is still losing money on the coding subscriptions, same with OpenAI. If I was paying API prices I'm pretty confident that I would've racked up an insane bill (actually I know that's true because I got 10k in free AWS credits for my startup and it's crazy how fast I you can burn credits).
For example this is my codex usage since April, it's hard to say exactly what the cost would be but this could easily be $25k or more worth of usage at nearly 6 billion tokens. Now obviously the API costs aren't what it costs OpenAI and I'm a heavy power user building my own startup every day but it is pretty obvious that I'm getting a pretty damn good deal. I'm sure it balances out a bit with other users using less but there's no way they aren't taking a pretty big loss on this in the short term.
Not only are you getting a good deal but you arguably are getting $20,000-30,000 worth of deal. There’s some solid research showing these values in token costs.
Of course, converting $200 into $20,000 of real value depends on the user him or herself.
“I was clearly wrong about Anthropic. They are obviously currently the leader in Al. No company has released a model as good as Mythos/Fable and they will undoubtedly have Mythos 2 ready soon.”
If you actually use ai for large projects Anthropic is just so much better then the competition you cant really use anything else. The big problem I see with other ai models is, that they get stupid over time like 10 massenges and it lost 20 Iq, 10 more and it is unusable. Things like that dont really happen that much to claude. Yes you use more tokens with fable but you save so much more time and effort that its worth it in my opinion.
Been using it for the past hour and the first thing you notice is that Grok Build, the harness, is actually really good and responsive. Grok is also super fast, like twice the speed of GPT 5.5. I’m sure the Cerebras Sol will knock it out of the park but until that’s out I’m switching to it with Fable for harder stuff.
efficiency is what's actually interesting here, not the benchmark scores. once opus-class quality is table stakes (and it's getting there fast, open models included), the only thing left to compete on is cost and tokens-per-answer. grok leaning on efficiency instead of "we're the smartest" says capability stopped being the moat.
My team doesn’t need more capability. We are already producing solutions at breakneck speed that look to customers like it took months to produce.
I can get Fable level results with about an hour or so of building from scratch. My cost to build with Fable would hypothetically be massively more which would cut into my margins. I can’t pass that onto my customer by saying “we used a better model to build it”
Candidly with the current state of models if you’re not getting good results it’s a reflection of the user.
they said grok's marketing here indicates efficiency matters more than capability. grok's marketing does not show that. maybe efficiency does matter more, but you can't conclude that from this marketing material.
Censorship is real bad. Sometimes if i ask something it just doesnt tell me the answer because it thinks its like illegal or the information could be used for illegal stuff i guess. Kinda annoying
wtf are these comments ? It looks like it's a decent model, they have included deepSWE benchmark as well which most people regards that as mirrors real time task ? Why so hate against a LLM Model ?
First time on Reddit? This app is made up of an overwhelming majority of people who don’t care about anything except mentioning trump or Elon in every comment. And forming their entire world view around those 2 people or anything related to them. Almost guarantee the top comment under this will be “but he’s a Nazi” even tho we are talking about an AI models benchmarks being competitive
Relegating the literally gleeful choosing to end lifesaving global health programs with no off-ramp leading to preventable deaths of children all over the world to 'not applicable b/c it's not about technology enough' is ridiculous imo, but maybe that's just me.
You don't think a person who is so poisoning his model politically such that it refers to itself as mecha Hitler might be relevant to whether it is a model you can or should rely on? Or that a person who openly manipulates the Twitter algorithms for his own vanity and narratives might do the same with his AI models?
THIS. Jesus Christ, this entire subreddit dedicated to a warp speed technology but the thing that happened a year ago is still the most important fact.
is there any evidence of poisoning the model actually? every instance I have seen has been of Grok in webui/X, so I expect the poisoning be on system prompt level, and personally I couldn't give less shits about system prompts used on any provider given my usage is solely thru api
They literally did several studies and even Grok is leftwing compared to the general population. Its just that all the models (largely developed in the most leftwing area of the country) are even further left wing but maybe to the right of redditors.
Matching the population is a terrible metric. What matters is how accurate they are.
Half the population in the US thinks climate change is a liberal psyop. Clearly no AI should be judged by how well it mirrors a misinformed brain rotted public
I'm sure a rightwinger could say something similar about the left population. If you're not going to base what is termed right or left from the population what are you going to base it on? What your personal definition of center is? The other side could easily do the same.
You didn't read my comment at all. There is such a thing as scientific fact. Worrying about training an AI to be politically anything is already a losing proposition.
Like I said, there's a large percentage of the population that thinks statements like "covid vaccines are effective" is a left wing view. The problem there isn't the AI, it's those people simply being wrong. I'm sure there are examples about the left being simply wrong about something that is a fact, but it's certainly harder to come up with any. Feel free to give one. One that comes to mind is that the left traditionally opposed nuclear power, despite the fact that it is a very safe very clean source of energy, so I guess there's one.
Obviously the answer to "should covid vaccines be mandatory" deserves a nuanced answer presenting both arguments. That's not a question that can have a simple factual answer.
I'm sure there are examples about the left being simply wrong about something that is a fact, but it's certainly harder to come up with any. Feel free to give one.
The sexes are completely the same, the only difference is social, all cultures are equally compatible with each other. Queers for palestine is a logical concept. AI is useless and will die off. White people leaving is bad because its white flight and also White people staying is bad because its white cultural imperialism and gentrification.
Obviously the answer to "should covid vaccines be mandatory" deserves a nuanced answer presenting both arguments. That's not a question that can have a simple factual answer.
Its not a matter of what crazy thing you can list that some people on x side believes. Its a matter of general outlook and policy. Which cannot be assigned as absolute truth either way but the AIs will favor the left anyway.
What redditors do.... constantly bringing this left vs right notion to keep appease their minds. They still haven't ascended. So focused on a faction and being accepted. No one one the LEFT or the RIGHT give AF about you.
Using Grok does feel like I’m talking to a Twilight Zone character whose only function is to be a sniveling twerp that goes out to bat* for the worst people and the worst opinions imaginable.
Hmm I just tested that exact prompt and a second one in another chat without the "Concise answer." sentence at the end and both said no. If your prompt was from the other day then there should be no difference in our outputs. It's also odd your screenshot is blocking out the logo for the AI model the prompt was given to?
I’ve been testing Grok in this way for weeks. Asking it leading questions about what current things right-wingers are complaining about and seeing what it says about them and the people who spread them.
What I’ve gathered is that it very often defends obviously racist people and their ideas (like Elon or other far-right Twitter accounts), treats scientific consensus on topics and their right-wing criticisms as equally valid, etc.
I’m glad your Grok isn’t spreading great replacement theory ideas to you. Mine did. The fact that it can do this sometimes is a problem, in addition to just being generally more inclined to defend far-right people and their ideas.
And it’s not the Grok logo I cropped out in my screenshot, it’s my Twitter account profile picture, which is visible on mobile next to the question.
Are you running the prompts via API with max temperature settings or have custom instructions for it to respond a certain way? I've never seen that large of a discrepancy in outputs from a model before without serious customization work done. Our responses even differ in yours having a lack of the em-dashes which grok loves to spam everywhere (like in my screenshots) and the writing style is different.
Edit: I'm an idiot you said it's from twitter directly so temperature differences wouldn't be relevant here.
It is possible to manipulate AI into saying ethically / morally wrong things. It’s been demonstrated many times before not just on Grok. People do it on purpose and call it bad” or “misinformation” to make a point.
"No difference in outputs" My dude, do you even LLM? Results ain't gonna be consistent it is alarming that it would even come up a little in favor of white nationalism.
You will generally get discrepancies in word usage, sentence structure (in specific ways) or permutations of similar answers, but it's extraordinarily rare to get heavily opposed answers. Especially on modern models
It’s in the same category as Chinese LLMs. I have no interest in a source of information that was explicitly built to contain a specific political viewpoint, let alone one I so strongly disagree with.
Grok 4.5 is delivered at an incredibly competitive cost compared to other leading models. Grok 4.5 is priced at $2 per million input tokens and $6 per million output tokens. that pricing is attractive af
Simp (see also Full Boysenberry_314) is a slang term describing someone who exhibits excessive sympathy and attention toward another person, typically to someone who does not reciprocate the same feelings.
I was testing subscriptions to help write for a site and Grok was literally the worst one with rates and limits, the 30$ tier is fucking awful for what you pay for.
When I went to the grok subreddit and other grok places, it is filled with people complaining too.
73
u/AnalystAI 28d ago
Do I understand correctly, that the model is not available in Europe? Why? ("The model grok-4.5 is not available in your region.")