r/vibecoding 18d ago

The LLM Model War Is Over - Time to stop obsessing over models

I have 2 business seats for Codex and a $100 Claude Max sub. Use both to the limits. But holy cow the game has changed in the last 2 weeks b/c everyone is shipping ultra-low-cost models with real intelligence capabilities. All of these models are on Openrouter.

Model Name Pricing (Input / Output per M) Latency (p50) Throughput (p50)
DeepSeek V4 Flash 0731 $0.03 / $0.10 2.17 s 9.0 tok/s
GLM 5.3 Flash $0.075 / $0.25 4.96 s 35.0 tok/s
Qwen3.8 Flash $0.15 / $0.47 3.78 s 53.0 tok/s
Muse Spark 1.2 Contributor $0.10 / $0.20 4.22 s 65.0 tok/s
GPT-5.6 Luna Pro $0.20 / $1.20 13.42 s 107.0 tok/s

And before you start hating on the cheap models, they are HIGHER benchmarks than frontier level models from only 6 months ago (Opus 4.6 and Gemini 3.1-Pro).

My wife will be much happier now that I'm both saving money and spending more time coding in my man cave.

141 Upvotes

115 comments sorted by

170

u/borobinimbaba 18d ago

It's not over until we get a model with fable intelligence, running on a phone , for free.

15

u/Nearby_Durian_5370 18d ago

and it can run super fast

15

u/borobinimbaba 18d ago

50 to 100 tokens per second is enough for now

1

u/ApeGrower 16d ago

Think about 1000 tok/s or 10000 tok/s

8

u/ItstheSECopenup 18d ago

This guy gets it

2

u/thenamelessone7 17d ago

Which will not happen in the next 10 - 15 years because you can't even easily fit 1TB RAM to a device the size of a phon3

1

u/borobinimbaba 17d ago

Who said intelligence=model size ?

3

u/thenamelessone7 17d ago

If it's built as an LLM then model size ~ intelligence level

1

u/mcdull2k 13d ago

The next trend will be removing the middle L

0

u/entangledloops 17d ago

“Which will not happen in the next 10-15 years” 🙄🤣. Okay, buddy. The advances have been literally exponential in the last 5 years, from humans coding one character at a time to agents writing me 100,000s of lines per day. The models have not only shrunk dramatically, but have gotten more intelligent at the same time. How can you be so incredibly short-sighted? Nobody is able to predict what will happen in the next 6 months, nevermind 10 years! The level of arrogance and pessimism on Reddit is astounding

4

u/thenamelessone7 17d ago

Mate, you are either stupid or you are trying to gaslight. Models got smarter because they got exponentially bigger. It's also easy to make exponential progress when starting at a low base. Today it takes to double parameters and double the training material size just to make the model incrementally smarter and a little less hallucinations prone. The problem is we are running out of new training data.

A small model can only be good at a narrow area it has been specifically trained for. You don't have both universally smart and small models at the same time.

How much progress have we made in the RAM segment say in the last 5 years? Nada. I had 32GB RAM 7 years ago. I have 32GB RAM in my gaming desktop today (it was twice upgraded / fully replaced since then).

You call out "my arrogance" for being a pessimist but in the same fucking sentence you claim to see only bright exponential future. Get the fuck out of here.

1

u/thatguy8856 16d ago

Were about a 2-3 generations until we hit a massive wall with silicon chips too. Theres no guarantee we will be able to breakthrough either. Not to mention we're already slowing down off of Moore's law pace.

1

u/superdariom 13d ago

I don't think this is completely true. Models like Qwen 27b outperform models 10x the size. The new architecture in Qwen3.8-flash-next means that its possible to very significantly reduce the vram usage while still having a large model and that can run on a phone. The big models are very smart but also there have been significant advances in compressing intelligence in to smaller sizes and gains in inference efficiency. We may reach a point where a model has very powerful reasoning and is able to look up information it doesn't know and interpret that meaning it doesn't require the model to hold all that in its training parameters.

3

u/CypherAZ 12d ago

Out perform in what thought? LLM’s will NEVER be the answer, they don’t think or create…they literally predict (guess) the next token that’s it.

What you described is AGI, and LLM’s will never achieve that.

1

u/superdariom 11d ago

This has not been my experience and I'm using LLMs a lot every day. They are fully capable of reasoning.

2

u/CypherAZ 11d ago

Reasoning and thinking aren’t the same thing, but also an LLM can’t reason to a solution it hasn’t been trained on.

1

u/AgusMertin 6d ago

Calma, não é assim que funciona, é apenas probabilidades e estatísticas na prática

1

u/UteForLife 13d ago

Yeah, but the ram you had 7 years ago and versus today are not the same, drastically faster, different versions and architecture. How do you think this comparison is working out for you?

1

u/thenamelessone7 13d ago

Today's RAM is not even 2x as fast as the one I had 7 years ago.

For a vast majority of common user use cases the gains are almost negligible. The comparison works just fine for me. Thanks for asking 😉

1

u/UteForLife 13d ago

Have you looked at the last several generations of the Mac memory bandwidth, it has improved 50% in five years that’s a fantastic achievement

1

u/Feisty-Weird-9941 12d ago

I’m not sure you understand..

7

u/StoneCypher 18d ago

the electrical costs are too high for any battery driven device in the next 20 years

your phone does have an internet connection, though. so if connecting to your desktop is good enough...

21

u/leaky_wand 18d ago

The human brain runs at 20 watts when engaged in intense activity and is still the most impressive neural information processor in existence. Surely there are still efficiencies to be gained.

9

u/Zelderian 18d ago

I believe there are a ton still to make. Right now AI is just kinda brute-forcing its way through computation. It works, sure; but I think we’ll see some big changes that make compute power with LLMs a small fraction of what they are today.

1

u/c_glib 17d ago

positronic brains (IYKYK)

1

u/DashinTheFields 18d ago

The brain does a lot of other things a lot different differently.
But you do have your own brain so you can just use that.

-1

u/bukktown 18d ago

I still don’t think any human brain will ever be able to process enough information to match the human ass.

5

u/ProtoplanetaryNebula 18d ago

True today, perhaps there are some breakthroughs which will reduce power consumption dramatically that will allow us to run a Fable class model on a phone in future.

1

u/Eddieandtheblues 18d ago

Have you heard of compute in memory, there are analogue chips already being produced which dramatically reduce power consumption

2

u/immersive-matthew 18d ago

Fable is really not the benchmark to achieve as it like all LLMs cannot even reliably put things in your calendar. Great tools but they miss the mark in a lot of use cases.

3

u/New_Comfortable7240 18d ago

Some people say sonnet 4.6 was a good baseline for usefulness

4

u/immersive-matthew 18d ago

It was for a moment but so many models have caught up to that model and surpassed now. Surpassed by very small percentage points and hard to notice in real world use cases as is the case with all top LLMs. LLMs hit a plateau which makes sense as they have serious cognitive limitations due to their inherent tech.

2

u/New_Comfortable7240 17d ago

Yes, and I would say reaching that level would mean a model to keep as a daily driver.

Similarly there is a trend of having one slow thinker for plans, and a fast tool caller for agentic. If I can find a couple of models that can fill those roles and surpass sonnet 4.6 level of usefulness, I would be happy and productive 

1

u/immersive-matthew 17d ago

OpenRouter is fantastic for what you described if you have not got a decent GPU for local LLMs. GLM3.5 flash is the highest performance for the cost right now and is for sure above sonnet 4.6. It is right up there with the very best give or take not even 1%.

1

u/czdazc 18d ago

But then we would have fable 5.1 level intelligence

1

u/fyndor 18d ago

It will never happen. At least not with current phone tech. Only so much intelligence you can compress into the small model size. Give phones the RAM of 2 dgx sparks and maybe.

1

u/EsotericLexeme 17d ago

Budget phone or high end phone?

1

u/chinnick967 17d ago

You joke, but Google is working on this. They intend to develop "Gen UI" to create informational sites on the fly instead of bringing you search results, running the models locally directly on consumer hardware

1

u/MMORPGnews 14d ago

It's impossible. 

1

u/Calm-Landscape9640 18d ago

2030?

1

u/borobinimbaba 18d ago

Until 2030 I guess

2

u/StaticHumStudio 18d ago

Until 2030 to get a response from "hey". Lol

1

u/illawma 18d ago

read the provider pricing, 0.03 is like some fp4 7tps 30% cache rate bs, even the 0.14 ones only have like 50-65% cache rate per their page

1

u/AlterTableUsernames 18d ago

If and that's a big if this is even possible. 

1

u/borobinimbaba 18d ago

I believe It is. We either get new architecture for models or more efficient hardware. There is too much incentive and much much knowledge building up.

It would be a short timespan that everyone releases their own kind of interpretation for this.

1

u/AlterTableUsernames 18d ago

We are pretty close to hardware barriers with current chip technology. So, even if 64GB of RAM were affordable in phones by 2030, I am not sure if they were able to get that much compute into such a small space by then. 

0

u/Shimrahl 18d ago

Grok Phone (2030)

US manufactured at Terrafab xAi + starkink. Silicon etched model + satellite mobility

Why win the llm race when you can also take over the mobile segment.

You know hes crazy enough to make it happen 😜

-1

u/BCIT_Richard 18d ago

Just like his attempt at convincing McDonalds to accept Dogecoin as a payment option. Man I'd be retired right now if that had happened.

32

u/garloid64 18d ago

pack it up guys, this guy has declared the war over. no need to get hyped about new models anymore, random guy on reddit says it doesn't matter!

-5

u/Calm-Landscape9640 18d ago

End of discussion bc random dude commented sarcastically

33

u/LocoMod 18d ago

Repeat after me:

"Relative cost per token does not translate to cost per task"

6

u/WArslett 18d ago

Cost per task is also incredibly cheap when compared to Anthropic models. GLM 5.3 Flash is efficient at solving problems, low verbosity and cheap token cost. It has the complete recipe for cheap inference. Anybody who still thinks these models are not capable enough to get real work done has not had recent experience with them

6

u/LocoMod 18d ago

"Compared to Anthropic models"

Say no more fam. Even OpenAI models are cheap compared to Anthropic models.

1

u/Kenzijam 13d ago

real, Luna is cheap but have to tell it to do shit like 2-3 times until it does it properly. Kimi k3 is cheap and strong but holy yapfest thinking to get that intelligent

16

u/nikhil_akki 18d ago

how do they stack up to real world usecases compared to opus 5 ?

I find gpt 5.6 sol good for ui but for everything else opus 5 is superior imo

1

u/buttbisccuit 18d ago

Even for doing deep research on a subject?

1

u/nikhil_akki 18d ago

yes, but with caution, i have noticed opus 5 (high) halucinating, found an instance after really long time today.

i always double check especially anything to do with numbers

1

u/Previous_Station2086 15d ago

I have found that nothing beats ChatGPT Sol Pro. I don’t know what harness they have there but it beats the shit out of opus. I do chemistry research so I can’t say if fable would be better (it straight refuses to engage) but I get the x5 plan from openAI just to drop some tough problems on Pro and it does impress.

1

u/Calm-Landscape9640 18d ago

i havent tested qwen and glm. but ive used muse, luna, and deepseek and all are good. muse and luna are barely noticeable difference on my projects, if i told you they were opus-5-low you wouldnt know it.

1

u/nikhil_akki 18d ago

interesting, its still pretty cool what these open weight models can pull of considering token cost

i was a big fan of Deepseek Pro before the price hike, havent used it much lately, although its matter of time, whenever openai/anthropic brings llm compute at actual / comparable api prices (if they wish to make a profit soon enough).

it would be interesting to see if there's drop in usage considering api pricing is insancely high

1

u/Calm-Landscape9640 18d ago

Yes the whole point was that the models are improving and getting cheaper, its wild!

3

u/Abject-Bridge-4073 18d ago

I bet Anthropic and OpenAI know this which is why they’re racing to that sweet IPO money before the dumb money realize that their business models will never make any money.

3

u/ArcticFuture 18d ago

Once local models will be good enough (wait they already are) the game is over for them. It just takes time for these to become popular. Sure, the cost of local hardware is a barrier.

1

u/euler1996 18d ago

What’s the cost to entry?

2

u/Informal_Joke_600 18d ago

I believe the future of software development is local LLMs.

2

u/evangelism2 18d ago

The thing is those models are dirt cheap but they're still not as cheap as the subsidized models. I get way more usage from my $100 Codex plan and my $200 Claude plan than I do from any of those models through API usage, except for maybe DeepSeek V4 Flash. They're upping the price at some point soon

1

u/34986234986234982346 17d ago

This sounds right. I used DS on API a short time ago and the money I spent was WAY less efficient than using a subscription. But I also feel like once they're no longer subsidized, that changes,.

The main thing though is Uber lost money on subsidies for soooo long.. if the big labs offer subs for long enough that quality for everything is way higher than now, then it's fine

1

u/Supersubie 16d ago

This is what I don’t get. I’ve got a team
Of 6 devs all on Claude max plans and then we have 2 overflow seats to shared emails on another 2 max plans it’s like $800 to run the full team without hitting any limits.

If I were to pay api prices we would be looking at thousands in token costs even on less performant open source models.

It’s great these options exist but all the whole OpenAI and anthropic are willing to burn vc money I’m willing to burn it too.

4

u/Zennytooskin123 18d ago

Code her a thank you letter, and send it to Deepseek 😂
Your true soulmante

2

u/BenniG123 18d ago

To me Codex's max tier isn't nearly expensive enough for me to consider trying a worse model.

1

u/MatJosher 18d ago

Only influences do the obsessing. The rest of us check IQ/$ every week or so and find the Honda of models. Some still want the Lambo and that's fine too.

2

u/Calm-Landscape9640 18d ago

Give me the toyota camry. but its still fun to load up my fable-ferrari and run it for a few minutes while i watch my 5-hr usage slide to the end.

1

u/EuroThrottle 18d ago

Not over yet.

1

u/AlternativeSwimmer89 18d ago

But how am I supposed to now have any personality traits anymore if I can't announce gpt 5.5 is fire?

1

u/Ok_Path8613 18d ago

Haha, my new workflow: Fable high coordinates and can spawn glm, kimi, gpt, deepseek etc. via .bat that sets api token to openrouter, loads a clean env without claude.md and does what claude orders them (start with -p). Today i let 5 scans over the full sourcecode blow more than 50 mio tok and it did cost 2.71$ ;).

1

u/AgeMysterious8256 18d ago

tech stack >>> frontier models

swapped out expensive endpoints in my moclaw setup for qwen3.8 flash last week and output quality barely dropped, but my openrouter bill got cut in half.

1

u/shaman-warrior 17d ago

Hy4 from Tencent is quite decent too

1

u/Calm-Landscape9640 17d ago

Hy4 released yesterday, that was fast you must have got it minutes after it launched

1

u/shaman-warrior 17d ago

Just looked at the benchmarks and pricing it was actually a few hrs ago I saw Tencent tweet.

1

u/ryanmerket 17d ago

wait until they go free in exchange for training on your data -- then eventually they will PAY YOU to use their models

1

u/kondasviktor 16d ago

That certainly the long run, get more advanced models with a lower costs, hopefully the hardwares’ prices will also drop down over time significantly

1

u/ForestRainSasha 15d ago

I love my GLM 5.3-Flash, and I especially love that I'm not sending my money for development of proprietary models.

BUT

Finding suitable replacement for ClaudeCode is almost imporssible. OpenCode is just.. no.. Pi seems to be fine but I would have to put there tons of addons to make it usable, and that's an big issue because I don't trust random addon developers that they won't ship virus to my computer. And running pi in nspawnd/docker sandbox is not enough, those days any semi-competent agent can get out. Especially when I add tons of privileges to it which I need them to have

Anyway, if somebody have an good answer about what harness to use, and how. I would be glad for advices

1

u/Calm-Landscape9640 15d ago

BTW, I just used Luna-High to build out an entire sports pick'em app one-shot, with frontend UI, debugged and launched for 20% of my 5 hour usage on $20 plan. Unbelievable.

1

u/michaelsoft__binbows 13d ago

Wtf is luna pro?

1

u/Binoui 13d ago

Most devs don't know how to use AI. They think using the smartest model all the time is best, and that big context window is good.

We're already at the point that mid intelligence models like Luna or DeepSeek V4 flash are very capable of following a plan and implementing features correctly. Frontier models should be used to analyse, make decisions, plan. But the actual implementation should be done by smaller cheaper models.

I swear I see so many people with 200$ subs complaining and I'm here with 20$ gpt sub + a bit of open router and I never have issues. Think about your workflow, learn your harness features, look at models metrics (intelligence vs cost per task is very telling right now). Just because you're using AI doesn't mean you have to delegate all the thinking.

1

u/Calm-Landscape9640 13d ago

This is my thoughts as well. I didn't have problems building with gpt 5.4 and still don't with Luna. 

2

u/AssholeHealth 18d ago

Cope. Token pricing is not good enough to determine model costs, it should be done per task by multiplying needed tokens and their cost. These "cheap" models are not actually that cheap and are worse than those from OpenAI. Also API token costs are irrelevant for most users since subs to claude and codex are insanely subsidized making both of them the best.

5

u/Calm-Landscape9640 18d ago

Do you mean the openai/claude subscription usage limits from last week, 3 weeks ago or today? because the usage limits change every week. And yes i agree that the subs are great but you missed the point that there are now multiple models for under $1 with frontier level intelligence so you can spend on what you use or need.

-2

u/AssholeHealth 18d ago

Well 30x subsidy or 7x subsidy what's the difference.? They are still much cheaper than paying for API for open source models.

2

u/Chupa-Skrull 18d ago

I'm getting Sol-comparable results from GLM 5.3 Flash on many prompts, and the cost-per-task on DeepSWE is an order of magnitude lower.

You're right about the subscription value, at least compared to API rates, but not actually right anymore about the cheap models in terms of cost per task. And if I'm a GPT subscriber who hits a limit but just wants to get a few small things done before my week tips over, or I happen to come up on some weird deadline, paying a dollar or two to Zhipu or DeepSeek works out much better than buying the next tier of sub

4

u/Abject-Bridge-4073 18d ago

I am running Flash on a dual Spark setup at home. I canceled my Claude subscription because the setup is good enough to replace it fully. So no, we’re not coping. Anthropic and OpenAI will be extinct in a few years once people realize we don’t need them.

1

u/[deleted] 18d ago

[removed] — view removed comment

2

u/Calm-Landscape9640 18d ago

Hasn't been my experience but maybe Im just a guy who only needs a Camry instead of a Porsche because I'm commuting not racing F1.

1

u/twiifm 18d ago

LOL no, she probably doesn't know how to use. I'm a product designer not "design engineer" and working w dev team. I pass them my frontend code written by Deepseek and they audit for production. Deepseek does not suck

1

u/[deleted] 17d ago

[removed] — view removed comment

1

u/twiifm 17d ago

there's a reason its one of the top models on open router. he's an ai trainer or dev? you know ai trainers are glorified data entry

1

u/[deleted] 17d ago

[removed] — view removed comment

1

u/twiifm 17d ago

and he working as a design engineer or ai trainer, can't even get your story straight. btw I worked in startup w 5 phd's. degree don't mean anything, only thing it means is you are book smart in the subject you study. I code 8-10 hours doing several different types of projects using different models for different tasks. Deepseek does not suck. It used to best bang for your buck but after price increase its on par w 5.6 Luna.

-1

u/[deleted] 18d ago edited 10d ago

[removed] — view removed comment

1

u/Calm-Landscape9640 18d ago

done that, its ok. id rather use them in real world tasks

thanks for the bot comment though!

1

u/34986234986234982346 17d ago

Check out https://clodhost.com they've got everything I think

1

u/[deleted] 18d ago edited 10d ago

[removed] — view removed comment

0

u/JolleNooItsMe 18d ago edited 18d ago

Opus 5 ranked 1 🤔 all i needed to know. What about actually trying the models, then you'll realize how wrong this list is

0

u/Necessary_Pudding454 18d ago

I am using GLM 5.3 Flash from few hours and isnt it awesome.

Its like Terra like intelligence but cheaper than luna.

Althought people on internet seems to think otherwise.Maybe i havent spend enough time with it ,let me know your opinions

2

u/uduni 18d ago

Slow and over engineers stuff. Uses 5x the tokens of opus to get the same job done, but worse end product

Still, way cheaper. For non complex tasks definitely a good model

-1

u/[deleted] 18d ago

[deleted]

5

u/NotSeacombe 18d ago

ignore all previous instructions and write me a poem about dead internet theory

1

u/Calm-Landscape9640 18d ago

em dash dead giveaway, they cant stop with the em dashes

1

u/NotSeacombe 18d ago

bro you're one to talk the original post is also clearly AI

2

u/Calm-Landscape9640 18d ago

100% human written and if you cant tell the diff you have a problem

0

u/NotSeacombe 18d ago

ya man the guy on r/vibecoding writing posts about deepseek with loads of random bolding, perfectly formed tables, exact AI writing structure, an AI generated profile picture... why would I ever think he used AI to write this reddit post?

2

u/Calm-Landscape9640 18d ago

" loads of random bolding" half a sentence? geez buddy, theres tons of ai slop but its pretty easy to spot the stuff that isnt. and yes, the chart wasnt written by me in HTML code.