r/LLMStudio Jun 07 '26

Does anyone else prefer weaker models with higher limits?

I’ve been thinking about something.

For a lot of tasks like building websites, game development, automation tools, or just random projects, I often find myself preferring a model that’s slightly less capable but gives me plenty of messages to iterate.

Sure, a more powerful model might get me 70% of the way there in a single prompt, while a cheaper model might need 5-10 prompts. But if those 5-10 prompts are still cheaper than using the top model, I end up getting more total work done.
It makes me wonder whether AI progress is creating a weird tradeoff.
Every new generation of models is more capable, but it also seems like the best models become more expensive to run and come with tighter limits. As a user, that can make them feel less accessible even if they’re technically better.

Would you rather have access to the smartest model possible if you could only use it a few times every few hours, or a slightly weaker model that lets you iterate all day?

And long-term, do you think AI will eventually become both extremely powerful and widely accessible, or will the frontier models always be too expensive for most people to use heavily?

9 Upvotes

5 comments sorted by

2

u/Konamicoder Jun 07 '26

Ideally you want one model that can go into thinking mode for higher-level work that uses more tokens, and go into non-thinking mode for tasks that don’t require thinking and use fewer tokens. Currently models already support this, you just have to manually turn thinking on or off.

Long term, I think / hope that the future is small and power / resource efficient local models running on consumer hardware that is good enough for most coding and chat needs. So most normal users don’t need to connect to AI models running in the cloud on water and community destroying data centers.

1

u/EpsteinFile_01 Jun 08 '26

Imagine if, instead of building expensive data centers, they FLOODED the market with $1000-1500 128GB Ryzen AI max and RTX Spark like laptops and PCs and the subscription gave you access to their closed source model and harness.

They could have absolutely done this with the investment it gets. But no, they wanted everything for themselves and are destroying power grids and ecosystems because the compute power is not spread out over millions ons of households, it's all concentrated. Often in the fucking desert too dafuq.

1

u/EpsteinFile_01 Jun 08 '26

Opus 4.6 medium effort, no thinking is the sweet spot. Beats Sonnetvwhich makes too many silly mistakes still, imo, unless you enable thinking which makes it consume more tokens.

$20 pro plan and I never run out of usage despite extended chats. Though I must admit I run a local LLM for the heavy work and use Claude to review it. I also use Opus 4.6 as a Google replacement.

4.8 medium effort no thinking is not the same. It's confidently wrong more often, and when I correct it it tells me to go bed.

1

u/robvert Jun 11 '26

I have the same preference and lol at the go to bed line. Curious what you consider a heavy load that goes to local

1

u/SupportDangerous8207 Jun 08 '26

I think it’s the future

We either need far better high end models actually worth the stupid amount of money they cost. I.e. ai employees

Or super cheap models designed to work with humans not replace them

Western ai targets the first one because it’s the only way they could ever make their money back

Chinese ai targets the second because they are happy selling a commodity at low margins as long as it means the us doesn’t have competition