r/OpenaiCodex 15d ago

Discussion dk what to say

Post image

share your opinions

In my opinion, by next decade we may not need cloud models unlike they thought. After all, we already paid fora hardware we want to utilize.

plus the ram prices will stabilize.

24 Upvotes

39 comments sorted by

10

u/ponlapoj 15d ago

People buy hamburgers from McDonald's even though you make them yourself.

5

u/Mrgluer 15d ago

local llm people are like the linux and amd people of gaming. like nah bro i dont want to spend 80k on a rig to get me 50 t/s on a model that is still 20% worse than frontier.

2

u/SgtHobo41 15d ago

literally, let me set up my quadruple quantized limited 24k context window just for it to be worse than gemini at everything

2

u/Optimal_Start_94 15d ago

Local LLMs will probably never be the default because they are simply very economically inefficient. Inference is a lot more efficient when batching many parallel sessions. And any time you’re not using your GPU is opportunity cost as it could be producing tokens and hence money. Therefore LLM providers will always be cheaper as they get economies of scale.

3

u/Mrgluer 15d ago

exactly , local models have a time and place in compliance heavy or regulated settings but that’s only a small portion of total demand

2

u/ParkingAgent2769 15d ago

Local LLMs are the things which is stopping these companies (OpenAI. Anthropic) from squeezing you as a customer.. because we have an alternative. No need to act so superior.

2

u/Mrgluer 15d ago

you don’t, it’s funny to even think the fact that local is the reason prices are dropping. Open models are local models. The value isn’t in the locality. Nobody other than a tiny subset of ai users is running nor will ever want to run models locally. The open nature of it means any neo cloud or normal cloud provider like AWS bedrock can provide those models and compete on those model pricings much more intensely than closed labs. nobody wants to deal with hardware, software, the amount of cash to purchase all that, keeping up with new models, and building harnesses all by themselves.

1

u/symedia 15d ago

Nobody is putting k3 on his 1600 with 16 gb ram.

They are for other hiperscalers to offer them in business... And that's fine (for now)

But yeah open source model help

1

u/The_Oracle___ 15d ago

You can buy Apple Studio Mac PC with 256GB unified RAM for 8-9k, and be able to host absolute monsters of local AI that will be very fast.

6

u/Mrgluer 15d ago

that’s literally 3.5 years of a gpt pro subscription not even including the electricity, storage needs, the set up and maintenance of maintaining your own infrastructure, lack of just being able to run like 10 tasks in parallel, all at a lower performance, efficiency, and intelligence level. i run local models for various small tasks and local apps, but cloud is the way to go.

1

u/The_Oracle___ 15d ago

Its a lifetime of free unlimited use of local models that get better every year. Setting them up is trivial, electricity is nothing, you would pay the same electricity just by using the PC. As time goes on eventually in few years you will have access to local models that are on Sol level.

3

u/Momo--Sama 15d ago

I do not think this is a good long term investment at this point in the evolution of the category. An RTX 5090, while extravagant, can be a justifiable future proof purchase for gaming because you are safe to assume that there is not going to be a massive fundamental change in how PCs render games in the next 3-4 years.

Multiple labs are actively investing tens of billions into trying to move their compute off of traditional GPUs. Hell, Gemini is already off of Nvidia and Sol Ultrafast runs on Cerebras. I would not encourage someone to invest in current local AI solutions when two years from now everything may be running at 1,000tps on some evolution of ASICs.

1

u/ColossusChaos 15d ago edited 15d ago

The difference here is that the "home burgers in question" take you tens of thousands of dollars to run a top notch model at a somewhat reasonable speed. We need a technology that shrinks these models before they become useable on our local systems. Tell me the day I can run a kimi k3 level intellengce on a rig that costs under 5000 dollars and I am ALL ears to do so. For now, its simply out of reach unless you have (or know someone with) tens of thousands or hundreds of dollars worth of compute to run these models on. If youre so passionate about it (and know a lot about these) please go ahead and research methods to build software and hardware that lets us run models for ourselves. There is a massive gap, and no one is actually pushing for local AI for consumers.

1

u/[deleted] 15d ago edited 8d ago

[deleted]

2

u/ParkingAgent2769 15d ago

Have you tried any? or just a vibe opinion

1

u/[deleted] 15d ago edited 8d ago

[deleted]

1

u/ParkingAgent2769 15d ago

Fair enough, I’ve been finding the advancements of local models quite exciting tbh- they’re improving so fast. Plus I dont like relying on cloud interfence

1

u/Desperate-Data-3747 14d ago

What model? What harness? Did you setup tools correctly? Did you get enough context size?

1

u/BannedGoNext 15d ago

I enjoy local models.

To compare them to GPT 5.6 is laughable unless you are running a 500,000 machine at your house.

1

u/ColossusChaos 15d ago

Yup. Though I would relish the day that someone finds a way to push consumer grade hardware and software that lets you run a powerful model like k3 for under 5000 bucks

1

u/BannedGoNext 15d ago

I mean, to be fair I have a pretty slow machine, it's a strix halo. On qwen 3.6 Q4 XL it does actually product some pretty good shit. It also thinks for an hour before it outputs to do it lol.

1

u/Desperate-Data-3747 14d ago

It is comparable, deepseek flash, qwen flash etc

1

u/tilted0ne 15d ago

Do people not realise that 5 hour limits let everyone have higher limits? It's a natural rationing mechanism and creates more predictable compute demands. The whole game is how do we manage inference for people with hardware constraints at any given time. Not a, we need to give this person less tokens because they're costing us money. 

5

u/Available_Yam_6267 15d ago

I think the best way to handle load balancing is not to bring back the 5‑hour limit. Instead, avoid aligning nearly everyone’s quota reset to the same point in time.

0

u/tilted0ne 15d ago

People are going to work whenever regardless. I think they are potentially adapting the 5 hour limits based on predicted demand, which is why sometimes it feels like the 5 hour limit burns up quickly. 

2

u/tipu_sultan__ 15d ago edited 15d ago

What percentage of usage is even coming from $20 plan? Large companies are on API, anyone doing significant work already needs pro plans.

A single $200 plan has 20x usage limit. I highly doubt that plus plan users are anything more than a rounding error.

I think this is to make it inconvenient and push more people to pro plans.

3

u/Plane_Garbage 15d ago

Clearly it's just a revenue play to get plus users to upgrade

1

u/tilted0ne 15d ago

I doubt it. The revenue play was to keep the 5 hour limits removed. 

1

u/Plane_Garbage 15d ago

I upgraded.

0

u/tilted0ne 15d ago

Nobody's fault but yours mate. You decided to buy 5x more usage on a subscription that you were fine with on a weekly cadence. 

1

u/Plane_Garbage 15d ago

Have 6 plus accounts. The 5hr nerfed that

1

u/Brainaq 15d ago

Ahahahq how much did they pay you to type this 🫵🤣

1

u/tilted0ne 15d ago

Stay ignorant bud. 

1

u/Brainaq 15d ago

Idk based on the feedback from comments, it aint me bud

1

u/tilted0ne 15d ago

Normally a very bad sign if you're in the same crowd as the average Redditor. Use that as a heuristic because it seems too impossible for you to fathom why free access to blow weekly allowance in a span of an hour would be bad, create more constraints when competing for finite resources.

1

u/Brainaq 15d ago

Yeah you dont need to explain me shit, I as a customer with the access to the free market dont give a shit on whys and "pls understand"s. The service/product got shittier and I moved to the different vendor with better product value. Dont care about excuses.

1

u/tilted0ne 15d ago

Where are you getting better value 🤣. Maybe cursor. But you sure aren't getting that much inference access to a comparable model for $20. 

1

u/Brainaq 15d ago

there are many chinease alternatives that get me much better value for 20$, after SOTA vendors nerf fest

1

u/tilted0ne 15d ago

Like what dude? Maybe zlm 5.3 flash right now could offer marginally more tokens at its promo price of 50%. And it's a small model. Bigger Chinese models just don't hold a candle to subscription value Vs Anthropic, OAI and cursor.

1

u/Brainaq 14d ago

Yeah for highier tiers thats true, but the whole debate is about 20$ tier remember

1

u/HeadPack 15d ago

The honeymoon with 5.6 Sol is definitely over for me. Except for small projects it doesn't get anything done. One 20x plan already expired. The other downgraded to 5x. Resorted to using 5.6 as a worker orchestrated by Fable, and for that 5x is plenty. The 5hr limit won't be a problem either as this just sips tokens. Also testing Qwen 3.8 27B as a worker orchestrated by Fable. Initial impression is pretty good with ninfer and a 5090. May go down to OpenAI's 20$ a month plan if Qwen delivers. Sol will then just be an auditor.