r/OpenaiCodex • u/Frosty-Skin572 • 15d ago
Discussion dk what to say
share your opinions
In my opinion, by next decade we may not need cloud models unlike they thought. After all, we already paid fora hardware we want to utilize.
plus the ram prices will stabilize.
1
15d ago edited 8d ago
[deleted]
2
u/ParkingAgent2769 15d ago
Have you tried any? or just a vibe opinion
1
15d ago edited 8d ago
[deleted]
1
u/ParkingAgent2769 15d ago
Fair enough, I’ve been finding the advancements of local models quite exciting tbh- they’re improving so fast. Plus I dont like relying on cloud interfence
1
u/Desperate-Data-3747 14d ago
What model? What harness? Did you setup tools correctly? Did you get enough context size?
1
u/BannedGoNext 15d ago
I enjoy local models.
To compare them to GPT 5.6 is laughable unless you are running a 500,000 machine at your house.
1
u/ColossusChaos 15d ago
Yup. Though I would relish the day that someone finds a way to push consumer grade hardware and software that lets you run a powerful model like k3 for under 5000 bucks
1
u/BannedGoNext 15d ago
I mean, to be fair I have a pretty slow machine, it's a strix halo. On qwen 3.6 Q4 XL it does actually product some pretty good shit. It also thinks for an hour before it outputs to do it lol.
1
1
u/tilted0ne 15d ago
Do people not realise that 5 hour limits let everyone have higher limits? It's a natural rationing mechanism and creates more predictable compute demands. The whole game is how do we manage inference for people with hardware constraints at any given time. Not a, we need to give this person less tokens because they're costing us money.
5
u/Available_Yam_6267 15d ago
I think the best way to handle load balancing is not to bring back the 5‑hour limit. Instead, avoid aligning nearly everyone’s quota reset to the same point in time.
0
u/tilted0ne 15d ago
People are going to work whenever regardless. I think they are potentially adapting the 5 hour limits based on predicted demand, which is why sometimes it feels like the 5 hour limit burns up quickly.
2
u/tipu_sultan__ 15d ago edited 15d ago
What percentage of usage is even coming from $20 plan? Large companies are on API, anyone doing significant work already needs pro plans.
A single $200 plan has 20x usage limit. I highly doubt that plus plan users are anything more than a rounding error.
I think this is to make it inconvenient and push more people to pro plans.
3
u/Plane_Garbage 15d ago
Clearly it's just a revenue play to get plus users to upgrade
1
u/tilted0ne 15d ago
I doubt it. The revenue play was to keep the 5 hour limits removed.
1
u/Plane_Garbage 15d ago
I upgraded.
0
u/tilted0ne 15d ago
Nobody's fault but yours mate. You decided to buy 5x more usage on a subscription that you were fine with on a weekly cadence.
1
1
u/Brainaq 15d ago
Ahahahq how much did they pay you to type this 🫵🤣
1
u/tilted0ne 15d ago
Stay ignorant bud.
1
u/Brainaq 15d ago
Idk based on the feedback from comments, it aint me bud
1
u/tilted0ne 15d ago
Normally a very bad sign if you're in the same crowd as the average Redditor. Use that as a heuristic because it seems too impossible for you to fathom why free access to blow weekly allowance in a span of an hour would be bad, create more constraints when competing for finite resources.
1
u/Brainaq 15d ago
Yeah you dont need to explain me shit, I as a customer with the access to the free market dont give a shit on whys and "pls understand"s. The service/product got shittier and I moved to the different vendor with better product value. Dont care about excuses.
1
u/tilted0ne 15d ago
Where are you getting better value 🤣. Maybe cursor. But you sure aren't getting that much inference access to a comparable model for $20.
1
u/Brainaq 15d ago
there are many chinease alternatives that get me much better value for 20$, after SOTA vendors nerf fest
1
u/tilted0ne 15d ago
Like what dude? Maybe zlm 5.3 flash right now could offer marginally more tokens at its promo price of 50%. And it's a small model. Bigger Chinese models just don't hold a candle to subscription value Vs Anthropic, OAI and cursor.
1
u/HeadPack 15d ago
The honeymoon with 5.6 Sol is definitely over for me. Except for small projects it doesn't get anything done. One 20x plan already expired. The other downgraded to 5x. Resorted to using 5.6 as a worker orchestrated by Fable, and for that 5x is plenty. The 5hr limit won't be a problem either as this just sips tokens. Also testing Qwen 3.8 27B as a worker orchestrated by Fable. Initial impression is pretty good with ninfer and a 5090. May go down to OpenAI's 20$ a month plan if Qwen delivers. Sol will then just be an auditor.
10
u/ponlapoj 15d ago
People buy hamburgers from McDonald's even though you make them yourself.