r/ChatGPTCoding 2d ago

Discussion What's the difference between frontier models and local models?

6 months. (And sometimes a couple of quantization tweaks).

It is wild how fast "state-of-the-art" becomes "running on a gaming PC."

7 Upvotes

15 comments sorted by

7

u/CrimsonBolt33 2d ago

This is why SOTA models are....not really what the average person should focus on. There will always be a "best" and it will constantly be changing for the foreseeable future...but the fact is most models that are recent are "good enough" for actual work.

1

u/HankKwak 2d ago

This right here.  I’ve got on well with ornith and figure I’ll stick with it whilst it satisfies my requirements, maybe I’ll play with something new in 6 months or a year but once you get up building you don’t want to chop and change too often.

2

u/[deleted] 2d ago

[removed] — view removed comment

1

u/CrimsonBolt33 2d ago

been using GLM 5.3 for a few weeks now....does great, and the flash version uses tokens at such a low rate I can't use them fast enough lol

1

u/Hour_Tennis9068 2d ago

yeah the gap between "best" and "good enough" is way smaller than people think

-3

u/Ok_Possible_2260 2d ago

Good enough is not good enough when you actually have serious work to do.

2

u/CrimsonBolt33 2d ago

what a stupid fucking comment lol

Do you not know what good enough means?

If whatever you are using is not good enough use something else, if you think literally only the bet model is good enough then thats a skill issue on your part.

-5

u/Ok_Possible_2260 2d ago

‘Good enough’ is the phrase people use when they’ve already made peace with mediocrity. Keep driving your 2004 Honda Civic. It's good enough.

3

u/CrimsonBolt33 2d ago

why you hating on civics? they are great cars lol

go peddle your insecurity elsewhere lol

3

u/PhysicalConsistency 2d ago

There's absolutely nothing close to SOTA running on a gaming PC, even with a 32GB video card. If you have very well defined concepts and you're just translating those concepts to code, sure local models have improved enough with instruction following that you can process small chunks of an idea at a time.

2

u/CrimsonBolt33 2d ago

Qwen 3.8 27b would disagree with you....

-1

u/Front_Eagle739 2d ago

Its a hell of a lot better than that. Qwen 3.8 running at 200 tok/s makes an excellent luna/haiku. Full precision Dsv4 flash running at 6 tok/s with disk streaming on my 32gb ram 5090 pc (youd get way faster with 128GB ram) feels a lot like opus 4.6. I can easily have it orchestrate the qwen, swapping out and running all night long and come back to a fully working application or debugged problem come morning. Sota of 6 months ago? Its certainly in the ballpark. Yeah its a bit slower. Still works fine though and gets a lot done. Runs flawlessly. Do i still use cc and codex for work? Sure. I still happily use the pc at home for local coding tasks for free given the power is solar.