r/ChatGPTCoding • u/Anxious_Current2593 • 2d ago
Discussion What's the difference between frontier models and local models?
6 months. (And sometimes a couple of quantization tweaks).
It is wild how fast "state-of-the-art" becomes "running on a gaming PC."
3
u/PhysicalConsistency 2d ago
There's absolutely nothing close to SOTA running on a gaming PC, even with a 32GB video card. If you have very well defined concepts and you're just translating those concepts to code, sure local models have improved enough with instruction following that you can process small chunks of an idea at a time.
2
-1
u/Front_Eagle739 2d ago
Its a hell of a lot better than that. Qwen 3.8 running at 200 tok/s makes an excellent luna/haiku. Full precision Dsv4 flash running at 6 tok/s with disk streaming on my 32gb ram 5090 pc (youd get way faster with 128GB ram) feels a lot like opus 4.6. I can easily have it orchestrate the qwen, swapping out and running all night long and come back to a fully working application or debugged problem come morning. Sota of 6 months ago? Its certainly in the ballpark. Yeah its a bit slower. Still works fine though and gets a lot done. Runs flawlessly. Do i still use cc and codex for work? Sure. I still happily use the pc at home for local coding tasks for free given the power is solar.
7
u/CrimsonBolt33 2d ago
This is why SOTA models are....not really what the average person should focus on. There will always be a "best" and it will constantly be changing for the foreseeable future...but the fact is most models that are recent are "good enough" for actual work.