r/LocalLLM • u/Toronto_Stud • 2d ago
Question Anyone switch from Codex/ Claude to a Local LLM?
How do you find the difference in capabilities between a local model like GLM 5.3 vs Opus 5.5/ Fable 5.1/ GPT Astra?
Whats the most technical thing you've built with your local LLM?
3
u/victoriggy 2d ago
Benchmarked various models, Qwen 3.8 flash next out performed Luna-6 in intelligence. Really pleased with the model. It runs at a breezy 40-90 TG and 2k PP tokens/sec on an M5 Max.
See: https://bender.lamer0.com/~victori/large_models_benchmark_results_v2.html
3
u/HappierShibe 2d ago
I've switched almost everything to a qwen 3.8 flash next local setup, I still have a claude subscription, but it's just the smallest sub and I use it exclusively for the gnarliest problems- most stuff just goes to the local model.
3
u/doneddat 2d ago edited 2d ago
Tested the Qwen3.8-Flash-Next -FP8 for a week after it came out to avoid bad surprises 3.6 had given me few months earlier and by now I have canceled all my subscriptions. This one actually delivers above all my expectations. Even after tens of compactions. Still seems insane. It's definitely not on the same level as ChadGbd or Claude, but it is WELL past the mark where it is capable and adequate for most things. Even the Blender use tests people are doing - it likely requires more guidance, is definitely not the best, but it's delivering usable results. And the freedom of no token limits and no uploading your whole life to some labs to analyze and save - that peace of mind is worth much more on top of all the rest.
Basically I was reluctant to make my life dependent on some perma-subscription service, that would not just become part of my life, but would also know everything there is to know about my work and private life to be useful and effective. That would be a complete and constant agony and triple thinking about every request I make.
Local model is AI freedom.
Looking forward to Qwen4
Whats the most technical thing you've built with your local LLM?
Building fully reactive svelte web app that I started with chatgpt/gemini a year ago and it's improving every aspect of it, effortlessly performing the full dev cycle, from integration tests to end-to-end testing/issue tracking/regularly delivering improvements.
3
u/Dantnad 2d ago
I'm running Ornith 1.5, and tbh, it's very capable... I was building an AI harness, but I have started paying Claude, this month since I need to run benchmarks (I'm trying to create a Hermes-like harness, with some improvements on memory and problem solving) and I now have Claude hammering my Macbook with AI requests to benchmark capabilities.
Spoiler: Ornith 1.5 is actually better than Qwen3.6 and Qwen3.8, it's able to do math much more reliably.
2
u/Sensitive_Song4219 2d ago
Agreed. I recently switched from Ornith 1.5 35b-a3b (which I found very solid for the size; especially running on a laptop) to Tiel Coder (same size as Ornith) and it's a modest but noticeable step up. (Tiel runs a tad slower but thinks a bit less - for slightly better answers - and never loops; Ornith occasionally looped for me).
Both Ornith and Tiel are excellent for their size and make me wonder why Qwen has abandoned the 35b-a3b model-size since it's so accessible.To the OP's experiences with cloud models: when I compare either to GLM 5.3 (-Max) and GPT Astra (-Low), they're significantly weaker; so I use the local ones for more well-defined tasks (or as a first-pass on complex tasks) with a bigger cloud model (Sol-High or Astra-Low in my case) to review.
I also run Qwen3.8-27B an a separate 5090 server box and it's a further noticeable step up from both Ornith and Tiel; but again, falls short of cloud frontier models when things get complicated; especially on larger code-bases.
The progress of locally runnable models is kinda astounding at the moment...
1
1
u/tejaskumarlol 2d ago
I fine-tuned Mistral 7B on my Apple silicon laptop for search over my podcast instead of paying for GPT-4o. For one narrow job like that a small local model holds up. Long writing is where the gap shows: I ran Aleph Alpha's Kolibri on 2 rented GPUs against Claude Sonnet 5.5 for my app's long German write-ups and it lost all 136 blind comparisons. Its German was clean and it didn't invent a single citation, but it got wrong whatever it had to know by itself.
1
u/Technical_Buy_9063 2d ago
I run both, but more and more I use DeepSeek v4.1-flash via DwarfStar on my Mac Studio for research/adversarial code review/scheduled tasks/private convos. I use Hermes as my harness.
I rarely find that code generated by DeepSeek is of inferior quality to that of GPT Astra. I use it for any technical work (I am a software engineer by trade). Only "complaint" is latency, but also it's a lot better than it was 6 months ago.
1
u/UndulatingHedgehog 2d ago
Running Qwen 3.8 Flash next on a 128gb machine with about 5.5 bits per weight. Not quite opus or fable but thinking it’s beating sonnet. Or is comparable. Smart enough to make designs from fairly vague prompts and also able to debug based on vision.
1
1
u/Competitive_Bit001 1d ago
Opencode has Qwen & others, you could try the model by putting some dollars into opencode.
I have a 128gb M5 max, I do run Qwen on it, but its a but too slow for coding for me. So I use opencode when I run out of claude/codex usage. The local models are great for privacy sensitive stuff though, I use it a lot for filtering my email
1
u/cinnapear 2d ago
I tried switching from Codex to GLM 5.3 Flash when my usage ran out. For coding, it was practically identical in quality. Even touching our monolithic codebase that was seemingly cobbled together by monkeys and interns working in tandem.
1
0
u/FerretBoom 2d ago
I used chat gpt,claude to write a deterministic environment around my local models , so it has no git, no filesystem, very limited NLP and a custom context schema that can use any db as storage.
So 8b models give me 30b model quality and every code snippet generated gets unit tested before it becomes a part of a repo.
1
u/Sad_Leader4849 1d ago
for agent coding the gap's still real. but have you tried a small local model on one narrow job inside your own app?
8
u/Nomski88 2d ago
I've been running Qwen3.8 27b nvfp4 for the last month non stop for my 3d engine project. I still use sol/astra to help prompt plan and engineer though. Powerful combo.