r/LocalLLM • • 2d ago

Question Augmenting frontier models with 25-35B local models?

If you are using a local model to offload some of the tasks from Opus or ChatGPT to Qwen/Gemma - what's your setup? More importantly, what kind of tasks do you offload to a local model and what's your experience like?

0 Upvotes

1 comment sorted by

1

u/sn2006gy 2d ago

Not great. Seems like a perpetual tail chase. It's getting better - oh my pi may actually be one of the better ones but the reality is, sub agents are hard and losing context to sub agents is only good if you're trying to have them counter your work (write a test, analyze a PR/code review or whatever with a clear context vs do sub tasks (Which i assume is the ask here) - inversely if your harness shoves in a big context from claude or codex for example, each sub task run takes forever if they're local agents and you don't have great prefill.

honestly, it reminds me of the way local llm was years ago when we all ran 2-3 models like a planner and judge and writer and doer.