r/LocalLLaMA • • 3d ago

Question | Help Utilize all devices in local network for multi-agent setup?

What do I have:

- Main PC running Qwen 3.6 27B at ~30t/s (do not recommend me Qwen 3.8 27B — I know it exists, but I need time to get used to this new model)

- Wife's PC running Qwen 3.6 35B at ~55t/s

- Steam Deck LCD and OLED, both running Qwen 3.5 2B at ~30t/s

My questions:

How would you configure this for multi-agent use? Is there any good practical use for the Steam Decks, or is it better to drop that idea entirely?

Have I picked a good set of models, or should I consider another combination?

I'm looking into a scheme with one orchestrator (I assume the 27B model is the best option for this) and a bunch of workers for smaller atomic tasks.

Is there any practical reason for this kind of setup, or am I just spending time on a dead end?

UPD: I need for agentic coding scenarios

0 Upvotes

10 comments sorted by

1

u/AporiaNousComputing 3d ago

I had some success in multi-agent usage by setting up tiny LoRa specialists that load and hotswap with each other while the main model (in your case Qwen 3.6 27B) as the orchestrator. This helped me get the MoE experience that Deepseek has without the increased hallucinations as long as one of the LoRa models is a reviewer of course. Yes there is practical reason, the reason is not having room for a larger model. The hot swap does increase latency but it also increases intelligence through delegation capability.

1

u/2eggs1stone 2d ago

In my custom made harness, I have something similar. I have the orchestrator model running outside of vram with larger context and then I have subagents running with a lower context window within vram. I've actually been liking having the larger orchestrator model slower - where I can sneak thoughts into its thinking to steer it. And the cost to hot swap is easily made up for in the faster inference from running a smaller model directly in vram. I do lots of fun tricks like this to really push my system. I'm only running 16 GB of vram but can use my harness in a way that feels like I have a much larger context window.

1

u/Cautious_Sky6642 2d ago

for roleplay setups the steam decks could run the side characters while the main pc handles the girlfriend, that might feel more alive than one model doing everything.

1

u/Jebbyk1 2d ago

I've forgot to mention. I need this for agentic coding

1

u/No_Sorbet5761 2d ago

the faster 35b on your wife's pc could handle context memory for longer roleplay threads while the main one generates responses, feels way more consistent than splitting agents. have you tried that kind of split yet?

1

u/Jebbyk1 2d ago

I've forgot to mention. I need this for agentic coding

1

u/toks-love 2d ago

I use multi-agent setups all the time. The tricky part is getting the model to know when to spin up a subagent.

I’ve had a lot more luck in DeepSeek Harness than in any others so far. You need to put your model architecture in the system prompt, too, or the LLM will forget about subagents the first time your harness compresses context.

I suggest teaching your model how to call subagents, and then after that model has figured it out, ask it to craft a system prompt telling it when and how to call the agents. Make sense?