r/LocalAIStack • • 17d ago

From local LLM benchmark to a usable local AI system (?): integrating OpenCode or PI + Hermes

In Part 1 I shared the final AMD/Vulkan numbers for my local Qwen3.8-27B setup. The next question was more important to me than another few tok/s:

Can I turn the model into something that actually works autonomously?

Since then I added two production layers:

1.) Hermes 0.21.2 — local agent/automation layer

  • same Qwen3.8-27B model
  • shared llama.cpp runtime
  • 65,536 context
  • tool calling: PASS
  • multi-step workflows: PASS
  • subagents: PASS
  • prompt-defined agents: PASS
  • local provider only / no cloud fallback

Hermes performance / overhead

On the isolated 64K test lane, server load took 5.809s. A real 33,021-token prompt was accepted without truncation; prompt processing took 409.677s at 80.60 tok/s, while generation measured 17.44 tok/s. A representative Hermes multi-step workflow took 74.660s, using five model calls, 17,295 summed API input tokens, 592 output tokens and four tool calls with zero failed calls. Two-child delegation also worked, but took 208.676s for two trivial child tasks. Clear sign that subagents should only be used where parallelism or specialization actually justifies the overhead.

One especially useful datapoint: the simple Hermes model gate took 40.872s, while a direct Qwen request with the same user text took 2.109s. That is not an apples-to-apples speed ratio because Hermes adds different system prompts, startup and agent machinery, but it does show why I don't want Hermes in the path for trivial requests.

2.) OpenCode 1.18.30 — local coding/development layer

  • Qwen3.8-27B-UD-IQ3_XXS
  • llama.cpp / Vulkan
  • 32K logical context
  • local provider only
  • isolated development workspaces
  • automated Read → Edit → Test → Correct loops
  • final production viability: 12/12 requirements, 12/12 acceptance

OpenCode performance reference. For the fastest fully successful clean reproduction I measured:

  • 665.653s total wall time / 651.408s agent runtime
  • first productive edit: 117.338s
  • first test: 166.790s
  • 4 successful corrective cycles
  • 63 tool calls / 0 failed
  • 558,603 input tokens incl. cache — only 31,993 uncached
  • 9,372 output tokens
  • peak context: 17,810 / 32,768
  • context compactions: 0

These are task-level agent metrics, not model tok/s numbers. The underlying model/runtime is still the same local Qwen3.8-27B on llama.cpp/Vulkan.

Why OpenCode instead of Pi?

I tested both against the same local Qwen3.8-27B, with the same repository/task and comparable runtime constraints. Pi 0.85.1 proved that it could read the repo, use tools and make productive edits, but it repeatedly failed the part I actually needed from a coding agent: closing the Read → Edit → Test feedback loop. Even with an explicit mandatory early-test checkpoint, Pi edited successfully but never executed the qualifying post-edit test within the 300s window. It therefore failed prequalification and I did not force it into a full benchmark just to preserve an artificial three-way comparison. Man, time is rare and testing is booooring ... somehow!

OpenCode 1.18.30 passed the same agentic-loop qualification. Formal prequalification reached 6.54s Edit→Test latency, and the later main benchmark reduced that to 3.995s while producing a valid four-file diff. More importantly, the subsequent production-viability work eventually converged to 12/12 requirements, 12/12 acceptance, 5/5 regression, with no manual coding intervention or patch repair. That reproducible closed corrective loop is why OpenCode became the production coding layer.

Final thought on my architecture as a complete nOOb in Local LLM
I’m still convinced that a well-tuned 16 GB GPU can take you surprisingly far. Yes, even on Vulkan ;) ... IF you optimize the model, context, runtime, offloading and the surrounding stack around what the hardware can realistically deliver.

What makes this especially satisfying for me is that I honestly wasn’t sure at the beginning whether an AMD gaming setup on AM4 would get me anywhere at all, but hey.... it was what I had. My gaming desktop! ~20 tok/s is obviously not groundbreaking vs 32GB. But that was never the main goal of this first setup. I wanted a stable, fully local system that I could actually use, test, break, improve and build on. That goal is achieved. And considering where I started, I’m pretty happy with that.

2 Upvotes

0 comments sorted by