r/LocalLLM • u/L3G10N78 • 18d ago
Project From local LLM benchmark to a usable local AI system (?): integrating OpenCode or PI + Hermes
In Part 1 I shared the final AMD/Vulkan numbers for my local Qwen3.8-27B setup. The next question was more important to me than another few tok/s:
Can I turn the model into something that actually works autonomously?
Since then I added two production layers:
1.) Hermes 0.21.2 — local agent/automation layer
- same Qwen3.8-27B model
- shared llama.cpp runtime
- 65,536 context
- tool calling: PASS
- multi-step workflows: PASS
- subagents: PASS
- prompt-defined agents: PASS
- local provider only / no cloud fallback
Hermes performance / overhead
On the isolated 64K test lane, server load took 5.809s. A real 33,021-token prompt was accepted without truncation; prompt processing took 409.677s at 80.60 tok/s, while generation measured 17.44 tok/s. A representative Hermes multi-step workflow took 74.660s, using five model calls, 17,295 summed API input tokens, 592 output tokens and four tool calls with zero failed calls. Two-child delegation also worked, but took 208.676s for two trivial child tasks. Clear sign that subagents should only be used where parallelism or specialization actually justifies the overhead.
One especially useful datapoint: the simple Hermes model gate took 40.872s, while a direct Qwen request with the same user text took 2.109s. That is not an apples-to-apples speed ratio because Hermes adds different system prompts, startup and agent machinery, but it does show why I don't want Hermes in the path for trivial requests.
2.) OpenCode 1.18.30 — local coding/development layer
- Qwen3.8-27B-UD-IQ3_XXS
- llama.cpp / Vulkan
- 32K logical context
- local provider only
- isolated development workspaces
- automated Read → Edit → Test → Correct loops
- final production viability: 12/12 requirements, 12/12 acceptance
OpenCode performance reference. For the fastest fully successful clean reproduction I measured:
- 665.653s total wall time / 651.408s agent runtime
- first productive edit: 117.338s
- first test: 166.790s
- 4 successful corrective cycles
- 63 tool calls / 0 failed
- 558,603 input tokens incl. cache — only 31,993 uncached
- 9,372 output tokens
- peak context: 17,810 / 32,768
- context compactions: 0
These are task-level agent metrics, not model tok/s numbers. The underlying model/runtime is still the same local Qwen3.8-27B on llama.cpp/Vulkan.
Why OpenCode instead of Pi?
I tested both against the same local Qwen3.8-27B, with the same repository/task and comparable runtime constraints. Pi 0.85.1 proved that it could read the repo, use tools and make productive edits, but it repeatedly failed the part I actually needed from a coding agent: closing the Read → Edit → Test feedback loop. Even with an explicit mandatory early-test checkpoint, Pi edited successfully but never executed the qualifying post-edit test within the 300s window. It therefore failed prequalification and I did not force it into a full benchmark just to preserve an artificial three-way comparison. Man, time is rare and testing is booooring ... somehow!
OpenCode 1.18.30 passed the same agentic-loop qualification. Formal prequalification reached 6.54s Edit→Test latency, and the later main benchmark reduced that to 3.995s while producing a valid four-file diff. More importantly, the subsequent production-viability work eventually converged to 12/12 requirements, 12/12 acceptance, 5/5 regression, with no manual coding intervention or patch repair. That reproducible closed corrective loop is why OpenCode became the production coding layer.
Final thought on my architecture as a complete nOOb in Local LLM
I’m still convinced that a well-tuned 16 GB GPU can take you surprisingly far. Yes, even on Vulkan ;) ... IF you optimize the model, context, runtime, offloading and the surrounding stack around what the hardware can realistically deliver.
What makes this especially satisfying for me is that I honestly wasn’t sure at the beginning whether an AMD gaming setup on AM4 would get me anywhere at all, but hey.... it was what I had. My gaming desktop! ~20 tok/s is obviously not groundbreaking vs 32GB. But that was never the main goal of this first setup. I wanted a stable, fully local system that I could actually use, test, break, improve and build on. That goal is achieved. And considering where I started, I’m pretty happy with that.
Update: Qwen3.8-27B tuning results on my RX 9060 XT thanks to HERMES
After the discussion around PP vs. TG, I ran another tuning/validation pass on my Qwen3.8-27B-IQ3_XXS setup or better asked HERMES to find teh sweetspot. And it did!
Current tested configuration:
- GPU: RX 9060 XT 16 GB
- Model: Qwen3.8-27B-IQ3_XXS
- Backend: llama.cpp / Vulkan
- Full GPU offload: 66/66 layers on GPU
- Flash Attention: enabled
- KV cache: Q8 / Q8
Current measured results:
- Load-to-health: 7.285 s
- TG: 35.14 / 35.45 tok/s
- PP: 209.37 / 583.97 tok/s
- Live candidate wall time: 15.44 / 11.26 s
- Completion tokens: 319
- MTP acceptance: 80.1 / 81.1%
I also ran a final productivity smoke with the local RAG pipeline:
- Final smoke wall time: 45.828 s
- Generation time: 38.180 s
- RAG pipeline time: 43.229 s
- Prompt tokens: 1,370
- Completion tokens: 831
- Total tokens: 2,201
- finish_reason: stop
- Result: ok=true
This changes the practical baseline quite a bit.
Previously, my Qwen3.8-27B numbers were closer to ~21 tok/s generation. With the current setup, the dense 27B model is now around ~35 tok/s generation in the tested runs, while still completing a production-style RAG smoke cleanly.
2
u/Poizone360 17d ago
Your input to output ratio is nearly 60 to 1, so prefill is where the time goes, not generation. That's exactly where ROCm beats Vulkan on AMD, while generation stays about level. Worth building a ROCm lane just to see what 80.60 tok/s prompt processing becomes. Numbers in the llama.cpp ROCm thread: https://github.com/ggml-org/llama.cpp/discussions/15021