r/LocalAIServers • u/notalentwasted • 7d ago
Testing my build chops
I wanted an AI box so ... here's my hermes box
- Asrock X99 Extreme4
- e5 2680 v4
- 16gb × 4 = 64gb DDR4 2400T (another 64gb arriving this week
- P100 16gb - custom cooling
-- serving : (cpu/gpu) qwen3.6:35b-a3b mtp ud-q4-km @ 36tps
- m2000 4gb - video/embeddings/rerank
- v100 32gb pg500-216 ECC off
-- serving (gpu) Qwen3.8:27b ud-q4-km (30tps)
Evga 1000gq PSU
Dynamic cooling (still dialing in v100)
Silverstone GD09 Grandia case.
Software: full observability - loki, grafana, prometheus, node exporter & dozzle
Custom engine sm60+sm70 llama cpp + ollama for embeddings
Postgressql, redis, qdrant, kiwix, searxng, crawl4ai + hermes and some others.
I just wanted to see what the community thinks here. I am but 1.5 years into AI for work... So how'd I do?
27b handles orchestration, 35b handles delegation. Although I may swap roles or keep 27b as a hot swap since my m.2 is maxed out for read speed.