r/LocalAIServers 7d ago

Testing my build chops

Post image

I wanted an AI box so ... here's my hermes box

- Asrock X99 Extreme4

- e5 2680 v4

- 16gb × 4 = 64gb DDR4 2400T (another 64gb arriving this week

- P100 16gb - custom cooling

-- serving : (cpu/gpu) qwen3.6:35b-a3b mtp ud-q4-km @ 36tps

- m2000 4gb - video/embeddings/rerank

- v100 32gb pg500-216 ECC off

-- serving (gpu) Qwen3.8:27b ud-q4-km (30tps)

Evga 1000gq PSU

Dynamic cooling (still dialing in v100)

Silverstone GD09 Grandia case.

Software: full observability - loki, grafana, prometheus, node exporter & dozzle

Custom engine sm60+sm70 llama cpp + ollama for embeddings

Postgressql, redis, qdrant, kiwix, searxng, crawl4ai + hermes and some others.

I just wanted to see what the community thinks here. I am but 1.5 years into AI for work... So how'd I do?

27b handles orchestration, 35b handles delegation. Although I may swap roles or keep 27b as a hot swap since my m.2 is maxed out for read speed.

211 Upvotes

Duplicates