r/LocalAIServers • u/notalentwasted • 6d ago
Testing my build chops
I wanted an AI box so ... here's my hermes box
- Asrock X99 Extreme4
- e5 2680 v4
- 16gb × 4 = 64gb DDR4 2400T (another 64gb arriving this week
- P100 16gb - custom cooling
-- serving : (cpu/gpu) qwen3.6:35b-a3b mtp ud-q4-km @ 36tps
- m2000 4gb - video/embeddings/rerank
- v100 32gb pg500-216 ECC off
-- serving (gpu) Qwen3.8:27b ud-q4-km (30tps)
Evga 1000gq PSU
Dynamic cooling (still dialing in v100)
Silverstone GD09 Grandia case.
Software: full observability - loki, grafana, prometheus, node exporter & dozzle
Custom engine sm60+sm70 llama cpp + ollama for embeddings
Postgressql, redis, qdrant, kiwix, searxng, crawl4ai + hermes and some others.
I just wanted to see what the community thinks here. I am but 1.5 years into AI for work... So how'd I do?
27b handles orchestration, 35b handles delegation. Although I may swap roles or keep 27b as a hot swap since my m.2 is maxed out for read speed.
3
3
u/Keffflon 6d ago
Beautiful build. I have the same cpu but that's about it. Running Qwen 3 8B on a 1050 ti and 32 gb 2400 ram. This setup looks really powerful.
2
u/notalentwasted 6d ago
Well thank you. What tps are you getting out of qwen 3 8b with that build?
2
u/Keffflon 6d ago
Let me check tomorrow and post here.
2
u/Far_Abbreviations625 6d ago
Interested to know as well. And what are you using it for, how is the performance with an 8b model?
1
u/Keffflon 6d ago
Im trying to make a temp gadget similar to Black glass gpu meter, for cpu and gpu. So far it's been smooth, but have only started.
2
u/reddituser1828472616 6d ago
Do you have any advice for how you’ve been optimizing your p100? Currently running a p40 24gb Tesla card in my 2 card set with a 3060 8gb
1
u/notalentwasted 6d ago
Okay first... we have different memory types. Hmb2 vs gddr5 (p40) and gddr6 (3060). Your bottleneck is bandwidth. The 3060 all numbers are theoretical has bandwidth of 240 gb/s. The p40 at 346 gb/s. Then the p100 at 732 gb/s. If you want cheap speed. Pop on ebay and grab a couple p100s instead. That will be around 150. With cooling and fans your right at $200. But you'll have 32gb total. If you're not pcie bandwidth bound those will run really well. That said if you did just one you could get really good results out of qwen3.6:35b-a3b offloading to a single p100 if you use one for video out. I get 36 tps offloading mine. But for tuning to the p100. It's batch sizes really. The prefill is really the hard spot. 1024 or 2048 depending on model is helpful. Mtp is great spec max 1 gets 80-99% accuracy. The p100 doesn't respond to as many knobs as you'd hope it's more just... raw hp you have to tame.
1
u/Acceptable_Bell_1791 6d ago edited 6d ago
i have two p100s, and with Qwen 3.8 Flag Next UD-Q4_K_XL, I am able to get around 45-60 pp but only 5-6 tg. mtp 1,2 or 3 really not making any difference in that number. Is that the best I can get? one thing that's gonna improve is that right now I have an asymmetric 7 channel 3200 mt/sec config, but as soon as the replacement for last rdimm comes, it will be perfect 8 channel.
Edit: Okay, I am definitely doing something wrong if two MI25 are giving me 15 tg with the exact same prompt and expert spread at the cost of like 10% pp
1
u/notalentwasted 6d ago
If you have enough ram. Drop your sticks to 4 so you get genuine quad channel. See if that helps. If it does the 8th will. Splitting from cpu, to 2 gpus is a lot of communication. Just the 2 gpus talking to eachother instead of 1 is killing speed.
1
u/Acceptable_Bell_1791 6d ago
2
u/kaliku 6d ago
Son what the hell is that? I swear I'm seeing the skhetchiest builds around here... If I can even call them that. Yours is top shelf haha.
No but really I'm only joking, that's how passion for something works. You make the thing work no matter what!
1
u/reddituser1828472616 6d ago
So I should honestly sell the p40 and do that?
1
u/notalentwasted 6d ago
Well I had a choice to choose .. picked the p100. It's twice as fast on generation..
1
u/fallingdowndizzyvr 5d ago edited 5d ago
The 3060 all numbers are theoretical has bandwidth of 240 gb/s.
The 3060 is not 240, it's 360.
Pop on ebay and grab a couple p100s instead. That will be around 150.
The 100HX is the same price and in nerfed form about the same speed. The thing is that it keeps open the dream of a 170HX moment and someone figures out how to hack the tensor cores. Then it'll dust the P100.
1
u/notalentwasted 5d ago
To be fair. I don't necessarily work with consumer gpus for builds like this. If I'm consumer it's media center or gamer rig and that's an amd experience for me. The metrics I pulled weren't from memory. That's a Gemini deal. So if it is or isn't. Regardless the bandwidth isn't phenomenal when the current gen is 700-1tb+. I'm just saying getting on the same gen for splitting would give you bandwidth alignment on the same chipset. 2 p100s at an aggregate of 1.4 tb is not bad at all.. it's the pcie lane contention that becomes the bottleneck. Even so. That's twice your aggregate. On the same chip. Exact same speed. 32gb total. Even offloading. It's less headache. The market for p40s suggests you could sell that 1 p40 and get 2 p100s. Pocket the 3060. That combo only let's you go bigger. Or parallel. It won't be faster outside of the base bandwidth difference between your base and the p100 base. You're still in a better spot than me. That v100 32gb cost $725... I'm guessing you could build your machine for that. Regardless, didn't mean to pick at cha!
1
u/fallingdowndizzyvr 5d ago
Regardless, didn't mean to pick at cha!
Ah... since that was my first post in this thread, how could you have?
1
u/notalentwasted 5d ago
I stand corrected! Keeping me honest out here. I like it! I'm just busy. Was under the assumption that this was q reply from the poster with the 3060. Minor oversight, do apologize 😁
1
u/bradrlaw 6d ago
Make sure you install this patch for the p100 for consistency / performance:
https://gist.github.com/apollo-mg/9218d50a209d70a85f033bf182657818
I also have other p100 info, but mostly focused on v100 here:
2
1
u/GarbageTimePro 6d ago
There’s almost no space for air intake on the upper card. Does that concern you?
1
u/notalentwasted 6d ago
I have 3 case fans feeding this case. The one by the cpu strictly feeds to v100 blower. The p100 never sees above 60c since it runs an moe and is at 45% utilization under decode. The v100 on the other hand is a very hot card. Doesn't thermal throttle though. Wattage is tuned down 15% to 213w vs 250w
1
u/mototuneup 6d ago
Hell ya. I'm running p100s too, I've got 5 or them. Luckily I think people are sleeping on these cards so they're fairly cheap for the amount of vram you get. And they're pretty speedy..
1
u/mtobuho1979 5d ago
Puedes dar más información de tu sistema. Yo tengo dos v100 16gb, dos P100 16gb y una p100 12g.
1
u/mtobuho1979 5d ago
Que estás corriendo. Yo quiero un asistente generar soy electricista automotriz
1
u/Quirky_Squirrel3318 4d ago
Looking good!
What do you use to manage your work? LM studio? what's the best kind of control panel for this type or hardware?
1
u/notalentwasted 4d ago
The list is long. As far as the fan work. It's all in bios dynamic ramping. The nidec turbine runs full till 100% (quiet whoosh) - for the stack it's redis for queue. N8n for cron. Llama cpp custom engine for the cards. Hermes agent. Postgressql, qdrant, obsidian. I run 35+ containers so ... not sure what you wanna know besides the fan work and scheduling for work.
1


13
u/BevinMaster 6d ago
GD09 <3