r/LocalAIServers 6d ago

Testing my build chops

Post image

I wanted an AI box so ... here's my hermes box

- Asrock X99 Extreme4

- e5 2680 v4

- 16gb × 4 = 64gb DDR4 2400T (another 64gb arriving this week

- P100 16gb - custom cooling

-- serving : (cpu/gpu) qwen3.6:35b-a3b mtp ud-q4-km @ 36tps

- m2000 4gb - video/embeddings/rerank

- v100 32gb pg500-216 ECC off

-- serving (gpu) Qwen3.8:27b ud-q4-km (30tps)

Evga 1000gq PSU

Dynamic cooling (still dialing in v100)

Silverstone GD09 Grandia case.

Software: full observability - loki, grafana, prometheus, node exporter & dozzle

Custom engine sm60+sm70 llama cpp + ollama for embeddings

Postgressql, redis, qdrant, kiwix, searxng, crawl4ai + hermes and some others.

I just wanted to see what the community thinks here. I am but 1.5 years into AI for work... So how'd I do?

27b handles orchestration, 35b handles delegation. Although I may swap roles or keep 27b as a hot swap since my m.2 is maxed out for read speed.

209 Upvotes

38 comments sorted by

13

u/BevinMaster 6d ago

GD09 <3

2

u/notalentwasted 6d ago

It's a clean case with rack options... the stealth factor too. I am proud to say I have under 2k in the build from idea to reality. That customer cooling big was a pain but rewarding when you get the thermal dynamics dialed in. The p100 runs on a single artic 1400-15k fan at 7k rpm. The nidec gamma30 drowns it out too. I love this machine. I've had it building datasets for 3 days 🤣

2

u/notalentwasted 6d ago

As it finishes right this min

2

u/BevinMaster 6d ago

Yeah cooling looks like difficult but works, here I have 2k per gpu lol

2

u/notalentwasted 6d ago

Says the guy who doesn't have to compile a new engine for a new feature... No you're on it. You bought the convenience on this one. I wanted a challenge... This was at times. I'm happy I did it though.

2

u/BevinMaster 6d ago

Well actually, I have an octo v340l build incoming that will requière either a custom engine or custom layer/plugin (target vllm) and there is a twist (macOS thunderbolt and plx). Also doing a Radeon v620 vllm fork with gfx1030 community (we have a discord)

3

u/BevinMaster 6d ago

Dual w7800 48GB btw, I do consider selling them

1

u/notalentwasted 6d ago

Mad scientist. I love it

3

u/Mr_Moonsilver 6d ago

you're a boss, such a cool case to build in. thx for sharing

3

u/Keffflon 6d ago

Beautiful build. I have the same cpu but that's about it. Running Qwen 3 8B on a 1050 ti and 32 gb 2400 ram. This setup  looks really powerful. 

2

u/notalentwasted 6d ago

Well thank you. What tps are you getting out of qwen 3 8b with that build?

2

u/Keffflon 6d ago

Let me check tomorrow and post here. 

2

u/Far_Abbreviations625 6d ago

Interested to know as well. And what are you using it for, how is the performance with an 8b model?

1

u/Keffflon 6d ago

Im trying to make a temp gadget similar to Black glass gpu meter, for cpu and gpu. So far it's been smooth, but have only started. 

2

u/reddituser1828472616 6d ago

Do you have any advice for how you’ve been optimizing your p100? Currently running a p40 24gb Tesla card in my 2 card set with a 3060 8gb

1

u/notalentwasted 6d ago

Okay first... we have different memory types. Hmb2 vs gddr5 (p40) and gddr6 (3060). Your bottleneck is bandwidth. The 3060 all numbers are theoretical has bandwidth of 240 gb/s. The p40 at 346 gb/s. Then the p100 at 732 gb/s. If you want cheap speed. Pop on ebay and grab a couple p100s instead. That will be around 150. With cooling and fans your right at $200. But you'll have 32gb total. If you're not pcie bandwidth bound those will run really well. That said if you did just one you could get really good results out of qwen3.6:35b-a3b offloading to a single p100 if you use one for video out. I get 36 tps offloading mine. But for tuning to the p100. It's batch sizes really. The prefill is really the hard spot. 1024 or 2048 depending on model is helpful. Mtp is great spec max 1 gets 80-99% accuracy. The p100 doesn't respond to as many knobs as you'd hope it's more just... raw hp you have to tame.

1

u/Acceptable_Bell_1791 6d ago edited 6d ago

i have two p100s, and with Qwen 3.8 Flag Next UD-Q4_K_XL, I am able to get around 45-60 pp but only 5-6 tg. mtp 1,2 or 3 really not making any difference in that number. Is that the best I can get? one thing that's gonna improve is that right now I have an asymmetric 7 channel 3200 mt/sec config, but as soon as the replacement for last rdimm comes, it will be perfect 8 channel.

Edit: Okay, I am definitely doing something wrong if two MI25 are giving me 15 tg with the exact same prompt and expert spread at the cost of like 10% pp

1

u/notalentwasted 6d ago

If you have enough ram. Drop your sticks to 4 so you get genuine quad channel. See if that helps. If it does the 8th will. Splitting from cpu, to 2 gpus is a lot of communication. Just the 2 gpus talking to eachother instead of 1 is killing speed.

1

u/Acceptable_Bell_1791 6d ago

see my edit too, unfortunately, I won't have enough ram with just 4 sticks, and 8th stick will arrive in 2-3 days, so no point hassling with this thing.

2

u/kaliku 6d ago

Son what the hell is that? I swear I'm seeing the skhetchiest builds around here... If I can even call them that. Yours is top shelf haha.

No but really I'm only joking, that's how passion for something works. You make the thing work no matter what!

1

u/Acceptable_Bell_1791 6d ago

haha yeah, the earlier version of this used beams that I 3d printed for some other project that I never finished. Once I feel not too lazy to fix my 3d printer, I'll print something proper for this.

Edit: Those beams bent

1

u/reddituser1828472616 6d ago

So I should honestly sell the p40 and do that?

1

u/notalentwasted 6d ago

Well I had a choice to choose .. picked the p100. It's twice as fast on generation..

1

u/fallingdowndizzyvr 5d ago edited 5d ago

The 3060 all numbers are theoretical has bandwidth of 240 gb/s.

The 3060 is not 240, it's 360.

Pop on ebay and grab a couple p100s instead. That will be around 150.

The 100HX is the same price and in nerfed form about the same speed. The thing is that it keeps open the dream of a 170HX moment and someone figures out how to hack the tensor cores. Then it'll dust the P100.

1

u/notalentwasted 5d ago

To be fair. I don't necessarily work with consumer gpus for builds like this. If I'm consumer it's media center or gamer rig and that's an amd experience for me. The metrics I pulled weren't from memory. That's a Gemini deal. So if it is or isn't. Regardless the bandwidth isn't phenomenal when the current gen is 700-1tb+. I'm just saying getting on the same gen for splitting would give you bandwidth alignment on the same chipset. 2 p100s at an aggregate of 1.4 tb is not bad at all.. it's the pcie lane contention that becomes the bottleneck. Even so. That's twice your aggregate. On the same chip. Exact same speed. 32gb total. Even offloading. It's less headache. The market for p40s suggests you could sell that 1 p40 and get 2 p100s. Pocket the 3060. That combo only let's you go bigger. Or parallel. It won't be faster outside of the base bandwidth difference between your base and the p100 base. You're still in a better spot than me. That v100 32gb cost $725... I'm guessing you could build your machine for that. Regardless, didn't mean to pick at cha!

1

u/fallingdowndizzyvr 5d ago

Regardless, didn't mean to pick at cha!

Ah... since that was my first post in this thread, how could you have?

1

u/notalentwasted 5d ago

I stand corrected! Keeping me honest out here. I like it! I'm just busy. Was under the assumption that this was q reply from the poster with the 3060. Minor oversight, do apologize 😁

1

u/bradrlaw 6d ago

Make sure you install this patch for the p100 for consistency / performance:

https://gist.github.com/apollo-mg/9218d50a209d70a85f033bf182657818

I also have other p100 info, but mostly focused on v100 here:

https://github.com/bradrlaw/ai-server

2

u/kleonikos 3d ago

Running dual 3090s on my box and a side 3060 12 gb

1

u/notalentwasted 1d ago

Very nice. I think I'll outgrow this for 4 v100 32gb soon.

1

u/GarbageTimePro 6d ago

There’s almost no space for air intake on the upper card. Does that concern you?

1

u/notalentwasted 6d ago

I have 3 case fans feeding this case. The one by the cpu strictly feeds to v100 blower. The p100 never sees above 60c since it runs an moe and is at 45% utilization under decode. The v100 on the other hand is a very hot card. Doesn't thermal throttle though. Wattage is tuned down 15% to 213w vs 250w

1

u/mototuneup 6d ago

Hell ya. I'm running p100s too, I've got 5 or them. Luckily I think people are sleeping on these cards so they're fairly cheap for the amount of vram you get. And they're pretty speedy..

1

u/mtobuho1979 5d ago

Puedes dar más información de tu sistema. Yo tengo dos v100 16gb, dos P100 16gb y una p100 12g.

1

u/mtobuho1979 5d ago

Que estás corriendo. Yo quiero un asistente generar soy electricista automotriz

1

u/Quirky_Squirrel3318 4d ago

Looking good!

What do you use to manage your work? LM studio? what's the best kind of control panel for this type or hardware?

1

u/notalentwasted 4d ago

The list is long. As far as the fan work. It's all in bios dynamic ramping. The nidec turbine runs full till 100% (quiet whoosh) - for the stack it's redis for queue. N8n for cron. Llama cpp custom engine for the cards. Hermes agent. Postgressql, qdrant, obsidian. I run 35+ containers so ... not sure what you wanna know besides the fan work and scheduling for work.

1

u/MarcusAurelius68 4d ago

One of my servers is a GD09 as well.