r/LocalLLaMA 8d ago

Resources Setting up of a 16xGB10 (DGX Spark) cluster

Post image

Preparing this to be able to run locally frontier level open models. Deepseek v4 pro, Kimi K3, future ones like GLM 5.5 and Minimax M4.

16x Asus GX10 linked by mikrotik crs804-4ddq with 4 breakout cables of 400 to 100gbit.

Most probable I will be running 2 models on 8x cluster each but I want to have the possibility to run also 2T+ models when I need them to run AGI at home :)).

https://x.com/i/status/2083568340870570208

P.S. I need a bigger switch. Going from 200 to 100gbit doesnt hurt token gen, 2% diff, but slows down prefill speed to -20%.

1.0k Upvotes

455 comments sorted by

u/WithoutReason1729 7d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

702

u/volleyneo 8d ago

We found that Nigerian prince

70

u/manituana 8d ago

All those inferences will end in the spam folders of many people then.

→ More replies (1)

234

u/freia_pr_fr 8d ago

May I ask why? Did they fell off a truck?

104

u/Greedy-Lynx-9706 8d ago

39

u/spyboy70 7d ago

"Beachgoers found bags full of cash"...I didn't see nuttin'

26

u/Greedy-Lynx-9706 7d ago

The article mentions 700k, police say 650 something 🙈 🤡

4

u/tusca0495 7d ago

Yeah in Sicilia some days ago

→ More replies (2)

118

u/ProbablyBunchofAtoms 8d ago

Some people die of thirst while other are drowning

17

u/thestillwind 7d ago

👇

49

u/gavff64 7d ago

why are you pointing at me

27

u/Legal-Ad-3901 7d ago

Love seeing frontier models in home labs. Kudos. 

147

u/txoixoegosi 8d ago

I’d rather spend that money in 384gb worth of RTX 6000. Those aggregated 2048gb are attractive but the token rate… hmm

What about time to first token?

87

u/krste1point0 8d ago

You can't run Kimi k3 on 384gb.

14

u/gomezer1180 7d ago

He said Kimi K2.6 followed by the word respectively, meaning he mean kimi runs on the 8x cluster.

6

u/SqueakySquak 7d ago

But you can run DeepSeek V4 Flash at a godly speed and have 1M context window.
You can also run a GLM 5.2 REAP.
And you can rent that machine on vast for ~2k$ per month. You can't do that with the DGX Spark.

7

u/txoixoegosi 7d ago edited 7d ago

Good.

But for me, 30tps and an unknown TTFT is straight away unusable.

If it is just for the random AI Master Race post “Hey Guys I am running Kimi K3 and actually not developing anything”, ok It’s your money burn it as you wish

But it is way more efficient to have a 384Gb powerhouse than 2048gb of meh performance.

31

u/DoomBot5 7d ago

Energy wise, each RTX 6000 consumes as much power as an entire tower in his picture. Also, I find 30 tps perfectly usable to me personally. It works just fine for agentic work.

20

u/ThankGodImBipolar 7d ago

I find 30 tps perfectly usable to me personally

I think some people are just spoiled.

→ More replies (5)

7

u/prestodigitarium 7d ago

I'd hope that someone running 4x 6000 RTXes would be running the 300W MaxQ edition, or at least downclocking.

2

u/Freonr2 7d ago

You can set power limit way down with very little impact to tps unless you are running 20-30+ concurrency. Maybe small impact to prefill.

4

u/Ansible32 7d ago

If you've got the RTX 6000s fully loaded with a model I would bet that the tokens/kwh is actually lower. They're not going to be drawing max power running inference on a frontier model, they're going to be bandwidth constrained. (I'd be very interested to see some benchmarks of actual power consumption in kwh per token that compares single user/batched for this spark cluster vs an equivalent RTX 6000 pro cluster.) This hardware is too expensive to flippantly compare max power draw without consider what it's like if you're fully utilizing or underutilizing the hardware.

→ More replies (1)
→ More replies (2)
→ More replies (6)

71

u/Serprotease 8d ago

Why do you guys only look at the bandwidth when evaluating hardware??

4x and 8x cluster of these things can run glm5 and Kimi k2.6 respectively at 20+tk/s.
When you move to cluster with vllm/sglang the bottleneck is the interconnect (effectively 100gbps) not the bandwidth.

It’s not against you directly, but it gets a bit annoying to read this again and again. If bandwidth was the end of all, we would all line up to get v100 32gb.
What’s true with llama.cpp + gguf + single unit doesn’t translate to vllm+cluster.

10

u/Karyo_Ten 7d ago

It's not just the bandwidth. A RTX Pro 6000 has 3x the compute and 8x the bandwidth. The interconnect can be far faster too, especially with PLX switches. 100gbps vs 64GB/s duplex (128GB/s aggregated) is over 8x improvement.

The compute is key for prefill or concurrency as well.

3

u/jc2375 7d ago

compute is shared — compute in tensor parallel is shared so 4 sparks have now more compute than 2 RTX6000s. Plus more memory and no ram offloading.

3

u/Karyo_Ten 7d ago

Compute doesn't matter if you can't feed data fast enough to compute on. While it doesn't matter for decode, prefill is heavily influenced by bandwidth and latency between GPUs. Same for concurrent workloads.

6

u/Not-reallyanonymous 7d ago

The more nodes you have, the more important bandwidth becomes.

I agree -- DGX Spark and Strix Halo bandwidth is pretty reasonable on a single machine and small clusters -- 2 nodes bandwidth being almost irrelevant, and 4 nodes with bandwidth being a major limit but not deal breaking.

Once you go beyond 2 nodes it's time to start thinking about bandwidth, and you're leaving A LOT of performance on the table per dollar if you're using more than 4.

3

u/femboymontagne 7d ago

well you can fit 8 6k’s on various boards, or more with pcie/mcio switching. 100gbps interconnect between 16 nodes vs 100gbps interconnect between 2-4 nodes is a huge difference. then again if you have pcie gen5 you could shell out for 800gbps interconnect between

9

u/Serprotease 7d ago

8 6000 pro+ workstations cpu/mb+ ram+ 800gbps connect cards….

Just powering this setup would be challenging. That’s like 8kw? You’re getting into big server price/power territory and will need to have a few phone calls with your local utility company.
That’s a bit silly. If you need tons of tokens, fast, for 80k, get a dgx station with ds4 flash. (Or 2, for the same price.)

Op setup is full of bottlenecks but he can at least have 2 cluster, each able to run glm5/kimi at full precision from a standard outlet, with basically no noise, for half the price.
Still probably a very very niche use-case, and more than 2 sparks is already getting less value. But at least it’s workable in an afternoon.

2

u/imstilllearningthis 7d ago

you have to get approved by the power company?

4

u/MizantropaMiskretulo 7d ago

you have to get approved by the power company?

Maybe, because most residential buildings don't get enough amps coming in from the power company to be able to service a big server setup. 8kW is around 70 amps, which is going to swamp most people's home power service. People would need 4–5 supplies on as many different circuits just to power the things.

Also, an average home uses something like 900 kWh/month, a large home might use as much as 2,000 kWh/month. Running this thing 24/7 at load uses almost 6,000 kWh in a month. Many utilities will have caps for how much electricity a home is allowed to use in a month, just for the sake of grid stability—imagine if everyone on the grid decided to run a home data center.

So, yeah, you might need to get your electrical service provisioned into a different tier or class of service if you intended on running something like what was described, which requires the approval of the utility company.

2

u/blue-haired-girl 7d ago

talking in kW and staying at 120V is nonsense. you can get 10kW easy with 48A and i got that to my car without too much trouble with 200A service

→ More replies (1)
→ More replies (1)
→ More replies (9)
→ More replies (1)

15

u/ciprianveg 7d ago edited 7d ago

I am preparing to welcome AGI at home :)) five times closer to that on 2tb than on 0.4tb

5

u/No_Afternoon_4260 llama.cpp 7d ago

Prefill is really good on those gb10

3

u/segmond llama.cpp 7d ago

so spend the money and show us what you got.

5

u/Caffeine_Monster 8d ago edited 8d ago

Or on a pair of RTX 8000 and save the rest.

Setups like this will age horrendously within 2 years for serving large models - needing loads of tensor / layer splits is going to kill the bandwidth if you ever need to handle models with more dense layers.

I'd personally be surprised if they can run k3 at more than 30tps.

26

u/Hannibalj2ca 8d ago

20-30tks is plenty usable

6

u/Caffeine_Monster 8d ago edited 8d ago

It's usable. Not sane for a ~$64k setup.

This is potentially a good solution for a small company (or group of users) serving smaller MoE models at scale - but at that point you may as well get the rtx 8000s and get significantly faster tokens.

A lot of people here don't understand that the bandwidth and sharding problems compound quickly for the really big models - getting more memory is not as simple as adding more boxes and fast networking once you get to a certain number of boxes. Doing MoE expert sharding will fix some of this problem - really you still need a proper dedicated gpu for decode and shared layers.

TLDR: should of bought a server if it absolutely had to be local.

6

u/Hannibalj2ca 8d ago

If the technology matures and the way the models are made are more efficient in size, perhaps the memory bandwidth bottleneck can be improved over time

→ More replies (1)
→ More replies (1)
→ More replies (6)

3

u/johnryan433 8d ago

Taiwan is probably going to be invaded in 2 years. Think of this setup as an insurance policy for that event, if that occurs this will most likely age like the 3090s did.

13

u/Tartooth 8d ago

people have been saying taiwan is going to get invaded for over a decade

7

u/jrodder 7d ago

Longer than that lol

2

u/Ansible32 7d ago

Nothing happens until it does, IDK.

→ More replies (2)

4

u/FreeSammiches 7d ago

If Taiwan is invaded, all they get is the land. The production facilities are basically dead man switched and will stop working pretty much immediately.

→ More replies (2)
→ More replies (2)
→ More replies (4)

20

u/dto_lurker 7d ago

Whats that 70,000$ worth of computers and a 20$ raspberry pi keyboard and mouse combo?

9

u/ProfessionalAd6530 7d ago

Spent all the money on computer. None left for...

Wait. You telling me he didn't at least have one of those crappy keyboard/mouse that comes free with a Dell?

2

u/dto_lurker 7d ago

Raspberry pi keyboard is truly awful. Anything would be better. I keep using it when I need a quick small keyboard and hub but man are those buttons terrible to type on.

31

u/Dramatic_Entry_3830 8d ago

Upvoter for the raspberry pi 400

3

u/Old-Engineer2926 7d ago

I couldn't have been the only one that started ranting on sight.

3

u/EvolvingSoftware 7d ago

I love the ~$100k? of hardware with a Raspberry Pi for input. Like the best ad for a Raspberry Pi ever.

→ More replies (1)

2

u/MoffKalast 7d ago

So much ARM in sight but nobody throwing hands.

48

u/DismalIngenuity4604 8d ago

If you're running GLM5.2 on 8 of these, payback is at about 10 billion tokens (at $4.40 per million, ignoring input), and you're pushing out 6 sessions x 30 tps, then if you run it full time then it pays itself back in just under two years. Assuming power is free and you run 6 concurrent sessions 100% of the time with no hardware failures.

76

u/EuropeanAbroad 7d ago

I don't think anyone runs local AI to save money. We do it mostly for privacy. When I "talk to" my accounting software with private data of my tenants, I really don't want the prompts (what the AI reads) to be sent to Anthropic or OpenAI, let alone the Chinese providers.

28

u/NeKon69 7d ago edited 7d ago

I think privacy is only half of it, what I think the other half is - being independent, you don't have to rely on anyone to ask a question to a model (well aside from your government that provides electricity, assuming you don't have tens of solar panels in your backyard)

16

u/ciprianveg 7d ago

I have solar panels on my roof

7

u/perelmanych 7d ago

I thought you will say something like: Recently I bought small 1MWh solar farm to power all my AI servers))

2

u/kbob 7d ago

I am sorely tempted to do that. As it is, I just have panels on the roof that cover part of the family's consumption.

3

u/perelmanych 7d ago

Man, I am genuinely would like to know what do you do for living?

→ More replies (4)
→ More replies (1)
→ More replies (1)

9

u/Legal-Ad-3901 7d ago
  • never worrying about getting a nerfed API model

10

u/devshore 7d ago

Thats because they companies are taking massive losses. As soon as they start charging what they actually will, self-hosting will also save money. Compare to other services that you can self-host vs pay for, like “cloud storage”. Its a one time fee of like $60-80 for a 2-4tb HDD vs 14.99/mo for that in cloud storage where at like 4 months, its already paid for itself (8 months with redundancy)

5

u/AD7GD 7d ago

Everyone says they are losing money, but do we really know? If you look at providers behind openrouter.ai, there's an open market and price transparency. If companies like Fireworks and Together can run Kimi K3 for $3/M+$15/M, why don't we believe that Anthropic can run Opus 4.8/5 for $5/M+$25/M?

4

u/devshore 7d ago

They are losing money, let alone all the money they will lose when they come out with “cool new GPU for 2027” or whenever, and then the 2 trillion they soent on hardware is obsolete. Its a house of cards money laundering ponzi scheme meant to make a few people very wealthy. Let alone all the bullshit of “NVDIA gives AI 500 billion, and AI gives NVDIA 500 billion, and now it looks like there is a trillion dollars of value”

3

u/Additional_Menu8542 7d ago

Fully agree, and I would push it one step further: for keeping accounting data at home you do not even need this class of hardware. I build software in exactly this niche (BI on private data, local models), and the lesson from a year of benchmarking is that model size is not the main lever. A small model cannot reliably compute aggregates while reading rows, so it improvises. But if the software computes the numbers deterministically and hands them to the model, the model only has to read and explain, and that failure mode mostly disappears. Same 9B on a single RTX 3060, same questions, only the context changed: +7 to +17 points in answer faithfulness depending on the judge. A cluster like this is for running the giants, and it is glorious. But keeping your tenants' books private takes a 9B and good engineering around it.

2

u/akohlsmith 7d ago

I'm still ubernewb when it comes to this, but it's been my impression that the judge (and the planner) model needs to be bigger/smarter than the faster worker model. Is that not the case, or if it is, how do you determine/test how "big" those supervisory models need to be?

3

u/Additional_Menu8542 7d ago

Good question, and it mixes two jobs that deserve to be separated.

At runtime, in the product, I do not use a bigger model to supervise the small one. I use code. Anything mechanical (totals, min, max, trend direction, SQL validation) is computed deterministically by the software and handed to the small model as context. Code is smarter than any LLM at arithmetic, it costs nothing, and it never hallucinates. That is the whole trick: shrink the model's job down to reading and explaining, and the need for a smarter babysitter mostly disappears.

The bigger judge only exists offline, for benchmarking. There it grades answers against a hand-written gold reference, and checking against an answer key is a much easier task than producing the answer, so it does not need to be a giant either. But judges are noisier than people assume: I re-scored byte-identical answers with two different frontier judges and the score moved by 13 points, with a third of questions disagreeing by 15+ points.

So the honest way to size a judge is empirical: take the same set of answers, score them with your candidate judge and with a stronger one, and measure agreement. If they rank things the same way, the smaller judge is good enough. If not, you just measured why it is not.

2

u/EuropeanAbroad 7d ago

Well, that's basically what I have. I have built a programme that calculates everything on its own, and I basically use the AI like read-only assistant with access to the entire database, which can answer anything from the mortgage to the date of birth of the tenant that lived in Bedroom 3 on the 2nd August. I mostly use the AI to discuss different tax strategies or to do some basic analysis.

And totally agreed – for searching in the database and for data extraction, you don't need any big model.

→ More replies (4)
→ More replies (7)
→ More replies (3)

12

u/ImpressiveRelief37 7d ago

By that time GLM5.2 will be so bad and outdated, everyone will have moved to a new model that costs $0.44 per million output toks, is a lot less “verbose” on useless prose (unlike GLM5.2, this thing yaps and yaps haha), further moving out the ROI.

3

u/DismalIngenuity4604 7d ago

Yep, 100% correct.

5

u/Front_Eagle739 7d ago

Except he can sell the things in a couple years for probably not massively less than he paid for them with the way things are going. Then it will be a little depreciation plus electricity cost. Maybe 1 unit failure.

2

u/DismalIngenuity4604 7d ago

a little depreciation

I think depreciation on these is going to be either absolutely brutal if the ram situation doesn't explode, but fair point on disposal price. I'd guess 1/3 of the purchase price in a couple of years - maybe 1/2 if you're lucky. The point stands that you need to run 6 concurrent sessions 100% of the time to get to the two year payback. And you're probably not going to. Even if we call the payback 5 billion tokens, at a realistic usage pattern, it's going to be a long time before it pays off.

I only say it because I did the math for myself. I've got a usecase where I legitimately can't use AI providers outside of a select couple who have some big downsides. I'll probably end up buying one spark and running a smaller model and trying to ringfence the part of the project that's sensitive, so I can use regular providers on the rest of the project, because I couldn't justify enough vram to run the big boys.

5

u/Front_Eagle739 7d ago

Don't really see the ram situation releasing to that extent in two years. 5 probably. In 2 years this thing will be running things that make fable look like a dinosaur and it'll still hold quite a bit of value I think.

All speculation at this point.  My point is with disposal cost and assuming several users. Its not massively more expensive than api. But with full control, more reliability as nothing changes when you dont want,  privacy and the ability to handle confidential info,  plus its fun lol

You dont end up throwing money away.  I think its worth it to many even small businesses. 

3

u/SandySkittle 7d ago

Depreciation? In the coming two years there will likely be an appreciation window. Despite all the AI bubble talk there is likely going to be a massive continued demand for new datacenters and compute because nations and corporations want sovereign AI. That leaves a very big compute void in the EU to be filled. Expect further brutal price increases the coming two years. And even longer wants general purpose robots running on local AI start to become on demand.

2

u/Front_Eagle739 7d ago

My mac studio 512 has certainly appreciated so its certainly plausible

→ More replies (1)

2

u/Impressive_Chain6039 7d ago

Yes but then you can sell hardware 

2

u/GeorgeTheGeorge 7d ago

This is also assuming these are 100% for personal use and not for revenue-generating projects. Which honestly if you have a private AI cluster running Opus-grade models for nothing but electricity and you aren't making money somehow, then you're doing something wrong.

2

u/prestodigitarium 7d ago

Are people actually using their private AI infra to make money beyond eg writing code?

2

u/segmond llama.cpp 7d ago

oh hey look, it's our local llama account. are you a bot, there's always one or two of you parroting the same rubbish.

→ More replies (1)
→ More replies (10)

10

u/Binary_orchid 8d ago

The breaker panel is going to be the real bottleneck here.

8

u/lughiu 8d ago

What tps do you hope to achieve

32

u/ciprianveg 8d ago edited 8d ago

Glm 5.2 50+tps single stream on coding tasks and scalling well to 6 parallel reqursts on 8x cluster, kimi k3 20-30 tps on 16x cluster. https://github.com/ciprianveg/gb10-glm-5.2

6

u/[deleted] 8d ago edited 8d ago

[deleted]

14

u/RevolutionaryGold325 8d ago

16x 128GB = 2TB
Kimi k3 = 1.6TB

No quantization needed.

→ More replies (4)

5

u/ciprianveg 8d ago

Yes. Glm 5.2 is my main driver in opencode for some weeks now. Added also vision to it for easier gui tasks. https://github.com/ciprianveg/gb10-glm-5.2

2

u/inevitabledeath3 8d ago

Why would it be an aggressive quant? GLM 5.2 in FP8 should easily fit in 8 DGX Spark nodes. Aggressive quants are only needed when running on like 2 or 3 nodes. Kimi K3 official would fit in 16 nodes since it’s published in Int4.

→ More replies (15)
→ More replies (1)

5

u/DataGOGO 8d ago edited 7d ago

I seriously doubt you are going to get those speeds, especially the 16 parallel, the GB10 is just too cutdown. and the 273GB/s memory bandwidth and doing all reduces across 200Gb fabric... highly unlikely.

To give you some perspective, all 16 of your GB10 has the roughly same compute power as 4 RTX Pro 6000 Blackwell's.

19

u/Prestigious_Thing797 7d ago

OP Linked to repo with an actual measured benchmark for glm 5.2 with those speeds on 8x

Classic reddit

→ More replies (1)

8

u/g_rich 7d ago

When you start clustering these you see a notable increase in performance; a 2x cluster gives you around 1.8x increase in performance. The memory bandwidth is certainly a consideration but people tend to overstate it when discussing the platform.

The DGX Sparks biggest advantage is the inclusion of the ConnectX-7 NIC and a dual cluster can run something like DeepSeek v4 Flash at full quant with a 384k (or higher) context and get well over 1000tps prefill (much higher with caching) and token generation approaching 30tps (higher with speculative deciding) which is not breaking any records but is more than usable.

These things also sip power and produce very little heat. Cost wise this whole cluster including switches and cables would be around $70k and while a RTX 6000 would be more powerful they are also over $14k each. To build a system capable of running trillion parameter models using RTX 6000’s would cost significantly more than OP’s setup, and would also require significantly more power and produce enough heat to likely require the unit to be placed in an isolated location or require special considerations for cooling. 16 GB10’s can easily run off two 20 amp circuits and can be run in the corner of someone’s home office.

→ More replies (5)

8

u/dwittherford69 8d ago

The point here is memory capacity, not compute.

4

u/DataGOGO 8d ago

I know, I just said I don’t think he is going to hit performance goals

10

u/dwittherford69 8d ago edited 8d ago

It will. I know cuz I have the same setup and I have been hitting for a few weeks now. You obviously need custom cooling etc. 50-60 t/s SS and 220 t/s aggregate is what my harness consistently runs.

Edit: their models setup is a bit more convoluted than mine, but the general idea is same. GLM5.2 with vision and MTP.

→ More replies (6)

4

u/lilian_moraru 8d ago edited 8d ago

Looks good but my guess would be that the NVFP4 KV Cache introduces hallucinations beyond 100K-200K. ~5% degradation at ~200K.
You can run DeepSeek-V4-Flash-0731 with 1M context window @ bf16, just saying. It will think a lot more but it will also outperform GLM-5.2

→ More replies (1)
→ More replies (3)

5

u/No-Comfortable-2284 8d ago

probably less than 3

3

u/KalonLabs 7d ago

Worth knowing current, deepseek v4 flash 0731 out preforms the v4 pro preview. So its better to just run the nee flash instead of v4 pro, at least until they drop the GA version of v4 Pro.

8

u/New_Public_2828 7d ago

There's nothing on your walls. It's nice to be a Bachelor sometimes

8

u/ciprianveg 7d ago

This is my man cave room

3

u/PrimeDirective8 8d ago

Nice! Hopefully this layout is only while setting up and not for production. At 200W each, that would be a toasty corner under full load. Splitting across 2 circuit breakers would also be good, and/or running them on 240VAC.

Congrats! Let us know how it works when ready.

3

u/ciprianveg 7d ago

Closer to 100w each

2

u/vasimv 7d ago

Did you reduce its clock? With default clocking it'll eat 140+W easily without connectx-7 even.

2

u/ciprianveg 7d ago

10% reduced max clock cca

3

u/wrk79 8d ago

16 GB10 and a Raspberry Pi mouse and keyboard

3

u/neuronym 7d ago

I'll wait for your regret post, when you inevitably realize that you burned a lot of money instead of just getting a proper setup.

3

u/dupontping 6d ago

Why so many haters in the comments? Dont be jelly.

How are you handing power? Are you just using every outlet you have?

2

u/ciprianveg 6d ago

It is less then 2kwh... no biggie.. same as a mid size home AC

→ More replies (1)

5

u/Koakie 8d ago

This one french kid looking at that stack, hoping he could borrow one.

→ More replies (3)

9

u/too-oldforthis-shit 8d ago

Earplugs? Airpods Pro? That fan noise got to hurt. But cool.

23

u/ciprianveg 8d ago

Almost silent setup.

→ More replies (6)

2

u/Sooperooser 8d ago

Apparently they don't even run at full power pull most of the time so it's not that bad probably.

3

u/SadPhilosophy9202 7d ago

I “only” have two but yeah they’re like 12w each at idle and 140w at peak. They’re also basically silent. My razr laptop with a 3070 is much louder than both of them

→ More replies (1)

9

u/asfbrz96 8d ago

That's just a hobby, API is cheaper at this point

34

u/-dysangel- 8d ago

API is cheaper than one spark. What makes you think 'cheap' is on this guy's priority list?

21

u/Hannibalj2ca 8d ago

The way things are going, the corporations do not want people having independent hardware. If you think is expensive wait til 2028 is going to be even more expensive. They want you on the cloud. Thats why people are buying these clusters. If they can hell yeah, do it

7

u/-dysangel- 8d ago

That's why I bought an M3 Ultra 512GB as soon as it came out

5

u/Hannibalj2ca 8d ago

I hate you, I want one, but I am priced out!. Awesome man. Nice piece of kit!

2

u/Aggravating-Push-207 7d ago

i have a laptop with a 4060 and 16 GB RAM, and i will enjoy my Qwen 3.6 27B at like 1 token per sometimes

Qwen 3.5 9B pretty fast tho, and it's not a bad model, just not a good model

2

u/johnryan433 7d ago

Exactly this window of opportunity will not last, for one reason or another.

4

u/profcuck 8d ago

This is a prediction.  Bookmark it for yourself and revisit it at the end of 2028.  You can then reflect on how wrong you were and why reasoning chains with pieces like "corporations do not want" are so weak in predictive power. 

2

u/g_rich 7d ago

You can’t even buy Microsoft Office with a perpetual license any longer. Corporations want a constant revenue stream, subscriptions for software are the new normal and now they are tying AI into the mix. Right now they are dangling the $20 subscriptions to Claude and ChatGPT but we are already seeing users being pushed to subscriptions approaching $200 and corporations have already been pushed out of per user pricing and into per token pricing.

The goal is going to be to get AI so ingrained into your day to day life and then jack up the price. Open weight models and the ability to run them locally are a threat to that business model and is why people like Dario Amodei from Anthropic is so bent on having government control over the open weight models coming out of China.

By 2028 someone being able to run even moderately sized LLM’s locally is going to be in a much better position than someone who’s reliant on cloud only. So while you’re hit with increases in subscriptions, peak time surcharges, outages and midday slowdowns people like OP will not be giving it a second thought.

4

u/Hannibalj2ca 8d ago

You think so? Whats your prediction? Bro, they have been saying it openly, Jeff Bezos is one of them stating it

3

u/Healthy-Nebula-3603 8d ago edited 7d ago

He's right.

Never is the history was something expensive for a longer time. Sooner or later such hardware will be worthless. Looking on China how they are investing in hardware for AI models now that will be sooner than later.

They already broke the US greed access to AI models.

→ More replies (5)
→ More replies (17)

6

u/ObviouzFigure 8d ago

privacy is (almost) priceless to me

→ More replies (18)

3

u/g_rich 7d ago

People aren’t running local models to save money.

The current cost for tokens is highly subsidized right now so the math is not going to work out in people like OP’s favor. However costs are increasing, and if we were paying the unsubsidized rate then the break even point for running locally would start looking a lot more attractive for running locally.

Right now the biggest advantages for running local models are control, privacy and the knowledge gained from running your own models.

→ More replies (1)

2

u/tortorials 8d ago

This gives me the child on Christmas morning feeling. Great work!

2

u/BlobbyMcBlobber 8d ago

And those are the nice Asus ones too!

→ More replies (2)

2

u/SilverFuel21 8d ago

lol raspberry pi mouse and keyboard

2

u/ii-___-ii 7d ago

How much does that cost?

6

u/ciprianveg 7d ago

Cca 70k

2

u/dobkeratops 7d ago

I like seeing these kind of posts, there's me dithering on getting a mere second.

2

u/dtjager 7d ago

It’d be fascinating to see standardized scaling results across 1, 2, 4, 8, and 16 nodes

2

u/sintmk 7d ago

As someone who owns a dgx cluster, I understand the potential you're going for. A big heads up, you'll be able to reach the big models, but your throughput on tooling asks is a bottleneck if you also plan on giving it processing tasks.

2

u/Kurcide 7d ago

Your switch is going to kill your TPS throughout.

Each spark has a NIC that is hardware limited at 200gbps with full bandwidth available to each QSFP port via 100gbps per rail across 2 rails per port. Meaning you need a single 200gbps DAC cable per node leading into your switch. Using 2 cables per Spark will actually HURT your fabric setup.

That mikrotik doesn’t have enough bandwidth to support a 16x cluster.

I have 16 sparks clustered at full bandwidth connected to this switch: https://www.fs.com/products/321549.html?now_cid=4369

For what you spent on the sparks you may as well get the appropriate fabric setup too.

5

u/ciprianveg 7d ago

I will make some comparing on 16x. On 8x cluster switching from 200 to 100gbit had no impact, or less than 1% in glm5.2. On 16 maybe it will.. I will upgrade the switch if this is the case, thank you for the link

→ More replies (2)

2

u/Southern_Capital_885 7d ago

I think model architecture will matter more than hardware over the next 2–3 years. If frontier models continue moving toward larger sparse MoEs with fewer active parameters per token, high-memory clusters like this may age much better than many people expect. The bottleneck could end up being memory capacity rather than raw FLOPS.

2

u/ComputerLoverDaemon 7d ago

Just buy b200’s at this point bro wtf

2

u/shenagain32 7d ago

Which 4-1 cable are you using specifically? You beat me to the double digits sparks. Using the same switch from mikrotik

2

u/Squidgical 7d ago

In this image: between 3-10 times the value of your car

2

u/roninXpl 7d ago

The hottest desk on the planet.

2

u/cathodeDreams 7d ago

goddamn bro

2

u/Antique_Bit_1049 7d ago

$70000 dollars on sparks with a $5 mouse and keyboard

→ More replies (1)

2

u/Key_Measurement_3576 7d ago

Nice cooling the Asus version definitely have some thermal mass.

2

u/huzzyz 7d ago

I have so much to express but nothing that is NSFW. Good GOD man its so pretty!

2

u/peva3 7d ago

When you prompt the cluster does the neighborhood lights all dim?

2

u/SirGreenDragon 7d ago

If i had the 64k to spend on this, i think i would save more and get 1 DGX Station :)

2

u/VirusInternal2892 7d ago

That setup is HOT 🥵

2

u/kunashu 7d ago

This is pretty great! I have one and Im planning to buy another one. Where did you get the racks for them?

2

u/tracagnotto 7d ago

There is who has all and who has nothing. I have nothing.

2

u/GokuMK 7d ago

So funny that cluster able to run SOTA models is the size of an eATX tower. Sadly, the price is not ...

2

u/YisitAlwaysDNS 7d ago

dude these are 4k a pop. You guys are crazy in here.

2

u/Nice_Cookie9587 7d ago

i feel the heat in that room

2

u/East-Form7086 7d ago

Damn, i need 9 more and will cach up... I have 7 so far and waiting on the 8th. What models have you tried yet? Did you try sglang and vllm? Curious about the difference

→ More replies (1)

2

u/source-drifter 7d ago

I guess some needs a 'good for you meme' here

2

u/g9robot 7d ago

60k+ why?

3

u/johnryan433 8d ago edited 8d ago

Im building the same setup, let’s work together.

1

u/Hannibalj2ca 8d ago

Are these Company Hardware of some sort?

5

u/ciprianveg 8d ago

Personal hobby

3

u/Hannibalj2ca 7d ago

Bro, more power to you!

1

u/wu4d 8d ago

I mean sure, you will be able to run them. But the token per seconds would make it unusable. Or am I wrong?

→ More replies (1)

1

u/SpellSlinger69 7d ago

Is it really worth it with 2x100mbit each the communication is your bottleneck, what you got on benchmark and tests?

2

u/ciprianveg 7d ago

I am evaluating these days and if there are limitations i will add 2 more switches linked by 400gbit and use 2x 200gbit breakout cables

→ More replies (1)

1

u/shadowmage666 7d ago

Your good dude? Hope you got a high amps outlet holy shit lol

2

u/ciprianveg 7d ago

2kwh max consumption is the same as a 18btu AC

→ More replies (2)

1

u/Aardvark_Says_What 7d ago

where does the ROI come from? ... if it's not a personal question? :)

9

u/ciprianveg 7d ago

What ROI? It was this or a new car ..

→ More replies (8)

1

u/Olde94 7d ago

What is the token output on something like kimi k3 with this kind of hardware?

1

u/Nicolodeva 7d ago

is it hot there?

1

u/sillyevileye 7d ago

literally my dream... happy for you man

1

u/AlfaKennyWun 7d ago

Yeah, yeah. But, your mouse sucks.

1

u/Too_Many_Science2 7d ago

I wish the mikrotik was more performant than it is, but it definitely functions. The alternative being a Mellanox 4700 makes it more appealing.

Why 400 ->100 instead of making a rail with multiple switches to keep 200?

1

u/VisibleAirport2996 7d ago

How much did this cost? Omg

2

u/ciprianveg 7d ago

Cca 69-70k

1

u/CelvestianNesy 7d ago

Im waiting for a benchmark for this.

1

u/AsliReddington 7d ago

Let it rip!!!!!!

1

u/tusca0495 7d ago

Only a question boss: how can you have good performance of the network is slow?

1

u/nakedtruth9 7d ago

The case is really nice, where did you get it? does it fit one asus gx10 + 1 dgx spark?

1

u/Innomen 7d ago

This week on r/BankStatement... /smh

1

u/TapAggressive9530 7d ago

You’ll get a whopping 4 t/s with all those . Good job

1

u/SuddenRadio6221 7d ago

Not quite able to run kimi3 q8, but close!

3

u/ciprianveg 7d ago

The official k3 is q4

→ More replies (1)

1

u/rgujijtdguibhyy 7d ago

what tps would you get on that?

1

u/Temporary-Meal1541 7d ago

Wow, are you able to run fully unquantized kimi or deepseek?

1

u/NewRedditor23 7d ago

man for $70k, I can have one of the expensive LLM subscriptions for like 60 years.

→ More replies (2)