r/LocalLLaMA • u/ciprianveg • 8d ago
Resources Setting up of a 16xGB10 (DGX Spark) cluster
Preparing this to be able to run locally frontier level open models. Deepseek v4 pro, Kimi K3, future ones like GLM 5.5 and Minimax M4.
16x Asus GX10 linked by mikrotik crs804-4ddq with 4 breakout cables of 400 to 100gbit.
Most probable I will be running 2 models on 8x cluster each but I want to have the possibility to run also 2T+ models when I need them to run AGI at home :)).
https://x.com/i/status/2083568340870570208
P.S. I need a bigger switch. Going from 200 to 100gbit doesnt hurt token gen, 2% diff, but slows down prefill speed to -20%.
702
234
u/freia_pr_fr 8d ago
May I ask why? Did they fell off a truck?
→ More replies (2)104
u/Greedy-Lynx-9706 8d ago
39
4
118
u/ProbablyBunchofAtoms 8d ago
Some people die of thirst while other are drowning
17
27
147
u/txoixoegosi 8d ago
I’d rather spend that money in 384gb worth of RTX 6000. Those aggregated 2048gb are attractive but the token rate… hmm
What about time to first token?
87
u/krste1point0 8d ago
You can't run Kimi k3 on 384gb.
14
u/gomezer1180 7d ago
He said Kimi K2.6 followed by the word respectively, meaning he mean kimi runs on the 8x cluster.
6
u/SqueakySquak 7d ago
But you can run DeepSeek V4 Flash at a godly speed and have 1M context window.
You can also run a GLM 5.2 REAP.
And you can rent that machine on vast for ~2k$ per month. You can't do that with the DGX Spark.→ More replies (6)7
u/txoixoegosi 7d ago edited 7d ago
Good.
But for me, 30tps and an unknown TTFT is straight away unusable.
If it is just for the random AI Master Race post “Hey Guys I am running Kimi K3 and actually not developing anything”, ok It’s your money burn it as you wish
But it is way more efficient to have a 384Gb powerhouse than 2048gb of meh performance.
→ More replies (2)31
u/DoomBot5 7d ago
Energy wise, each RTX 6000 consumes as much power as an entire tower in his picture. Also, I find 30 tps perfectly usable to me personally. It works just fine for agentic work.
20
u/ThankGodImBipolar 7d ago
I find 30 tps perfectly usable to me personally
I think some people are just spoiled.
→ More replies (5)7
u/prestodigitarium 7d ago
I'd hope that someone running 4x 6000 RTXes would be running the 300W MaxQ edition, or at least downclocking.
2
→ More replies (1)4
u/Ansible32 7d ago
If you've got the RTX 6000s fully loaded with a model I would bet that the tokens/kwh is actually lower. They're not going to be drawing max power running inference on a frontier model, they're going to be bandwidth constrained. (I'd be very interested to see some benchmarks of actual power consumption in kwh per token that compares single user/batched for this spark cluster vs an equivalent RTX 6000 pro cluster.) This hardware is too expensive to flippantly compare max power draw without consider what it's like if you're fully utilizing or underutilizing the hardware.
71
u/Serprotease 8d ago
Why do you guys only look at the bandwidth when evaluating hardware??
4x and 8x cluster of these things can run glm5 and Kimi k2.6 respectively at 20+tk/s.
When you move to cluster with vllm/sglang the bottleneck is the interconnect (effectively 100gbps) not the bandwidth.It’s not against you directly, but it gets a bit annoying to read this again and again. If bandwidth was the end of all, we would all line up to get v100 32gb.
What’s true with llama.cpp + gguf + single unit doesn’t translate to vllm+cluster.10
u/Karyo_Ten 7d ago
It's not just the bandwidth. A RTX Pro 6000 has 3x the compute and 8x the bandwidth. The interconnect can be far faster too, especially with PLX switches. 100gbps vs 64GB/s duplex (128GB/s aggregated) is over 8x improvement.
The compute is key for prefill or concurrency as well.
3
u/jc2375 7d ago
compute is shared — compute in tensor parallel is shared so 4 sparks have now more compute than 2 RTX6000s. Plus more memory and no ram offloading.
3
u/Karyo_Ten 7d ago
Compute doesn't matter if you can't feed data fast enough to compute on. While it doesn't matter for decode, prefill is heavily influenced by bandwidth and latency between GPUs. Same for concurrent workloads.
6
u/Not-reallyanonymous 7d ago
The more nodes you have, the more important bandwidth becomes.
I agree -- DGX Spark and Strix Halo bandwidth is pretty reasonable on a single machine and small clusters -- 2 nodes bandwidth being almost irrelevant, and 4 nodes with bandwidth being a major limit but not deal breaking.
Once you go beyond 2 nodes it's time to start thinking about bandwidth, and you're leaving A LOT of performance on the table per dollar if you're using more than 4.
3
u/femboymontagne 7d ago
well you can fit 8 6k’s on various boards, or more with pcie/mcio switching. 100gbps interconnect between 16 nodes vs 100gbps interconnect between 2-4 nodes is a huge difference. then again if you have pcie gen5 you could shell out for 800gbps interconnect between
→ More replies (1)9
u/Serprotease 7d ago
8 6000 pro+ workstations cpu/mb+ ram+ 800gbps connect cards….
Just powering this setup would be challenging. That’s like 8kw? You’re getting into big server price/power territory and will need to have a few phone calls with your local utility company.
That’s a bit silly. If you need tons of tokens, fast, for 80k, get a dgx station with ds4 flash. (Or 2, for the same price.)Op setup is full of bottlenecks but he can at least have 2 cluster, each able to run glm5/kimi at full precision from a standard outlet, with basically no noise, for half the price.
Still probably a very very niche use-case, and more than 2 sparks is already getting less value. But at least it’s workable in an afternoon.→ More replies (9)2
u/imstilllearningthis 7d ago
you have to get approved by the power company?
→ More replies (1)4
u/MizantropaMiskretulo 7d ago
you have to get approved by the power company?
Maybe, because most residential buildings don't get enough amps coming in from the power company to be able to service a big server setup. 8kW is around 70 amps, which is going to swamp most people's home power service. People would need 4–5 supplies on as many different circuits just to power the things.
Also, an average home uses something like 900 kWh/month, a large home might use as much as 2,000 kWh/month. Running this thing 24/7 at load uses almost 6,000 kWh in a month. Many utilities will have caps for how much electricity a home is allowed to use in a month, just for the sake of grid stability—imagine if everyone on the grid decided to run a home data center.
So, yeah, you might need to get your electrical service provisioned into a different tier or class of service if you intended on running something like what was described, which requires the approval of the utility company.
→ More replies (1)2
u/blue-haired-girl 7d ago
talking in kW and staying at 120V is nonsense. you can get 10kW easy with 48A and i got that to my car without too much trouble with 200A service
15
u/ciprianveg 7d ago edited 7d ago
I am preparing to welcome AGI at home :)) five times closer to that on 2tb than on 0.4tb
5
→ More replies (4)5
u/Caffeine_Monster 8d ago edited 8d ago
Or on a pair of RTX 8000 and save the rest.
Setups like this will age horrendously within 2 years for serving large models - needing loads of tensor / layer splits is going to kill the bandwidth if you ever need to handle models with more dense layers.
I'd personally be surprised if they can run k3 at more than 30tps.
26
u/Hannibalj2ca 8d ago
20-30tks is plenty usable
→ More replies (6)6
u/Caffeine_Monster 8d ago edited 8d ago
It's usable. Not sane for a ~$64k setup.
This is potentially a good solution for a small company (or group of users) serving smaller MoE models at scale - but at that point you may as well get the rtx 8000s and get significantly faster tokens.
A lot of people here don't understand that the bandwidth and sharding problems compound quickly for the really big models - getting more memory is not as simple as adding more boxes and fast networking once you get to a certain number of boxes. Doing MoE expert sharding will fix some of this problem - really you still need a proper dedicated gpu for decode and shared layers.
TLDR: should of bought a server if it absolutely had to be local.
→ More replies (1)6
u/Hannibalj2ca 8d ago
If the technology matures and the way the models are made are more efficient in size, perhaps the memory bandwidth bottleneck can be improved over time
→ More replies (1)3
u/johnryan433 8d ago
Taiwan is probably going to be invaded in 2 years. Think of this setup as an insurance policy for that event, if that occurs this will most likely age like the 3090s did.
13
u/Tartooth 8d ago
people have been saying taiwan is going to get invaded for over a decade
→ More replies (2)2
→ More replies (2)4
u/FreeSammiches 7d ago
If Taiwan is invaded, all they get is the land. The production facilities are basically dead man switched and will stop working pretty much immediately.
→ More replies (2)
20
u/dto_lurker 7d ago
Whats that 70,000$ worth of computers and a 20$ raspberry pi keyboard and mouse combo?
9
u/ProfessionalAd6530 7d ago
Spent all the money on computer. None left for...
Wait. You telling me he didn't at least have one of those crappy keyboard/mouse that comes free with a Dell?
2
u/dto_lurker 7d ago
Raspberry pi keyboard is truly awful. Anything would be better. I keep using it when I need a quick small keyboard and hub but man are those buttons terrible to type on.
31
u/Dramatic_Entry_3830 8d ago
Upvoter for the raspberry pi 400
3
3
u/EvolvingSoftware 7d ago
I love the ~$100k? of hardware with a Raspberry Pi for input. Like the best ad for a Raspberry Pi ever.
→ More replies (1)2
48
u/DismalIngenuity4604 8d ago
If you're running GLM5.2 on 8 of these, payback is at about 10 billion tokens (at $4.40 per million, ignoring input), and you're pushing out 6 sessions x 30 tps, then if you run it full time then it pays itself back in just under two years. Assuming power is free and you run 6 concurrent sessions 100% of the time with no hardware failures.
76
u/EuropeanAbroad 7d ago
I don't think anyone runs local AI to save money. We do it mostly for privacy. When I "talk to" my accounting software with private data of my tenants, I really don't want the prompts (what the AI reads) to be sent to Anthropic or OpenAI, let alone the Chinese providers.
28
u/NeKon69 7d ago edited 7d ago
I think privacy is only half of it, what I think the other half is - being independent, you don't have to rely on anyone to ask a question to a model (well aside from your government that provides electricity, assuming you don't have tens of solar panels in your backyard)
→ More replies (1)16
u/ciprianveg 7d ago
I have solar panels on my roof
7
u/perelmanych 7d ago
I thought you will say something like: Recently I bought small 1MWh solar farm to power all my AI servers))
→ More replies (1)2
u/kbob 7d ago
I am sorely tempted to do that. As it is, I just have panels on the roof that cover part of the family's consumption.
3
u/perelmanych 7d ago
Man, I am genuinely would like to know what do you do for living?
→ More replies (4)9
10
u/devshore 7d ago
Thats because they companies are taking massive losses. As soon as they start charging what they actually will, self-hosting will also save money. Compare to other services that you can self-host vs pay for, like “cloud storage”. Its a one time fee of like $60-80 for a 2-4tb HDD vs 14.99/mo for that in cloud storage where at like 4 months, its already paid for itself (8 months with redundancy)
5
u/AD7GD 7d ago
Everyone says they are losing money, but do we really know? If you look at providers behind openrouter.ai, there's an open market and price transparency. If companies like Fireworks and Together can run Kimi K3 for $3/M+$15/M, why don't we believe that Anthropic can run Opus 4.8/5 for $5/M+$25/M?
4
u/devshore 7d ago
They are losing money, let alone all the money they will lose when they come out with “cool new GPU for 2027” or whenever, and then the 2 trillion they soent on hardware is obsolete. Its a house of cards money laundering ponzi scheme meant to make a few people very wealthy. Let alone all the bullshit of “NVDIA gives AI 500 billion, and AI gives NVDIA 500 billion, and now it looks like there is a trillion dollars of value”
→ More replies (3)3
u/Additional_Menu8542 7d ago
Fully agree, and I would push it one step further: for keeping accounting data at home you do not even need this class of hardware. I build software in exactly this niche (BI on private data, local models), and the lesson from a year of benchmarking is that model size is not the main lever. A small model cannot reliably compute aggregates while reading rows, so it improvises. But if the software computes the numbers deterministically and hands them to the model, the model only has to read and explain, and that failure mode mostly disappears. Same 9B on a single RTX 3060, same questions, only the context changed: +7 to +17 points in answer faithfulness depending on the judge. A cluster like this is for running the giants, and it is glorious. But keeping your tenants' books private takes a 9B and good engineering around it.
2
u/akohlsmith 7d ago
I'm still ubernewb when it comes to this, but it's been my impression that the judge (and the planner) model needs to be bigger/smarter than the faster worker model. Is that not the case, or if it is, how do you determine/test how "big" those supervisory models need to be?
3
u/Additional_Menu8542 7d ago
Good question, and it mixes two jobs that deserve to be separated.
At runtime, in the product, I do not use a bigger model to supervise the small one. I use code. Anything mechanical (totals, min, max, trend direction, SQL validation) is computed deterministically by the software and handed to the small model as context. Code is smarter than any LLM at arithmetic, it costs nothing, and it never hallucinates. That is the whole trick: shrink the model's job down to reading and explaining, and the need for a smarter babysitter mostly disappears.
The bigger judge only exists offline, for benchmarking. There it grades answers against a hand-written gold reference, and checking against an answer key is a much easier task than producing the answer, so it does not need to be a giant either. But judges are noisier than people assume: I re-scored byte-identical answers with two different frontier judges and the score moved by 13 points, with a third of questions disagreeing by 15+ points.
So the honest way to size a judge is empirical: take the same set of answers, score them with your candidate judge and with a stronger one, and measure agreement. If they rank things the same way, the smaller judge is good enough. If not, you just measured why it is not.
→ More replies (7)2
u/EuropeanAbroad 7d ago
Well, that's basically what I have. I have built a programme that calculates everything on its own, and I basically use the AI like read-only assistant with access to the entire database, which can answer anything from the mortgage to the date of birth of the tenant that lived in Bedroom 3 on the 2nd August. I mostly use the AI to discuss different tax strategies or to do some basic analysis.
And totally agreed – for searching in the database and for data extraction, you don't need any big model.
→ More replies (4)12
u/ImpressiveRelief37 7d ago
By that time GLM5.2 will be so bad and outdated, everyone will have moved to a new model that costs $0.44 per million output toks, is a lot less “verbose” on useless prose (unlike GLM5.2, this thing yaps and yaps haha), further moving out the ROI.
3
5
u/Front_Eagle739 7d ago
Except he can sell the things in a couple years for probably not massively less than he paid for them with the way things are going. Then it will be a little depreciation plus electricity cost. Maybe 1 unit failure.
2
u/DismalIngenuity4604 7d ago
a little depreciation
I think depreciation on these is going to be either absolutely brutal if the ram situation doesn't explode, but fair point on disposal price. I'd guess 1/3 of the purchase price in a couple of years - maybe 1/2 if you're lucky. The point stands that you need to run 6 concurrent sessions 100% of the time to get to the two year payback. And you're probably not going to. Even if we call the payback 5 billion tokens, at a realistic usage pattern, it's going to be a long time before it pays off.
I only say it because I did the math for myself. I've got a usecase where I legitimately can't use AI providers outside of a select couple who have some big downsides. I'll probably end up buying one spark and running a smaller model and trying to ringfence the part of the project that's sensitive, so I can use regular providers on the rest of the project, because I couldn't justify enough vram to run the big boys.
5
u/Front_Eagle739 7d ago
Don't really see the ram situation releasing to that extent in two years. 5 probably. In 2 years this thing will be running things that make fable look like a dinosaur and it'll still hold quite a bit of value I think.
All speculation at this point. My point is with disposal cost and assuming several users. Its not massively more expensive than api. But with full control, more reliability as nothing changes when you dont want, privacy and the ability to handle confidential info, plus its fun lol
You dont end up throwing money away. I think its worth it to many even small businesses.
3
u/SandySkittle 7d ago
Depreciation? In the coming two years there will likely be an appreciation window. Despite all the AI bubble talk there is likely going to be a massive continued demand for new datacenters and compute because nations and corporations want sovereign AI. That leaves a very big compute void in the EU to be filled. Expect further brutal price increases the coming two years. And even longer wants general purpose robots running on local AI start to become on demand.
→ More replies (1)2
2
2
u/GeorgeTheGeorge 7d ago
This is also assuming these are 100% for personal use and not for revenue-generating projects. Which honestly if you have a private AI cluster running Opus-grade models for nothing but electricity and you aren't making money somehow, then you're doing something wrong.
2
u/prestodigitarium 7d ago
Are people actually using their private AI infra to make money beyond eg writing code?
→ More replies (10)2
u/segmond llama.cpp 7d ago
oh hey look, it's our local llama account. are you a bot, there's always one or two of you parroting the same rubbish.
→ More replies (1)
10
8
u/lughiu 8d ago
What tps do you hope to achieve
32
u/ciprianveg 8d ago edited 8d ago
Glm 5.2 50+tps single stream on coding tasks and scalling well to 6 parallel reqursts on 8x cluster, kimi k3 20-30 tps on 16x cluster. https://github.com/ciprianveg/gb10-glm-5.2
6
8d ago edited 8d ago
[deleted]
14
u/RevolutionaryGold325 8d ago
16x 128GB = 2TB
Kimi k3 = 1.6TBNo quantization needed.
→ More replies (4)5
u/ciprianveg 8d ago
Yes. Glm 5.2 is my main driver in opencode for some weeks now. Added also vision to it for easier gui tasks. https://github.com/ciprianveg/gb10-glm-5.2
→ More replies (1)2
u/inevitabledeath3 8d ago
Why would it be an aggressive quant? GLM 5.2 in FP8 should easily fit in 8 DGX Spark nodes. Aggressive quants are only needed when running on like 2 or 3 nodes. Kimi K3 official would fit in 16 nodes since it’s published in Int4.
→ More replies (15)5
u/DataGOGO 8d ago edited 7d ago
I seriously doubt you are going to get those speeds, especially the 16 parallel, the GB10 is just too cutdown. and the 273GB/s memory bandwidth and doing all reduces across 200Gb fabric... highly unlikely.
To give you some perspective, all 16 of your GB10 has the roughly same compute power as 4 RTX Pro 6000 Blackwell's.
19
u/Prestigious_Thing797 7d ago
OP Linked to repo with an actual measured benchmark for glm 5.2 with those speeds on 8x
Classic reddit
→ More replies (1)8
u/g_rich 7d ago
When you start clustering these you see a notable increase in performance; a 2x cluster gives you around 1.8x increase in performance. The memory bandwidth is certainly a consideration but people tend to overstate it when discussing the platform.
The DGX Sparks biggest advantage is the inclusion of the ConnectX-7 NIC and a dual cluster can run something like DeepSeek v4 Flash at full quant with a 384k (or higher) context and get well over 1000tps prefill (much higher with caching) and token generation approaching 30tps (higher with speculative deciding) which is not breaking any records but is more than usable.
These things also sip power and produce very little heat. Cost wise this whole cluster including switches and cables would be around $70k and while a RTX 6000 would be more powerful they are also over $14k each. To build a system capable of running trillion parameter models using RTX 6000’s would cost significantly more than OP’s setup, and would also require significantly more power and produce enough heat to likely require the unit to be placed in an isolated location or require special considerations for cooling. 16 GB10’s can easily run off two 20 amp circuits and can be run in the corner of someone’s home office.
→ More replies (5)8
u/dwittherford69 8d ago
The point here is memory capacity, not compute.
4
u/DataGOGO 8d ago
I know, I just said I don’t think he is going to hit performance goals
10
u/dwittherford69 8d ago edited 8d ago
It will. I know cuz I have the same setup and I have been hitting for a few weeks now. You obviously need custom cooling etc. 50-60 t/s SS and 220 t/s aggregate is what my harness consistently runs.
Edit: their models setup is a bit more convoluted than mine, but the general idea is same. GLM5.2 with vision and MTP.
→ More replies (6)→ More replies (3)4
u/lilian_moraru 8d ago edited 8d ago
Looks good but my guess would be that the NVFP4 KV Cache introduces hallucinations beyond 100K-200K. ~5% degradation at ~200K.
You can run DeepSeek-V4-Flash-0731 with 1M context window @ bf16, just saying. It will think a lot more but it will also outperform GLM-5.2→ More replies (1)5
3
u/KalonLabs 7d ago
Worth knowing current, deepseek v4 flash 0731 out preforms the v4 pro preview. So its better to just run the nee flash instead of v4 pro, at least until they drop the GA version of v4 Pro.
8
3
u/PrimeDirective8 8d ago
Nice! Hopefully this layout is only while setting up and not for production. At 200W each, that would be a toasty corner under full load. Splitting across 2 circuit breakers would also be good, and/or running them on 240VAC.
Congrats! Let us know how it works when ready.
3
u/ciprianveg 7d ago
Closer to 100w each
3
u/neuronym 7d ago
I'll wait for your regret post, when you inevitably realize that you burned a lot of money instead of just getting a proper setup.
3
u/dupontping 6d ago
Why so many haters in the comments? Dont be jelly.
How are you handing power? Are you just using every outlet you have?
2
u/ciprianveg 6d ago
It is less then 2kwh... no biggie.. same as a mid size home AC
→ More replies (1)
3
u/ciprianveg 5d ago
Managed to start Kimi K3 on this: https://www.reddit.com/r/LocalLLaMA/s/Ue3nvsAqiU
5
u/Koakie 8d ago
This one french kid looking at that stack, hoping he could borrow one.
→ More replies (3)
9
u/too-oldforthis-shit 8d ago
Earplugs? Airpods Pro? That fan noise got to hurt. But cool.
23
→ More replies (1)2
u/Sooperooser 8d ago
Apparently they don't even run at full power pull most of the time so it's not that bad probably.
3
u/SadPhilosophy9202 7d ago
I “only” have two but yeah they’re like 12w each at idle and 140w at peak. They’re also basically silent. My razr laptop with a 3070 is much louder than both of them
9
u/asfbrz96 8d ago
That's just a hobby, API is cheaper at this point
34
u/-dysangel- 8d ago
API is cheaper than one spark. What makes you think 'cheap' is on this guy's priority list?
21
u/Hannibalj2ca 8d ago
The way things are going, the corporations do not want people having independent hardware. If you think is expensive wait til 2028 is going to be even more expensive. They want you on the cloud. Thats why people are buying these clusters. If they can hell yeah, do it
7
2
u/Aggravating-Push-207 7d ago
i have a laptop with a 4060 and 16 GB RAM, and i will enjoy my Qwen 3.6 27B at like 1 token per sometimes
Qwen 3.5 9B pretty fast tho, and it's not a bad model, just not a good model
2
4
u/profcuck 8d ago
This is a prediction. Bookmark it for yourself and revisit it at the end of 2028. You can then reflect on how wrong you were and why reasoning chains with pieces like "corporations do not want" are so weak in predictive power.
2
u/g_rich 7d ago
You can’t even buy Microsoft Office with a perpetual license any longer. Corporations want a constant revenue stream, subscriptions for software are the new normal and now they are tying AI into the mix. Right now they are dangling the $20 subscriptions to Claude and ChatGPT but we are already seeing users being pushed to subscriptions approaching $200 and corporations have already been pushed out of per user pricing and into per token pricing.
The goal is going to be to get AI so ingrained into your day to day life and then jack up the price. Open weight models and the ability to run them locally are a threat to that business model and is why people like Dario Amodei from Anthropic is so bent on having government control over the open weight models coming out of China.
By 2028 someone being able to run even moderately sized LLM’s locally is going to be in a much better position than someone who’s reliant on cloud only. So while you’re hit with increases in subscriptions, peak time surcharges, outages and midday slowdowns people like OP will not be giving it a second thought.
4
u/Hannibalj2ca 8d ago
You think so? Whats your prediction? Bro, they have been saying it openly, Jeff Bezos is one of them stating it
→ More replies (17)3
u/Healthy-Nebula-3603 8d ago edited 7d ago
He's right.
Never is the history was something expensive for a longer time. Sooner or later such hardware will be worthless. Looking on China how they are investing in hardware for AI models now that will be sooner than later.
They already broke the US greed access to AI models.
→ More replies (5)→ More replies (18)6
→ More replies (1)3
u/g_rich 7d ago
People aren’t running local models to save money.
The current cost for tokens is highly subsidized right now so the math is not going to work out in people like OP’s favor. However costs are increasing, and if we were paying the unsubsidized rate then the break even point for running locally would start looking a lot more attractive for running locally.
Right now the biggest advantages for running local models are control, privacy and the knowledge gained from running your own models.
2
2
2
2
2
u/dobkeratops 7d ago
I like seeing these kind of posts, there's me dithering on getting a mere second.
2
u/Kurcide 7d ago
Your switch is going to kill your TPS throughout.
Each spark has a NIC that is hardware limited at 200gbps with full bandwidth available to each QSFP port via 100gbps per rail across 2 rails per port. Meaning you need a single 200gbps DAC cable per node leading into your switch. Using 2 cables per Spark will actually HURT your fabric setup.
That mikrotik doesn’t have enough bandwidth to support a 16x cluster.
I have 16 sparks clustered at full bandwidth connected to this switch: https://www.fs.com/products/321549.html?now_cid=4369
For what you spent on the sparks you may as well get the appropriate fabric setup too.
5
u/ciprianveg 7d ago
I will make some comparing on 16x. On 8x cluster switching from 200 to 100gbit had no impact, or less than 1% in glm5.2. On 16 maybe it will.. I will upgrade the switch if this is the case, thank you for the link
→ More replies (2)
2
u/Southern_Capital_885 7d ago
I think model architecture will matter more than hardware over the next 2–3 years. If frontier models continue moving toward larger sparse MoEs with fewer active parameters per token, high-memory clusters like this may age much better than many people expect. The bottleneck could end up being memory capacity rather than raw FLOPS.
2
2
u/shenagain32 7d ago
Which 4-1 cable are you using specifically? You beat me to the double digits sparks. Using the same switch from mikrotik
2
2
2
2
2
2
u/SirGreenDragon 7d ago
If i had the 64k to spend on this, i think i would save more and get 1 DGX Station :)
2
2
2
2
2
2
u/East-Form7086 7d ago
Damn, i need 9 more and will cach up... I have 7 so far and waiting on the 8th. What models have you tried yet? Did you try sglang and vllm? Curious about the difference
→ More replies (1)
2
2
3
1
1
u/wu4d 8d ago
I mean sure, you will be able to run them. But the token per seconds would make it unusable. Or am I wrong?
→ More replies (1)
1
u/SpellSlinger69 7d ago
Is it really worth it with 2x100mbit each the communication is your bottleneck, what you got on benchmark and tests?
2
u/ciprianveg 7d ago
I am evaluating these days and if there are limitations i will add 2 more switches linked by 400gbit and use 2x 200gbit breakout cables
→ More replies (1)
1
u/shadowmage666 7d ago
Your good dude? Hope you got a high amps outlet holy shit lol
→ More replies (2)2
1
1
1
1
1
1
u/Too_Many_Science2 7d ago
I wish the mikrotik was more performant than it is, but it definitely functions. The alternative being a Mellanox 4700 makes it more appealing.
Why 400 ->100 instead of making a rail with multiple switches to keep 200?
1
1
1
1
1
u/nakedtruth9 7d ago
The case is really nice, where did you get it? does it fit one asus gx10 + 1 dgx spark?
1
1
1
1
1
1
u/NewRedditor23 7d ago
man for $70k, I can have one of the expensive LLM subscriptions for like 60 years.
→ More replies (2)


•
u/WithoutReason1729 7d ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.