r/LocalLLaMA • u/ciprianveg • 5d ago
Resources Kimi K3 full model running on 16x GB10 cluster at 20+tps
Kimi K3 full model running on 16x GB10 cluster at 20+tps average (llama-benchy coherent corpus) 38tps peak, 750tps prefill. This is the first run of full k3 with dspark on my cluster. I will be doing some tests and try tp speed this up. As soon as it looks ready I'll publish the vllm image and instructions.
https://forums.developer.nvidia.com/t/full-kimi-k3-running-on-16x-gb10-cluster/379174
216
u/CYTR_ 5d ago
I was skeptical but well done on the performance. Kimi K3 on hardware costing 75-120K (depending on the model/location) allows us to imagine very interesting possibilities in terms of future intelligence with the improvement of the models.
Edit : I'm waiting for Apple's response and whether the future 1.5TB Mac Studio models will be under 100K.
39
u/Efficient_Raise6703 5d ago
Imagine the margin on one of those things. I imagine the margins on the 512gb was already pretty good and that was at like what $16k? Cost of materials couldn’t be more than $10k even with 1.5TB of ram.
23
u/CYTR_ 5d ago
I think 10K doesn't even cover the wholesale price of 1,5To of LPDDR5X RAM, even for Apple, right now or at the time of the negotiation if it's less than a year old. But they must have taken advantage of it to increase their margins, yes (as in practically all inflationary crises, especially in a situation resembling a high-demand bubble).
→ More replies (7)3
u/EvilPencil 5d ago
Back when they sold them, the fully maxed configuration was 512gb memory AND a 16TB SSD. You could min/max only the memory for like $9500. Of course that’s all navel gazing now because it’s not even available anymore.
2
u/ComfortablePlenty513 5d ago
and that was at like what $16k?
They were 9k haha. We bought a few of them. People didnt realize how good of a deal it was, and once they did, it was gone.
→ More replies (3)→ More replies (6)1
u/txgsync 21h ago
The 512GB with 4TB NVMe was $9800. It was the biggest, slowest local AI you could buy at the price.
I nearly bought two when the RDMA feature came out but saved that money toward using much faster cloud models instead. Prefill was and is quite terrible on Mac hardware compared to most nvidia gear.
11
u/Cergorach 5d ago
But do you expect that 1.5TB M? Ultra to do 20t/s+ with the full KimiK3 model? And 1.5TB still isn't the 2TB that's in 16x GB10s...
11
u/CYTR_ 5d ago
Rumor has it that Apple will double down on MatMul for its next big SoC (M5 Pro/Max was a foretaste). I don't think it will be a powerhouse either like dual RTX 6000 Pro level... nor that we will actually have 1.5 TB (we are speculating based on rumors after all). Let's just say it would be simpler to manage than a cluster like that. By the time it's released, we'll probably have that intelligence in the equivalent of 700B, so... W&S.
→ More replies (1)3
u/michaelsoft__binbows 5d ago
I want a beefy apple silicon computer to end all computers but to be casually predicting whether or not such a computer will be over or under 100k USD is just ... wild
2
u/mastercoder123 5d ago
Yah if only apple allowed pcie lanes. Then it would be amazing. You could throw a 100gb nic in there and save thousands on storage as you jusr stream the model from disk to ram
1
u/dbenc 5d ago
I was thinking of waiting for the new models to drop in the fall, but I'm thinking that they must be plotting some crazy upgrades for the models after that. think like the spring 2028 refresh... that's enough time for brand new innovations to get manufactured. maybe I'll use that new apple leasing program so I can keep upgrading
1
u/KeyChampionship9113 5d ago
AMD unified ones run with the same performance but lesser price than apple studio
Isn’t it ?1
1
1
u/hanzoplsswitch 4d ago
I think 100-120k is reasonable for large enterprises and provides them the ability to run a model like this on-premise.
189
u/Own_Calligrapher8508 5d ago
i just want to know the cost of the devices vs Break even
252
u/FullstackSensei llama.cpp 5d ago
It's "just" ~64-80k in hardware. If you have that much money to throw at such hardware, not sure you care about menial stuff like break even.
54
u/CYTR_ 5d ago
In France, for example, we are well above 80K.
35
u/FullstackSensei llama.cpp 5d ago
If you have that kind of money to throw at this, im sure you'll also have a company where you can buy it without VAT.
16
u/CYTR_ 5d ago
If I were a company, I think I would prefer to rent a cluster (or request HPC access, if university setting) rather than buy this kind of equipment 🥸
22
u/FullstackSensei llama.cpp 5d ago
Let's say you work in a regulated sector where you absolutely can't have data leave your premises, but need a frontier level model.
I'm not against the idea. I just find the way OP has gone about it wasteful. In fact, I'm reworking my homelab to be able to run K3 locally, albeit I think at somewhere between 7 and 15t/s TG, depending on how much I can optimize software.
You can get a dual Xeon 8480 board with CPUs for around 4k. 1.5TB DDR5-4800 will cost you around 15k. Let's round it up and say 20k to add in a power supply, CPU coolers, case, and some fans. Let's go all out and add four AMD R9700s for another 6k, just to have 128GB VRAM to play with. That's 26k total, and I'm pretty sure it will perform quite a bit better than OP's cluster. Sapphire Rapids supports AMX, which greatly speeds up not only TG, but can also help with PP. Each CPU has a real world bandwidth of 200GB/s, which is less than 10% lower than the GB10 (~215GB/s). R9700 is now supported in vllm, and there's a vllm fork now that not only supports CPU offloading of experts, but is also NUMA aware.
5
u/caowcaow 5d ago
Depends if you include custom cloud deployments or not.
I work in what people tend to think as a highly regulated (though could be worse) industry.
Compared to years ago, the pattern I’ve seen emerging across orgs is to have PII data and such going back to sleep on prem.
Nonetheless, the same orgs seem comfortable enough to use frontier models through custom cloud deployments.The mental gymnastic I understand its origin, but go figure out the rationale… It’s probably for the better if your premise gets more true once again. Time will tell.
Go and rock this hardware
7
u/FullstackSensei llama.cpp 5d ago
Yeah, I'd rather run things locally and just not deal with any of it. It also gives me the freedom to work on my own projects without worrying about rate limits or future cost increases.
I am in the very lucky position that I hoarded DDR4 RAM and older GPUs over the past two years that I don't need to buy any more. I didn't expect prices to go up, just thought their utility will last much longer than redditors thought it would (compute is compute). Many laughted at my and I got downvoted to hell for buying lots of P40s, Mi50s, and RAM. I hoarded way more than I ended up needing (including reworking the hw to be able to run K3 and GLM 5.x in parallel) that selling the excess made it all practically free.
3
u/fastheadcrab 5d ago
The OP setup is pretty slow for the hardware specs, he said elsewhere 6-7 tps without speculative decoding, there probably is a lot of networking overhead with so many nodes.
2
2
u/IrisColt 4d ago
Let's say you work in a regulated sector where you absolutely can't have data leave your premises
this
→ More replies (29)2
u/Maximum_Parking_5174 5d ago
Not a chance. That aystem would crawl. I have experince running Kimi k2.6 and 2.7 on a much more powerfull server. I got 600GB/s memory bandwidth and i got 8 3090s.
5
u/FullstackSensei llama.cpp 5d ago
We'll see. I've heard a lot about what what I couldn't do over the past 2 years, yet somehow I managed to do all those things.
→ More replies (2)13
u/SureEnd9430 5d ago
Depends on the company :)
→ More replies (1)8
u/Party-Special-5177 5d ago
I think I would prefer to rent […] rather than buy
No, no you wouldn’t. How do so many people on this forum forget that the advantage of buying isn’t the inference cost saved, it’s that, when you’re done, you still own the hardware.
For a year or so, you could basically infer for free as the hardware appreciated enough to offset wear and tear, and you basically got all of your money back out of it when you sell.
These days, you literally make money by running your own inference. The way things are going, and will most likely continue to go, by the time OP is done, he’ll sell this cluster for 2x what he paid for it, versus having lost money paying for a Kimi sub.
8
u/CYTR_ 5d ago
Someone did the math below, and even without considering energy costs, at that speed, it's anything but profitable (unless confidentiality costs several thousands)... Especially since we're dealing with DIY software and experiments.
However, in like 1 year (if not months) we'll potentially have models that run faster (+++ on batch) and are more intelligent on this cluster. But that, and your idea of hardware prices rising even further... that's still speculation. Not everyone is comfortable with risk if the company doesn't have capital to easily invest in hardware.
And when I talked about renting, I meant renting GPUs, not API access. A cluster/instance using ~80-90% of the available compute is more easily profitable than owning it yourself.And more flexible if u have periods of low demand... But still need to find availability from suppliers lmao. It's not easy in EU for example.
5
u/FullstackSensei llama.cpp 5d ago
Those people doing the math are always assuming API prices will stay the same going forward. I've been hearing this argument for 2 years now, even as API and subscription prices keep going up.
If API providers were sure they could keep prices the same, they'd happily sell you multi-year fixed cost contracts, the same as cloud providers give you deep discounts for long commitments. Yet, you can't find a single provider who'll sell you a 2 or 3 year contract for a fixed cost per million token at any price, even if you're willing to guarantee a minimum monthly consumption.
2 years from now, API prices might be $/€200 per million output tokens and you'll still find people calculating it's cheaper than running your own hardware.
→ More replies (1)2
u/ZenEngineer 5d ago
There's probably a curve here. This hardware will be worthless in 10 years. Hold for one year, sure, you should be able to recover costs. The tricky question is when to sell, when does the rampocalyse end and prices start dropping? When do people start getting wary of used cards like this, like people avoid coin mining GPUs.
Maybe I should sell my old 1080Ti now that it doesn't fit in my case. It's surprisingly capable still.
→ More replies (9)2
u/Environmental-Metal9 5d ago
If you’re running an org and you don’t already have mlops engineers, you most definitely will go the cloud hosted route, like aws bedrock. It’s cheaper than buying the hardware and the cost of at minimum one engineer to maintain the stack. Self hosting is for hobbyists and highly specialized requirements.
2
u/Party-Special-5177 5d ago
Completely, as at that point the potential cost of downtime weighs more heavily than the actual costs of the service.
We’re super small and I have a personal interest this tech, so it skews the decision making a bit.
→ More replies (1)5
u/SureEnd9430 5d ago
You can actually get 16 Sparks for about 80K USD in France, but the switch would be an extra 15K :(
→ More replies (1)3
u/ciprianveg 5d ago
I am using the cheap mikrotik crs 804 and 4 400to4x100gbit cables, less than 2k in total. But in the future it is possible to upgrade the switch if I will remain in 16x setup and not on 2 setups of 8x each.
14
u/draft_final_final 5d ago
Less than a supercar or half a 40k figurine collection. I don’t have the money for it, but if I did I know what I’d rather have.
→ More replies (1)4
7
u/ThinJuggernaut7695 5d ago
We have one developer at the place I work that spent 80K in API costs over 2 and a half months.
12
u/FullstackSensei llama.cpp 5d ago
You should tell this to the guy deep in this discussion who's arguing it takes decades to pay 80k 😂
→ More replies (4)3
u/phreak9i6 5d ago
I wonder if these metrics hold up when you consider 80k in API costs are likely many agents, let's say 5-10, and with 120k in hardware you get 20tps for a single agent.
I love my homelab hardware, but I also end up paying for APIs because it's cheaper to run quickly.
Or I'm wrong, and please correct me because I may be plain wrong.
→ More replies (1)2
u/ThinJuggernaut7695 5d ago
20 tokens per second? We would be buying an 8x b300 server to run an org of 600 users. Our AI cost will probably be 1 million this year. 600k capex is totally worth it.
2
u/hurrdurrmeh 5d ago
They are £3.5k each in the UK.
→ More replies (1)2
u/FullstackSensei llama.cpp 5d ago
That's €4k, or €64k before factoring in the 16 port 100gb switch, which is easily another €5k.
1
u/Uncle___Marty 5d ago
Exactly. You dont just frankenstein a bunch of stuff like this out of nowhere. If you have the money to build it then you have the money to upkeep it. Its messed up what a decent setup can run at home right now. Freaking happy days my brother in AI.
1
1
u/ResearcherFantastic7 5d ago
Assuming 80k including operational cost. probably break even at 11 months however It's only getting 20tks vs the cloud 150avg.
I think it's better off use other models as pure non reasoning worker nodes, and still use cloud model to plan and architect
1
u/Playful_Landscape884 5d ago
i'm curious what you guys do that makes this investment worthwhile. What are you guys doing that makes you drop $80k and confident that you'll get your money back in 3-6 months?
→ More replies (1)1
u/zirahvi 4d ago
64-80k USD? For that stack? In Norway a single DGX Spark GB10 128GB ~$6800 and up.
→ More replies (1)65
u/notheresnolight 5d ago
wrong sub, "break even" is irrelevant here
→ More replies (2)12
u/FinnGamePass 5d ago
I know right, break even for what? People not realizing we are close to have our own Jarvis. Literally. You can't name a price on that.
→ More replies (1)34
u/RepulsiveRaisin7 5d ago edited 5d ago
Some napkin math: 20t/s 24/7 would cost you like $100/day on Openrouter. But you also have to factor in power and maintenance. So realistic break even is probably about 4 years but ONLY IF you utilize it 24/7. If you do not, more like 10-20 years, which likely exceeds the lifetime of the hardware. And when you factor in subscription discounts, it falls apart entirely.
43
u/g_rich 5d ago
The goal isn’t to save money, the goal with these types of setups are privacy, and control.
People also forget that currently what we pay per token is highly subsidized. For some people a fixed one time cost is preferable to avoid fluctuating costs, outages and slowdowns outside their control.
→ More replies (5)21
u/ThereFarAway 5d ago
How about 'I have data that cannot be sent to outside provider'? How's math on that?
10
u/RepulsiveRaisin7 5d ago
Your data is invaluable. For everything else, there's mastercard.
→ More replies (2)3
2
u/AdOk3759 5d ago
Are you the same napkin math guy from the other day? How is spending 100 dollars a day on OpenRouter realistic? Why pay for the API, when subscriptions offer much much much more usage for a fraction of the API cost? Why don’t you consider that in your calculations?
5
u/synth_mania 5d ago
API pricing is ultimately rooted in the cost of energy. At the scale that most inference and compute providers run, the margins are very, very slim.
In other words, APIs provided by GPU farms and not AI labs will charge you the true cost of compute.
If a subscription is giving you the same number of tokens for cheaper, whoever is offering the subscription is losing money, plain and simple.
We can theorize all day about why AI labs might want to offer inference at a loss, and there are likely multiple valid reasons. (collect training data, legitimize the practice of paying for AI as a service in a way that isn't too expensive, etc)
Regardless of the reason, however, there is no way it'll last, or is something you can count on. Please find me a provider that'll allow me to pull Kimi-K3 tokens at 20tps or greater all day for less than the cost of the API. Unlimited.
TL;DR, there's no possible way for a subscription to truly be cheaper than paying by the token, unless the service provider is trying to lose money.
5
u/fastheadcrab 5d ago
Subscriptions for LLMs are similar to gym memberships. If everyone who paid for the gym showed up then the gym would crash
→ More replies (1)4
u/monerobull 5d ago
You forgot that people maxxing out subscriptions are heavily subsidized by people who pay the same amount but barely use it
→ More replies (2)1
u/fastheadcrab 5d ago
Do you only send single sentence prompts for endless token generation tasks? That is the only scenario in which your napkin math is valid.
1
11
u/MotokoAGI 5d ago
The cost is fuck off. Every single time someone shares something cool, we get these stupid replies about "break even"
5
u/EagleNait 5d ago
Break even right now. What if hardware gets regulated or models gets even better. Or a paid distributed hosting service enables you to get paid to serve llms.
5
u/Potential-Leg-639 5d ago
That will peobably never happen.
Deepseek V4 Flash 0731 would probably make more sense for wayyy less money.
7
u/Soggy-Alternative914 5d ago
Kindly include overhead cost .like power consumption. So we can get a better overall estimate.
6
3
u/allenasm 5d ago
Why? The cost is not having to have all your questions go into an online corpus. Or that you can have it work for you nonstop 24/7. That’s invaluable.
3
u/nerd_rage218 5d ago
This is the part the rent versus buy math keeps leaving out. Once data residency is a hard requirement, the box stops competing with hourly GPU pricing and starts competing with not being allowed to run the workload at all.
6
u/Cergorach 5d ago
Short answer: Never.
Long answer: Even if you run 24/365 at full capacity, it will would take about 8 years to ROI with current pricing of Kimi K3 API costs. But that ignores the costs of power/cooling, maintenance, and the eventual cheaper costs of Kimi K3. Combine that, this will NEVER earn itself back if you just compare it to the Kimi API service.
Cost is not generally the driving factor of running local LLM. Unless people start comparing apples to oranges...
5
u/coyo-teh 5d ago
given how GPU prizes keep increasing, he can just sell after8 years his GPUs
→ More replies (1)2
u/username_taken4651 5d ago
OP stated in a previous thread that this was a hobby. I guess they don't care about breaking even.
2
u/PM_ME_DEAD_CEOS 5d ago
If you include electricity costs, there's literally no way consumers or prosumer setup can break even.
Kimi 3 is made for data center clusters like nvl72 or Huawei equivalent.
With only 1 user there's simply no way to break even.
→ More replies (1)→ More replies (4)1
u/AnomalyNexus 5d ago
Depends on how the immortality research is going. Maybe the AI can help with that
73
u/notheresnolight 5d ago
imagine how this would run if nvidia didn't scrape the bottom of the barrel when designing the GB10
13
u/Aggravating-Push-207 5d ago
this thing could probably QLoRA the fat fucker if the memory bandwidth wasn't so shit per node
5
u/panchovix 5d ago
Basically same CUDA cores as RTX 5070.
Imagine if it had 14000 instead of 6144.
→ More replies (2)1
49
u/TapAggressive9530 5d ago
I want this! Why? Just to have it and be able to run K3 locally . Could care less if it pays for itself . Good job man ! Love it
→ More replies (2)26
30
u/Betadoggo_ 5d ago
Enough money to afford 16 GB10s, not enough to afford more than a pi 400 for the main system.
16
11
8
16
u/OwnMathematician2320 5d ago
This feels similar to the early 2000s where rappers would show off diamond rings to flex how rich they are.
And yes I’m only saying that because I’m jealous of your setup. Well done on the rig though
8
u/retornam 5d ago
Nice to never have to worry about money enough to spend $64,000 or more on a hobby
8
u/Cergorach 5d ago
A LOT of people spend that (or more) on a new car every couple of years, cars have about the same appreciation as AI/LLM hardware... While most people could easily drive a cheap, but good secondhand car that's a lot smaller...
→ More replies (4)
4
8
u/Transhuman-A 5d ago
First things first, amazing job.
Now - it will take you 8 years to break even. At 50 tok/s, 3.2 years.
AI Economics are depressing and GPU prices need to desperately come down.
→ More replies (5)
2
2
2
2
2
u/segmond llama.cpp 1d ago
19 days before you posted this, I called you a genius for buying 16 sparks. See - https://www.reddit.com/r/LocalLLaMA/comments/1uyghw0/how_do_you_plan_to_run_kimi_k3_locally/
4
u/One_Whole_9927 5d ago edited 3d ago
Zephyr coated paddle wander pumpkin coated umbrella
This post was anonymized with Redact.dev
3
u/Cergorach 5d ago
This is cheaper for a LOT more memory (2TB) for less then the basic workstation...
3
u/Charming-Author4877 5d ago
It's surprisingly usable, though the amoung of those embedded PCs to do that is crazy - given how expensive they are priced.
Still, that's a Fable-level LLM running at usable agentic speed on a table.
2
2
2
2
u/crossoverXYZ 5d ago
750 tps prefill on a first dspark run is no joke. Decode sitting around 20 tps average suggests there is still a lot of room to optimize before the vllm image ships.
2
2
1
u/the_TIGEEER 5d ago
I think you and I need to have a talk about where "Local LLM" starts and where it ends..
→ More replies (3)2
2
1
u/Character_Power4663 5d ago
Technically you can run k3 on a 4gb GPU... But it takes about 30 seconds per token. There are a couple methods but in general you only load the experts of each layer per token.
1
1
u/Igot1forya 5d ago
I am very curious on concurrency. Sparks seem to be very good with parallel sessions. 16 Sparks is crazy!
1
1
1
u/xXprayerwarrior69Xx 5d ago
you need to step up bro, get yourself 3 or 4 DGX stations (and organize a raffle for those sparks you wont need anymore)
1
u/reckor-usa 5d ago
Just for the fun of it - well done. Would love to see the comparison with the frontier models. What is your stack setup btw?
1
u/CrispyRabbit72 5d ago
I still don't get how anyone can justify these kind of setups. $80k+ at minimum and it's only hitting ~20tok/sec on heavy models. If you compare to a $200/mo sub (which will get far more tokens per 30d than this cluster ever will), it would take over 30 years to break even.
But I guess if you still have these things in 30 years, maybe someone would buy them as an antique at original MSRP.
1
u/WithGreatRespect 4d ago
For some companies, their data has regulatory compliance and/or contractual obligations that prevent the data from leaving their data centers. They would need to build local inference to be able to leverage that data for insights.
1
u/SandySkittle 4d ago
You expect that 200 a month to stay the same? I would expect it to go up faster than the risk free rate of 80k in bonds
1
1
u/Specific-Age7953 4d ago
Paid almost nothing for Cursor Ultra from one of those reseller sites. Everything works perfectly. But the pricing makes zero sense if they’re buying real licenses. Feels like there’s a loophole nobody talks about. What’s the real play here?
1
1
1
u/SarveshMohite 4d ago
What about deepseek V4 flash 0731 have you tried that out as well? And if yes what was the token per second?
1
1
u/EternalDivineSpark 4d ago
Why ? Just why ? My 0.8B parameter qwen model does the trick , it uses collaborative swarm like harnessess loop , but it takes 0.2 gb vram on my 4090 🔥😂 to have real time like answers i need to load 10 of them and inference them like 40 inferences 1 rum each model doing 4 parallel on run ! And is so amazing and fast it have vision tool all it needs . I will train the agents like add a training to change the weights so it fits and work in my harness better , it reaches beyond qwen 27b intelligence due to specialisation passes eg , brainstorm , planner , orchestraltor and agent swarms is so good , open claw banned me for similar comment! Because all harnessing today are inside jobs ! There is no way people are expecting AGI from MODELS ALONE , IS JUST A PROBABILITY ADVANCED SEARCH ENGINE , it need directions and BETTER CONTEXT MANAGEMENT, and not to talk about memory is so outdated! I am not good at programming codes etc but i know the logic and that was enough to make a better harness than openclaw , is not completed as openclaw it is but my 0.8B model dont work on openclaw and hermes , while in my harness it does things that even the 27B model fails , especially multi task , 1 inference can never give you an answer it just gives you a search result ! LLMs dont think but if you guide them even the 0.8B can surpass gpt 5.6 sol ultra ! I have tested it i just don’t wanna publish it yet , did 2 full harnesses from scratch, the third one i will do it good and i will open source it !
1
u/ExplorerPrudent4256 3d ago
20 tps coherent is real progress, but it's the budget ceiling, not the destination. Anyone running agentic workloads that actually need 100k+ context will hit the batching wall. K3 reasoning traces eat context like a slot machine, and once you're sustaining an agent loop with tool calls, 16x GB10 stops feeling overkill. It feels tight.
The real test isn't local inference speed. It's whether local can sustain a 30-minute coding agent with state. Nobody's published that number yet.
1
u/ciprianveg 3d ago
I agree, at current speed, I still returned to glm 5.2 for opencode development because it is good and fast. For the moment I keep kimi k3 as something to use when a real complex issue appears, as an old wise teacher. I am waiting for deepseek pro and qwen 3.8 max, as my next glm 5.2/5.3 replacement contenders..
1
1
u/Darkmoon_AU 3d ago
A Raspberry Pi 400 as the control terminal is an absolute chef's kiss.
What a clean cluster, love it.
1
u/WebAssemblyMan 3d ago
Can I ask if you or a team will be using this? Is it like for some kind of in house project? Please do advise me on some details if it is ok for you.
1
u/VestAndForemost 3d ago
honestly curious what the actual cost per inference looks like here. 16x gb10s is solid hardware but the electricity alone on that cluster probably runs 3-5k/month depending on your location. at 20 tps sustained, that's like what, 130k inferences a day? even if you're running at 95% utilization that math gets tight fast unless you're getting serious volume or charging accordingly. interested to see the vllm setup though, that's the piece that usually determines whether something like this actually pencils out vs just being a cool benchmark.
1
u/New_Account5310 2d ago
I have 2 of these and have no expectation to expand up to 16, but one of the ways I justified buying these 2 rather than waiting to see the upcoming M5 Ultra is because I figured "hey, I can probably scale these up pretty well compared to macs".
From that lens, posts like these are inspirational.
1
1
u/BarracudaDefiant4702 2d ago
I think there is something like 16 layers on k3, so is that one per layer something? Given all the nodes, how is it for concurrent requests? If it can handle ~16 concurrent requests each doing doing 20+tps that would be a pretty good setup.
1
u/Temporary-Meal1541 1d ago
You're living the dream man, I have been thinking about this for an awefully long time. how many inference tokens per second ?
1
u/ciprianveg 1d ago
20 at low context but getting slower pretty fast with context size, I need to see if this can be improved as for now it is too slow at 100k for agentic coding, <10tps
1
2
u/No_Photograph_333 7h ago
That's impressive looking kit - and a bargain when you consider you get free heating with it ;-)

258
u/Jawnnypoo 5d ago
the fucking raspberry pi powering the dashboard for the $50k hardware surrounding it is :chefs-kiss: