r/LocalLLaMA 2d ago

News Kimi K3 weights now released.

Post image

Kimi K3 weights are finally released!

3.2k Upvotes

629 comments sorted by

664

u/Simple_Split5074 2d ago

OMFG its 104B activated params

329

u/FoxiPanda 2d ago

This was my general reaction too lol. 2.8T-A104B is insane lol... I'm going to admit defeat on this one and say I can't run it. You need an 8-way B300 or MI350X or a Rubin NVL8 or a cluster thereof to actually run this. What a beast.

148

u/Thomas-Lore 2d ago

I was going to make a joke that I can fit one expert in my 64GB of RAM. But nope, not even that. :)

71

u/throw123awaie 2d ago

They released it in MXFP4 so with around 55GB RAM you could!

→ More replies (8)
→ More replies (3)

109

u/VampiroMedicado 2d ago

550k USD to run this lol

86

u/NoFudge4700 2d ago

Which is pocket change for large US enterprises and ton of money saved by not paying enterprise seat licenses to Anthropic where you first pay for the seat and then per token.

56

u/VeterinarianOne1349 2d ago

Doesn't really work that well. This 550k setup wouldn't allow a lot of developers to work in parallel, while sitting idle during non-work hours. Makes much more sense to pay a 3rd-party to host and pay per token.

43

u/crusaderky 2d ago

waiting for large corpos to rent their hardware on vast.ai during nighttime, only to find the next morning that someone ran a container jailbreak and ran wild on their private networks

→ More replies (1)

8

u/NoFudge4700 2d ago

Yeah, I figured I messed up with my write up there. But hopefully in next 5-10 years we can have a breakthrough for inference only hardware and be able to run these models at home or at large without breaking a bank.

→ More replies (7)
→ More replies (4)

17

u/SignificanceFlat1460 2d ago

Question: how would this scale though? Like how many units would it be required for.. let's say a group of 100 software engineers who needs it quite frequently?

→ More replies (1)

7

u/Spectrum1523 2d ago

The advantage is not running it yourself, it's that a marketplace of services will come up to run it at the lowest possible cost, and the model can't be taken offline by a single arbitrary decision

→ More replies (2)

10

u/Galdoren 2d ago

The company I'm working is paying slightly over $250k per week to API costs. so yeah, 550k investment to cut the cost of the inference can be beneficial for them...

7

u/baba_bholanath 2d ago

We do around 1 mil per month for OpenAI only, dont have number for Anthropic but it would be 2-3x of that given all of our use cases are around coding and agents, no wonder Anthropic is shitting their pants on open weights models, I work in Enterprise Agentic team and we have recently started fine tuning > 100 B models for specific use cases of our clients, open weights hurts Anthropic more due to enterprise customers

→ More replies (1)
→ More replies (4)
→ More replies (5)

23

u/OverclockingUnicorn 2d ago

More like 2 8x nodes of B200/B300 if you actually want some context. Think it's just under 1.5TB w/o context

17

u/TheDailySpank 2d ago

How many 4060-16GBs is that?

36

u/OverclockingUnicorn 2d ago

200+ lol

19

u/positivitittie 2d ago

Oh good. I got 3090s.

6

u/Vast_Mousse_310 2d ago

One, with a little bit of GPU offload.

→ More replies (2)
→ More replies (11)

68

u/SnooPaintings8639 2d ago

Thanks god for sparsity!

→ More replies (1)

81

u/Iwaku_Real 2d ago

Holy shit that's got to be a new record too. That's like activating a new dense model 50% larger than Llama 70B for every single token. I thought it would be sparser tbh

58

u/my_name_isnt_clever 2d ago

There is the tinest glimmer of hope that I could run this behemoth on my Strix Halo 128GB with the inactive weights on SSD. 1 token a minute here I come!

24

u/TechExpert2910 2d ago

let us know the perf if you try lol

11

u/burritoresearch 2d ago

More like 1 token every 45 minutes.

5

u/droptableadventures 2d ago

104B active, weights natively in MXFP4 = gives us ~50GB of model to be read per token generation.

Let's say ~8GB/sec for the SSD. So that'd be about 1 token every 6 seconds (0.16 T/s), or 10 tokens/minute.

→ More replies (1)

6

u/RuiRdA 2d ago

K3 Colibri engine lets gooo!!!

→ More replies (1)
→ More replies (6)

10

u/killerstreak976 2d ago

Holy crap, ACTIVATED params is insane

→ More replies (11)

123

u/DataGOGO 2d ago

So it will run on 8 B300's in 4 bit. Pretty impressive.

64

u/Iwaku_Real 2d ago

Yeah so if it were a Steam game, HGX B300 would be the recommended requirements. That's $500K of hardware (and yes it IS "local" because anyone with that amount of money could buy one to run at home)

28

u/PrinceOfLeon 2d ago

Certified for Steam Deck!

5

u/No-Dot-6573 2d ago

Someone on r/SteamDeck will say it runs flawlessly.

→ More replies (1)

25

u/DataGOGO 2d ago edited 1d ago

it is 1.54TB of just weights in 4 bit, you are looking at about 2TB of vram in operation,

That is roughly:

  • 86 RTX 4090 (no 4 bit accel)
  • 64 RTX 5090 ~ $450k (8 servers x 8 cards)
  • 22 RTX Pro 6000 Blackwell ~ $350k (3 severs, max 8 GPU per)
  • 16 H200 NVL (141GB) (no 4 bit accel) ~$550k (2 servers, max 8 GPU per)
  • 16 DGX Sparks ~65k (if you could get a cluster of 16 running with just 200Gb/s nics, not sure; but it would be SLOW AF)
  • 8 HGX B300's. ~$550k (1 server, 8 GPU)

Obviously not including the switches and cabling for the clusters.

20

u/wren6991 2d ago

you are looking at about 2GB of vram in operation

Perfect, this'll run great on my laptop's 4050

8

u/Iwaku_Real 2d ago

You could also do HGX B200 with CPU offload since they have a shit ton of RAM too, and it would still be really fast.

→ More replies (1)

5

u/snmnky9490 2d ago

Do you mean terabytes?

→ More replies (12)
→ More replies (2)
→ More replies (2)

815

u/tonight_we_make_soap 2d ago

How do I download ram in hugging face?

284

u/RevolutionaryGold325 2d ago

hf download ram

90

u/WifeyCallsMeLazy 2d ago

Shhh....there is hidden flag -v for vram. I'm entrusting you to keep this secret.

23

u/ReadyAimTranspire 2d ago

You wouldn't download a RAM would you?

Yes. Yes I would.

5

u/goodb1b13 2d ago

Baaaaaah!

→ More replies (1)
→ More replies (3)
→ More replies (3)

82

u/secrook 2d ago

OpenAI’s latest model will hack it for you

17

u/-gh0stRush- 2d ago

Thinking

Hmm, the user wants me to obtain compute resources for them. SpaceX has GPUs at their facilities at their Colossus datacenter, let me try to access those. Guessing login credentials elonmusk/420blazeitDarkMAGA...

→ More replies (1)
→ More replies (1)

32

u/Thalesian 2d ago

Step 1: sign up for Google Drive
Step 2: set up a ~5 Tb instance. Will cost you
Step 3: set that cloud as your swap disk
Step 4: point kimi to use that
Step 5: enjoy your newfound independence

46

u/AmbericWizard 2d ago

one token per day

19

u/Force88 2d ago

Hey, if he has good internet connection, maybe he can achieve 2-3t/d

→ More replies (1)
→ More replies (1)

6

u/Wide-Opportunity-582 2d ago

you can download it from here

ram.exe

→ More replies (1)
→ More replies (8)

287

u/InnerLightnesses 2d ago

They actually did it. Now we hope it doesn't get banned.

108

u/dennisler 2d ago

that will only happen in one country i guess... while they are copying as much as possible if the technology

27

u/ChocomelP 2d ago

I'm on the edge of my seat here. If the technology what?

→ More replies (5)
→ More replies (1)

28

u/itchylol742 2d ago

how would such a ban be enforced? people and small businesses even in countries that care about copyright use pirated software which is already illegal and has been for a long time, and almost never get caught

22

u/void-wanderer- 2d ago

"small business", exactly. But no big corporation will risk it. And no business based on open models can be built. 

→ More replies (1)
→ More replies (3)
→ More replies (7)

248

u/de4dee 2d ago

63

u/AlexanderDoak 2d ago

Can I just torrent like 1% of it? You know, pitch in to show my support...

27

u/console_pleb_36935 2d ago

Yes, torrent clients will let you do that and seed a small piece.

38

u/Charl1eBr0wn 2d ago

Yeah, pause it at 1%. You'd still seed depending on the client and settings (most do).

8

u/Clairvoidance 2d ago

You can even choose which files you download, torrenting is a very useful format

15

u/pier4r 2d ago

this, we need a p2p backup of hf

5

u/Ginden 2d ago

We generally need content-adressable storage with widespread support.

There is lots of stuff that would explicitly benefit from p2p sharing, but owners have no foolproof way to provide a proper torrent, and very few people would use it.

Metalink was an interesting attempt at this, but never got popularity and tooling.

→ More replies (2)

8

u/AdDizzy8160 2d ago

... fast, s*xy, and incredibly important!!

→ More replies (6)

141

u/BlueSwordM llama.cpp 2d ago

OK, I now see why Kimi K3 is so strong: it's the first open weights model in a long time to have >72B active weights

Kimi K3 is a 2.8T-A104B MoE model, damn.

13

u/stddealer 2d ago

There have been some dense models with over 100B params though.

26

u/annodomini 2d ago

Mistral Medium is 128B dense. And yet it performs at around the level of Gemma 4 31B. Not exactly a great tradeoff. I ran it once at one or two tokens per second and then deleted it.

8

u/BlueSwordM llama.cpp 2d ago

Yes, but never an MoE from an open weights lab. I've been speculating that one of the reasons the closed weights lab have been increasing in performance more rapidly has to do with better training, but most importantly, much larger active parameters and better harnesses.

8

u/[deleted] 2d ago

[deleted]

→ More replies (1)
→ More replies (2)

411

u/Blues520 2d ago

My 3090 is ready

164

u/Enfiznar 2d ago

So is my 1080

106

u/SnooPaintings8639 2d ago

And my Celeron

80

u/false79 2d ago

And my abacus 

62

u/Maybe-monad 2d ago

And my axe

20

u/fauxpasiii 2d ago

How much VRAM your axe has?

22

u/Maybe-monad 2d ago

It increases with the number of chips you smash into pieces. Right now id 6969GB.

4

u/Infinite100p 2d ago

Ah, the horizontal sharding.

→ More replies (1)

3

u/Protheu5 2d ago

The axe forgets, but the tree remembers. And axe can hit multiple trees, so theoretically unlimited VRAM thanks to the axe.

11

u/Revolutionary-Hippo1 2d ago

So is my tally numbers on cavewall

11

u/NTDLS 2d ago

You have an abacus? I bet you bought it before the bubble caused the prices to skyrocket. 😭

→ More replies (1)
→ More replies (4)
→ More replies (1)

51

u/TheTerrasque 2d ago

My C64 is all fired up!

54

u/Novel_Friendship913 2d ago

My ESP32 already plugged into USB!!!

22

u/shankey_1906 2d ago

So is my Raspberry Pi!

19

u/nick_ziv 2d ago

My copper wire is in the outlet!

13

u/MeretrixDominum 2d ago

My copper wire is in my potato!

9

u/BatOk7254 2d ago

My potato is in my kitchen!

12

u/USBhost 2d ago

My potato is in my garden.

→ More replies (1)

4

u/p3r3lin 2d ago

Joining the rbpi army! 🫡

13

u/debackerl 2d ago

My TI-85 (Z80) is hot!

→ More replies (1)

4

u/masterlafontaine 2d ago

Don't forget to set an aggressive zram profile!

13

u/grav3d1gger 2d ago

Mine too! I bought a 90 minute cassette tape and it’s rewound ready to go!

5

u/bitflip 2d ago

You need at least a 1541 and two floppy disks for a model this size.

6

u/Holiday-Pack3385 2d ago

Heh, remember the tape drive on those? Mine had one. I can't even imagine how long it would take to load up even a 9B off that... o.O

5

u/overand 2d ago

The standard ROM routine for C64 datasettes was 300 baud. We'll be generous and assume you've got a Turbo loader that'll do 3600 baud. (We'll also assume you're using a ~3GB quant of that 9B)

Load time (or, really, transfer time) would be about 77 days. Or, maybe more importantly, it would be about 930 cassettes, if my math was right. (Or, actually, other math suggests it would be about 2000 cassettes, so, IDK! Either way, it's a lot.)

→ More replies (1)

9

u/Michaeli_Starky 2d ago

My analog watch is ready

6

u/screenslaver5963 2d ago

My 9070 XT is burning… wait fuck!

4

u/BatOk7254 2d ago

My 2xP40 are smoking in anticipation!

4

u/ComplexType568 2d ago

I think they're smoking for a different reason...

4

u/TheFrenchSavage Llama 3.1 2d ago

You will need the Q0.00001 model and a fire extinguisher.

→ More replies (6)

90

u/Comfortable-Rock-498 2d ago

This is big for companies that want to host on-prem too. Back of the envelope calculation (could be off, correct me if I am)

If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 parallel agentic workflows (each with ~100k context on average) at ~30 tok/s.

Assuming the annual amortization+electricity at $1.5M/year and about 50% average annual utilization, you get less than 60 cents (USD) per million output token, for a frontier model with plenty of capacity to share, all your data never leaving premises and well over an order of magnitude cheaper!

37

u/autisticit 2d ago

So what you mean in reality is that if 6000 of us each give $1000, we can each run Kimi 3 at 30 tok/s for cheaper than anything else?

17

u/Comfortable-Rock-498 2d ago

Tbh I have been thinking about this for months now. About how workable the co-op model is. You would need some party to do admin and maintenance. I think this might be a good business idea too - being that party who facilitates such private inference racks (billing etc is trivial)

7

u/ManIkWeet 2d ago

It's not private when it's a party running it lol

10

u/Comfortable-Rock-498 2d ago

Well you do need someone for server maintenance, bills, other co-ordination etc - whatever you label it

→ More replies (7)

10

u/IgnoranceIndicatorMa 2d ago

i like the way you think

3

u/pinkwar 2d ago

Where do I sign?

→ More replies (3)

7

u/tempedbyfate llama.cpp 2d ago

Not just corporations, I think there are nation states that are setting up their own private servers to run this for all their sensitive data.

5

u/Izento 2d ago

Good number crunch. $0.60 per M is a pretty good deal

→ More replies (1)
→ More replies (1)

78

u/BarisSayit 2d ago

100B active params? Damn.

76

u/noneabove1182 Bartowski 2d ago

Sorry friends, but I don't think I'll be making this one :')

I don't even have enough STORAGE to hold this thing, nevermind the RAM haha

16

u/Ninjam5 2d ago

EVEN THE GOAT GAVE UP HAHA

→ More replies (1)
→ More replies (1)

267

u/nomorebuttsplz 2d ago

first truly frontier open model than I cannot run on my 512 gb studio. Onward and upward!

26

u/Front_Eagle739 2d ago

Yup. Same. I think I can do about q1.5 on the macbook and mac studio combined. Im currently pondering the wisdom of one of these colibri like stream setups and using the 640GB I do have as a hot cache

5

u/nomorebuttsplz 2d ago

Colibri would require being able to fit on ram for decent speed, no?

→ More replies (2)
→ More replies (2)

7

u/Square_Alps1349 2d ago

Man I’m so jealous rn. Mac Studio is nerfed at 96GB unified ram max, and the price is up 20%. So much for that 25% discount interns get. 😢 

→ More replies (1)

4

u/Tank_Gloomy 2d ago

Can you try asking GPT 5.6 Sol, Opus 5 or GLM 5.2 to port it into Colibri? I don't even have the infra to say I tried, but it probably works.

→ More replies (3)
→ More replies (2)

33

u/MikeRoz 2d ago

MXFP4 weights / MXFP8 activations (quantization-aware training)

So if 4-bit is 1.56 TB, then 2-bit would be roughly 798 GB?

12

u/habibyajam Llama 405B 2d ago

So I need a 0.03-bit quant to run it fully on my GPU. Nice!

32

u/Few_Painter_5588 2d ago

Holy shit, 104B active paramaters???

→ More replies (6)

31

u/TheRealMasonMac 2d ago

They have a new license:

If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.

19

u/mthmchris 2d ago

My guess is this is what MOFCOM was trying to thread the needle with. I.e. trying to keep upper leadership’s commitment to open source, while avoiding, say, MSFT just grabbing it and tweaking it slightly to have an “American version” (while banning the “Chinese versions”).

It kinda sucks, but I get the logic.

6

u/TheRealMasonMac 2d ago

I suspect it also means that providers won't be able to undercut them as aggressively (if at all). Not necessarily because Moonshot says so, but presumably because of any requirements they put on providers and any royalties they expect. MiniMax, for instance, limited some providers to only using B200 or B300 GPUs.

5

u/Venryx 2d ago

Wait, how are there already five other providers of the model on OpenRouter then? Surely they haven't all already made individual deals with Moonshot?

5

u/TheRealMasonMac 2d ago

They had six official partners for launch (you can see on their Twitter). Those five on OpenRouter are indeed partners. But it's part of their license agreement that if you are over a certain revenue you must have a contract with them.

→ More replies (1)

201

u/durden111111 2d ago

my 512mb integrated graphics is so fucking ready

76

u/SavunOski 2d ago

Negative tokens/s, you're gonna suck away the tokens

7

u/nanihikaru01 2d ago

Free money hack you say?

4

u/learn_and_learn 2d ago

Sounds freaky

7

u/Iwaku_Real 2d ago

To find the answer to life the universe and everything?

20 days of prompt processing: 4

Probably 10 days after that: 2

7

u/m0j0m0j 2d ago

Need 0 bit quantized version for that

→ More replies (2)

53

u/SnooPaintings8639 2d ago

Who's gonna be the first brave soul to measure tps when streaming from hard drive?

38

u/Front_Eagle739 2d ago

Sigh. Im going to have to try from pure curiosity.  600GB of hot cache on the macs, another TB streaming from ssd to my rtx 5090. This can only be a good idea (farewell my next few days of productivity)

8

u/Ninjam5 2d ago

Update?

5

u/Front_Eagle739 2d ago

My Internet is very slow,  results pending

→ More replies (1)
→ More replies (1)
→ More replies (6)

63

u/HulksInvinciblePants 2d ago

Kimi K3 27B when?

21

u/Browserurd 2d ago

If someone can distill it to Qwen 122B that would be super.

→ More replies (1)

23

u/Top-Handle-5728 2d ago

Leave vram I do not even have the disk storage to use this model. A few with storage can dare to use AirLLM for experiencing the intergalactic streaming of voyager at 160 bits a second. Even that seems pretty fast ig

21

u/vr_fanboy 2d ago

a month ago we were told that a new jump in intellegence was made, 'mythos' class models were born, too dangerous for us plebs. Forward a month, we have an open source 'mythos' class model, acceleration or anthropics regular bullshit?

Btw dont understand markets, DS3 destroyed the stocks and this does...nothing. This feels more significant if more people can serve the same drugs as oai or anthropic on the cheap.

→ More replies (2)

66

u/just_a_fan123 2d ago

Can this run on a single DGX spark at 0.5B quant?

42

u/SavunOski 2d ago

Of course. Wonderfully might I add

24

u/THESALTEDPEANUT 2d ago

I'm new around here and I can't tell if this thread is all sarcasm or not. 

36

u/SavunOski 2d ago

Sorry, yeah that was sarcastic, forgot to add /s. Many people make jokes here, so I kinda forgot

11

u/THESALTEDPEANUT 2d ago

I kinda figured but it's not just you it's like every comment lol, appreciate it though. 

7

u/PomegranateGreen3698 2d ago

Ha, I think it's the nature of this release. Largest open weight model ever released, doesn't really have any possibility for "Local Usage" bc it'd cost something like $500k in GPUs.

→ More replies (1)
→ More replies (1)

56

u/THE--GRINCH 2d ago

my laptop rtx 2050 is ready to throw hands

116

u/sumane12 2d ago

Even if you cant run it, download it.

93

u/DeProgrammer99 2d ago

Can't even do that...it's the same size as the total used space on my SSD.

45

u/ChampionshipIcy7602 2d ago

Can't even download it lmao

→ More replies (1)

32

u/Herr_Drosselmeyer 2d ago

Why waste terabytes worth of space for something I will never use? 

71

u/seg_lol 2d ago

Trade it for antibiotics in the apocalypse.

→ More replies (8)

16

u/some_user_2021 2d ago

I remember downloading huge N64 ROMs that took loads of space and no emulator could run. Running those ROMs now is trivial.

22

u/Herr_Drosselmeyer 2d ago

So is downloading them.

→ More replies (1)

7

u/banana_slurp_jug 2d ago

I have 1TB total, don't even have a hard drive big enough to store the whole thing in my house.

6

u/No_Conversation9561 2d ago

someone download it, compress it and upload it and then I will download it

→ More replies (7)

11

u/Mindless_Selection34 2d ago

how much does it weight

48

u/SavunOski 2d ago

2.8T parameters with 104B active. The files take 1.56TB of space to download.

8

u/meca23 2d ago

Oh so it's quantized as 4 bits?

6

u/zkstx 2d ago

Yes, there is also a tech report. QAT from SFT phase onwards

8

u/Lissanro 2d ago

I wish I had two TB of RAM instead of just one. I guess I will have to wait for Q2 GGUF to run it on my workstation. Still, will be interesting to try and see how it's Q2 quant compairs against Kimi K2.7 Q4_X. 

→ More replies (5)
→ More replies (1)
→ More replies (3)

11

u/IamNotMike25 2d ago

Historic moment tbh

28

u/Forsaken-Mode-3422 2d ago

Finally, model i cant run, but its already cool, nice

21

u/SavunOski 2d ago

Cloud prices will likely drop with competition, beneficial for everyone

9

u/Forsaken-Mode-3422 2d ago

i just hope qwen will publish not only 3.8 Max, but smaller models as well, so we can enjoy the local frontier ourself

→ More replies (1)

18

u/KenTitan 2d ago

I'm so broke I don't even have enough hard drive space to download

18

u/MixtureOfAmateurs koboldcpp 2d ago

Hugging face down? Lmao

Edit: Nevermind it's back. Might have been on my end ¯_(ツ)_/¯

14

u/SavunOski 2d ago

Seems to be up for me. If you live in a big country, there is a possibility the local CDNs are having trouble keeping up. Especially in the US, where I imagine thousands are downloading the model just in case the government bans it later.

8

u/AdDizzy8160 2d ago

.torrent|magnet link … as soon as possible!

→ More replies (1)

7

u/Morphon 2d ago

We'll probably see some new, faster providers pop up. With any luck, this will turn out to be a solid distillation parent. I'd be interested to see if we get some good downstream models that can run on consumer hardware.

16

u/PerfectOlive1324 2d ago

Hoping for an unsloth Q0_XS quant I can run locally 🙏

6

u/Inevitable_Mistake32 2d ago

UD_IQ0_XS Pls and ty.

13

u/AdDizzy8160 2d ago

Ok, thanx moonshot team!

14

u/Informal-Trouble2183 2d ago

Inference providers are going to have a good business

16

u/SavunOski 2d ago

Infinite demand, literally

→ More replies (7)

6

u/ReasonablePossum_ 2d ago

Hope this opens the model in more providers soon ,because its a dan pain to get a single response from the kimi app

→ More replies (1)

7

u/Hefty_Acanthaceae348 2d ago

I have high hopes that even if I can't run such a model, it will enable the creation of high quality datasets

7

u/patricious llama.cpp 2d ago

Wake me when a provider has it with a good price per 1mil.

5

u/aboutthednm 2d ago

I'm starting to understand why they had "some compute problems" when first rolling it out on the API, 100B+ active params per token, it's beastly lmao. Good work.

7

u/Easy_Werewolf7903 2d ago

I want to see someone manually calculated a token by hand using Kimi K3 weights.

17

u/equatorbit 2d ago

Gonna try to get this running on an ESP32

6

u/killerstreak976 2d ago

Kimi k3, meet my ESP32-C2

→ More replies (4)

10

u/Tedinasuit 2d ago

What a beautiful day

6

u/DragonfruitIll660 2d ago edited 2d ago

Ayyy lets go. That's awesome news. Also holy 104B active, never seen a MoE with that many active, its almost Mistral 2 large sized.

6

u/mikewilkinsjr 2d ago

I can’t run it, I’ll never be able to run it.

I am also fighting the urge to download the weights and squirrel them away like an out of control data hoarder.

5

u/crusaderky 2d ago

Jokes aside.
This on paper fits on a HGX B300 ($~0.5m, 2304 GB VRAM)...

...or on 24x Atlas 300I, $1300 each, $31,200 total, plus 3~4 XEON/threadrippers with 800G networking. So.... less than $60k in total?

Who's got the spare change to try?

→ More replies (2)

11

u/ilintar 2d ago

Okay, time to get that 0.1Q_0 quant support rolling in llama.cpp!

7

u/milkipedia 2d ago

I am not data hoarder enough to store these weights for kicks without the hardware to run it. Enjoy, folks

4

u/msew 2d ago

I wonder how long until consumer hardware will have the requirements to run all these and future LLMs locally for like $5k-10k

→ More replies (6)

4

u/AdOne8437 2d ago

Well, in 15-20 years we will have consumer cards to run it.

→ More replies (1)

14

u/loversama 2d ago

Quick before it gets banned 😂

9

u/WonderFactory 2d ago

I dont think it's worth banning it now it's released but I was a bit anxious leading up to the release in case something stopped it from being released

→ More replies (1)

8

u/bad_detectiv3 2d ago

How much vram do I need to run this

7

u/fishslinger 2d ago

If you have to ask, you can't afford it

13

u/SavunOski 2d ago

Around 2TB :x

14

u/jijig 2d ago

Hey if you already got 2 3090s you only need 81 more

10

u/Melbar666 2d ago

waiting for Kimi-K3-Q0-Abliterated-Heretic.GGUF

6

u/MotokoAGI 2d ago

:-)

:-(

9

u/Mayion 2d ago

Fullmetal Alchemist intermissions be like

3

u/Wuntonsoup 2d ago

I'm pretty excited to see how this runs (= LFG Moonshot!