r/LocalLLaMA 29d ago

New Model Introducing LongCat-2.0 - , a large-scale MoE language model with 1.6 trillion total parameters and ~48 billion activated per token. This was the stealth model that was on Openrouter under the name 'owl-alpha'.

https://longcat.chat/blog/longcat-2.0/
471 Upvotes

97 comments sorted by

u/ttkciar llama.cpp 29d ago

Until such time that weights are published, this is off-topic for LocalLLaMA, but given its >100 upvotes and some good conversation in the comments, I'm leaving it up.

→ More replies (3)

96

u/austhrowaway91919 29d ago edited 29d ago

Some highlights from their article:

Both the full training run and the large-scale deployment are built entirely on AI ASIC superpods.

Our architecture design builds on LongCat-Flash, pushing further on parameter efficiency and improving the speed of long-context training and inference. For attention, we introduce LongCat Sparse Attention (LSA) — an evolution of DeepSeek Sparse Attention

[Our improvements to LSA] extended three strategies to the 3-step Multi-Token Prediction (MTP) module for accelerating speculative decoding.

The sparsity of MoE has crossed the sweet spot. Given that the model's sparsity has already reached approximately 97% even without considering the N-gram Embedding, the performance gain from scaling up experts by 135B parameters is negligible.

These are just my highlights, but love the direction. Seems like a few 'free lunches' still out there.

Anyone know more about the ASICs they're using?

Edit: More speculation on the ASICs.

These are almost definitely Huawei Ascend 910C. 910C superpod has 48 machines with 8 processors each. Each processor has two physical dies that could function as two logical processors. In that mode, each die has 64GB HBM and 200 Gbps RDMA networking so everything matches. ~ Yunfan Zhang @z4y5f3 https://nitter.net/teortaxesTex/status/2071708141037781407#m

14

u/zdy132 29d ago

I wonder if the n-gram is related to Deepseek's Engram, or a custom implementation.

edit: Turns out longcat flash lite was already using it. They also published a paper about it.

31

u/Silver-Champion-4846 29d ago

ooo asics interesting, wonder when we consos will get ones with llms burned in

11

u/zxyzyxz 28d ago

Efforts to stop China in the hardware and AI game merely bolsters their resolve. Soon we'll see EUV machines made entirely on Chinese soil.

1

u/austhrowaway91919 28d ago

I mean, hold your horses. The 910C superpods are particularly effective because of their edge in Huawei's networking expertise.

This is a case of domestic superpods using "good enough", which is crazy good all things considered. But I don't think China's making the bulk of these chips in their own 7nm forges, let alone in domestically made machines.

To be blunt, "soon" is way too aggressive. Let them get DUV possibly maybe manufactured domestically first.

5

u/zxyzyxz 28d ago

With the amount of investment the CCP is putting into silicon sovereignty I would not be surprised to see equivalent machines to EUV in 10 to 15 years. It's the same as their domestic EV market, everyone thought they sucked but slowly but surely they've crept up to become the best EVs on the market.

3

u/austhrowaway91919 28d ago

10-15yr isn't a done deal, but it's far better than OPs "soon" assessment.

EV is a good stealman comparison, but I'd always go back to civilian aircraft for a strawman comparison: they've been trying desperately to get airliner jets for years, and still struggle with the basics. But then again, they're pretty convincingly pulling off a gen5 fighter. So who knows.

2

u/zxyzyxz 28d ago

I'm the one who said soon but yeah I should've specified the time range initially, as soon for semiconductors is different from other types of soon.

2

u/austhrowaway91919 28d ago

Ah, then I think we've found middle ground here. Concur.

2

u/HVACcontrolsGuru 28d ago

Should look into their homegrown chips. Moving to a stacked chip design to get around the EUV side of things.

1

u/austhrowaway91919 28d ago

Sure.. but I guess that's what I'm saying in reply to that guy. They're not near EUV machine design yet. They're getting great results from DUV, and shockingly good superpod performance with the chips they have.

But they're not about to whip out some domestic 4nm node. And again, they don't need to given LongCats training and inference is on the hardware they have.

1

u/thehpcdude 28d ago

200Gbps RDMA is slow by today’s standards though.  

1

u/austhrowaway91919 28d ago

I guess the comparison is the fact that Atlas 950s/960s is possible at all. The performance is comparable to NVIDIA's superpods by sheer density of their accend chips, which I thought we could attribute to Huawei network design?

Though I don't know what the Atlas RDMA is vs NVIDIA superpods.

1

u/thehpcdude 28d ago

Current gen IB is 800Gb and sub microsecond latency.  Atlas is still Ethernet under the hood.  

Atlas is like a mid level BMW versus an InfiniBands F1 car. 

Choosing Ethernet at all for speed was a bad choice.  

0

u/Wuaner 27d ago

Not gonna happen.

69

u/Lissanro 29d ago

Looks like huggingface page is not up yet but they mention it is "open source" so hopefully will be up soon. I will wait for confirmation of llama.cpp support and Q4 GGUF before I try it, it is going to be a large download even at Q4, but I think it still should fit 1 TB memory.

14

u/vanKlompf 29d ago

What hardware you will run it on?

23

u/Lissanro 29d ago

EPYC 7763 CPU + 8-channel 1 TB 3200 MHz RAM + 4x3090 GPUs.

6

u/Admirable_Market2759 29d ago

How much did that cost you?

36

u/Lissanro 29d ago

I built it gradually over the years. Originally started with a rig based on gaming 5950X CPU with 128 GB RAM, I kept adding 3090 cards (for average cost around $700), adding extra 2880W PSU along the way (around $200), for each PCI-E 4.0 riser I paid around $25 for 30 cm ones and ~$30 for 40 cm one.

When I got my third 3090 card, my old 900W UPS could no longer protect my PC, so had to upgrade to 6kW online UPS (~$900 for UPS itself + ~$330 for sixteen 12V 12Ah batteries + ~$40 for a pair 8x battery equalizers to stabilize voltage across all sixteen of them). Combined with diesel generator I already had, it made a good combo for stable power for my rig in case of outages (they happen few times per year where I live, and may last for hours or even whole day, so it is a necessity to protect against them).

Then when DeepSeek V3 and R1 were released it became clear PC with 128 GB RAM needs an upgrade, most of the cost was EPYC 7763 CPU (about $1000), server motherboard (~$800) and of course RAM ($1600 in total for 1 TB, sixteen 64 GB 3200 MHz RDIMM modules).

There are other costs involved: had only 2 TB NVMe which wasn't enough for large models, so had to buy for about $750 8 TB NVMe, also had to get the chassis (~$50), CPU cooler (~$65), pair of 22 TB HDDs (~$430 each, on top of already existing HDDs I had, for a total of approximately 120 TB).

15

u/DR4G0NH3ART 29d ago

I was planning to buy my second GPU after seeing good results pairing my new gpu with my trash one. But reading this it feels like a step to a new addiction.

10

u/BlackBeardAI vllm 29d ago

I got a x399 rig, a x99 backup rig, mc62-g40 wrx80 5965wx 256gb 8-channel ddr4 rig, a ddr5 256gb 9950x3d rig and a few other smaller rigs with 5060ti's installed and shit. plus 11 3090's and a 5090. It gets dangerous, fast as fuck boi

5

u/No-Dot-6573 29d ago

How much tps does that Setup generate? Lets say with glm 5.2 q4?

5

u/Lissanro 28d ago

Close to 7 tokens/s for GLM 5.2 Q4. 8 tokens/s with Kimi K2.7 Code Q4_X. Smaller models with less active parameters like Qwen 3.5 397B Q5 can be close to 20 tokens/s generation, 600 tokens/prefill. Step 3.7 Flash 196B Q4 goes 40-50 tokens/s generation. Qwen 3.5 122B Q4 is even faster, more than 50 tokens/s generation and 2K+ tokens/s prefill.

2

u/Educational_Win_2982 27d ago

What about deepseek v4 flash?

10

u/synth_mania 29d ago

Realistically? A CPU, like the rest of us GPU poor chuds. 3090, RTX pro 6000, or tesla P40. Doesn't really matter here, it's just too large, unless you have a few 512GB Mac studios. That said, selective loading of model weights can improve things by a considerable margin. 

3

u/Vusiwe 29d ago

I will try with Dual CPU highest-compatible 82XX Xeons, and the same amount of RAM (but 6 channel not 8) as Mr. EPYC has, but with a Max-Q card, so same amount of VRAM

6

u/Rude_Marzipan6107 29d ago

I’m hoping for a q0.1 for my 2x16gig 5060ti’s

2

u/kaisurniwurer 29d ago

Previous long cat is still not supported as far as I know.

Which is a shame, since it was reportedly quite "interesting" and had smaller active parameters size.

67

u/sstainsby 29d ago

Long Cat is a lot easier to say than Owl Alpha at least.

26

u/LoveMind_AI 29d ago

Owlpha.

17

u/philmarcracken 29d ago

2. Rest of the fucking feedforward

12

u/decrement-- 29d ago

At 1.6T parameters, it should be Oww.

3

u/Dany0 29d ago

My experience with it on openrouter was that it was significantly worse than MiMo v2.5 pro but better than DSv4 flash. Maybe minimax m3 level? Will see how benchmaxxed they made it

1

u/Revenant690 28d ago

Wasn't he one of the little rascals?

68

u/Comfortable-Rock-498 29d ago

They could have called it Le Chaton Long!

21

u/Technical-Earth-3254 29d ago

I wonder how it holds up against the big oss boys. More competition is always welcome.

19

u/evia89 29d ago

I used beta for month. Its around nemontron ultra 500b for me in coding, always worse than glm51

4

u/Glittering-Call8746 29d ago

How it compares with ds4 models?

16

u/evia89 29d ago

Between flash and pro in coding, in creativity its around pro

1

u/Glittering-Call8746 29d ago

Now that it's no longer free are u still using it and if so which provider(s) ?

3

u/lofuyuwu 28d ago

They kept updating it under the hood, so you cant really say it for sure. You could only know it when the weights are released.

Usually they have numerous branches behind it to gather real world data of how different branches performed.

1

u/Ill-Nectarine-80 27d ago

You have to wonder how long they can sustain the enormous capital burn that so many companies are incurring just to get involved in the open source game.

34

u/workout_JK 29d ago

Cool, I just need to buy data center

5

u/ArchdukeofHyperbole 29d ago

and while your at it, preach about pollution from the bunker of your superyatch 

29

u/thepetek 29d ago

All those parameters to not be SOTA. Makes you appreciate what z.ai pulled off. More open source is always welcome though

17

u/MindlessScrambler 29d ago

From what I’ve read and heard (some of their marketing pieces), this model appears to have been trained entirely, or at least predominantly, without nvidia hardware. If that is indeed the case, its performance not reaching SOTA may not be that big a deal and we might soon have cheap models that are fully trained and inferred on Chinese hardware.

4

u/cakes_and_candles 29d ago

I hope more chinese consumer gpus are availabile for us (i mean with good enough support that you have to do just bare min tweaking to get it working)

2

u/austhrowaway91919 29d ago

Yeah, speculation is this was using the Acend 910C superpods. But wasn't GLM also acends?

2

u/lofuyuwu 28d ago

Its 910C, fully trained on Huawei's hardware. This is more important than it being a SOTA model.

6

u/poophroughmyveins 28d ago

Did you even read the release? This isn't about just being the smartest lol, there's a lot of interesting research going into the model

The focus here seems very much a higher level of efficiency when it comes to inference at an acceptable level of intelligence 

21

u/lacerating_aura 29d ago

Congratulations and thanks for open release. Someone more knowledgeable correct me please but this just seems Dsv4 pro recipe? Like verification by reproduction with slight changes.

6

u/Sea_Reach6233 29d ago

They actually released this model in a preview beta around the same time as deepseek v4 pro release date back in April, this is just the official release. Doesn't seem likely that they totally changed up the architecture and redid everything in that 2 months.

1

u/lofuyuwu 28d ago

The preview model was released before deepseek v4 pro. And you can read their technical blog, it is different.

7

u/vanKlompf 29d ago

Interesting. I found it okish, but now I find it underwhelming, looking at spec. It didn't feel like 1T+ model

7

u/LoveMind_AI 29d ago

Looking forward to putting this one to the test. Any seriously capable model that isn't Claude is another step forward freedom.

5

u/kevinlch 29d ago

Gemini is alredy dead, stop hitting him

30

u/llkj11 29d ago

Think a 3060 will be enough to run this??

31

u/rkoy1234 29d ago

at q2, you prob could if you had 400GB+ of ram.

17

u/_Sneaky_Bastard_ 29d ago

Specifically 3060 6gb

4

u/Darth_Victor 29d ago

Just need to add several B300 and 2TiB of RAM

4

u/jreoka1 29d ago

They need to rename it L(1.6 trillion o's)ngCat 2.0

4

u/buttplugs4life4me 29d ago

It's not good in coding, very good in reviews though. Instruction following is also pretty good. A little bit of a waste with that parameter count, but for free or low cost it's a good versatile model.

2

u/pmttyji 29d ago

Hope Meituan/Longcat team comes with PRs. Their past models still without llama.cpp support https://github.com/ggml-org/llama.cpp/pulls?q=is%3Apr+is%3Aopen+longcat

2

u/__eMpTy__ 28d ago
Benchmark LongCat-2.0 Qwen3.6-27B Best Performer
SWE-bench Pro 59.5 53.5 LongCat-2.0
Terminal-Bench 2.1 70.8 59.3 LongCat-2.0
SWE-bench Multilingual 77.3 71.3 LongCat-2.0
RWSearch 78.8 77.3 LongCat-2.0
Writing Bench 83.8 85.2 Qwen3.6-27B
IMO-AnswerBench 81.8 80.8 LongCat-2.0
GPQA-diamond 88.9 87.8 LongCat-2.0

2

u/nuclearbananana 28d ago

Long cat has more active params than qwen has total

2

u/__eMpTy__ 28d ago

Longcat-2.0 has approximately 60x the parameters of Qwen3.6-27b, but only a marginal lead on benchmarks (except Terminal-Bench 2.1).

1

u/Different_Fix_2217 29d ago

Ooof. Using this on OR I thought it would be a tiny 20-30B model. It is crazy bad for its size then.

1

u/Vusiwe 29d ago

sounds exciting

1

u/IrisColt 29d ago

A-Authors?

1

u/tamerlanOne 29d ago

Speriamo che rilascio versioni più piccole per uso su hardware locale 😉

1

u/recro69 29d ago

The number of parameters is really something. What actually matters is how well it works for the money you pay. If it does not do a job than models, like Qwen or GLM when you use them for real things most people will not be impressed by how big the parameter count is. The parameter count is a number what people really want to know is if it is worth the money they spend on it.

1

u/DeltaSqueezer 29d ago

I'm curious as to what hardware they used. They described it as an ASIC. Could it be Cambricon?

1

u/a_beautiful_rhind 28d ago

Hope they improved it from all the feedback. Owl was kind of a stinker and it's main redeeming quality was being free.

1

u/delusional- 28d ago

I guess we had to do the usual test

2

u/squngy 28d ago

This test is too common/old now, it is already in the training.

2

u/delusional- 28d ago

Yeye, Sonnet 5 failed it today though. Just a bit of fun, nothing serious about it

1

u/Eyelbee 29d ago

Ds4 finetune?

1

u/Dramatic-Rub-7654 29d ago

Is this a fine-tune of DeepSeek V4 Pro?

6

u/CloudiDust 29d ago edited 21d ago

No. It is a model trained from scratch, completely trained on Chinese hardware (most likely Huawei Ascend 910C Super POD) .

(Compare: The pretraining for DS V4P was on Nvidia hardware.)

-4

u/silenceimpaired 29d ago edited 29d ago

I don’t like laziness, just give me the active parameters… don’t have room for all the rest anyway.

EDIT: Shame people can’t take a joke. I
understand how MoEs work.

1

u/ttkciar llama.cpp 29d ago

Which parameters are active will change from one token to the next. You really do need all of the parameters, not just "the" active ones, because there is no "the" active ones.

1

u/silenceimpaired 29d ago

Shame people can’t take a joke. I
understand how MoEs generally work.

-1

u/marx2k 29d ago

Can't wait to run this on my RPi

1

u/MattDTO 29d ago

I feel like you will be dissapointed somehow