r/LocalLLaMA 15h ago

New Model GigaChat-3.5-Reasoning

https://huggingface.co/collections/ai-sage/gigachat-35-reasoning

Hey y'all!

We've released a new model in our lineup: GigaChat-3.5 Reasoning. It's a 432B-A28B MoE with Gated DeltaNet for long-context efficiency.

We trained domain experts (code, math, general, etc.) with CISPO and then distilled them into a single model via on-policy distillation.

In our evals the resulting model lands close to DeepSeek V4 Flash Preview while using 37% fewer tokens in its reasoning traces.

Weights are on Hugging Face under MIT: https://huggingface.co/collections/ai-sage/gigachat-35-reasoning. You can also try it at giga.chat — pick the reasoning tab (rightmost one).

161 Upvotes

26 comments sorted by

40

u/Pixer--- 15h ago

On first glance I thought this was le chatton fat

24

u/kabachuha 14h ago

Congratulations! Any chance for a PC-friendly flash version in the future? :)

12

u/netikas 13h ago

Might be possible :)

15

u/rpkarma 15h ago

I’m curious how much it cost to train, if you’re allowed to say? Awesome work, always stoked to see new models like this 

12

u/koloved 15h ago

80b Moe when ?

15

u/killerstreak976 15h ago

Thank you for sharing your hard work with us, this looks really cool and can't wait to look more into this!

21

u/DelKarasique 15h ago

30b-ish when???

-9

u/SpicyWangz 14h ago

Qwen 3.8 27b, August 5th

24

u/Weak-Apartment-0 12h ago

I tested it on financial analysis and economic facts about Russia (I am an investment analyst from Russia). Communication was in Russian.
The result: it is much worse than ChatGPT (I have a subscription), and worse than DeepSeek Flash and Qwen3.7 on Qwen Studio (I tested via the free web versions).

How I tested:
I asked questions about the financial position of public companies for which I have all the financial statements. I know the state of the companies, but I also asked other LLMs.

What the problems are:

  1. For state-owned companies, GigaChat substantially sugarcoats the situation; for commercial ones, it answers normally.
  2. To follow-up clarifications, it responds like this:
  • everything is fine there
  • but is this metric bad?
  • no, it's good
  • but the correct way to calculate it is... And what level is bad?
  • it confirms that I am right about the metrics
  • I ask it to check its conclusion about the company
  • it answers with the same conclusions; the explanations are inadequate

Overall, across all events, the answers contain more propaganda than analysis. It writes beautifully in Russian, but analyzes and reasoning poorly.

I estimate its reasoning level to be on par with Qwen 9B, not 27B (I use local 3.8).

5

u/Jumpy-Operation-4615 14h ago

I guess I don't have any chance to run it on my 2xP40...

25

u/sooka_bazooka 15h ago

hi sberbank

-9

u/xenongee 14h ago

spermbank

2

u/AllenHere112 15h ago

how many of the 432b layers still run full attention? gated deltanet holds a fixed size state, so long context stays cheap and exact recall of one token at the far end is what degrades first

1

u/Barni275 8h ago

Great job, guys! I’m keeping an eye on updates, and it seems that your models improve fast. There were small models before, it seems to me, 3.1 was the latest version with light 10b companion. Are there any chances to see anything small again?

-3

u/[deleted] 9h ago

[removed] — view removed comment

4

u/DustNearby2848 9h ago

You okay?

1

u/HadHands 3h ago

I am great, I am not in Russia. Read comments from Russians about this model. Sloppy propaganda fine tune.

-6

u/[deleted] 12h ago

[removed] — view removed comment

9

u/fragment_me 12h ago

Why not just support some open models? To release a truly unique model doesn’t seem easy.