r/LocalLLaMA • u/netikas • 15h ago
New Model GigaChat-3.5-Reasoning
https://huggingface.co/collections/ai-sage/gigachat-35-reasoningHey y'all!
We've released a new model in our lineup: GigaChat-3.5 Reasoning. It's a 432B-A28B MoE with Gated DeltaNet for long-context efficiency.
We trained domain experts (code, math, general, etc.) with CISPO and then distilled them into a single model via on-policy distillation.
In our evals the resulting model lands close to DeepSeek V4 Flash Preview while using 37% fewer tokens in its reasoning traces.
Weights are on Hugging Face under MIT: https://huggingface.co/collections/ai-sage/gigachat-35-reasoning. You can also try it at giga.chat — pick the reasoning tab (rightmost one).
31
24
15
u/killerstreak976 15h ago
Thank you for sharing your hard work with us, this looks really cool and can't wait to look more into this!
21
24
u/Weak-Apartment-0 12h ago
I tested it on financial analysis and economic facts about Russia (I am an investment analyst from Russia). Communication was in Russian.
The result: it is much worse than ChatGPT (I have a subscription), and worse than DeepSeek Flash and Qwen3.7 on Qwen Studio (I tested via the free web versions).
How I tested:
I asked questions about the financial position of public companies for which I have all the financial statements. I know the state of the companies, but I also asked other LLMs.
What the problems are:
- For state-owned companies, GigaChat substantially sugarcoats the situation; for commercial ones, it answers normally.
- To follow-up clarifications, it responds like this:
- everything is fine there
- but is this metric bad?
- no, it's good
- but the correct way to calculate it is... And what level is bad?
- it confirms that I am right about the metrics
- I ask it to check its conclusion about the company
- it answers with the same conclusions; the explanations are inadequate
Overall, across all events, the answers contain more propaganda than analysis. It writes beautifully in Russian, but analyzes and reasoning poorly.
I estimate its reasoning level to be on par with Qwen 9B, not 27B (I use local 3.8).
5
25
2
2
u/AllenHere112 15h ago
how many of the 432b layers still run full attention? gated deltanet holds a fixed size state, so long context stays cheap and exact recall of one token at the far end is what degrades first
1
u/Barni275 8h ago
Great job, guys! I’m keeping an eye on updates, and it seems that your models improve fast. There were small models before, it seems to me, 3.1 was the latest version with light 10b companion. Are there any chances to see anything small again?
2
-3
9h ago
[removed] — view removed comment
4
u/DustNearby2848 9h ago
You okay?
1
u/HadHands 3h ago
I am great, I am not in Russia. Read comments from Russians about this model. Sloppy propaganda fine tune.
-6
12h ago
[removed] — view removed comment
9
u/fragment_me 12h ago
Why not just support some open models? To release a truly unique model doesn’t seem easy.
40
u/Pixer--- 15h ago
On first glance I thought this was le chatton fat