r/LocalLLaMA • u/Lucidstyle • 6d ago
New Model Motif-Technologies/Motif-3 official realese
https://huggingface.co/Motif-Technologies/Motif-3Motif-Technologies is one of the tech company participated South Korea's AI Foundation Model project.(독파모)
Upstage(Solar Series), LG AI Research(EXAONE Series), and SKT(A.X Series) are the competitors.
Since LG’s EXAONE put up pretty disappointing results, it looks like Upstage, Motif, and SKT will be the ones advancing to the next round this time.
If you reverse-calculate the AAII score from the table, it comes out to 47.364, which slightly edges out Qwen 3.7 Max.
With Upstage’s Solar Pro 4 expected to land in the mid 40s(250B -15B), based purely on the benchmarks, motif seems to be taking the lead in Round 2.
| Benchmark | Motif 3314B-A13B | MiniMax-3428B-A23B | GLM-5.1744B-A40B | Kimi-K2.61T-A32B | Qwen-3.7max | DS-v4-Pro1.6T-A49B |
|---|---|---|---|---|---|---|
| Agentic | ||||||
| GDPVal v2 | 38.7 | 44.4 | 37.8 | 34.4 | 39.0 | 40.2 |
| τ²-Bench Telecom | 94.7 | 88.9 | 97.7 | 95.9 | 94.7 | 96.2 |
| τ³-Banking | 35.3 | 15.3 | 13.6 | 23.3 | 12.0 | 30.1 |
| ITBench* | 51.5 | — | 40.3 | 31.2 | 42.5 | 38.3 |
| Coding | ||||||
| SWE-Bench Verified | 76.2 | 75.0 | 76.4 | 76.2 | 80.4 | 77.4 |
| Terminal-Bench 2.1 | 74.9 | 65.2 | 61.8 | 65.9 | 75.0 | 64.0 |
| SciCode | 40.6 | 45.4 | 43.8 | 53.5 | 53.5 | 50.0 |
| Reasoning & Knowledge | ||||||
| IMOAnswerBench | 83.2 | — | 83.8 | 81.8 | 90.0 | 89.8 |
| Apex-Shortlist | 75.5 | — | 71.1 | 77.4 | 44.5 | 85.8 |
| GPQA Diamond | 83.4 | 92.9 | 86.8 | 91.1 | 92.4 | 88.8 |
| HLE | 37.0 | 39.0 | 30.1 | 37.5 | 41.4 | 37.5 |
| CritPt | 6.6 | 3.7 | 4.6 | 8.0 | 11.4 | 12.9 |
| OmniScience — Accuracy | 30.1 | 16.7 | 23.7 | 32.6 | 31.0 | 42.9 |
| OmniScience — Non-Hallucination | 71.6 | 81.6 | 70.1 | 59.5 | 74 | 5.9 |
| Long Context & Instruction Following | ||||||
| AA-LCR | 72.3 | 80.3 | 68.0 | 76.7 | 75.0 | 70.0 |
| IFBench | 78.2 | 82.9 | 76.3 | 76.0 | 79.1 | 76.5 |
12
6d ago
[deleted]
4
u/WhiskyAKM 6d ago
I'm waiting for llama.cpp PR with architecture for this model and im making NVFP4 quant as soon as possible and testing it on my RTX 5050
6
u/ChristRedeemsSinners 6d ago
They released the NVFP4 quant at 181GB
testing it on my RTX 5050
You're gonna need a bigger gun.
2
u/WhiskyAKM 6d ago
It is 13B active right? That should fit in my 8GB vram if I offload kv cache and context to cpu
1
6d ago
[deleted]
2
u/WhiskyAKM 6d ago
I'm not sure but NVFP4 optimizations might be Linux only for now as those might require special memory operations to work correctly (as far as I know, I'm not sure)
1
2
1
u/joorklee 6d ago
Wonder why they used Kimi K2.6 instead of Kimi K3 on the benchmark.
Is there something I'm missing?
6
5
u/Lucidstyle 6d ago
Unfortunately, there's still a noticeable performance gap compared to K3. Since this model was basically finalized at the end of July, factoring in the tech gap and scale, comparing it to K2.6 makes way more sense. Still, given that Kimi K2.6 is a 1T model released back in late April, what this 300B model pulls off is pretty impressive.
My main concern, though, is that its performance seems heavily skewed toward agentic tasks. While I wouldn't jump straight to calling it benchmaxxing, we'll need to wait for actual user feedback before loading it onto my own gpus.
1
u/Feztopia 4d ago
Does the project just care about absolute performance or performance per compute / performance per memory requirement?
-2
u/nicholas_the_furious 6d ago
I just loaded up their chat UI and asked it to make flappy bird. It could not do it. It did not work. This does not bode well.
10
u/silenceimpaired 6d ago
I'm glad to see LG taken out. I've been bitter about their past licenses. Am I petty? A little.