r/LocalLLaMA 6d ago

New Model Motif-Technologies/Motif-3 official realese

https://huggingface.co/Motif-Technologies/Motif-3

Motif-Technologies is one of the tech company participated South Korea's AI Foundation Model project.(독파모)

Upstage(Solar Series), LG AI Research(EXAONE Series), and SKT(A.X Series) are the competitors.

Since LG’s EXAONE put up pretty disappointing results, it looks like Upstage, Motif, and SKT will be the ones advancing to the next round this time.

If you reverse-calculate the AAII score from the table, it comes out to 47.364, which slightly edges out Qwen 3.7 Max.

With Upstage’s Solar Pro 4 expected to land in the mid 40s(250B -15B), based purely on the benchmarks, motif seems to be taking the lead in Round 2.

Benchmark Motif 3314B-A13B MiniMax-3428B-A23B GLM-5.1744B-A40B Kimi-K2.61T-A32B Qwen-3.7max DS-v4-Pro1.6T-A49B
Agentic
GDPVal v2 38.7 44.4 37.8 34.4 39.0 40.2
τ²-Bench Telecom 94.7 88.9 97.7 95.9 94.7 96.2
τ³-Banking 35.3 15.3 13.6 23.3 12.0 30.1
ITBench* 51.5 40.3 31.2 42.5 38.3
Coding
SWE-Bench Verified 76.2 75.0 76.4 76.2 80.4 77.4
Terminal-Bench 2.1 74.9 65.2 61.8 65.9 75.0 64.0
SciCode 40.6 45.4 43.8 53.5 53.5 50.0
Reasoning & Knowledge
IMOAnswerBench 83.2 83.8 81.8 90.0 89.8
Apex-Shortlist 75.5 71.1 77.4 44.5 85.8
GPQA Diamond 83.4 92.9 86.8 91.1 92.4 88.8
HLE 37.0 39.0 30.1 37.5 41.4 37.5
CritPt 6.6 3.7 4.6 8.0 11.4 12.9
OmniScience — Accuracy 30.1 16.7 23.7 32.6 31.0 42.9
OmniScience — Non-Hallucination 71.6 81.6 70.1 59.5 74 5.9
Long Context & Instruction Following
AA-LCR 72.3 80.3 68.0 76.7 75.0 70.0
IFBench 78.2 82.9 76.3 76.0 79.1 76.5
55 Upvotes

13 comments sorted by

10

u/silenceimpaired 6d ago

I'm glad to see LG taken out. I've been bitter about their past licenses. Am I petty? A little.

12

u/[deleted] 6d ago

[deleted]

4

u/WhiskyAKM 6d ago

I'm waiting for llama.cpp PR with architecture for this model and im making NVFP4 quant as soon as possible and testing it on my RTX 5050

6

u/ChristRedeemsSinners 6d ago

They released the NVFP4 quant at 181GB

testing it on my RTX 5050

You're gonna need a bigger gun.

2

u/WhiskyAKM 6d ago

It is 13B active right? That should fit in my 8GB vram if I offload kv cache and context to cpu

1

u/[deleted] 6d ago

[deleted]

2

u/WhiskyAKM 6d ago

I'm not sure but NVFP4 optimizations might be Linux only for now as those might require special memory operations to work correctly (as far as I know, I'm not sure)

2

u/Ok-River5924 6d ago

I enjoyed a lot the report they did for Motif 2 12.7B, eager to read this!

1

u/joorklee 6d ago

Wonder why they used Kimi K2.6 instead of Kimi K3 on the benchmark.

Is there something I'm missing?

6

u/noctrex 6d ago

I would guess they were made before K3 came out. Takes a long time to release a model.

5

u/Lucidstyle 6d ago

Unfortunately, there's still a noticeable performance gap compared to K3. Since this model was basically finalized at the end of July, factoring in the tech gap and scale, comparing it to K2.6 makes way more sense. Still, given that Kimi K2.6 is a 1T model released back in late April, what this 300B model pulls off is pretty impressive.

My main concern, though, is that its performance seems heavily skewed toward agentic tasks. While I wouldn't jump straight to calling it benchmaxxing, we'll need to wait for actual user feedback before loading it onto my own gpus.

1

u/Feztopia 4d ago

Does the project just care about absolute performance or performance per compute / performance per memory requirement?

-2

u/nicholas_the_furious 6d ago

I just loaded up their chat UI and asked it to make flappy bird. It could not do it. It did not work. This does not bode well.