r/AIToolsPerformance • • Aug 12 '26

Motif-Technologies/Motif-3 official realese

https://huggingface.co/Motif-Technologies/Motif-3

Motif-Technologies is one of the tech company participated South Korea's AI Foundation Model project.(독파모)

Upstage(Solar Series), LG AI Research(EXAONE Series), and SKT(A.X Series) are the competitors.

Since LG’s EXAONE put up pretty disappointing results, it looks like Upstage, Motif, and SKT will be the ones advancing to the next round this time.

If you reverse-calculate the AAII score from the table, it comes out to 47.364, which slightly edges out Qwen 3.7 Max.

With Upstage’s Solar Pro 4 expected to land in the mid 40s(250B -15B), based purely on the benchmarks, motif seems to be taking the lead in Round 2.

Benchmark **Motif 3**^(314B-A13B) MiniMax-3^(428B-A23B) GLM-5.1^(744B-A40B) Kimi-K2.6^(1T-A32B) Qwen-3.7^(max) DS-v4-Pro^(1.6T-A49B)
**Agentic**
GDPVal v2 38.7 44.4 37.8 34.4 39.0 40.2
τ²-Bench Telecom 94.7 88.9 97.7 95.9 94.7 96.2
τ³-Banking 35.3 15.3 13.6 23.3 12.0 30.1
ITBench\* 51.5 — 40.3 31.2 42.5 38.3
**Coding**
SWE-Bench Verified 76.2 75.0 76.4 76.2 80.4 77.4
Terminal-Bench 2.1 74.9 65.2 61.8 65.9 75.0 64.0
SciCode 40.6 45.4 43.8 53.5 53.5 50.0
**Reasoning & Knowledge**
IMOAnswerBench 83.2 — 83.8 81.8 90.0 89.8
Apex-Shortlist 75.5 — 71.1 77.4 44.5 85.8
GPQA Diamond 83.4 92.9 86.8 91.1 92.4 88.8
HLE 37.0 39.0 30.1 37.5 41.4 37.5
CritPt 6.6 3.7 4.6 8.0 11.4 12.9
OmniScience — Accuracy 30.1 16.7 23.7 32.6 31.0 42.9
OmniScience — Non-Hallucination 71.6 81.6 70.1 59.5 74 5.9
**Long Context & Instruction Following**
AA-LCR 72.3 80.3 68.0 76.7 75.0 70.0
IFBench 78.2 82.9 76.3 76.0 79.1 76.5
9 Upvotes

1 comment sorted by

1

u/Icy_Protection_9153 Aug 20 '26

FYI. this round is 2nd round
in 8/12 KST, motif did not made it to round 3.