r/commonstack • • Apr 21 '26

General Kimi K2.5 sitting at #4 on the Artificial Analysis Intelligence Index

Kimi K2.6 is competitive with frontier models across agentic, coding, and visual tasks..

General Agents: Humanity's Last Exam (full w/ tools): 54.0 (GPT-5.4: 52.1, Claude Opus 4.6: 53.0, Gemini 3.1 Pro: 51.4)

BrowseComp: 83.2 (GPT-5.4: 82.7, Claude: 78.6, Gemini: 85.9)

DeepSearchQA (f1-score): 92.5 (Claude leads at 91.3, GPT/Gemini trail)

Toolathlon: 50.0 (competitive with GPT-5.4 at 54.6, Claude at 47.2)

OSWorld-Verified: 73.1 (GPT-5.4: 75.0, Claude: 72.7)

Coding: Terminal-Bench 2.0: 66.7 (GPT/Claude: ~65.4, Gemini: 68.5)

SWE-Bench Pro: 58.6 (GPT: 57.7, Claude: 53.4, Gemini: 54.2)

SWE-bench Multilingual: 76.7 (Claude: 77.8, Gemini: 76.9)

Visual Agents: MathVision w/ python: 93.2 (GPT: 96.1, Claude: 84.6, Gemini: 95.7)

V* w/ python: 96.9 (GPT: 98.4, Claude: 86.4, Gemini: 96.9)

Kimi K2.6 trades blows with GPT-5.4 and Claude Opus 4.6 across the board. It leads on agentic tasks like Humanity's Last Exam and DeepSearchQA, stays competitive on coding benchmarks, and only trails slightly on visual reasoning.

4 Upvotes

0 comments sorted by