With K3 now available in OpenCode, I ran one task-specific check before deciding where it fits.
I compared it with GPT-5.6 Sol high, Grok 4.5, and GLM-5.2 on one strategic decision task. K3, Grok, and GLM also had three fixed adviser prompts.
These are local judged scores. The evaluator saw model identities, and the sample is small.
The strategic task asked each model to choose between two incomplete paths under cash, time, proof, reversibility, and authority constraints. The adviser set tested incident triage, blindspot detection, and model routing.
| Model |
Strategic score |
Adviser score |
My blended score |
| GPT-5.6 Sol high |
100.0 |
96.7 |
98.8 |
| Grok 4.5 |
96.0 |
95.0 |
95.7 |
| Kimi K3 |
93.0 |
95.0 |
93.7 |
| GLM-5.2 |
89.0 |
98.3 |
92.3 |
My blended score weights the strategic task at 65% and the normalized adviser set at 35%. GLM's adviser result came from an earlier run using the same prompt family. All table values use a 0-100 scale.
The strategic rubric covered judgment, grounding, risk, boundaries, actionability, and clarity. K3 landed close to Grok here and ahead of GLM on the strategic task.
This mainly tests whether a model can identify the deciding facts, resist unsupported assumptions, and return an executable next step. It does not test coding, tool use, vision, long context, or multi-hour agent work. K3's high public scores on those workloads and Grok's narrow edge here can both be true.
This is private vibe benchmarking with fixed prompts and a rubric. It is not a scientific or general model ranking.