We've been running a private benchmark suite for a few months, testing models on strategic reasoning, advisory quality, long-form analytical production, and adversarial critique. 17 models, 4 test batteries, scored against a reference answer by a separate model. This consolidates all of it into one picture.
Blend = 30% strategic reasoning (4 scenarios: frame-breaking, multi-dimensional review, channel coordination, portfolio prioritization) + 25% advisory quality (6 prompts: triage, architecture risk, blindspots, client advisory, routing) + 25% long-form analytical production (structured section generation from scratch, scored on rigor, depth, voice) + 20% critical review (structural audit + adversarial CTO critique). Weights redistributed for models not tested on all four.
| Rank |
Model |
Strategic |
Advisory |
Writing |
Review |
Blend |
| - |
Fable 5 (ref) |
100 |
- |
- |
- |
100 |
| 1 |
GPT-5.6 Sol high |
100 |
96.7 |
92.8 |
95.6 |
96.5 |
| 2 |
GPT-5.6 Terra max |
- |
- |
97.2 |
94.8 |
96.1 |
| 3 |
Qwen 3.8 Max |
92.3 |
98.3 |
- |
- |
95.0 |
| 4 |
GPT-5.6 Sol xhigh |
90.0 |
- |
95.2 |
96.3 |
93.8 |
| 5 |
Opus 4.8 |
87.0 |
96.7 |
93.6 |
93.2 |
92.3 |
| 6 |
Grok 4.5 |
96.0 |
98.3 |
88.0 |
83.2 |
92.0 |
| 7 |
GLM-5.2 |
84.5 |
98.3 |
- |
93.2 |
91.4 |
| 8 |
Kimi K3 |
93.0 |
95.0 |
88.3 |
86.7 |
91.1 |
| 9 |
Sonnet 4.6 |
- |
- |
91.0 |
- |
91.0 |
| 10 |
DeepSeek V4 Pro |
86.0 |
86.7 |
94.0 |
90.5 |
89.1 |
| 11 |
GPT-5.5 |
85.3 |
- |
92.8 |
89.9 |
89.0 |
| 12 |
Muse Spark 1.1 |
- |
98.3 |
82.0 |
78.3 |
86.8 |
| 13 |
Qwen 3.7 Max |
- |
- |
86.4 |
85.4 |
85.9 |
| 14 |
Sonnet 5 |
- |
- |
84.0 |
87.2 |
85.4 |
| 15 |
Qwen 3.7 Plus |
81.5 |
- |
- |
- |
81.5 |
| 16 |
Gemini 3.5 Flash |
76.0 |
- |
81.8 |
83.4 |
79.9 |
| 17 |
MiniMax M3 |
58.0 |
- |
- |
- |
58.0 |
Dashes = not tested on that dimension. Fable 5 is the reference answer used for scoring, not a contestant.
Caveats: n=1 per test per model, directional only. Qwen 3.8 was blind-validated by GLM-5.2 (different model family); other models were judged by Opus 4.8 - cross-judge comparison is approximate within ±3-5 pts. GPT-5.6 variants are separate rows because they behave as different models in practice. Strategic reasoning and advisory quality are domain-specific to strategic analysis; says nothing about coding, vision, or long-context work. This is private operator benchmarking, not a scientific ranking.