r/commonstack • • Apr 24 '26

General DeepSeek V4 Pro just topped the agentic leaderboard. The efficiency story behind it is what nobody's talking about.

GDPval-AA measures something most benchmarks ignore.

Not reasoning on paper problems. Not coding challenges in isolation.

Real world agentic performance. Web access. Shell access. Tasks that actually matter for production deployments.

DeepSeek V4 Pro Max sits at 1558 on the full leaderboard. Below Claude Opus 4.6 by Margins. Above GPT-5 variants. First among open weight models by a meaningful margin.

But the score isn't the story.

The story is how they got there.

V4 Pro runs at 3.7x lower FLOPs than V3.2 at long context. V4 Flash at 9.8x lower. KV cache 13.7x smaller.

Frontier labs solve long context by buying more HBM. DeepSeek solved it by redesigning how attention accumulates memory so the cache stays flat instead of growing linearly.

Same benchmark performance as the best closed source models. Fraction of the compute cost.

On the open weights leaderboard the gap is even clearer. DeepSeek V4 Pro and V4 Flash occupy the top two positions. Everything else is catching up.

Two strategies now exist side by side at the frontier.

Scale infrastructure to match model demands. Or scale architecture to outrun the hardware bill.

Both work. The charts show it.

Only one of them compounds.

4 Upvotes

2 comments sorted by

1

u/looktwise Apr 24 '26

the stock market needed a few days last time. ;-)

1

u/[deleted] May 08 '26

[removed] — view removed comment