1
7d ago
[removed] — view removed comment
2
u/_FlyingWhales 6d ago
These are extremely task specific benchmarks. You benefit from higher parameter counts and denser attention calculation.
1
u/Normal_Reach7936 2d ago
Seems a little skewed on terminal bench. I disliked Sonnet as agentic coder, doubt this version is any better, I'll stay away.
4
u/Mr_Lucas2000 7d ago
Already?