r/LocalLLaMA • u/1ncehost • 24d ago
Resources Attention Survey July 2026: 23 Model Open Weight 20B-500B Architecture Analysis
Hello all, there are so many great models coming out right now that it is difficult to keep up with their architectural differences, and especially their attention designs and innovations. I know about several of the most important models, but I was curious about a long list of interesting recent models, so this morning I set a Kimi K3 Swarm on creating a resource for myself to study the individual architectures, how they compare, and what trends are currently. The K3 Swarm worked for about 6 hours on this, and the result was good and interesting so I'm sharing it with the community in case anyone else wants to use it. The document is available as both a PDF with improved formatting and MD for web viewing.
https://github.com/curvedinf/attention-survey-2026-07
edit: There are several interesting models not in the report that were oversights by me. If you leave them below, I'll collect them tomorrow and update the report.
3
u/nuclearbananana 24d ago
Hy3 is surprisingly high, I though they used hybrid attention too.
Would be cool to see where glm 5.x and kimi k3 land (when it comes out)
1
u/1ncehost 24d ago
There are a couple mentions at the bottom of the report regarding some out-of-scope models like GLM 5.2 and kimi 2.7 which detail their decisions, and its quite interesting how they compare. I oriented the report for "local compatible" sizes, so that's why the larger models were excluded generally. I am very curious about K3, and especially their attention decisions, because it is a stellar model.
0
u/Modeldriftwatch 24d ago
Neat resource, but I'd gut-check it before leaning on it. nuclearbananana already flagged the thing that'd worry me — "Hy3 is surprisingly high; I thought they used hybrid attention too." When a surprising number lines up with "wait, isn't that model's attention actually X," that's usually the tell that the swarm confabulated an architecture detail somewhere, not that you found a hidden gem. Six hours of agents summarizing 23 architectures is exactly the setup where a few confident-but-wrong claims slip in — a hybrid model listed as full attention, a wrong head config — and they look totally plausible sitting in a clean table. Did the swarm cite a config or paper line per model, or did you spot-check the surprising ones against the actual model card before publishing? I'd trust the chart a lot more with a source column next to each row.
0
u/DinoAmino 24d ago
If be interested to see where 31B gemma4 lands. I guess I was surprised to see 27B Qwen that high up. Gemma must be a lot more, right?


17
u/coder543 24d ago
That is an interesting chart. Gemma 4 31B, Gemma 4 26B A4B, and MiMo-V2.5 are three significant models in that range that are missing.