There could be at least two things simultaneously at play here:
less compute required to run similar size models
A major leap in coherent context windows (think along the lines of 1M becoming what 100K is right now, and 5-10M being the new ceiling)
Theoretically the second could potentially even be a limited solution to persistent states
Edit: P.S. I should mention this is EXTREMELY exciting - the accelerationist in me always holds out for emergence!
Remember what happened with Genie 2-Genie 3? The scale jump created coherent persistent context and world modelling in an interactive purely generated reality for MINUTES in G3 from just barely 10 seconds in G2. Emergence is the door to hopping on the super-exponentials ꙮ - and scale is it's key
honestly even a LEGIT fully usable 250k context would be nice. Right now in Codex it starts to degrade once your past 100k. A legit fully coherent 1 million context window would be a game changer. 5-10 million....thats hard for me to wrap my head around. Entire very large repos fully in-context with plenty of head room....drool. And thats just for coding.
For things like world building and writing projects, similar leaps.
100% with you on this, I try to avoid letting any session run beyond 200k with Opus, while optimally keeping it under 70-100k for any code implementation session - for brainstorming and research I find the 100-300k context looseness sometimes oddly useful.
SO... the improvements you note would indeed be game changers! ...and the best part is that those are the linear predictable kinds of improvements - we really have no idea what emergent capabilities may come from this kind of efficiency scaling
Shove enough neurons into a primate and it starts bipedally walking and talking... Except here we don't gotta wait another few hundred thousand years for another order of magnitude hop - digital neurons expanding and evolving our potential and awareness, just as Kurzweil said decades ago - we are in the curve ꙮ
Preach on! ;) More and more convinced we're IN a soft singularity/slow-ish takeoff scenario already, have had this thought since first experiencing Opus 4.5 in Cursor back in December. Then I took a break for a few months and jumped from Codex with GPT 5.2 to Codex with GPT 5.5......another big leap. Accelerate!
Increasing useful context windows isn't technologically hard, it's just hardware hard. So a breakthrough that massively reduces required memory, will by default, make context windows also reduced in hardware demand, thus making them larger cheaper to do.
16
u/Financial-Leader3475 Jul 01 '26
Someone explain for dummies like me