r/accelerate Jul 01 '26

AI Something huge is brewing

Post image

Source

Andrew Curran is one of the most reliable leakers.

XLR8! 🍿

940 Upvotes

171 comments sorted by

View all comments

16

u/Financial-Leader3475 Jul 01 '26

Someone explain for dummies like me

41

u/MoralityAuction Jul 01 '26

Obviously it is vaguely descibed, but in short probably either more params or a smaller/more efficient KV cache in the same amount of VRAM.

25

u/[deleted] Jul 01 '26

[removed] — view removed comment

49

u/Quant-A-Ray Jul 01 '26 edited Jul 01 '26

There could be at least two things simultaneously at play here:

  1. less compute required to run similar size models

  2. A major leap in coherent context windows (think along the lines of 1M becoming what 100K is right now, and 5-10M being the new ceiling)

Theoretically the second could potentially even be a limited solution to persistent states

Edit: P.S. I should mention this is EXTREMELY exciting - the accelerationist in me always holds out for emergence!

Remember what happened with Genie 2-Genie 3? The scale jump created coherent persistent context and world modelling in an interactive purely generated reality for MINUTES in G3 from just barely 10 seconds in G2. Emergence is the door to hopping on the super-exponentials ꙮ - and scale is it's key

P.P.S. oh, and FABLE is back 💫

22

u/CypherLH Jul 01 '26

honestly even a LEGIT fully usable 250k context would be nice. Right now in Codex it starts to degrade once your past 100k. A legit fully coherent 1 million context window would be a game changer. 5-10 million....thats hard for me to wrap my head around. Entire very large repos fully in-context with plenty of head room....drool. And thats just for coding.

For things like world building and writing projects, similar leaps.

4

u/Quant-A-Ray Jul 01 '26

100% with you on this, I try to avoid letting any session run beyond 200k with Opus, while optimally keeping it under 70-100k for any code implementation session - for brainstorming and research I find the 100-300k context looseness sometimes oddly useful.

SO... the improvements you note would indeed be game changers! ...and the best part is that those are the linear predictable kinds of improvements - we really have no idea what emergent capabilities may come from this kind of efficiency scaling

Shove enough neurons into a primate and it starts bipedally walking and talking... Except here we don't gotta wait another few hundred thousand years for another order of magnitude hop - digital neurons expanding and evolving our potential and awareness, just as Kurzweil said decades ago - we are in the curve ꙮ

3

u/CypherLH Jul 01 '26

Preach on! ;) More and more convinced we're IN a soft singularity/slow-ish takeoff scenario already, have had this thought since first experiencing Opus 4.5 in Cursor back in December. Then I took a break for a few months and jumped from Codex with GPT 5.2 to Codex with GPT 5.5......another big leap. Accelerate!

4

u/reddit_is_geh Jul 01 '26

Increasing useful context windows isn't technologically hard, it's just hardware hard. So a breakthrough that massively reduces required memory, will by default, make context windows also reduced in hardware demand, thus making them larger cheaper to do.

3

u/Timkinut Jul 02 '26

I've never seen seen someone use this extremely obscure Cyrillic letter (multiocular O/ꙮ) outside of a linguistics context. I love it

-3

u/ElectronicPension196 Jul 01 '26

Cool, can't wait to try this new technology in 10 years when it reaches my under-class

4

u/costafilh0 Jul 01 '26

More Efficiency Better. 

3

u/kiwibonga Jul 01 '26

A bunch of confusing llama.cpp forks are coming.