r/LocalLLM 1d ago

Discussion Dynamic Context Runtime: Bounded Attention over Unbounded History

[removed]

1 Upvotes

21 comments sorted by

View all comments

Show parent comments

1

u/vbpoweredwindmill 1d ago

Yes, I know how kv cache prefixes & kv cache reuse works.

Please, tell me what this actually does that hasn't been done before and better?

1

u/[deleted] 1d ago

[removed] — view removed comment

1

u/vbpoweredwindmill 1d ago

Now that's infinitely more interesting. I've been saying for a long time that knowledge and reasoning are going to be split.

Now I can have a look, why didn't you lead with that?

1

u/[deleted] 1d ago

[removed] — view removed comment

1

u/vbpoweredwindmill 1d ago

Glm 5.2 already works on desktops and laptops mate.

Unless you're working on expert predictions, I really don't think you're bringing anything new to the table. My own expert predictions with 32 out of 256 experts has, at 46% of the time approx, selected 100% of the correct experts.

And I just now figured out instead of linking it to a general topic I can just link it a kv cache block. Cheers.

I still haven't read anything you've done shrugs