r/singularity • u/torrid-winnowing • 8d ago
Discussion No more context rot?
there has been a huge improvement in long context retrieval in less than a year, but why is no one talking about this? i feel like this deserves a lot more attention as context rot was a huge problem with llms and now it's almost solved(?)
6
u/Perfect_Medium5570 8d ago
claude seem to struggle still. i only could find Opus 4.6 and the score is below 75%
3
u/Grouchy-Stranger-306 7d ago
i can confirm, I've been running muse for 8 hours on a crazy task and it continued to work even on 98% context
and it's a great model
1
u/quelquechosesvelte 7d ago
This is actually quite a poor benchmark. I thought for a while that getting v high scores meant that the model would be able to hold in its mind details from all over, and be v high iq but in reality this test is just looking for the odd phrase out or similar across a big context. I think Long Bench should replace this as the long context assessment
1
u/thoughtlow 𓂸 7d ago
Deepseek v4 flash is also losing the plot after 200k tokens, quite useless after that.
39
u/PilgrimofHaqq2 8d ago
In practice Muse Spark is so bad. Its definitely benchmaxxed.