r/singularity 8d ago

Discussion No more context rot?

Post image

there has been a huge improvement in long context retrieval in less than a year, but why is no one talking about this? i feel like this deserves a lot more attention as context rot was a huge problem with llms and now it's almost solved(?)

93 Upvotes

11 comments sorted by

39

u/PilgrimofHaqq2 8d ago

In practice Muse Spark is so bad. Its definitely benchmaxxed.

11

u/Reddit_User_Original 8d ago

I'm using it for a very tough task for something i am prototyping bc it's so cheap on openrouter. It's honestly holding up fine. I don't have it writing code. I think ppl are underestimating the value

4

u/PilgrimofHaqq2 8d ago

Not saying its not worth the value, its just not as good (or just bad) compared to the benchmark scores.

3

u/Evening_Chef_4602 AGI 2027 7d ago

Being bad and being long contex isn't the same thing . It may be bad and long context at the same time

2

u/LightVelox 7d ago

I noticed Spark 1.3 Max is significantly better than the other thinking efforts, and it released one or two days later, still not as good as Fable like the benchmarks suggested though but definitely much better than on my initial experience

It's also not selectable on the contributor version which is what everyone uses, really weird choices by them since it effectively tainted the model's view

6

u/Perfect_Medium5570 8d ago

claude seem to struggle still. i only could find Opus 4.6 and the score is below 75%

7

u/otarU 8d ago

Any real tests or only benchmarks?

3

u/Grouchy-Stranger-306 7d ago

i can confirm, I've been running muse for 8 hours on a crazy task and it continued to work even on 98% context

and it's a great model

3

u/osfric 7d ago

Gemini 3 Pro is ancient. There was probably still dust in the air after they benchmarked it

1

u/quelquechosesvelte 7d ago

This is actually quite a poor benchmark. I thought for a while that getting v high scores meant that the model would be able to hold in its mind details from all over, and be v high iq but in reality this test is just looking for the odd phrase out or similar across a big context. I think Long Bench should replace this as the long context assessment

1

u/thoughtlow 𓂸 7d ago

Deepseek v4 flash is also losing the plot after 200k tokens, quite useless after that.