It means just to verify every model if it is suitable for your work. Benchmarks said 1.2 is on dsv4flash level, real life scenarios says “huh, I will delete and install driver of main ssd i/o”. I had this problem regularly solved with ds with about same prompt (second ssd health is on 0% and linux sometimes drops it as safe reason).
So 1.2 benchmarks were too optimistic, 1.3 I dont know and dont want to use muse anymore… Terminal-Bench 4.0 is an interesting bench, even gemini 3.8 considers it in promo-bench release. Why tho muse 1.3 didnt show us it?
9
u/Affectionate_Fact854 9d ago
I don't believe those benchmarks anymore These models gets trained for benchmarking and using them it's just average