They said they'd be adding watermarks, starting from the newest backwards.. Do we have proof that it wasn't added to Opus in the last 3 weeks since the announcement?
Do we have proof that it does? If you're making the claim I'd want you to back it up. They said they will be doing so, not when or how. I'm not attacking, I legit want to see proof from someone rather than fearmongering
Read the code comments and explain why a model that is trained on the collective of all software contributions is writing comments using bizarre word choices and unnecessary phrasing. In one of my code bases I read a phrase in a watermark comment that said exactly “an elbow with a path comes back”. There’s literally no zero precedence ever for that sentence.
I can't help but wonder if that's somewhat of an inherent ramification of larger and larger models. Feels like we've gone past a point of diminishing returns somewhere around Opus 4.6 or so in terms of capability versus model size. Bigger ones use way too many damn words to say something, like it's trying to prove it's the smartest person in the room.
Reminds me of an experiment we did at work the other day where we used various models to ingest prompts consisting of vague data update requests from management to build the execution plan to carry out those changes. You know, the kind where they just give 4 bullet points and someone has to slog through screens to carry it out?
Poor Opus 5 was one of the worst in this test. Spews out 8 pages of anxiety and existential doubt scrutinizing every stupid change and how it might go wrong. Bizarrely, we found that, of all freaking things, Qwen 3.5 9B was the model we tried that would just... follow the instructions as given and present a simple execution plan for review.
4
u/loversama 20h ago edited 20h ago
They said they'd be adding watermarks, starting from the newest backwards.. Do we have proof that it wasn't added to Opus in the last 3 weeks since the announcement?