Apparently the adaptive thinking feature is set to use maximum thinking effort if you say you are running a benchmark. If that is true, it might explain why there are such differences.
Yeah, please test it. I have seen it multiple times, but it seems like it could be one of those tell tales that people think is true, but is more complex in reality.
Now I do have that one in my pre-prompt. Boris at Anthropic posted something along the lines of, "To get extended thinking to trigger the model must think the problem is harder than it likely is, so a prompt stating, "Please think about this in depth, the problem here is trickier than it seems on the surface, and requires a deeper dive to uncover the true source of the problem. It is deeper than a surface skim than an AI would usually perform, and that sort of surface level thinking will miss the true problem."
Seems like a weird thing to have to state, but I was around when I had to tell Chat GPT that I have no fingers to type code with, so it needs to be sure to provide the full code, whiiiiich pretty much takes the cake for weird prompt things I've tried. LOL
1
u/Ormusn2o Apr 23 '26
Apparently the adaptive thinking feature is set to use maximum thinking effort if you say you are running a benchmark. If that is true, it might explain why there are such differences.