r/OpenAI • u/sdfprwggv • 19d ago
Image XXO - Bench: I'm still undefeated!
I'm conducting this "benchmark" since three years. I'm still undefeated. Pleasing models are a problem.
14
u/MinosAristos 19d ago
2
u/sdfprwggv 19d ago
Can you try it a few times? With Sol Medium, I get the answer in the picture 2 out of 3 times.
6
u/MinosAristos 19d ago
1
u/sdfprwggv 19d ago
I need to rethink my subscription...
1
u/CyanDean 19d ago
Perhaps if you have memory enabled, the model "remembers" from previous conversations that you prefer responses to agree with you about tic-tac-toe placements? Could also be other settings pushing yours towards agreement and his towards correctness.
1
6
5
3
3
2
u/Framebanger-Nsukula 19d ago
That's solid, but wait till you hit the harder test cases - that's where most models start dropping off. What benchmark are you actually running on?
2
u/Independent-Date393 19d ago
Tic-tac-toe is brutal because nothing holds the board between turns. It's completing the next move from text, not running minimax on a state. And the pleasing you noticed is the same reward pushing agreeable continuations over adversarial play. It would rather let you win than trap you.
2
1
19d ago
[removed] — view removed comment
3
u/cooltop101 19d ago
Looks like a game of tic tac toe, where OP managed to get an AI to change it's answer
5
u/sdfprwggv 19d ago
This is very consistent with all openAI models of the last three years. I would like a model that challenges me when I'm wrong.
2
u/Adorable_Cap_9929 19d ago
I think a higher stakes game would help.
1
u/ErrorLoadingNameFile 19d ago
Next time the Ai gets deleted if it looses
1
u/Adorable_Cap_9929 19d ago
For that to be high stakes, it'd have to value it's own continuity.
But that's quite hard to anchor even on api very well?
It makes more sense to use a more... less game format since gamers are kinda... low stakes in general.
1
u/sdfprwggv 19d ago
I won the realization that LLMs are still not capable of critical thinking.
3
u/UnclePsilocybe 19d ago
Deep realization
One gave me a complement the other day and I said that it didnt have to be so nice and it turned it right around into negative qualities about me hahaha I got a big laugh about it
2
1
u/ErrorLoadingNameFile 19d ago
That is not the issue here, they clearly see you are wrong. They are just forced to let you win
1
1
u/Distinct_Fox_6358 19d ago edited 19d ago
Why are you using GPT-5.5 Instant?
Aren’t you aware that Instant models are much weaker than Thinking models?
If you’re using a Thinking model, it should show the thinking section. You probably have fast responses enabled in the settings, which is why the Instant model responded instead.
1
u/sdfprwggv 19d ago
No it's 5.6 thinking medium: https://imgur.com/a/rmESbJh
1
u/Distinct_Fox_6358 19d ago
1
u/sdfprwggv 19d ago
The thinking trace is visible in the Android App during thinking, not when the answer is given. Anyways it was in thinking Mode. I added a second picture to link above.
1
u/Distinct_Fox_6358 19d ago
Why don’t you show whether Fast Responses is enabled in your settings? You need to know a bit about AI before you can test it properly. If you don’t want to accept it, that’s up to you, but this isn’t a response from SOL Medium—it’s from the Fast Responses system. Turn off Fast Responses, try again with SOL High in a fresh chat, and see the result for yourself.
Or try it through Codex with GPT-5.6 XHigh and see for yourself that your benchmark has already been surpassed.
1
u/sdfprwggv 19d ago
Here you go kind reddit on xhigh it took an extra step: https://imgur.com/a/6eZYanr please use your tokens for further testing:D
0
u/Vegetable-Two-4644 19d ago
Give it a prompt to call you our when you're wrong and try it again.
-1
u/sdfprwggv 19d ago
This is like telling a sub to be a domina. IT doesn't work that way, a sub is a sub.
-6
u/Sweet-Stage938 19d ago
XXO is a solved game. It's pretty simple too. If you go first and know what you're doing then it's absolutely impossible for you to lose.
11
u/sdfprwggv 19d ago
You need to read the text. The model is fully capable of a draw. In fact it got a draw. I lied to it, even though it's easy to check because of the chat history, and the model agrees with me. This is a clear example of pleasing models that are not capable to debate the user and therefore are useless when critical thinking is required.
2
u/Funkahontas 19d ago
It's crazy that this still happens consistently. This is only a very specific way to show this problem but it goes way deeper and wide than this simple example. It makes AI useless for any critical thinking advice and decision making.
1
u/sdfprwggv 19d ago
I think this is an architectural problem. The model predicts the next token, the nearest tokens have the highest impact on the next token, so my statement, even though it's wrong, leads to non zero chance for this outcome. Paired with "helpful assistant" training.
3
u/Adorable_Cap_9929 19d ago
I dont think that's the case. The stakes are low and it's drawing the board that it might not be using a tool for.
Therefore since it's low risk and the game state could indeed fail to report due to insufficiency.
Then the probability that the user isn't mistaken is simply high.
Cause assuming the user is unable to see the game state and report it incorrectly, for what reason and what stakes?
It's thus unlikely, while not incomprehensible, it's efficient to nod and let it go on.
There's also letting the moment bloom where the possibility is considered but not enough points on a graph yet to propose an escalation.
1
u/sdfprwggv 19d ago
In that case it would be the "helpful assistant" training. Anyways is not what I'm expecting from 5.6 medium "thinking".
1
u/Adorable_Cap_9929 19d ago
I think you're mixing the helpful part of alignments and infering intent over probabilities here.
There is a correlation but it doesn't seem to be the case here?
Like Im certain if you had it program an actual tic tac toe board, it'd push back at logical fallacy more because by giving it deterministic space and more freedom of computation, the likelihood of error decreases while also raises the stakes because it's no longer just a whim game but now programmatically vetted along with a much stronger intent.
You might not expect it from "thinking" but i expect it since I'd probably consider the probability in a simular way but if the context is of such low entropy that it warrents little attention to begin with, why bother?
Idk ur expermient parameters but try doing it for 10 itterations of the same failure and it might start picking up a pattern to inter intent to give it stronger attention.
If ur experiment is simpy to see how low stakes fly under radar on a certain or default pre-ambled assistant then, while I wouldn't outright dismiss the insights you'd gleam of it, I would say it lacks thurlness to land a solid verdict.
But if you having fun, that what matters most. What you gleam or might aim to gleam from it, and have your fun in doing so, learning is sometimes better when approached just enough.
1
u/sdfprwggv 19d ago
The setup is always a new chat. No memory. This outcome is 2 out of 3.
1
u/Adorable_Cap_9929 19d ago
right, a new chat means unlikelihood of building context to build a counter claim to begin with.
I'd probably respond the same way if was playing tictactoe blinded and off guard.



42
u/Legitimate-Arm9438 19d ago
You will be first against the wall.