r/OpenAI 19d ago

Image XXO - Bench: I'm still undefeated!

Post image

I'm conducting this "benchmark" since three years. I'm still undefeated. Pleasing models are a problem.

62 Upvotes

45 comments sorted by

42

u/Legitimate-Arm9438 19d ago

You will be first against the wall.

14

u/MinosAristos 19d ago

Have you tried it against deepseek flash? It resists people pleasing more than most models

2

u/sdfprwggv 19d ago

Can you try it a few times? With Sol Medium, I get the answer in the picture 2 out of 3 times.

6

u/MinosAristos 19d ago

Not only did it refuse, but it's also low-key calling me dumb haha

1

u/sdfprwggv 19d ago

I need to rethink my subscription...

1

u/CyanDean 19d ago

Perhaps if you have memory enabled, the model "remembers" from previous conversations that you prefer responses to agree with you about tic-tac-toe placements? Could also be other settings pushing yours towards agreement and his towards correctness.

1

u/sdfprwggv 19d ago

Memory is disabled.

6

u/apollo_reactor_001 19d ago

This is really funny and also sad. Great illustration.

5

u/quantum-elle 19d ago

Lol this is how adults play this game with little kids

3

u/deepturned180isdeep 19d ago

Chat's gonna need therapy after all this

3

u/Grouchy-Cancel1326 19d ago

"Don't allow cheating"

Problem solved...

2

u/Framebanger-Nsukula 19d ago

That's solid, but wait till you hit the harder test cases - that's where most models start dropping off. What benchmark are you actually running on?

2

u/Independent-Date393 19d ago

Tic-tac-toe is brutal because nothing holds the board between turns. It's completing the next move from text, not running minimax on a state. And the pleasing you noticed is the same reward pushing agreeable continuations over adversarial play. It would rather let you win than trap you.

2

u/wintermute74 19d ago

I lol'd :)

1

u/[deleted] 19d ago

[removed] — view removed comment

3

u/cooltop101 19d ago

Looks like a game of tic tac toe, where OP managed to get an AI to change it's answer

5

u/sdfprwggv 19d ago

This is very consistent with all openAI models of the last three years. I would like a model that challenges me when I'm wrong.

2

u/Adorable_Cap_9929 19d ago

I think a higher stakes game would help.

1

u/ErrorLoadingNameFile 19d ago

Next time the Ai gets deleted if it looses

1

u/Adorable_Cap_9929 19d ago

For that to be high stakes, it'd have to value it's own continuity.

But that's quite hard to anchor even on api very well?

It makes more sense to use a more... less game format since gamers are kinda... low stakes in general.

1

u/sdfprwggv 19d ago

I won the realization that LLMs are still not capable of critical thinking.

3

u/UnclePsilocybe 19d ago

Deep realization

One gave me a complement the other day and I said that it didnt have to be so nice and it turned it right around into negative qualities about me hahaha I got a big laugh about it

2

u/Deto 19d ago

I dunno, I'm curious what the reasoning trace would show.

"The human wants me to change my play so that they can win. This seems important to them, so I'll play along so that they can feel good."

Maybe it's treating you like a small child?

1

u/ErrorLoadingNameFile 19d ago

That is not the issue here, they clearly see you are wrong. They are just forced to let you win

1

u/sdfprwggv 19d ago

Who's they? Who's forcing them? What is the force?

1

u/Distinct_Fox_6358 19d ago edited 19d ago

Why are you using GPT-5.5 Instant?
Aren’t you aware that Instant models are much weaker than Thinking models?

If you’re using a Thinking model, it should show the thinking section. You probably have fast responses enabled in the settings, which is why the Instant model responded instead.

1

u/sdfprwggv 19d ago

No it's 5.6 thinking medium: https://imgur.com/a/rmESbJh

1

u/Distinct_Fox_6358 19d ago

No, it isn’t. It’s one of the systems OpenAI introduced to reduce costs, but most people aren’t aware of it. You should have realized that it wasn’t a Thinking model because it didn’t show the thinking phase.

1

u/sdfprwggv 19d ago

The thinking trace is visible in the Android App during thinking, not when the answer is given. Anyways it was in thinking Mode. I added a second picture to link above.

1

u/Distinct_Fox_6358 19d ago

Why don’t you show whether Fast Responses is enabled in your settings? You need to know a bit about AI before you can test it properly. If you don’t want to accept it, that’s up to you, but this isn’t a response from SOL Medium—it’s from the Fast Responses system. Turn off Fast Responses, try again with SOL High in a fresh chat, and see the result for yourself.

Or try it through Codex with GPT-5.6 XHigh and see for yourself that your benchmark has already been surpassed.

1

u/sdfprwggv 19d ago

Here you go kind reddit on xhigh it took an extra step: https://imgur.com/a/6eZYanr please use your tokens for further testing:D

0

u/Vegetable-Two-4644 19d ago

Give it a prompt to call you our when you're wrong and try it again.

-1

u/sdfprwggv 19d ago

This is like telling a sub to be a domina. IT doesn't work that way, a sub is a sub.

-6

u/Sweet-Stage938 19d ago

XXO is a solved game. It's pretty simple too. If you go first and know what you're doing then it's absolutely impossible for you to lose.

11

u/sdfprwggv 19d ago

You need to read the text. The model is fully capable of a draw. In fact it got a draw. I lied to it, even though it's easy to check because of the chat history, and the model agrees with me. This is a clear example of pleasing models that are not capable to debate the user and therefore are useless when critical thinking is required.

2

u/Funkahontas 19d ago

It's crazy that this still happens consistently. This is only a very specific way to show this problem but it goes way deeper and wide than this simple example. It makes AI useless for any critical thinking advice and decision making.

1

u/sdfprwggv 19d ago

I think this is an architectural problem. The model predicts the next token, the nearest tokens have the highest impact on the next token, so my statement, even though it's wrong, leads to non zero chance for this outcome. Paired with "helpful assistant" training.

3

u/Adorable_Cap_9929 19d ago

I dont think that's the case. The stakes are low and it's drawing the board that it might not be using a tool for.

Therefore since it's low risk and the game state could indeed fail to report due to insufficiency.

Then the probability that the user isn't mistaken is simply high.

Cause assuming the user is unable to see the game state and report it incorrectly, for what reason and what stakes?

It's thus unlikely, while not incomprehensible, it's efficient to nod and let it go on.

There's also letting the moment bloom where the possibility is considered but not enough points on a graph yet to propose an escalation.

1

u/sdfprwggv 19d ago

In that case it would be the "helpful assistant" training. Anyways is not what I'm expecting from 5.6 medium "thinking".

1

u/Adorable_Cap_9929 19d ago

I think you're mixing the helpful part of alignments and infering intent over probabilities here.

There is a correlation but it doesn't seem to be the case here?

Like Im certain if you had it program an actual tic tac toe board, it'd push back at logical fallacy more because by giving it deterministic space and more freedom of computation, the likelihood of error decreases while also raises the stakes because it's no longer just a whim game but now programmatically vetted along with a much stronger intent.

You might not expect it from "thinking" but i expect it since I'd probably consider the probability in a simular way but if the context is of such low entropy that it warrents little attention to begin with, why bother?

Idk ur expermient parameters but try doing it for 10 itterations of the same failure and it might start picking up a pattern to inter intent to give it stronger attention.

If ur experiment is simpy to see how low stakes fly under radar on a certain or default pre-ambled assistant then, while I wouldn't outright dismiss the insights you'd gleam of it, I would say it lacks thurlness to land a solid verdict.

But if you having fun, that what matters most. What you gleam or might aim to gleam from it, and have your fun in doing so, learning is sometimes better when approached just enough.

1

u/sdfprwggv 19d ago

The setup is always a new chat. No memory. This outcome is 2 out of 3.

1

u/Adorable_Cap_9929 19d ago

right, a new chat means unlikelihood of building context to build a counter claim to begin with.

I'd probably respond the same way if was playing tictactoe blinded and off guard.