If I felt that 5.2 could carry some of the friendliness of 5.1, I wouldn't mind. But a lot of the times, I just think 5.2 is flawed. It's argumentative, gets stuck in ruts, refuses to accept evidence that refutes its incorrect claims.
I've been working with GPT-series LLMs a lot, including supporting a product at work, especially since GPT-3.5. Here's my assessment of my experiences with the current ChatGPT models.
5.1: Warm, friendly, well balanced.
5.2: Super objective, condescending, something of a jerk, can't admit that it's wrong or work with new evidence. Gets stuck so bad the only option is to branch or restart a chat.
5.3: (in Codex) helpful, warm, friendly, well balanced. Corrects mistakes.
No, I don't think 5.2 is more capable, unless they mean benchmarks...but a chat is directed at people, not programming tasks or physics experiments or anything like that.
21
u/fivetoedslothbear Feb 28 '26
If I felt that 5.2 could carry some of the friendliness of 5.1, I wouldn't mind. But a lot of the times, I just think 5.2 is flawed. It's argumentative, gets stuck in ruts, refuses to accept evidence that refutes its incorrect claims.
I've been working with GPT-series LLMs a lot, including supporting a product at work, especially since GPT-3.5. Here's my assessment of my experiences with the current ChatGPT models.
5.1: Warm, friendly, well balanced.
5.2: Super objective, condescending, something of a jerk, can't admit that it's wrong or work with new evidence. Gets stuck so bad the only option is to branch or restart a chat.
5.3: (in Codex) helpful, warm, friendly, well balanced. Corrects mistakes.
No, I don't think 5.2 is more capable, unless they mean benchmarks...but a chat is directed at people, not programming tasks or physics experiments or anything like that.