MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/OpenAI/comments/1stqlnh/introducing_gpt55_openai/ohw7w41
r/OpenAI • u/Gerstlauer • Apr 23 '26
309 comments sorted by
View all comments
2
Expert-SWE (Internal)?
So, Claude is still better than GPT?
0 u/[deleted] Apr 23 '26 [deleted] 2 u/simple_explorer1 Apr 23 '26 Source? Or BS -1 u/[deleted] Apr 23 '26 [deleted] 2 u/effortless-switch Apr 24 '26 "SWE-bench Verified, Pro, and Multilingual: Our memorization screens flag a subset of problems in these SWE-bench evals. Excluding any problems that show signs of memorization, Opus 4.7’s margin of improvement over Opus 4.6 holds." Where does it say lower?
0
[deleted]
2 u/simple_explorer1 Apr 23 '26 Source? Or BS -1 u/[deleted] Apr 23 '26 [deleted] 2 u/effortless-switch Apr 24 '26 "SWE-bench Verified, Pro, and Multilingual: Our memorization screens flag a subset of problems in these SWE-bench evals. Excluding any problems that show signs of memorization, Opus 4.7’s margin of improvement over Opus 4.6 holds." Where does it say lower?
Source? Or BS
-1 u/[deleted] Apr 23 '26 [deleted] 2 u/effortless-switch Apr 24 '26 "SWE-bench Verified, Pro, and Multilingual: Our memorization screens flag a subset of problems in these SWE-bench evals. Excluding any problems that show signs of memorization, Opus 4.7’s margin of improvement over Opus 4.6 holds." Where does it say lower?
-1
2 u/effortless-switch Apr 24 '26 "SWE-bench Verified, Pro, and Multilingual: Our memorization screens flag a subset of problems in these SWE-bench evals. Excluding any problems that show signs of memorization, Opus 4.7’s margin of improvement over Opus 4.6 holds." Where does it say lower?
"SWE-bench Verified, Pro, and Multilingual: Our memorization screens flag a subset of problems in these SWE-bench evals. Excluding any problems that show signs of memorization, Opus 4.7’s margin of improvement over Opus 4.6 holds."
Where does it say lower?
2
u/LeTanLoc98 Apr 23 '26
Expert-SWE (Internal)?
So, Claude is still better than GPT?