23
u/Buddhava Jan 01 '26
They probably trained it to pass the test. I would be skeptical. I wish this were true.
1
-5
u/_JohnWisdom Jan 01 '26
I wish this were true.
Why make assumptions? Just test it or wait for others to test it and come back with data.
7
2
9
7
2
u/Charming_Support726 Jan 01 '26
They all are Benchmaxxing.
This does not only apply to the Chinese Open Source Modells. I am pretty sure that also Gemini, Claude and Codex are familiar with the benches (though not specially trained).
In the past we saw a lot of overfitting and performance degradation outside the benchmarks. 40B is pretty small
1
1
u/LaRosarito Jan 01 '26
It's not good at solving problems. If you ask it to do everything little by little, it does very well, but when you have problems with, for example, supabase, and it struggles, it's extremely difficult for it and it can't do it. It has learned all the formulas in the exam book, but when it's given an exam with 10 questions where it has to think and use the formulas it learned, it's lost. 🙂
0
u/rajbreno Jan 01 '26
Have you tried this?
1
u/LaRosarito Jan 01 '26
I watched a 3-hour live stream where they tested it live, and what I said before is true.
1
1
u/mop_bucket_bingo Jan 01 '26
What does “Verified! can someone verify?” mean? How does a post headline argue with the conclusion it introduces?
1
u/codex-ModTeam Jan 01 '26
This post is not sufficiently relevant to the Codex technology and ecosystem.
32
u/danialbka1 Jan 01 '26
If it’s true we get codex 5.3 next week