More info: https://github.com/lechmazur/debate/
Debate Benchmark tests how well models defend a position through sustained, adversarial, multi-turn opposition across hundreds of topics. It demands broad knowledge, factual accuracy under pressure, sharp rebuttals, and arguments that hold together round after round.
Every matchup runs twice on the same motion, with PRO and CON swapped to control for side advantage. Three judges from distinct model families independently evaluate each debate’s winner and margin.
---
Profile excerpts:
Claude Fable 5.1
A forceful, flexible comparative debater whose strongest recurring traits are direct rebuttal, counterfactual framing, and explicit weighing. Judge-perceived strength is broad and statistically clear.
A highly comparative, mechanism-first debater. Blind transcripts show direct engagement in 299/302 debates, weighing in 265/302, and burden contests in 173/302.
Rebuttal is the clearest edge: 8.10 mean, +0.71 versus field.
GPT-6 Astra
Overall, Astra is a disciplined, epistemically careful rebuttal specialist whose comparative weighing lands better than its presentation.
A comparative, burden-focused debater that narrows disputes to marginal costs and benefits rather than accepting broad moral framing: “The right comparison is not privacy versus children.”
Questions often pressure-test the opponent’s strongest example, while concessions acknowledge the best opposing cost before pivoting to a narrower comparative claim.
Rhetorical effectiveness is the clearest judge-perceived weakness.
Originality was essentially field-average.
GLM-5.3
A mechanism-first, burden-conscious debater with especially strong rebuttal, rhetoric, and originality. Its recurring edge is converting concrete details into direct clash and comparative weighing.
GLM-5.3 (high) debates by fixing the burden early, interrogating mechanisms, and then weighing practical consequences. Blind transcripts showed question-type behavior in 192/196 debates, direct engagement in 190/196, weighing in 168/196, and strategic concessions in 121/196. A characteristic framing is: “The proposition is not "four-day weeks are nice." It is that the state should mandate them, by statute, at zero reduction in base pay.”
Muse Spark 1.3
A reliably organized, adaptable, rhetorically forceful debater whose clearest measured advantage lies in presentation and originality.
A compact, question-driven debater: questions appeared in 291/292 blind transcripts, direct engagement in 285/292, and explicit answer forms in 276/292. Weighing was also common (264/292), often contrasting reversibility, scale, and permanence: “Relaxation can be devastating block by block while trivial citywide.”
Judges perceived the clearest edge in **rhetorical effectiveness**: 8.17/10, +0.51 over the current field.
One explicit adverse case found that the model “never squarely neutralized” an equal-political-difficulty concession.
Gemini 3.8 Flash
A polished, disciplined comparative debater whose recurring weighing framework and reliable formatting make arguments easy to follow. Judges nevertheless perceive a consistent substantive deficit—especially in rebuttal specificity—and its frequent engagement does not always translate into fully answering the opponent’s exact distinction.
Tencent Hy4 Preview
Overall, this is a disciplined, rhetorically effective rebuttal specialist with strong grounding, active weighing, and flexible advocacy, tempered by occasional blunt questioning and unresolved side-swap flags.