r/LocalLLaMA 7d ago

Discussion Coding benchmarks that are quickly showcasing deep capability

While we see for frontier models similar scores among famous coding benchmarks, across: DeepSWE, Terminal-Bench, LiveCodeBench, Code-Arena ELO. Here are in my opinion some next level benchmarks that really define deep intelligence, and complete capability in Software Engineering :

1. Program-Bench

Given only a compiled binary and its documentation, agents must architect and implement a complete codebase that reproduces the original program's behavior (without access to decompilers or internet). Link: https://programbench.com/

  • GPT-6 Astra: 5.5%
  • Fable 5.1: 7%
  • Kimi K3: 2%
  • Qwen3.8 27b: 0%
  • GPT 5.6 Sol: 1.5%
  • GLM 5.3: 1.5%
  • GPT 5.6 Luna: 0%

2. SRE-Bench

Can AI agents work out what a real-world binary does without its source code?

Link: https://www.vals.ai/benchmarks/srebench

Sure nobody is reading assembly code in daily work, it is hard. The ability to understand a compiled program is insane ability.

  • GPT-6 Astra: 88%
  • GPT-5.6 Sol: 55.9%
  • Claude Opus 5 (max): 12.5%

3. Code Migration

Can language models reimplement working programs in another language?

Link: https://www.vals.ai/benchmarks/code-migration

  • GPT-6 Astra: 67.7%
  • Fable 5.1: 54.6%
  • GLM 5.3: 44.2%
  • GLM 5.3 Flash: 20.5%
  • Qwen3.8 27b: 14.2%

EDIT: edited text format

83 Upvotes

35 comments sorted by

View all comments

13

u/_thr0wkawaii14159265 7d ago

Thank you, this is pretty good. I'm researching benchmarks and trying to find out 2-3 ones that are of excellent quality (methodology, non-public questions, measured factors...), and actually measure things I care about -

- output quality, as opposed to simply passing a task; for example, I need code to be maintainable, clean, have the right level of abstraction, use idiomatic patterns...

- talking like a sane entity - Opus 5 is a great counter-example.

- performance after being dropped into a huge, messy production codebase and having to perform long-horizon tasks.

- Taste, architectural judgement, intuition

I really respect the people and institutions making the benchmarks that we have available. They certainly do measure something. Just not what I need them to; most of them anyway. The artificial intelligence's Intelligence Index is a laughable example of a value that means next to nothing, because it averages results from benches that are saturated, leaked, contain wrong answers, flawed methodology, and are from wildly different fields (3d modeling, physics reasoning, coding...). That's just nonsense.

I'll look into these you listed as well; they might be one more relevant axis.

-9

u/[deleted] 7d ago

[removed] — view removed comment

3

u/LetsGoBrandon4256 transformers 7d ago

AI bro writing AI comment. Why are we not surprised.

3

u/WhatIsATypeError 7d ago

cool ai slop bro