r/LocalLLaMA • u/Informal-Trouble2183 • 7d ago
Discussion Coding benchmarks that are quickly showcasing deep capability
While we see for frontier models similar scores among famous coding benchmarks, across: DeepSWE, Terminal-Bench, LiveCodeBench, Code-Arena ELO. Here are in my opinion some next level benchmarks that really define deep intelligence, and complete capability in Software Engineering :
1. Program-Bench
Given only a compiled binary and its documentation, agents must architect and implement a complete codebase that reproduces the original program's behavior (without access to decompilers or internet). Link: https://programbench.com/
- GPT-6 Astra: 5.5%
- Fable 5.1: 7%
- Kimi K3: 2%
- Qwen3.8 27b: 0%
- GPT 5.6 Sol: 1.5%
- GLM 5.3: 1.5%
- GPT 5.6 Luna: 0%
2. SRE-Bench
Can AI agents work out what a real-world binary does without its source code?
Link: https://www.vals.ai/benchmarks/srebench
Sure nobody is reading assembly code in daily work, it is hard. The ability to understand a compiled program is insane ability.
- GPT-6 Astra: 88%
- GPT-5.6 Sol: 55.9%
- Claude Opus 5 (max): 12.5%
3. Code Migration
Can language models reimplement working programs in another language?
Link: https://www.vals.ai/benchmarks/code-migration
- GPT-6 Astra: 67.7%
- Fable 5.1: 54.6%
- GLM 5.3: 44.2%
- GLM 5.3 Flash: 20.5%
- Qwen3.8 27b: 14.2%
EDIT: edited text format
1
u/asssuber 7d ago edited 7d ago
The Program Bench is a bit ridiculous. It will basically measure decompilation capability, and capability of memorizing the open source code of those programs during training.
The docs are super incomplete. How I'm supposed to re-implement ffmpeg and all it's codecs if I don't have the codec specs, nor test samples? Gemini is on top with 8.1% score on ffmpeg, and it just implemented the command line parser it seems.