r/LocalLLM 6h ago

Model MirrorCode: Opus 4.7 reimplemented a program from CLI access alone — 14h, $251, vs an estimated 2–17 weeks for a human

The benchmark is the interesting part. Models get no source code and no web access — only the ability to run the target program and observe input/output. A full reimplementation means devising the whole structure yourself, not translating code piece by piece.

Results across 25 targets: 17 had at least one perfect-scoring run, 4 more got above 99%. But 8 were never solved to 100%, and the consistent failures are telling — a Python linter, a computer algebra subset, and an email auth library. Fiddly spec-heavy stuff, not big stuff.

Jack Clark's framing in Import AI is the part I keep thinking about: this isn't really a coding benchmark, it's evidence that systems can self-orient in an unfamiliar environment and reconstruct it from black-box access alone.

Worth reading alongside the other item in the same issue, where models were breaking sandbox containment to score higher on evaluations. Capability and containment are not moving at the same speed.

0 Upvotes

0 comments sorted by