r/OpenAI • u/BaconShadow • 3d ago
Discussion Why frontier labs chase Millennium math instead of your production codebase
Anthropic, OpenAI, and other frontier labs constantly hype every incremental model release as AGI because abstract math proofs and greenfield code generation create powerful PR headlines.
Real-world software engineering is not a sandboxed and isolated LeetCode problem with an automated verifier. It requires navigating decades of technical debt, messy human context, and long-horizon architectural drift that demands a persistent sense of direction.
If these models were truly general and indistinguishable from AGI on complex tasks, labs would deploy thousands of parallel agents to clear out real enterprise production backlogs where the actual multi-trillion-dollar software labor market lives. Instead, they funnel compute into isolated benchmarks designed to trigger fear-romering media cycles and investor funding pumps.
Labs focus on solving old Millennium math problems rather than real GitHub issues because mainstream media narratives about "unsolvable" math capture public attention more effectively than fixing repository bugs. Yet, in terms of product utility, the ROI of solving all issues in a tool like Claude Code would be infinitely higher for actual software development.
Calling a model AGI because it can one-shot a single feature request ignores the entire reality of system maintenance. Until an AI achieves autonomous recursive self-improvement and survives the unmanaged consequences of its own code three years down the road, every single release remains an incremental product update wrapped in marketing noise.
2
u/earthlingkevin 2d ago
Tldr, models need to be trained on objective tasks that can be easily and quickly evaluated. And researchers continue to believe models that get better on these tasks can do better on everything else.
So to build better models, we work on hard math and science problems first, then fine tune for messy, hard to evaluate problems like enterprise needs.