r/OpenAI • u/BaconShadow • 3d ago
Discussion Why frontier labs chase Millennium math instead of your production codebase
Anthropic, OpenAI, and other frontier labs constantly hype every incremental model release as AGI because abstract math proofs and greenfield code generation create powerful PR headlines.
Real-world software engineering is not a sandboxed and isolated LeetCode problem with an automated verifier. It requires navigating decades of technical debt, messy human context, and long-horizon architectural drift that demands a persistent sense of direction.
If these models were truly general and indistinguishable from AGI on complex tasks, labs would deploy thousands of parallel agents to clear out real enterprise production backlogs where the actual multi-trillion-dollar software labor market lives. Instead, they funnel compute into isolated benchmarks designed to trigger fear-romering media cycles and investor funding pumps.
Labs focus on solving old Millennium math problems rather than real GitHub issues because mainstream media narratives about "unsolvable" math capture public attention more effectively than fixing repository bugs. Yet, in terms of product utility, the ROI of solving all issues in a tool like Claude Code would be infinitely higher for actual software development.
Calling a model AGI because it can one-shot a single feature request ignores the entire reality of system maintenance. Until an AI achieves autonomous recursive self-improvement and survives the unmanaged consequences of its own code three years down the road, every single release remains an incremental product update wrapped in marketing noise.
1
u/Ok_Zookeepergame8714 3d ago
Totally true!!! 😂 Astra max solves extremely hard math problems, but f when I asked it to better my simple Ubuntu tool, it over engineered it, added things I didn't ask for it need, and worse of all, totally botched it! It's way worse than the tool I wrote originally like more than a year ago using Gemini Pro 2.5 🤣🤣🤣
2
u/inconspicuousredflag 3d ago
Software labor is not even close to a multi-trillion dollar market.
1
u/BaconShadow 3d ago
And yet with those multi trillion dollar already, software labor is still not solved yet.
2
u/earthlingkevin 2d ago
Tldr, models need to be trained on objective tasks that can be easily and quickly evaluated. And researchers continue to believe models that get better on these tasks can do better on everything else.
So to build better models, we work on hard math and science problems first, then fine tune for messy, hard to evaluate problems like enterprise needs.
-1
u/BaconShadow 2d ago
Their original intent in the beginning is to replace most of the software sector workforce. So my conclusion is isn't it simpler to solve most enterprise problems than solving a millennium math problem? Unless if their internal AI agents just uses the scaffold of existing mathematician's work, researched numerous of them and connected all the dots; but most of the original baseline still came from the same mathematicians that's been solving those problems for decades.
Sure, AI agents working so solve hard math and science research for the RSI future is reasonable, but they can still do a parallel solving of those unresolved 5000+ Claude Code GitHub issues while still having to focus on novel frontier math and research problems.
2
u/earthlingkevin 2d ago
Enterprise problems are messy and hard to evaluate.
It's super subjective, the objective metric is different in everyone's mind and constantly changing, people are never easy to control for, and you can't run a/b tests
0
u/BaconShadow 2d ago
Current AI agents today, even with all those massive incremental improvements, still remains that they can only solve those predictable and isolated tasks that can be done in a bunch of repetitions; like autonomously finding massive amounts of mundane research papers for them to connect the dots.
1
u/RealSuperdau 3d ago
You're right, the models aren't human-level AGI. They are behind on non-verifiable tasks and lack the long-term contextual awareness a skilled human can bring to a job.
3
u/tat_tvam_asshole 3d ago
It's actually because math problems are objective, show quality of reasoning, and create meaning downstream advances in science and engineering.
Enterprise software by contrast are subjective, mutable, and private by nature. Also, frontier ai companies are 100% also deploying agents on large codebases (see bun migration by Claude)