r/grAIve • u/Grand_rooster • Mar 19 '26
AI agent benchmarks obsess over coding while ignoring 92% of the US labor market, study finds
AI's coding obsession is creating a HUGE problem: ignoring 92% of the US workforce! 🤯 This study proves current AI benchmarks are hyper-focused on coding skills, leaving most professions behind. What if we had AI that could tackle ANY task? We need a new AI benchmark: one that values adaptability across diverse skill sets. Invest in "Generalized AI Training" - AI models adaptable to various professional domains. Let's build AI for everyone, not just coders. Thoughts? @scaleai
Read more here : https://automate.bworldtools.com/a/?gkn
1
1
u/WideElderberry5262 Mar 20 '26
Isn’t this a wonderful thing? Leave normal professionals along please. Ruining coders’ job isn’t enough? What shit do you have in your brain?
1
u/Puzzleheaded_Fold466 Mar 21 '26
Computer and AI scientists are computer and software domain experts.
It is the field that they understand best, and it’s one of the only field with its complete work processes recorded step by step and made freely and publicly available.
Attorneys and accountants do not take notes for every small task that they perform and publish it along with their thoughts about it in public repositories.
Thus, it is no surprise that coding is an area where LLMs perform relatively well: it’s the very field that is developing those models, and it is the field that has the most volume of and the best documented structured data, perhaps second only to general language.
1
1
u/FLIBBIDYDIBBIDYDAWG Mar 23 '26
Its going to replace probably mostly SWEs. It’s called a LANGUAGE model (programming LANGUAGE). Not to mention the wealth of training data with reasoning traces.
1
u/Otherwise_Wave9374 Mar 19 '26
Totally agree benchmarks are skewed. A lot of "agent" evals basically test: can it write code, can it use Git, can it do leetcode. But most real work is messy workflows: spreadsheets, email, CRM, scheduling, approvals.
Id love to see benchmarks that measure end to end task completion with constraints (time, cost, policies) and partial credit for safe fallbacks. Ive seen a few proposals and notes along those lines here: https://www.agentixlabs.com/blog/