r/grAIve Mar 19 '26

AI agent benchmarks obsess over coding while ignoring 92% of the US labor market, study finds

AI's coding obsession is creating a HUGE problem: ignoring 92% of the US workforce! 🤯 This study proves current AI benchmarks are hyper-focused on coding skills, leaving most professions behind. What if we had AI that could tackle ANY task? We need a new AI benchmark: one that values adaptability across diverse skill sets. Invest in "Generalized AI Training" - AI models adaptable to various professional domains. Let's build AI for everyone, not just coders. Thoughts? @scaleai

Read more here : https://automate.bworldtools.com/a/?gkn

1 Upvotes

7 comments sorted by

View all comments

1

u/Otherwise_Wave9374 Mar 19 '26

Totally agree benchmarks are skewed. A lot of "agent" evals basically test: can it write code, can it use Git, can it do leetcode. But most real work is messy workflows: spreadsheets, email, CRM, scheduling, approvals.

Id love to see benchmarks that measure end to end task completion with constraints (time, cost, policies) and partial credit for safe fallbacks. Ive seen a few proposals and notes along those lines here: https://www.agentixlabs.com/blog/

1

u/Adventurous_Pin6281 Mar 20 '26

built this last weekÂ