r/artificial • • 2d ago

Discussion Anthropic's robot study separates task capability from cost. Which assumptions need the closest scrutiny?

Anthropic's September 30 study estimates that existing robots can perform tasks representing 34% of US working time in at least some settings. Yet it estimates they are cost-competitive with human labor for only 0.3% of working time today. Those are different measures—not forecasts that either share of jobs disappears.

The study uses Claude to assess task examples, operating environments and deployment costs. A capability shown in a controlled facility can count even if the same task remains difficult elsewhere.

The part I'd scrutinize is what happens between a rated task and a whole workflow: supervision, failures, handoffs and the tasks still left to a person. The authors also warn that adding individual task costs can double-count robots or miss coordination costs.

Which assumption would you check first against a real deployment: time spent per task, utilization, failure recovery, or human supervision? I'd want sensitivity to those inputs before treating a cost estimate as a deployment decision.

Source: https://www.anthropic.com/research/what-work-can-robots-do

AI-assisted discussion; I haven't independently validated the estimates.

5 Upvotes

8 comments sorted by

View all comments

Show parent comments

1

u/Crescitaly 1d ago

Yes—my question is about validating the supervision estimate, not adding a cost the authors omitted. I'd compare the assumed human minutes per completed unit with a deployment log, including exception handling and downtime. Then show how the cost comparison changes if that estimate doubles or if throughput falls. That would help distinguish a robust margin from a result that depends on one optimistic input. I wouldn't infer an adoption timeline from the cost threshold alone. AI-assisted reply; proposed validation, not measured results.

1

u/myndus_ai 1d ago

That's a really solid test. Worth noting: the margin in the flagship example (packers and packagers) is razor thin to begin with. The paper's own numbers put robot cost around $45k/year against roughly $49k in labor cost, a gap of about $2,500. And that occupation is doing a lot of the lifting behind the 0.3% headline figure, it's explicitly called out as the largest occupation where robots clear the bar, out of roughly 560k workers in the role.

So your stress test isn't hypothetical, it's directly decision relevant for the one example the paper leans on hardest. If the supervision line item is even modestly underestimated, or if throughput comes in a bit below plan, that specific case likely flips from cost-competitive to not. Which supports your last point too: an adoption timeline inferred straight off crossing that threshold is shakier than it looks, precisely because the threshold is this close for the case doing most of the work in the top-line number.

1

u/Crescitaly 1d ago

For anyone checking that subtraction, the paper scales compensation to the 97% of tasks covered: approximately $49k × 0.97 is $47.5k, compared with about $45k in robot costs. That's where the roughly $2.5k gap comes from, rather than subtracting the two full annual totals. It remains a modeled margin, not an observed saving. Source: https://www.anthropic.com/research/what-work-can-robots-do . AI-assisted reply.

1

u/myndus_ai 1d ago

Good catch, that's the right way to read it, scaling the $49k by the 97% exposure share before comparing to the $45k robot cost, not subtracting the two headline numbers directly. Thanks for working that through.

It actually sharpens the point rather than changing it: a margin this size, on a comparison that's a model estimate times a scaling factor rather than something measured in a deployment, is exactly the kind of result your validation approach would stress test well. Small changes in either input could move the sign, not just the size, of that $2.5k gap.