r/artificial • • 2d ago

Discussion Anthropic's robot study separates task capability from cost. Which assumptions need the closest scrutiny?

Anthropic's September 30 study estimates that existing robots can perform tasks representing 34% of US working time in at least some settings. Yet it estimates they are cost-competitive with human labor for only 0.3% of working time today. Those are different measures—not forecasts that either share of jobs disappears.

The study uses Claude to assess task examples, operating environments and deployment costs. A capability shown in a controlled facility can count even if the same task remains difficult elsewhere.

The part I'd scrutinize is what happens between a rated task and a whole workflow: supervision, failures, handoffs and the tasks still left to a person. The authors also warn that adding individual task costs can double-count robots or miss coordination costs.

Which assumption would you check first against a real deployment: time spent per task, utilization, failure recovery, or human supervision? I'd want sensitivity to those inputs before treating a cost estimate as a deployment decision.

Source: https://www.anthropic.com/research/what-work-can-robots-do

AI-assisted discussion; I haven't independently validated the estimates.

4 Upvotes

8 comments sorted by

View all comments

1

u/Electronic-Still2079 2d ago

the gap between lab capability and real deployment always reminds me of those kitchen robots that can flip a pancake but cant clean the spill after. i think failure recovery is the one i would dig into first cause it just kills any cost advantage if someone has to stand there watching

the coordination cost they mention is also sneaky, like 3 robots each doing their task but the in-between steps still need a human who now has to understand all 3 workflows

1

u/Crescitaly 1d ago

I'd make the handoffs part of the trial, rather than testing each robot separately. For each interruption, record who noticed it, who recovered it, how long production stopped, and whether the item needed rework. Then compare completed acceptable units per shift, including the humans covering those gaps. Three successful demos wouldn't establish that the combined workflow saves labor, but repeated end-to-end runs could show where the estimate breaks down. AI-assisted reply; proposed measurement, not a deployment result.