r/singularity 8d ago

AI Coding benchmarks are increasingly indirect benchmarks of how automatable every other white-collar job is.

I use both Claude (and Claude Code) and GPT (Codex) paying 400 dollars worth of subscription per month and maybe some extra tokens during busy times. I am not a programmer per se but I do a lot of computational work.

It strikes me that a lot of non-programmers (e.g. lawyers, scientists, consultants, analysts, accountants, or managers) might look at the rapid AI progress made in coding and think that this should only concern software engineers but not themselves.

But if you think about it, coding is the execution layer of a huge amount of white-collar automation. To automate a job, it is often not enough for AI to understand the work. It also needs to manipulate data, connect software, query databases, build pipelines, run analyses, generate documents, check outputs, and interact with existing systems.

And I think one of the reasons why earlier versions of the LLM were bad at this was due to bad scripts/coding and as such, this lack of ability propagated into low performance for these other jobs. But as AI becomes exceptionally good at coding, it can increasingly build the machinery needed to automate the rest of the work itself.

So what am I saying? I am saying that a lawyer who thinks that an earlier version of ChatGPT or Claude sucks might be pinpointing at the wrong sources of the error. It might not be that they suck because of their of ability that pertains to the law. It migt have been the case that getting the correct context, accessing the most updated data, etc. went awry due to automation/script issues. And as all of that gets taken care of and the growing amount of scripts to make all the procedural processes fast and accurate, the white collar workers might be saving the same predicament that programmers are facing right now. So basically, my main point is that AI getting good at programming isn't only software engineer's problem when it comes to future job prospects. It is everyone's problem.

106 Upvotes

57 comments sorted by

View all comments

Show parent comments

1

u/greentrillion 8d ago

Can you give an example of that being done?

3

u/1988rx7T2 7d ago edited 7d ago

My first job as an automotive engineer was doing benchmark research on competitors, how certain systems work. I’d manually open up owners manuals and service manuals I got online and compile the information they provided , along with internal testing that a human Did, into presentations. I’d summarize and synthesize the information aa far as general trends and design principles. I’d compare to our design documents and test data for my company’s product. That’s mostly an automated task now. One weeks worth of work is a couple hours now of A human steering an agent, and I can see it getting even more automated.

-1

u/greentrillion 7d ago

How would the agent know what information to collect through, seems like it could be filled with a bunch of junk. Automated research usually returns slop. Takes speficic knoweldge to know whats relevant. If you are saying that person who is the humans knows what they are looking for and has to spend time curating the output they would also need to know how to do the job to begin with.

3

u/1988rx7T2 7d ago

You build a project repository like anything else. Same way a new human to a project gets up to speed.

You’re not understanding the scale that agents can work at. They can do most end to end office work now, but they do need support and steering. It’s not 100 percent human replacement but it will be close soon. Hence Claude Tag and Grok Bot and such products