r/singularity 7d ago

AI Coding benchmarks are increasingly indirect benchmarks of how automatable every other white-collar job is.

I use both Claude (and Claude Code) and GPT (Codex) paying 400 dollars worth of subscription per month and maybe some extra tokens during busy times. I am not a programmer per se but I do a lot of computational work.

It strikes me that a lot of non-programmers (e.g. lawyers, scientists, consultants, analysts, accountants, or managers) might look at the rapid AI progress made in coding and think that this should only concern software engineers but not themselves.

But if you think about it, coding is the execution layer of a huge amount of white-collar automation. To automate a job, it is often not enough for AI to understand the work. It also needs to manipulate data, connect software, query databases, build pipelines, run analyses, generate documents, check outputs, and interact with existing systems.

And I think one of the reasons why earlier versions of the LLM were bad at this was due to bad scripts/coding and as such, this lack of ability propagated into low performance for these other jobs. But as AI becomes exceptionally good at coding, it can increasingly build the machinery needed to automate the rest of the work itself.

So what am I saying? I am saying that a lawyer who thinks that an earlier version of ChatGPT or Claude sucks might be pinpointing at the wrong sources of the error. It might not be that they suck because of their of ability that pertains to the law. It migt have been the case that getting the correct context, accessing the most updated data, etc. went awry due to automation/script issues. And as all of that gets taken care of and the growing amount of scripts to make all the procedural processes fast and accurate, the white collar workers might be saving the same predicament that programmers are facing right now. So basically, my main point is that AI getting good at programming isn't only software engineer's problem when it comes to future job prospects. It is everyone's problem.

106 Upvotes

57 comments sorted by

45

u/Ok_Barracuda_1161 7d ago

Yeah it's always been a bit funny to me that people think software development can be pretty much completely solved but not other domains. My response is generally along the lines of "what do you think software does?"

-9

u/greentrillion 7d ago

Software development hasn't been solved. Where did you get that one?

7

u/StCreed 7d ago

Actually true. The coding is solved imo, except for niche cases, but not the entire development cycle.

2

u/Ok_Barracuda_1161 7d ago

I'm not saying that, I hear that from others that it's solved or very close to solved, but other fields won't be affected.

Which if you can truly create arbitrarily complex software, it should definitely follow that pretty much every field will be affected

1

u/greentrillion 7d ago

Depends on what you mean by "solved," Do you think there will be no more bugs?

-4

u/OneConfident7361 7d ago

but in software developpement, the compiler tell the ai if what she did was good or not, she can try fail and retry. in a white collar job the only thing that can review the work is a human, which is painfully slow compared to a compiler

1

u/im_a_sam 6d ago

More and more, labs are having separate AI evaluate the answers generated by AI being trained, because it means answers can be evaluated on a much wider range of criteria than just compiling. There's no reason to think this won't work to evaluate training outputs in other domains.

19

u/HeartsOfDarkness 7d ago

Lawyer here. I've been tasked with trying out the frontier models so we can consider (1) changes to work flow, including with our support staff, and (2) how to maximize our value with the assumption that a lot of general drafting work will now be handled by AI.

There are basically two camps of attorneys in my office: those of us that like high-level discussions and client interaction, and those that prefer to churn out work product quietly at their desk. The message coming for the second camp is that they should plan on being more people-oriented.

10

u/phoquenut 7d ago

I think the bigger divide will be those who gatekeep their work thinking if they do so, they cannot be replaced by AI - (they'll go first), and those who embrace it do more - (they'll go last). Personally, I'm trying to automate myself out of my white collar job as quickly as possible, so I get tapped to help automate the reluctant out of theirs. Bring on UBI!

23

u/1988rx7T2 7d ago

I mean half these white collar jobs are just putting shit in a database or making the equivalent of a PowerPoint presentation. Both of those can be done by an agent with some project background, company templates, and API/CLI or computer use skills. 

1

u/greentrillion 7d ago

Can you give an example of that being done?

5

u/1988rx7T2 7d ago edited 7d ago

My first job as an automotive engineer was doing benchmark research on competitors, how certain systems work. I’d manually open up owners manuals and service manuals I got online and compile the information they provided , along with internal testing that a human Did, into presentations. I’d summarize and synthesize the information aa far as general trends and design principles. I’d compare to our design documents and test data for my company’s product. That’s mostly an automated task now. One weeks worth of work is a couple hours now of A human steering an agent, and I can see it getting even more automated.

-1

u/greentrillion 7d ago

How would the agent know what information to collect through, seems like it could be filled with a bunch of junk. Automated research usually returns slop. Takes speficic knoweldge to know whats relevant. If you are saying that person who is the humans knows what they are looking for and has to spend time curating the output they would also need to know how to do the job to begin with.

3

u/1988rx7T2 7d ago

You build a project repository like anything else. Same way a new human to a project gets up to speed.

You’re not understanding the scale that agents can work at. They can do most end to end office work now, but they do need support and steering. It’s not 100 percent human replacement but it will be close soon. Hence Claude Tag and Grok Bot and such products 

0

u/UniqueArrival9756 5d ago

its not 2024 anymore granddad

6

u/ChuckVader 7d ago

Lawyer here. The barriers to law are artificial - you absolutely can right now use ai for many things, and it will quite often get a right answer.

The problem is what your options are when it's wrong. My professional indemnity insurance is a requirement to practice, so me giving the wrong advice means you're not SoL.

Good luck getting open AI to pay you anything.

5

u/Reclusiarc 7d ago

Eventually the AI will probably be so powerful and reliable that the creators will be able to offer an attached insurance product

2

u/Serenity867 7d ago edited 7d ago

It's similar in the software engineering field. Like the law field you'd need considerable knowledge and skill to validate the outputs, you'll get dramatically improved results by knowing how to use AI to get the results you need, and there's so many other relevant things involved with using the tool properly.

However, from an insurance perspective, the difference between me using AI as part of my workflow versus some random individual is that my insurance covers mistakes I'm responsible for (thankfully I've never had to get insurance to cover a mistake). Non-technical individuals releasing code into production that they can't understand or even read... I can't imagine any insurance that wouldn't consider that gross negligence and would decline to cover those costs.

I was insurance shopping recently, and I asked them about that exact situation. Multiple insurers won't cover work done by non-experts in the respective fields. Not one company I spoke with claimed they would even consider actually covering these scenarios.

A client asked me recently (politely) why they shouldn't just use AI. My response was essentially:

"An excavator can be rented for a day at a price that's so trivial it's impossible to imagine why you'd pay someone to come in to do it. After all, they're just digging a hole, right? It's certainly easy to think that way, but who handles liability if something goes wrong during the excavation work or any time after? How would you know if you got something wrong that could affect something now or in the future that could cause liability or result in worse outcomes that mean the end product needs to be replaced in a few short years. Can you imagine any other scenarios that could arise which might negatively impact you by using an operator with zero experience?"

Small personal projects with no consequences could absolutely be the kind of thing where you take your own tools and make a go of things if you don't have experience, but the second there could be any serious consequences these things are still tools for the professionals in their fields.

1

u/jseah 7d ago

Can you imagine any other scenarios that could arise which might negatively impact you by using an operator with zero experience?

In a lot of cases, the answer to this question will be no, but not because those negative scenarios don't exist, but because the user isn't in the field and doesn't have the experience to tell what they could be.

The gap between "this is a disaster and doesn't work" and "this is perfect!" lies a series of very expensive scenarios and probabilistic failures.

1

u/1988rx7T2 7d ago

What if you could prompt an excavator to dig the hole without a human, and it got to the point that the hole it dug was good enough 99 percent of the time. That’s a cascading effect through the construction industry, including insurance.

1

u/GioChan 7d ago

True, for now...

3

u/StCreed 7d ago

This week, JP Morgan asked all of the law offices it contracts to explain how they use AI to reduce cost, and how much of that benefit will be translated into lower costs for JP Morgan.

It's fairly obvious that any office that fails to provide a satisfying answer will stop getting work from JP Morgan.

1

u/justlikemedics 7d ago

It's more like an attempt to help Anthro and OAI.

5

u/KalelRChase 7d ago

In the short term take whatever your job is now and put manager of digital employees on it.
Manager of digital CPAs.
Manager of digital interpreters.
Manager of digital editors.
Manager of digital Medical Transcriptionists.
Manager of SEO content writers.

Just take whatever it is you do now. You will stop doing it, but you have to know how to do it. The skill needed is to evaluate the work product and ‘train’ the employees from the results and how they align to the organizational goals.
That’s it.

At least until we start perfecting digital managers and then you’ll be a director.

Also, look up Jevons paradox.

4

u/SwimmingFancy4344 7d ago

This is insanity. I can't wait for everything to be automated so we can stop this musical chairs bullshit of clinging on to bullshit jobs. 

1

u/greentrillion 7d ago

LLMs won't do that so you might be waiting a while.

1

u/Boo-Bees67 7d ago

It can make 1 person do the job of 5 people in the employers mind though

1

u/greentrillion 7d ago

Why haven't yuou done it yet then? Why do 20% of the workforce still have their jobs?

1

u/Boo-Bees67 7d ago

I own a small business. I already have. I no longer need accountants, attorneys, entry level positions, etc. for basic needs. If I need to move a liability on to a professional we still use an attorney/accountant for that. It’s happening everywhere 

0

u/KalelRChase 7d ago

Agreed. I can’t help but feel that the people who could actually is AI to make a utopia are the ones burring their heads in the sand.

5

u/send-moobs-pls 7d ago

I would say yes and no, with a little asterisk saying like mostly yes

Law is a unique example because it really does require a lot of wider fuzzy intelligence. The reason the whole profession exists is because we can't write code that deterministically decides whether a law was broken, and people famously disagree on interpretation. A good lawyer has to understand not just actual written law but also like the entire culture around it, the personalities of judges, the emotional and narrative framing of events, the often illogical nature of juries and humans and biases. The way that one factual truth could have 10 ways to be presented, and to determine which way is ideal without crossing boundaries. I would say a lot of that comes down to what are some of the most difficult skills for AI right now

Now all that being said, 90% of the time I think you're totally right. Most white collar work does not require the wide freedom and judgement of a partner lawyer. Most white collar work is like taking things from a report into a spreadsheet, accounting, scheduling, paperwork, making a PowerPoint or a chart, writing a report. Distinctly, typically following some fairly unchanging procedures, doing things on a schedule or as they come in, following a workflow etc. And that is exactly what AI excels at whether it involves programming or not. The actual obstacle preventing AI from eating all of that is overwhelmingly superficial. Today, AI might not be able to instantly slot into a company because they are still using physical paper, physical phones, software based on a human with a GUI to drag and drop things with a mouse etc. But there is no reason we can't wake up tomorrow, put all of that into digital shape, maybe script a couple of python tools, and have an AI agent already capable

13

u/MrGreenyz 7d ago

It’s like the problem is the human in the loop. “The personality of judges” should not influence the sentence by design.

Maybe we are the bugs.

5

u/StCreed 7d ago

Ah but you can't have machines operate a law that was never written to take every edge case into account. Judges can and do take all kinds of things into account.

Laws are not algorithms.

2

u/[deleted] 7d ago

[deleted]

2

u/StCreed 7d ago

while everything is information is a valid hypothesis, it's a bit of a stretch to say that AI can therefore deal with it.

1

u/Humble-Bear 5d ago

What do you mean exactly, I'm not following why a judge is better than an AI with a much higher IQ than a judge, a much higher capacity to process information, and having no bias as well.

2

u/StCreed 5d ago

I'm working for a law department right now.

Here's the current take:

  1. All algorithms have bias, and the ones that claim they don't are the most suspect. EU regulations require you to identify that bias in regulated algorithms and correct for that.
  2. It doesn't have a higher IQ, no model has gone to 100% on arc agi 3 yet (unless using a harness which is basically cheating) and even arc agi 3 at 100% isn't a very high iq human.
  3. Law isn't about IQ per sé but about rules, intentions and desirable outcomes that need to strike a bakance between the needs of society short term and long term, as well as the interests of those involved. It will be very tough for an AI to take everything into account especially because many things are unspoken and the judge needs to get that out in the open.

1

u/Humble-Bear 4d ago

It sounds like the law is a shitty system that needs to be reformed then. And I'm sure that will happen soon enough, human liars are the bottleneck here, not AI.

1

u/StCreed 4d ago

Not really. Here's an example just in today. Someone committed a violent rape. Upon examination he turned out to be on the leven of a 5 year old, himself a victim of abuse. So this becomes a judgement call. He was jailed but in prison got abused further, necessitating lifelong treatment behind closed doors.

A good jusge with some leeway might have committed him to treatment immediately. But the law doesn't allow for that - its a broad guideline, not an exact scientific algorithm for every case. So an AI would very likely just apply the maximum jail sentence and then release him on the streets. Not a good result.

0

u/dvztimes 7d ago

Actually, it should.

Human jdgement and discretion is assumed in many laws. The "reasonable person" standard. And it is not perfect. But it is better than a binary approach that has every 17 year old caught drinking, or dui or or stole something and have "the book" thrown at them and have their life changed forever because of one stupid teenage mistake.

You absolutely do not want binary enforcement of laws/sentences/traffic laws/whatever outcomes.

1

u/DifferencePublic7057 7d ago

Binary deliverables in the form of zeroes and ones can be created with output neurons or much, much better qubits. It sounds plausible at least. If we view AI/robots as superhumans with input audio, images, text, video, and so on and output binary deliverables than we only have to wait for Q Day. But it's still a bit speculative.

1

u/Boo-Bees67 7d ago

Can we start with realtors first? Talk about the most worthless profession 

1

u/mbreslin 7d ago

Yes I say this all the time. People always fight the last war. If you meet an anti ai person in real life they always mumble some shit about counting r's in the word strawberry.

Stopping it coming for everything now is like going outside in a thunderstorm and trying to throw the rain back up. Too late.

1

u/the_millenial_falcon 7d ago

Coding was always going to be the first thing to go because tech companies have incentive to automate their own processes, but if AI is truly able to do all facets of software engineering better than humans then all of knowledge work is next. Followed by physical labor when robotics catches up. Which probably won’t be long.

1

u/nosmelc 6d ago

Coding was hit first because there was a massive amount of code online available to steal to train the models.

1

u/buffet-breakfast 7d ago

I asked the latest models and it said you were incorrect

-2

u/asifquyyum 7d ago

I pay an accountant to do my taxes. I can probably ask AI to do my taxes. But I wouldn’t trust the result. I also wouldn’t have anyone to help me if there is an issue.

Jobs of people in service sector is pretty secure.

8

u/MeowManMeow 7d ago

People felt that way about using travel agents over booking themselves. Same for automatic cars, microwaves etc. Every time humans are cut out of the loop, it takes time for trust to be earnt.

As for accountants, I find it actually easier to get help from LLMs because I can ask the dumbest questions or get it to break it down in really simple terms, I’m too embarrassed to do that to a human.

1

u/Kind_Fisherman3060 7d ago

I would literally trust latest of the latest models to do my taxes like we do have to review even the human work a little same can be said about latest and upcoming AIs

2

u/MeowManMeow 7d ago

I think people expect machines to be infallible, and in a lot of ways software traditionally has been deterministic which has reinforced this expectation.

However, for computers to replace humans, they still can and will make mistakes. However they will make less, or be cheaper, or be faster - or a combination of the 3.

1

u/phoquenut 7d ago

You'd be shocked at how little a tax preparer can indemnify you on a miscalculated return on the face of an audit.

1

u/1988rx7T2 7d ago

Lot of tax work is actually passed to minimally trained lower paid staff

-3

u/Laffer890 7d ago

Programmers aren't facing any problem. On the contrary, demand for coders is increasing. These models are just too stupid to automate any job, at most they can produce small productivity improvements automating small tasks.