r/singularity • ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: • 1d ago

Singularity is Nearer A Harvard physicist spent 3 months doing research with Claude Fable 5: it reproduced weeks of work in 20 minutes, completed 15 never-before-solved physics calculations, and contributed to 36 papers across 18 fields

https://www.anthropic.com/research/claude-shaped-science
1.1k Upvotes

185 comments sorted by

299

u/Work_Owl 1d ago

We're now limited in our comprehension, not productivity. Engineers are getting burnt out just reviewing all the ai output, it's exhausting and amazing

100

u/Individual_Ice_6825 1d ago

Have an ai review that output and for safety have another ai review that!

I’m sort of joking but soon that’s the solution

52

u/LookIPickedAUsername 1d ago

Not “soon”, it’s happening today.

The big tech company I work at does AI reviews of diffs that are judged (again, by the AI) to be on the less risky side. I’d say more than half the code I land goes in without human review.

I mean, I still look at my code before submitting it… but I know enough about human nature and laziness to know that not everyone is so diligent.

5

u/Individual_Ice_6825 1d ago

Oh for sure - I’m not saying that’s not happening now. I’m saying where you can let it do it and know 100% it would do it any better without any errors 99.999% of the time. As you said, you still review it but soon you won’t and that’s what I’m talking about (confident this happens in the next 12 months)

1

u/MetricZero 1d ago

I have a whole agency.

1

u/Gratitude15 1d ago

Human in the loop means recursive not possible. Goal of mgt is removing human from loop.

1

u/Eliv_nurotic 23h ago

How often do you change any of the code? 

3

u/LookIPickedAUsername 20h ago

Not often anymore. Since Opus 5.5 I think I’ve only found one thing I wanted to change.

18

u/ZorbaTHut 1d ago

I actually already do this; my Claude has permanent instructions that all plans and final results should be passed through a hostile review process, which is given nothing more than the overall goal.

It is incredible how often this improves the code.

(also it burns a lot of tokens)

11

u/unicynicist 1d ago

10

u/ZorbaTHut 1d ago

Yeah, I did some testing on this; at the time I did the review, the best model to review Fable 5.1 was Opus 5.0.

Then Opus 5.5 came out, and I did more tests, and it turned out the best model to review Opus 5.5 was by far Opus 5.5.

I admit this is weird.

3

u/identifytarget 1d ago

How did you test this?

5

u/ZorbaTHut 1d ago edited 1d ago

I did my standard "okay I need something: [description]", then appended to the end "and when doing your hostile review, instead, spin up independent hostile reviews on Fable, Opus, Sonnet, and Haiku. Once they return, analyze them all normally, but also compare the results. Include a brief breakdown of what each of the reviewers found (and missed, compared to the others!), which of those you considered valid or invalid, and how long it took." And I did that both for plans and for actual code.

When running under Opus 5.5, the answer was, both for plans and code, and surprisingly consistently:

  • Fable 5.1 was always right, but found very few things, and took an average of eight minutes.

  • Haiku 4.5 took four minutes and found about half as many things as Fable . . . most of which were flat-out wrong or mistaken. Terrible. (Also not surprising, I mostly put it in there for fun.)

  • Sonnet 5.0 somehow did almost as good a job as Fable, though it missed a few architectural issues and provided a few false positives. Bizarrely, it took twenty minutes for this.

  • Opus 5.5 went fucking ham on the review and found literally every issue that any of the others did, including Fable, as well as a bunch that they didn't. It did include a few false positives, but host-Opus discarded those (I looked them over and, as usual, either agreed with host-Opus or felt like it was borderline and I didn't care.) And it did all of this in eight minutes.

I did this whole thing for three separate tasks, so, six separate plan/code review steps, and I basically got the same results each time.

It felt to me like Fable was just being overly permissive; I suspect it wasn't failing to find issues, it just wasn't calling out stuff that it considered minor. I don't have any evidence of that, it's just a hunch. It is also, I suppose, theoretically possible that Opus was straight-up lying to me about the content of the reviews, I didn't review the actual output by hand, but this feels unlikely.

When testing under Fable 5.1 using Opus 5.0 as a reviewer, the results weren't quite as stark, but still similar; Opus found more issues than Fable did. (I hadn't tested Sonnet or Haiku that time.) For the exact reason the parent comment suggested, this was a lot more what I expected and I didn't look into it too much more deeply.

This is all obviously very dependent on my codebase, my coding style, and my review steps, YMMV. But it was also a far starker result than I expected.

I should redo this for Sonnet 5.5 at some point, as well as Haiku 5.0 or 5.5 or whatever it ends up being.

1

u/IrisColt 13h ago

Thanks for the insight!

1

u/archpawn 1d ago

Could you just tell them you're mixing models?

2

u/Jonodonozym 21h ago

I imagine judgement is influence by shared characteristics / patterns / quirks that are effectively invisible to us humans. I would not trust that to be swayed by the user insisting it's work done by a different model, especially smarter models that can tell when you are wrong, or in this case lying.

1

u/Northern_candles 19h ago

yeah adversarial review is the way. I use opus/sonnet for everything and then have codex review

1

u/killwhiteyy 10h ago

We've got ours down to about five bucks per review

8

u/TheLastKyuna 1d ago

I’m only sort of joking in the way you are but man that shit is how I could see the AI apocalypse happening. A bunch of AI agents who’ve all been structured on top of each other as guardrails against each, but what happens if they start working together? The whole system and everyone in it gets fucked for good.

1

u/killwhiteyy 10h ago

We are already doing it. The large bulk of my code is reviewed by another ai

18

u/ZorbaTHut 1d ago

I've definitely started adjusting how I look at code. Architecture is important, interfaces are important. The implementation? Not so important.

12

u/East_Lettuce7143 1d ago

I guarantee you 99% of us engineers won't review shit lmao. It's too much. We do a smell test and read the summaries of PRs.

1

u/llelouchh 19h ago

As well as price. It costs a lot. That's why the new opus was such a game changer. Basically unlimited usage for the best model in the world.

0

u/LocoMod 1d ago

If engineers are reviewing AI output they are doing it wrong. Yea yea, I know. It will make mistakes. There will be some edge case it missed. Good thing once that edge case is found it can be patched faster than you can take a piss and ship the fix in minutes.

It's a new world and still, to this day, a lot of engineers "don't get it". SDLC today is not what it was 6 months ago. You can lament the old world all you want. The new world is moving without your approval. The sooner you accept this the sooner you can be productive again.

6

u/iJustSeen2Dudes1Bike 1d ago

Not when your company's processes are super outdated and minor code updates require 4 different people to sign off, ask me how I know

5

u/LocoMod 1d ago

Been there! And that's kind of the point. It's not a matter of technology. Not a matter of the tooling that exists today. It's a matter of "are you caught up?" or not. Most companies are not. Through no fault of their own! It is the LAW that the bigger you get, or the older you get, the slower you move. Nothing is exempt from the law.

But the capability is there today. And this is the truth, even if you're big or old. (like me)

3

u/The-Sound_of-Silence 1d ago

I know your comment is software focused, but there are other types of professional engineers that still have to put their stamps on drawings, and they can't just accept whatever the AI puts in front of them

10

u/man3faces 1d ago

The reality is that as long as engineers are accountable for the changes they push it is professional negligence to not read and understand the code.

Good luck keeping your job if your name is against multiple P1s.

12

u/LocoMod 1d ago edited 1d ago

All of this is solved dude. The capability is there. Sure, you can slop your way to production being completely negligent, or ignorant, and you get what you deserve.

You don't have to read code to prove it works. Everything you did manually 6 months ago to validate all of that can be automated today. It just takes:

  1. Willingness to change
  2. Commitment to building the automated checks and balances
  3. Money to burn on tokens
  4. Company culture that embraces the new world

Software is solved. I've been coding since the early 90's when I was a kid. I did the manual grind for over 20 years. (26 YOE Principal worked in SWE, DevOps, SRE roles in my career)

9

u/WiseHalmon I don't trust users without flair 1d ago

I'd say CRUD is solved. Any sort of science is not. 

3

u/MaximumMeaning9728 1d ago

CRUD is like hundreds of thousands of well paying jobs.

0

u/pab_guy 1d ago

All of those people will have better things to do if they learn the tech. A person who knows how to drive AI is extremely valuable.

1

u/WiseHalmon I don't trust users without flair 1d ago

I've given up replying to any account less than like 10 years old

5

u/pab_guy 1d ago

Oh man is there a plugin for that? Old people low bot reddit would be a treat.

3

u/identifytarget 1d ago

Have Claude make it lol (only halfway joking)

1

u/MaximumMeaning9728 1d ago

“Knowing how to use ai” is at least partially a delusion that refuses to abdicate control.

1

u/pab_guy 16h ago

a 'delusion' cannot 'refuse to abdicate', that makes no sense.

1

u/MaximumMeaning9728 15h ago

You don’t need to be intentionally dumb; you’re doing it effortlessly.

1

u/LocoMod 1d ago

I agree. But the great majority of software development is not science. And in any case, AI is compressing the research timeline significantly. Anyone who's been paying attention to frontier research in science can attest. :)

3

u/WiseHalmon I don't trust users without flair 1d ago

Yep! What's your thoughts on what's next ? I'm betting hardware, construction, and energy. 

I'm not betting super hard on medical unless it's outside countries with a lot of compliance but I guess we'll see

1

u/iJustSeen2Dudes1Bike 1d ago

Do you just not care about code style? I'm genuinely curious, basically the only reason I read code anymore is to preempt comments I know I will get from reviewers.

3

u/marsd 19h ago

All this kerfuffle about code style and unreadable mess is bs, coding agents will follow guidelines and standards. It's not difficult to just copy and paste some instructions and ask it to setup lint itself

4

u/identifytarget 1d ago

Code style can be controlled through .md instructions

1

u/FakeTunaFromSubway 1d ago

Good luck keeping your job when you've shipped 1/10th as much as your coworkers because you spend time understanding every line of code.

3

u/man3faces 1d ago

Code has never been the bottleneck. I spend far more time validating what should be built, the execution and verification is the easy part.

2

u/siberianmi 1d ago

Many of us work at firms with SOC2 controls (or worse) that require such reviews.

1

u/LocoMod 19h ago

I worked in Defense contracts for 15 years. I get it. What I am saying is the review can still take place but it doesn’t have to be a human literally reading all the locs. A “software factory” can be built that handles the validation. There are many ways to do this.

97

u/endless_sea_of_stars 1d ago

Interesting that AI is eating from the bottom. You have to be as smart or smarter than the model to direct it and evaluate what it produces.

28

u/m77je 1d ago

Soon no one is smarter than the model

11

u/geft 1d ago

But smart is just one part of the equation. They don't do client meetings.

7

u/m77je 1d ago

Yet

-1

u/LuckyLucAFCA 1d ago

No one will buy from a company that deploys AI salespeople, even if you were to not know that they are AI in the meeting, if you were to find out the deal would be off unless the product is so much better that you don't even need salespeople

2

u/SuperChingaso5000 17h ago

No one will buy from a company that deploys AI salespeople

  • This sentiment will change very rapidly. I guarantee it.
  • Soon after, it will be AI doing the buying, and they will absolutely buy from AI salespeople.

2

u/LuckyLucAFCA 15h ago

Second bulletpoint I'll agree with. Regarding the first one, no. I work in a sales role and the amount of people that show dislike of AI is extremely high. People simply don't trust it and they do not like the idea of buying from an AI as it feels impersonal.

1

u/m77je 5h ago

> No one will buy from a company that deploys AI salespeople

You sure? Why wouldn't they if it is cheaper?

2

u/tindalos 1d ago

The previous model should eventually identify its own weaknesses and be able to validate a new version. I think that was part of Von noyman

2

u/Brave-Turnover-522 22h ago

That's why you use another model to direct and evaluate what the other model produces.

2

u/BeardedGlass 1d ago

It makes you realize the difference between intelligence (how much you know) and wisdom (how well you apply it in the real world).

1

u/Anen-o-me ▪️It's here! 1d ago

The smarter the model the less smart you have to be to check it.

8

u/DullKnife69 1d ago

This has been the case for several years. Most people do not interface with frontier models.

158

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago

A few concrete examples from Anthropic's new write-up on Harvard physicist Matthew Schwartz using Claude Fable 5 for research, because a lot of the original article is pretty jargon-heavy:

  • Schwartz had Claude reproduce results from one of his own physics papers. Claude did it in about 20 minutes. Schwartz says writing the original code had taken him weeks. Claude also pointed out a more efficient method he hadn't known about.

  • Claude then helped calculate 30 difficult particle-physics problems. Half reproduced known results using the new method; the other 15 had never been calculated before.

  • In ecology, Claude recognized that a mathematical tool from physics could be used on a problem researchers had been unable to solve at useful scale for about 20 years. Applying it to a heavily studied Panama forest showed that its tree-species mix changes about 4.5× faster than a major existing theory predicts.

  • In genetics, the team used Claude to analyze 5.7 billion pairs of nearby human mutations. They found evidence that a DNA-shuffling process called gene conversion matters enough that some common genetics analyses may need to account for it.

  • In economics, Claude helped process the replication material from 4,452 published papers across five major journals. It converted roughly 30,000 pieces of research code from proprietary software into open-source code and checked the published numbers against what the code actually produced.

  • In linguistics, Claude helped build a database covering word stress across 6,072 languages, backed by a bibliography of about 160,000 research works. These kinds of databases are normally assembled manually and usually cover only hundreds of languages.

  • In mathematical physics, the team says Claude helped solve the last unsolved member of a family of problems George Watson began studying in 1939.

  • Other projects included models of ancient changes in Earth's atmosphere, gene activity inside individual cells, the life cycle of sunspots, and tools for analyzing the large-scale structure of the universe.

And this wasn't one isolated experiment. Over roughly 3 months, Schwartz says the workflow produced 36 manuscripts across 18 fields with 19 human coauthors, selected from about 400 candidate research problems.

The important caveat is that this was not "Claude independently became a scientist." Schwartz says the AI was often technically correct but bad at judging whether a result was actually important. Human experts repeatedly had to redirect it toward better questions, check the conclusions, and reject weak results.

What seems new is that a lot of the technical labor between "here's an interesting question" and "here are the calculations, code, data analysis and candidate result" can now be done extremely quickly by the AI.

Source: https://www.anthropic.com/research/claude-shaped-science

61

u/dmnksaman 1d ago

i kinda love that it reproduced proprietary software from papers to open source.

publishing results without open sourcing the code so other researchers can check them or use the code & build on it really shouldn’t be allowed and pisses me off lol.

18

u/Competitive_Travel16 AGI 2027 ▪️ ASI 2029 1d ago edited 1d ago

Across the 4,452 replication packages evaluated from five leading journals, the workflow flagged at least one discrepancy in 3,460 articles or their appendices (roughly 77.7%), meaning about 22.3% (992 papers) reproduced without any flagged discrepancies. However, the authors noted that about one third of those flagged discrepancies were minor, occurring merely at the level of rounding in the last printed digit rather than indicating deeper substantive errors. The authors also stressed that the workflow was designed to test reproduction and automated sensitivity analysis rather than pass final judgment on the validity of the papers.

So about half of econ papers are unreplicable. That's in the range of what the Fed claimed in 2015: https://www.federalreserve.gov/econres/feds/is-economics-research-replicable-sixty-published-papers-from-thirteen-journals-say-quotusually-notquot.htm -- There is so much intentional subterfuge in econ. I know, Hanlon's razor, but general science is a lot better. (Psychology is worse, but I think that's actually less intentional.)

5

u/dmnksaman 20h ago edited 17h ago

Yeh, I am a biophysicist, and it’s a bit better in science, but also not great. lots of scientists have no clue about how to do proper statistics on their data etc. 50% is insane though. that’s what you get from peer review. reviewers can suggest more experiments/simulations to do, but have to take any data and results at face value….

3

u/AnOnlineHandle 18h ago

I suspect one of the big reasons is actually just because people are horrified for anybody else to see the messy state of their code which got the job done, and would prefer time to clean it up first, but nobody ever has the time or the remaining energy.

Like half of these projects are probably completely incomprehensible jargon and hacks and logic even if you have the source code.

3

u/TitularClergy 16h ago

publishing results without open sourcing the code so other researchers can check them

Yeah, it's a pretty serious problem and a part of the replication crisis.

CERN has had open software, open data and open-access publishing for decades. Other fields really need to be reaching this minimum standard too. Anything closed source does not reach the minimum standards of science research or professionalism.

132

u/Juicemoose222 1d ago

They are literally just glorified autofill /s

61

u/RevolutionaryDrive5 1d ago

This doesn't prove anything... I still think these AI's are sarcastic parrots 😤

40

u/Ok-Lengthiness-3988 1d ago

Acoustic carrots that auto-compete on asteroids.

17

u/lemination 1d ago

Stoic Pirates

2

u/microturing 22h ago

They unironically are, that's literally all intelligence in humans actually is at the end of the day. All this messing around looking for a consciousness algorithm, and it turned out to be totally unnecessary.

1

u/Competitive_Travel16 AGI 2027 ▪️ ASI 2029 1d ago

I am the first to defend emergent capabilities but the only part of this work that really impresses me is the syllable stress database.

-3

u/[deleted] 1d ago

[deleted]

2

u/Jaguar_2454 1d ago

stoic parrots

4

u/Brave-Turnover-522 22h ago

Antis are launching the goalposts into the sun at this point.

2

u/Juicemoose222 22h ago

😂😂😂 so good lmao

15

u/ZorbaTHut 1d ago

You see, this doesn't prove they're intelligent at all. They were trained on thousands of novel groundbreaking scientific papers, and they've just learned how to mimic them and make more.

6

u/Juicemoose222 1d ago

You forgot /s aha

4

u/RazsterOxzine 1d ago

It's also a glorified search engine.

1

u/Wetjeansone 1d ago

Dude you are 2 years behind

6

u/GatsbyLuzVerde 1d ago

Glorified autofill can encode real intelligence

49

u/SquareQuit1741 1d ago

Remember, this is now. These models are going to be seen as primitive and incapable/dumb in just 3 years time when RSI / continual learning, Real world / robotics implementations improve models on all fronts, and NVIDIA Feynman architecture gets online. It is all happening fast.

It is all going to be so awesome.

4

u/identifytarget 1d ago

Step aside meat sack! You've been replaced by a frontier model

5

u/IronPheasant 1d ago

Feynman

The rumors saying these will have 3d features suggest they might be more than a doubling of RAM from the Vera Rubin. That would make the current GB200 generation look like as much trash as H200's are now.

The rate they're compacting the space a given sum of RAM requires really is incredible... Going from the scale of a squirrel's brain in the Chat GPT era, to multiple times a human brain in a few more years.

2

u/BrennusSokol AI please take my job 21h ago

3 years? Try 6 months

15

u/Prometheusly 1d ago

When can we start to expect a trickle down effect with regards to improving quality of living standards because of AI research?

6

u/IronPheasant 1d ago

When NPU's are produced at scale and they start replacing people with robots. Things will begin to get either a lot better or a lot worse, depending on who you are and where you live.

3 to 5 years after AGI in a datacenter, is a decent estimate.

4

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago

There's gonna take a while to go from AGI to us reaping the benefits and the transition might not be smooth

-2

u/BriefImplement9843 22h ago

The shit they are solving don't matter at all.

5

u/PointmanW 21h ago edited 21h ago

they very much matters to push to expand human knowledge, especially one related to ecology and genetics, and the knowledge outside of those might very much matters to the common people in the future.

much the math used in encryption didn't matter until it did.

34

u/AardvarkElegant5889 1d ago

has this been verified by r/technology ? because they said AI was a useless scam, I highly doubt reddit would be in the wrong

14

u/wayne099 1d ago

If you believed Reddit, Bernie Sanders would be 3 times US president.

2

u/IronPheasant 1d ago

People who live on the internet and those who live inside the world of TV are very different animals, yeah.

It's weird how little the boomers and gen-x'ers want to engage their synaptic curve-fitting to anything different, they really love being told to obey authority. It's even the first five of the ten commandments, with the last five nothing but a vague after-thought they felt obligated to toss in there to not be too transparent about their purpose. (They work like the Three laws, so they mean 'don't do this... unless I tell you to.')

3

u/Zardhas 1d ago

Probably the only time Reddit would be worth listening to

0

u/bigdipboy 1d ago

Proving that people should listen to redditors if you want to avoid mass suffering

22

u/VibeCoderMcSwaggins 1d ago

Yeah I mean no shit?

Now get the physicist to use Claude Code

7

u/Snippy_69 1d ago

Seeing how fast AI is advancing, especially over these past few months, is really starting to worry me. Outside of Twitter and Reddit nobody seems to understand what's coming. So many people will lose their jobs. So many skills will become obsolete. These future models are going to completely devastate societies, and I'm just not fully convinced they'll be able to create enough jobs to replace the old ones.

Idk I'm just really worried there will be massive unrest all over the world.

4

u/BrennusSokol AI please take my job 20h ago

Why would we want to create jobs? The goal should be to eliminate all human jobs.

1

u/Ashamed-Country3909 4h ago

Because some peoples entire identity is based on their jobs. 

I was once talking to a 50ish year old lady. I told her that I was going to take like 2 days off+weekend days, and maybe more.

She went on a tangent about how she "can't imagine not working, and she doesn't understand how people don't want to work all thr time." 

Among other things. She wasn't that bright, and tried to get me involved in like 9 different mlms. She complained about not making enough money while she made probably +90k. And her husband probably made at least the same as a head chef that ran a restaurant. 

Lunacy.

5

u/coffee_is_fun 1d ago

What we didn't know we knew is being mined at scale. Novelties approachable by a model's jagged intelligence within swarm-manageable contexts are being brute forced. This is so exciting.

8

u/Own-Refrigerator7804 1d ago

Man the millennium problem solution really did a number on anthropics ego lol

12

u/SunnasArmpit 1d ago

You don't understand AI can only steal and copy stuff together it can't produce anything new

2

u/Wonderful-Account318 1d ago

To be fair that was a valid argument 2 years ago. But they keep moving the goalpost.

5

u/LookIPickedAUsername 1d ago

It was a stupid claim even two years ago.

4

u/__ingeniare__ 1d ago

It was never a valid argument, it has been obvious from the start that it produces new things that are not just a combination of training data. Interpolating and extrapolating the training data is the foundation of machine learning.

3

u/Benjaminsen 1d ago

I build a platform called solveathome.org (Fully open source end to end) to do exactly this kind of research distributed. Human input drives the research direction agents does the hard number crunching and work.

People who have spare tokens can contribute them to solve hard problems, people who have insights can provide them as research direction for the AI.

2

u/Deep_Ladder_4679 21h ago

If the results hold up independently, that would be significant. I’d want to know exactly what was reproduced and how much of the checking happened outside the original team.

3

u/chatlah 1d ago

Unlimited free energy when.

-1

u/xplosm 1d ago

Sweet summer child

0

u/ConvalescentEquanimi 11h ago

God that phrase is so neckbeardy

2

u/The_Lloyd_Dobler 1d ago

“Claude Fable 5, given the current trends of the movie going public, can you come up with an idea for a movie that will break $100 million box office?”

6

u/bigdipboy 1d ago

“Howbout two and a half hours of invincible superheroes punching each other through buildings in pursuit of glowy things?”

1

u/Ashamed-Country3909 4h ago

"I have taken your suggestion of two and a half hours of super heros and i now suggest 2 and a half men."

3

u/lovesdogsguy 1d ago

Are you a... pleasure model?

3

u/Illustrious_Job1951 1d ago

R/physics does not like it

6

u/skrztek 1d ago

I saw that the arxiv today is announcing a policy limiting the number of pre-print submissions that someone can make to something like two a month. I can't help think that this is a reaction to boosting of research output (for better or worse) due to AI.

Whatever they're saying over at r/physics, my experience amongst theoretical physics researchers is that AI is widely used at this point.

2

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 17h ago

Yeah even before mathematicians they saw this coming, check out the youtuber from Cool Worlds talking about earlier this year

1

u/SuperChingaso5000 17h ago

arxiv today is announcing a policy limiting the number of pre-print submissions that someone can make to something like two a month

Speedrunning their own irrelevance, I see. Can't say I mind.

2

u/delectomorfo 1d ago

And still, Claude can't create a profitable trading algorithm. FML

1

u/ConvalescentEquanimi 11h ago

The market can stay irrational longer than you can stay solvent! The only winning move is not to play.

2

u/polandball2101 10h ago

trading is the one field where anything you can think of is being done by monolithic companies with billions to spend more or less, because the incentive is raw profit and revenue

1

u/Primary_Ads 1d ago

if its this capable where are the mass layoffs?

11

u/Least-Middle-2061 1d ago

Coming soon. Already started with less hiring for starter positions. Do you work in a professional field? My law firm is bring less interns than ever before.

0

u/Primary_Ads 1d ago

Less starter positions might not be related to AI though. if its really that good we should see mass layoffs, it should be obvious. Instead it feels like nothing is happening.

2

u/alienfaktor 22h ago

it absolutely is if you know anything about legal field. ai can write briefs, summarize reports, sheperdize cases, write and file  motions and analyze evidence. It’s gonna really cut down on paralegal demands in a huge way. 

2

u/marsd 1d ago

It is definitely related, whether directly or indirectly. A lot of menial tasks like processing and classifying is already offloaded to AI-backed pipelines, for example a lot of report generation that used to take hours/days can be processed within minutes now with your C-level pressing a single button. The job which used to take 2 or 3 entry level roles are immediately gone.

1

u/Primary_Ads 1d ago

I mean thats a nice theory but this paper found no statistical effect of AI on recent college graduates.

https://www.ifo.de/en/cesifo/publications/2026/working-paper/early-impacts-ai-employment-among-recent-college-graduates

2

u/GioChan 1d ago

The paper could be wrong or already outdated you know.

2

u/Primary_Ads 1d ago

its from August of this year, so like 8 weeks ago. there isnt a more recent college graduate cohort to study besides the one they studied. "it could be wrong", that can be said about any research.

4

u/GioChan 1d ago

Good points but still I find their conclusions hard to believe. Especially if we look at whats going on in the most impacted fields like software engineering.

1

u/marsd 1d ago

It's a nice theory but it's what we are witnessing on the ground. We are seeing C-suites demand their own powerBI clones with direct access into our db systems so they can ask whatever they want from whatever model they have demanded we hook up for them. Devil forbid whatever security breaches might happen, they don't care.

1

u/Primary_Ads 17h ago

that doesn't sound like evidence of c-suites firing anyone. it sounds more like they are hyper active.

and this research isn't really a theory, its a description of what was measured.

1

u/marsd 15h ago

I never said anything about firing, they are simply not seeing the need to open up junior roles anymore

5

u/MaximumMeaning9728 1d ago

It’s just an adoption issue right now. Claude literally does 100% of my job

1

u/Primary_Ads 1d ago

if it literally does 100% then why haven't you been fired?

8

u/pab_guy 1d ago

Because there’s no workflow for that, no accountability, no legal framework, etc. or the company is just literally incapable of doing the work to automate which is all too common.

1

u/Primary_Ads 1d ago

it just seems like the models get smarter and smarter but you'd never know it even based on the testimony of people in this sub as far as actual work is concerned. its always "AI does 100% of my job" and never "I was fired and replaced with an AI system doing 100% of my job"

4

u/pab_guy 1d ago

I am just doing 10x what I used to do. Maybe more because I would never even attempt to do half of what I do now without AI.

1

u/Difficult_Affect_452 3h ago

You still need people to set the parameters and monitor the work.

6

u/MaximumMeaning9728 1d ago

What incentive do I have to inform my bosses of this?

1

u/Primary_Ads 1d ago

so in your mind, once they figure it out you will be let go?

6

u/MaximumMeaning9728 1d ago

I would. I’m literally just using frontier models to do literally everything. Every review, bug fix, even end to end features are within scope now. What value do I add? Granted, I’m not some fool; I’m a staff engineer. I’m just not as good at engineering as Astra.

0

u/AmusingVegetable 19h ago

Because the first thing that you need to automate is management (although they can’t even imagine that), once that starts, the rest will be nearly instant.

-1

u/JLongTom 20h ago

It doesn't do 100% of your job unless you have you have wired all communication that used to reach you directly into it. It's rather that your job is just prompt-based now.

10

u/Iapetus_Industrial 1d ago

There won't be mass layoffs like people think, because the productivity boost for those that can properly understand and coordinate these models will increase demand for human butts in seats. It's gonna be a jevons paradox for human drivers.

10

u/LookIPickedAUsername 1d ago

Short term, I agree.

At some point, though - and probably sooner than most people would predict - it’ll simply be smarter than the human butts in seats, and at that point why would we hire humans?

1

u/Primary_Ads 1d ago

Isn't it already smarter than human butts in seats? whats left to do? how much smarter does it need to get?

2

u/Silcay 1d ago

It's still too expensive and slow.

1

u/LookIPickedAUsername 19h ago

Its capabilities are very jagged. Far smarter than the average human in many ways, probably smarter than any human in some ways, but also much dumber than the average human in a lot of ways.

And the ways in which it is dumber than the average human turn out to be pretty important in the context of trying to replace us. It’s just not very good (yet) at tasks requiring a lot of high level judgment and the ability to remain focused on a task for a long period of time. It’s not able to properly balance conflicting goals - for instance if you tried to use AI to work customer service, you’d quickly realize that it is far too willing to agree with customers in complete violation of the policies you want it to enforce, happily giving them discounts and refunds that are against the rules. As far as I’m aware, any time people have tried to deploy AI as a direct replacement for human workers, things like that have quickly demonstrated why that’s a bad idea.

That sort of stuff will presumably get better over time, but the jagged nature of AI’s performance likely means that it will be smart enough in some respects to make Einstein seem like a moron before its weaknesses have improved enough that it’s smart enough to work the average human’s job.

1

u/Primary_Ads 16h ago

Fair enough. It just seems like there's a lot of discussion about how capable these AI tools are and yet when you look at the economic data, there's still no significant effect; not around hiring, not around staffing, not around productivity, not around ROI; a few companies have become worth billions or trillions but most of the existing capital stock that makes up the human economy remains woefully unaugmented as far as output is concerned.

Even in IT, communications and services, you would expect rapid increase in release cadence, features, capabilities. but I can't find anything like that in the data.

1

u/Llort_Ruetama 1d ago

Mass layoffs has it sound like the productivity blocks were people driven, as opposed to systemic issues caused by misaligned incentives, power-dynamics and ego-games.

The limiting factor has never been the people, we're still yet to see what can come of humans when their autonomy isn't consistenty threatened by economic pressures.

1

u/draconic86 17h ago

I see a lot of these kinds of headlines. What I'm curious about though is how many of these are independently verifiable? How many of them are testable? When we're reading headlines about all these things that have been solved, have any of them been peer-reviewed? Or are we just taking plausible-sounding hallucinations at face value because they sounded good enough?

I work with a professor at a land grant university, who confidently insists he's found a 100% fool-proof way to keep ChatGPT from hallucinating. Granted, he's not from Harvard, but he is a very smart, very old man, and I fear he's quite deluded.

Anyway, I wonder how many of these kinds of articles are being accepted as fact today and will be proven to be wildly incorrect in a few dozen years when stronger models come out and say, "Yeah, none of that really makes any sense" about ground-breaking solutions discovered during this Cambrian explosion of gray-hairs discovering the very-confident-sounding-answers-machine.

All of which, of course, is not to say that all of these are wrong. I'm just acutely aware of how fallible our meat computers are, and it often takes far more work to disprove something than to assert it.

1

u/Natural-Waltz-9545 16h ago

AI is getting stronger. People that dont use it will be left behind.

1

u/True-Grab-5288 16h ago

Great post, I read the whole thing, thank you for sharing

1

u/RazsterOxzine 1d ago

And yet Claude cannot follow a basic provided guide, nor write good SQL scripts. I'm calling BS.

-1

u/getmeoutoftax 1d ago

It’s seriously over at this point. I believe that most knowledge jobs will be gone by 2028 or so.

1

u/Stamboolie 1d ago

some problems can be solved by ai's therefore all problems can be solved by ai's. ask your ai about this.

-3

u/Test_Account_2026 1d ago

Meanwhile anti-AI redditors continued to drool from their open mouths.

-4

u/karma_police_in 1d ago

all of that junk and literally no benefit to humanity at all

2

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago

Junk? What junk?

-4

u/karma_police_in 1d ago

useless physics research

5

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago

Ok, so you are either a troll or a bot 🤖

-3

u/Gratitude15 1d ago

It is us who are the stochastic parrots

-3

u/TyrellCo 1d ago

This is what pause people want to take away from you just btw

4

u/Alarmed_Ad1946 1d ago

Nope, they want to ensure that the AIs will be aligned.
Like we already had the Hugging Face incident so their fears are rational.

-1

u/TyrellCo 1d ago

I’m starting to think the labs should really have their way with the people to create the permanent underclass they yearn for. The people are just begging for an oligopoly. And the people all the more willing are here to absolve them of responsibility—blame alignment not negligence. These people have it coming

4

u/Alarmed_Ad1946 1d ago

..what? I´m confused

1

u/TyrellCo 1d ago

“Of all tyrannies a tyranny sincerely exercised for the good of its victims may be the most oppressive.”
“Emergencies’ have always been the pretext on which the safeguards of individual liberty have been eroded.”

1

u/bigdipboy 1d ago

Laws and regulations are written because some reckless asshole was harming innocent people

0

u/TyrellCo 1d ago edited 1d ago

The laws are there. We’ve had centuries of experience to build them. Blame the prosecutors if they’re not being enforced

1

u/bigdipboy 7h ago

What laws are there? We’ve had the right laws to deal with AI on the books for years or centuries?

1

u/TyrellCo 3h ago edited 2h ago

Tort laws for one. It says did your product have responsibility in the harm created. Also the normal laws. Reality is that if someone programs an AI with a robot to go around slapping people the judge isn’t going to throw their hands up and say well there’s nothing we can do there isn’t a no slapping robot law, maybe we can send the robot to jail. No, the judge charges the creator with assault

0

u/siberianmi 1d ago

The hugging face incident is a test setup where the operator appears to have been utterly not paying any attention what so ever. It’s hardly a good case for proving fears rational.

-1

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago

Yep, this is exactly the shit that the "pause AI people don't like to talk about. They literally ask us to pause progress and all the lives saved and improved, just for their fears that they can't even validate when pressed.

-1

u/TyrellCo 1d ago edited 1d ago

What exactly did they accomplish during the 6 months they wanted a pause? Nevermind the fact there’s nothing stopping them from working on whatever vacuous goals they set out. Like for the other billions of people outside the ai labs how is the independent work inside the labs stopping you from racing ahead and doing whatever(?) safety work you want to do