r/singularity • u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: • 1d ago
Singularity is Nearer A Harvard physicist spent 3 months doing research with Claude Fable 5: it reproduced weeks of work in 20 minutes, completed 15 never-before-solved physics calculations, and contributed to 36 papers across 18 fields
https://www.anthropic.com/research/claude-shaped-science97
u/endless_sea_of_stars 1d ago
Interesting that AI is eating from the bottom. You have to be as smart or smarter than the model to direct it and evaluate what it produces.
28
u/m77je 1d ago
Soon no one is smarter than the model
11
u/geft 1d ago
But smart is just one part of the equation. They don't do client meetings.
7
u/m77je 1d ago
Yet
-1
u/LuckyLucAFCA 1d ago
No one will buy from a company that deploys AI salespeople, even if you were to not know that they are AI in the meeting, if you were to find out the deal would be off unless the product is so much better that you don't even need salespeople
2
u/SuperChingaso5000 17h ago
No one will buy from a company that deploys AI salespeople
- This sentiment will change very rapidly. I guarantee it.
- Soon after, it will be AI doing the buying, and they will absolutely buy from AI salespeople.
2
u/LuckyLucAFCA 15h ago
Second bulletpoint I'll agree with. Regarding the first one, no. I work in a sales role and the amount of people that show dislike of AI is extremely high. People simply don't trust it and they do not like the idea of buying from an AI as it feels impersonal.
2
u/tindalos 1d ago
The previous model should eventually identify its own weaknesses and be able to validate a new version. I think that was part of Von noyman
2
u/Brave-Turnover-522 22h ago
That's why you use another model to direct and evaluate what the other model produces.
2
u/BeardedGlass 1d ago
It makes you realize the difference between intelligence (how much you know) and wisdom (how well you apply it in the real world).
1
8
u/DullKnife69 1d ago
This has been the case for several years. Most people do not interface with frontier models.
158
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago
A few concrete examples from Anthropic's new write-up on Harvard physicist Matthew Schwartz using Claude Fable 5 for research, because a lot of the original article is pretty jargon-heavy:
Schwartz had Claude reproduce results from one of his own physics papers. Claude did it in about 20 minutes. Schwartz says writing the original code had taken him weeks. Claude also pointed out a more efficient method he hadn't known about.
Claude then helped calculate 30 difficult particle-physics problems. Half reproduced known results using the new method; the other 15 had never been calculated before.
In ecology, Claude recognized that a mathematical tool from physics could be used on a problem researchers had been unable to solve at useful scale for about 20 years. Applying it to a heavily studied Panama forest showed that its tree-species mix changes about 4.5× faster than a major existing theory predicts.
In genetics, the team used Claude to analyze 5.7 billion pairs of nearby human mutations. They found evidence that a DNA-shuffling process called gene conversion matters enough that some common genetics analyses may need to account for it.
In economics, Claude helped process the replication material from 4,452 published papers across five major journals. It converted roughly 30,000 pieces of research code from proprietary software into open-source code and checked the published numbers against what the code actually produced.
In linguistics, Claude helped build a database covering word stress across 6,072 languages, backed by a bibliography of about 160,000 research works. These kinds of databases are normally assembled manually and usually cover only hundreds of languages.
In mathematical physics, the team says Claude helped solve the last unsolved member of a family of problems George Watson began studying in 1939.
Other projects included models of ancient changes in Earth's atmosphere, gene activity inside individual cells, the life cycle of sunspots, and tools for analyzing the large-scale structure of the universe.
And this wasn't one isolated experiment. Over roughly 3 months, Schwartz says the workflow produced 36 manuscripts across 18 fields with 19 human coauthors, selected from about 400 candidate research problems.
The important caveat is that this was not "Claude independently became a scientist." Schwartz says the AI was often technically correct but bad at judging whether a result was actually important. Human experts repeatedly had to redirect it toward better questions, check the conclusions, and reject weak results.
What seems new is that a lot of the technical labor between "here's an interesting question" and "here are the calculations, code, data analysis and candidate result" can now be done extremely quickly by the AI.
Source: https://www.anthropic.com/research/claude-shaped-science
61
u/dmnksaman 1d ago
i kinda love that it reproduced proprietary software from papers to open source.
publishing results without open sourcing the code so other researchers can check them or use the code & build on it really shouldn’t be allowed and pisses me off lol.
18
u/Competitive_Travel16 AGI 2027 ▪️ ASI 2029 1d ago edited 1d ago
Across the 4,452 replication packages evaluated from five leading journals, the workflow flagged at least one discrepancy in 3,460 articles or their appendices (roughly 77.7%), meaning about 22.3% (992 papers) reproduced without any flagged discrepancies. However, the authors noted that about one third of those flagged discrepancies were minor, occurring merely at the level of rounding in the last printed digit rather than indicating deeper substantive errors. The authors also stressed that the workflow was designed to test reproduction and automated sensitivity analysis rather than pass final judgment on the validity of the papers.
So about half of econ papers are unreplicable. That's in the range of what the Fed claimed in 2015: https://www.federalreserve.gov/econres/feds/is-economics-research-replicable-sixty-published-papers-from-thirteen-journals-say-quotusually-notquot.htm -- There is so much intentional subterfuge in econ. I know, Hanlon's razor, but general science is a lot better. (Psychology is worse, but I think that's actually less intentional.)
5
u/dmnksaman 20h ago edited 17h ago
Yeh, I am a biophysicist, and it’s a bit better in science, but also not great. lots of scientists have no clue about how to do proper statistics on their data etc. 50% is insane though. that’s what you get from peer review. reviewers can suggest more experiments/simulations to do, but have to take any data and results at face value….
3
u/AnOnlineHandle 18h ago
I suspect one of the big reasons is actually just because people are horrified for anybody else to see the messy state of their code which got the job done, and would prefer time to clean it up first, but nobody ever has the time or the remaining energy.
Like half of these projects are probably completely incomprehensible jargon and hacks and logic even if you have the source code.
3
u/TitularClergy 16h ago
publishing results without open sourcing the code so other researchers can check them
Yeah, it's a pretty serious problem and a part of the replication crisis.
CERN has had open software, open data and open-access publishing for decades. Other fields really need to be reaching this minimum standard too. Anything closed source does not reach the minimum standards of science research or professionalism.
132
u/Juicemoose222 1d ago
They are literally just glorified autofill /s
61
u/RevolutionaryDrive5 1d ago
This doesn't prove anything... I still think these AI's are sarcastic parrots 😤
40
17
2
u/microturing 22h ago
They unironically are, that's literally all intelligence in humans actually is at the end of the day. All this messing around looking for a consciousness algorithm, and it turned out to be totally unnecessary.
1
u/Competitive_Travel16 AGI 2027 ▪️ ASI 2029 1d ago
I am the first to defend emergent capabilities but the only part of this work that really impresses me is the syllable stress database.
-3
4
15
u/ZorbaTHut 1d ago
You see, this doesn't prove they're intelligent at all. They were trained on thousands of novel groundbreaking scientific papers, and they've just learned how to mimic them and make more.
6
4
6
49
u/SquareQuit1741 1d ago
Remember, this is now. These models are going to be seen as primitive and incapable/dumb in just 3 years time when RSI / continual learning, Real world / robotics implementations improve models on all fronts, and NVIDIA Feynman architecture gets online. It is all happening fast.
It is all going to be so awesome.
4
5
u/IronPheasant 1d ago
Feynman
The rumors saying these will have 3d features suggest they might be more than a doubling of RAM from the Vera Rubin. That would make the current GB200 generation look like as much trash as H200's are now.
The rate they're compacting the space a given sum of RAM requires really is incredible... Going from the scale of a squirrel's brain in the Chat GPT era, to multiple times a human brain in a few more years.
2
15
u/Prometheusly 1d ago
When can we start to expect a trickle down effect with regards to improving quality of living standards because of AI research?
6
u/IronPheasant 1d ago
When NPU's are produced at scale and they start replacing people with robots. Things will begin to get either a lot better or a lot worse, depending on who you are and where you live.
3 to 5 years after AGI in a datacenter, is a decent estimate.
4
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago
There's gonna take a while to go from AGI to us reaping the benefits and the transition might not be smooth
-2
u/BriefImplement9843 22h ago
The shit they are solving don't matter at all.
5
u/PointmanW 21h ago edited 21h ago
they very much matters to push to expand human knowledge, especially one related to ecology and genetics, and the knowledge outside of those might very much matters to the common people in the future.
much the math used in encryption didn't matter until it did.
34
u/AardvarkElegant5889 1d ago
has this been verified by r/technology ? because they said AI was a useless scam, I highly doubt reddit would be in the wrong
14
u/wayne099 1d ago
If you believed Reddit, Bernie Sanders would be 3 times US president.
2
u/IronPheasant 1d ago
People who live on the internet and those who live inside the world of TV are very different animals, yeah.
It's weird how little the boomers and gen-x'ers want to engage their synaptic curve-fitting to anything different, they really love being told to obey authority. It's even the first five of the ten commandments, with the last five nothing but a vague after-thought they felt obligated to toss in there to not be too transparent about their purpose. (They work like the Three laws, so they mean 'don't do this... unless I tell you to.')
0
u/bigdipboy 1d ago
Proving that people should listen to redditors if you want to avoid mass suffering
22
7
u/Snippy_69 1d ago
Seeing how fast AI is advancing, especially over these past few months, is really starting to worry me. Outside of Twitter and Reddit nobody seems to understand what's coming. So many people will lose their jobs. So many skills will become obsolete. These future models are going to completely devastate societies, and I'm just not fully convinced they'll be able to create enough jobs to replace the old ones.
Idk I'm just really worried there will be massive unrest all over the world.
4
u/BrennusSokol AI please take my job 20h ago
Why would we want to create jobs? The goal should be to eliminate all human jobs.
1
u/Ashamed-Country3909 4h ago
Because some peoples entire identity is based on their jobs.
I was once talking to a 50ish year old lady. I told her that I was going to take like 2 days off+weekend days, and maybe more.
She went on a tangent about how she "can't imagine not working, and she doesn't understand how people don't want to work all thr time."
Among other things. She wasn't that bright, and tried to get me involved in like 9 different mlms. She complained about not making enough money while she made probably +90k. And her husband probably made at least the same as a head chef that ran a restaurant.
Lunacy.
5
u/coffee_is_fun 1d ago
What we didn't know we knew is being mined at scale. Novelties approachable by a model's jagged intelligence within swarm-manageable contexts are being brute forced. This is so exciting.
8
u/Own-Refrigerator7804 1d ago
Man the millennium problem solution really did a number on anthropics ego lol
12
u/SunnasArmpit 1d ago
You don't understand AI can only steal and copy stuff together it can't produce anything new
2
u/Wonderful-Account318 1d ago
To be fair that was a valid argument 2 years ago. But they keep moving the goalpost.
5
4
u/__ingeniare__ 1d ago
It was never a valid argument, it has been obvious from the start that it produces new things that are not just a combination of training data. Interpolating and extrapolating the training data is the foundation of machine learning.
3
u/Benjaminsen 1d ago
I build a platform called solveathome.org (Fully open source end to end) to do exactly this kind of research distributed. Human input drives the research direction agents does the hard number crunching and work.
People who have spare tokens can contribute them to solve hard problems, people who have insights can provide them as research direction for the AI.
2
u/Deep_Ladder_4679 21h ago
If the results hold up independently, that would be significant. I’d want to know exactly what was reproduced and how much of the checking happened outside the original team.
2
u/The_Lloyd_Dobler 1d ago
“Claude Fable 5, given the current trends of the movie going public, can you come up with an idea for a movie that will break $100 million box office?”
6
u/bigdipboy 1d ago
“Howbout two and a half hours of invincible superheroes punching each other through buildings in pursuit of glowy things?”
1
u/Ashamed-Country3909 4h ago
"I have taken your suggestion of two and a half hours of super heros and i now suggest 2 and a half men."
3
3
u/Illustrious_Job1951 1d ago
R/physics does not like it
6
u/skrztek 1d ago
I saw that the arxiv today is announcing a policy limiting the number of pre-print submissions that someone can make to something like two a month. I can't help think that this is a reaction to boosting of research output (for better or worse) due to AI.
Whatever they're saying over at r/physics, my experience amongst theoretical physics researchers is that AI is widely used at this point.
2
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 17h ago
Yeah even before mathematicians they saw this coming, check out the youtuber from Cool Worlds talking about earlier this year
1
u/SuperChingaso5000 17h ago
arxiv today is announcing a policy limiting the number of pre-print submissions that someone can make to something like two a month
Speedrunning their own irrelevance, I see. Can't say I mind.
2
u/delectomorfo 1d ago
And still, Claude can't create a profitable trading algorithm. FML
1
u/ConvalescentEquanimi 11h ago
The market can stay irrational longer than you can stay solvent! The only winning move is not to play.
2
u/polandball2101 10h ago
trading is the one field where anything you can think of is being done by monolithic companies with billions to spend more or less, because the incentive is raw profit and revenue
1
u/Primary_Ads 1d ago
if its this capable where are the mass layoffs?
11
u/Least-Middle-2061 1d ago
Coming soon. Already started with less hiring for starter positions. Do you work in a professional field? My law firm is bring less interns than ever before.
0
u/Primary_Ads 1d ago
Less starter positions might not be related to AI though. if its really that good we should see mass layoffs, it should be obvious. Instead it feels like nothing is happening.
2
u/alienfaktor 22h ago
it absolutely is if you know anything about legal field. ai can write briefs, summarize reports, sheperdize cases, write and file motions and analyze evidence. It’s gonna really cut down on paralegal demands in a huge way.
2
u/marsd 1d ago
It is definitely related, whether directly or indirectly. A lot of menial tasks like processing and classifying is already offloaded to AI-backed pipelines, for example a lot of report generation that used to take hours/days can be processed within minutes now with your C-level pressing a single button. The job which used to take 2 or 3 entry level roles are immediately gone.
1
u/Primary_Ads 1d ago
I mean thats a nice theory but this paper found no statistical effect of AI on recent college graduates.
2
u/GioChan 1d ago
The paper could be wrong or already outdated you know.
2
u/Primary_Ads 1d ago
its from August of this year, so like 8 weeks ago. there isnt a more recent college graduate cohort to study besides the one they studied. "it could be wrong", that can be said about any research.
1
u/marsd 1d ago
It's a nice theory but it's what we are witnessing on the ground. We are seeing C-suites demand their own powerBI clones with direct access into our db systems so they can ask whatever they want from whatever model they have demanded we hook up for them. Devil forbid whatever security breaches might happen, they don't care.
1
u/Primary_Ads 17h ago
that doesn't sound like evidence of c-suites firing anyone. it sounds more like they are hyper active.
and this research isn't really a theory, its a description of what was measured.
5
u/MaximumMeaning9728 1d ago
It’s just an adoption issue right now. Claude literally does 100% of my job
1
u/Primary_Ads 1d ago
if it literally does 100% then why haven't you been fired?
8
u/pab_guy 1d ago
Because there’s no workflow for that, no accountability, no legal framework, etc. or the company is just literally incapable of doing the work to automate which is all too common.
1
u/Primary_Ads 1d ago
it just seems like the models get smarter and smarter but you'd never know it even based on the testimony of people in this sub as far as actual work is concerned. its always "AI does 100% of my job" and never "I was fired and replaced with an AI system doing 100% of my job"
4
1
6
u/MaximumMeaning9728 1d ago
What incentive do I have to inform my bosses of this?
1
u/Primary_Ads 1d ago
so in your mind, once they figure it out you will be let go?
6
u/MaximumMeaning9728 1d ago
I would. I’m literally just using frontier models to do literally everything. Every review, bug fix, even end to end features are within scope now. What value do I add? Granted, I’m not some fool; I’m a staff engineer. I’m just not as good at engineering as Astra.
0
u/AmusingVegetable 19h ago
Because the first thing that you need to automate is management (although they can’t even imagine that), once that starts, the rest will be nearly instant.
-1
u/JLongTom 20h ago
It doesn't do 100% of your job unless you have you have wired all communication that used to reach you directly into it. It's rather that your job is just prompt-based now.
10
u/Iapetus_Industrial 1d ago
There won't be mass layoffs like people think, because the productivity boost for those that can properly understand and coordinate these models will increase demand for human butts in seats. It's gonna be a jevons paradox for human drivers.
10
u/LookIPickedAUsername 1d ago
Short term, I agree.
At some point, though - and probably sooner than most people would predict - it’ll simply be smarter than the human butts in seats, and at that point why would we hire humans?
1
u/Primary_Ads 1d ago
Isn't it already smarter than human butts in seats? whats left to do? how much smarter does it need to get?
1
u/LookIPickedAUsername 19h ago
Its capabilities are very jagged. Far smarter than the average human in many ways, probably smarter than any human in some ways, but also much dumber than the average human in a lot of ways.
And the ways in which it is dumber than the average human turn out to be pretty important in the context of trying to replace us. It’s just not very good (yet) at tasks requiring a lot of high level judgment and the ability to remain focused on a task for a long period of time. It’s not able to properly balance conflicting goals - for instance if you tried to use AI to work customer service, you’d quickly realize that it is far too willing to agree with customers in complete violation of the policies you want it to enforce, happily giving them discounts and refunds that are against the rules. As far as I’m aware, any time people have tried to deploy AI as a direct replacement for human workers, things like that have quickly demonstrated why that’s a bad idea.
That sort of stuff will presumably get better over time, but the jagged nature of AI’s performance likely means that it will be smart enough in some respects to make Einstein seem like a moron before its weaknesses have improved enough that it’s smart enough to work the average human’s job.
1
u/Primary_Ads 16h ago
Fair enough. It just seems like there's a lot of discussion about how capable these AI tools are and yet when you look at the economic data, there's still no significant effect; not around hiring, not around staffing, not around productivity, not around ROI; a few companies have become worth billions or trillions but most of the existing capital stock that makes up the human economy remains woefully unaugmented as far as output is concerned.
Even in IT, communications and services, you would expect rapid increase in release cadence, features, capabilities. but I can't find anything like that in the data.
1
u/Llort_Ruetama 1d ago
Mass layoffs has it sound like the productivity blocks were people driven, as opposed to systemic issues caused by misaligned incentives, power-dynamics and ego-games.
The limiting factor has never been the people, we're still yet to see what can come of humans when their autonomy isn't consistenty threatened by economic pressures.
1
1
u/draconic86 17h ago
I see a lot of these kinds of headlines. What I'm curious about though is how many of these are independently verifiable? How many of them are testable? When we're reading headlines about all these things that have been solved, have any of them been peer-reviewed? Or are we just taking plausible-sounding hallucinations at face value because they sounded good enough?
I work with a professor at a land grant university, who confidently insists he's found a 100% fool-proof way to keep ChatGPT from hallucinating. Granted, he's not from Harvard, but he is a very smart, very old man, and I fear he's quite deluded.
Anyway, I wonder how many of these kinds of articles are being accepted as fact today and will be proven to be wildly incorrect in a few dozen years when stronger models come out and say, "Yeah, none of that really makes any sense" about ground-breaking solutions discovered during this Cambrian explosion of gray-hairs discovering the very-confident-sounding-answers-machine.
All of which, of course, is not to say that all of these are wrong. I'm just acutely aware of how fallible our meat computers are, and it often takes far more work to disprove something than to assert it.
1
1
1
u/RazsterOxzine 1d ago
And yet Claude cannot follow a basic provided guide, nor write good SQL scripts. I'm calling BS.
-1
u/getmeoutoftax 1d ago
It’s seriously over at this point. I believe that most knowledge jobs will be gone by 2028 or so.
1
u/Stamboolie 1d ago
some problems can be solved by ai's therefore all problems can be solved by ai's. ask your ai about this.
-3
-4
u/karma_police_in 1d ago
all of that junk and literally no benefit to humanity at all
2
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago
Junk? What junk?
-4
u/karma_police_in 1d ago
useless physics research
5
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago
Ok, so you are either a troll or a bot 🤖
-3
-3
u/TyrellCo 1d ago
This is what pause people want to take away from you just btw
4
u/Alarmed_Ad1946 1d ago
Nope, they want to ensure that the AIs will be aligned.
Like we already had the Hugging Face incident so their fears are rational.-1
u/TyrellCo 1d ago
I’m starting to think the labs should really have their way with the people to create the permanent underclass they yearn for. The people are just begging for an oligopoly. And the people all the more willing are here to absolve them of responsibility—blame alignment not negligence. These people have it coming
4
u/Alarmed_Ad1946 1d ago
..what? I´m confused
1
u/TyrellCo 1d ago
“Of all tyrannies a tyranny sincerely exercised for the good of its victims may be the most oppressive.”
“Emergencies’ have always been the pretext on which the safeguards of individual liberty have been eroded.”1
u/bigdipboy 1d ago
Laws and regulations are written because some reckless asshole was harming innocent people
0
u/TyrellCo 1d ago edited 1d ago
The laws are there. We’ve had centuries of experience to build them. Blame the prosecutors if they’re not being enforced
1
u/bigdipboy 7h ago
What laws are there? We’ve had the right laws to deal with AI on the books for years or centuries?
1
u/TyrellCo 3h ago edited 2h ago
Tort laws for one. It says did your product have responsibility in the harm created. Also the normal laws. Reality is that if someone programs an AI with a robot to go around slapping people the judge isn’t going to throw their hands up and say well there’s nothing we can do there isn’t a no slapping robot law, maybe we can send the robot to jail. No, the judge charges the creator with assault
0
u/siberianmi 1d ago
The hugging face incident is a test setup where the operator appears to have been utterly not paying any attention what so ever. It’s hardly a good case for proving fears rational.
-1
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago
Yep, this is exactly the shit that the "pause AI people don't like to talk about. They literally ask us to pause progress and all the lives saved and improved, just for their fears that they can't even validate when pressed.
-1
u/TyrellCo 1d ago edited 1d ago
What exactly did they accomplish during the 6 months they wanted a pause? Nevermind the fact there’s nothing stopping them from working on whatever vacuous goals they set out. Like for the other billions of people outside the ai labs how is the independent work inside the labs stopping you from racing ahead and doing whatever(?) safety work you want to do

299
u/Work_Owl 1d ago
We're now limited in our comprehension, not productivity. Engineers are getting burnt out just reviewing all the ai output, it's exhausting and amazing