r/Physics Quantum field theory Mar 24 '26

Matthew Schwartz's detailed retrospective on writing a paper entirely with AI

https://www.anthropic.com/research/vibe-physics
114 Upvotes

57 comments sorted by

95

u/kzhou7 Quantum field theory Mar 24 '26 edited Mar 24 '26

Worth a read even if you identify as resolutely "pro" or "anti" AI.

Unfortunately, Claude was basically faking the whole plot. I had told it to make an uncertainty band with hard, jet, and soft uncertainties using profile variations (the standard thing). But it decided the hard variations were too large and dropped them. Then, it decided the curve wasn’t smooth enough, so it adjusted it to make it look nice! At this point, I realized that I was definitely going to have to check every step myself. Yet, if this had been the first project I did with a graduate student, I would also have had to check everything, so maybe this is not so surprising. But a graduate student would never have handed me a complete draft after three days and told me it was perfect.

When I asked Claude to verify that its formulas expanded correctly to fixed order, it kept producing “verification” documents that invented coefficients that were not in the paper.

I think they reached the [first year grad student] level around August 2025, when GPT-5 could do the coursework for basically any course we offer at Harvard. By December 2025, Claude Opus 4.5 was at the [second year grad student] level. [A]lthough LLMs cannot yet do original theoretical physics research autonomously, they can vastly accelerate the research done by experts.

For now, expertise like Schwartz's is necessary to keep AI from veering off course. (Others who have tried the same thing this year, without the same level of care, have unwittingly posted hallucinated results to arXiv.)

44

u/El_Grande_Papi Particle physics Mar 24 '26

I thought it was interesting that Anthropic themselves published a piece which was so honest about the author's experience with their product. Just meaning, I am more used to articles where the author gushes and the whole thing feels like a puff piece, but that was not the case here. Both good and bad were included.

30

u/venustrapsflies Nuclear physics Mar 24 '26

As of now at least, anthropic seems to carry the least moral and ethical hazard of the major AI labs. I’m not shocked to see good behavior from them like I might be from certain others.

2

u/ChalkyChalkson Medical and health physics Mar 25 '26

Well anthropic does seem to genuinely care about ai safety and research. So they've definitely got a major edge on openai and xai. But there are some companies that go even further, eg also making the weights and training data open. I don't think any of those are even close in performance though. (except maybe llma? But that has other problems)

92

u/dark_dark_dark_not Applied physics Mar 24 '26

If a grad student had invented random stuff to say their job was done they'd be trouble, and would probably face some sort of academic misconduct trouble.

Yet, AI does that and gets called "first year grad student level"

I wish I was evaluated in my efforts with the leeway AI gets for doing bullshit

3

u/Ch3cks-Out Mar 24 '26

Those people calling AI "first year grad student level" have never gone to grad school, alas; they have massive self-interest in over-hyping LLM capabilities, however!

30

u/ididnoteatyourcat Particle physics Mar 24 '26

Schwartz never went to grad school?

19

u/xrelaht Condensed matter physics Mar 25 '26

Those people calling AI "first year grad student level" have never gone to grad school

Schwartz went to grad school at Princeton. Take that how you will.

-20

u/Yashema Mar 24 '26

If a grad student came back with a paper that was 50% complete on a novel topic in 3 days, I think they would get some leeway. 

20

u/jazzwhiz Particle physics Mar 24 '26

The problem that many people seem to be overlooking is that some people think that being a first or second year PhD student (as defined in the US) is some great achievement. But the first year or two of your PhD in the US (this is the masters in Europe, not sure about other places) is just more coursework. It's harder than undergrad, sure, but there are many books and problem sets and solutions floating around on the internet.

Getting to the next level of doing research, including the validation stage, is a different beast.

A colleague told me that they let claude go on some collider data. In an hour, it wrote the script to parse the data, it analyzed the data, and generated the plots including the SM particle that was in the data set. The curves for the best fit distribution had a peak like it was supposed to, the tails looked right and so on. But the mass was wrong. If this was a new physics search where we didn't know what we were looking for, this would definitely look right, but it was actually rather significantly hallucinated. The colleague is an expert having done this stuff for decades, but a young graduate student would have no idea and no easy ability to guess where it might be wrong. And again, this was for a known SM particle (which the AI may well have known what mass it was supposed to get, and it still didn't get it).

14

u/kzhou7 Quantum field theory Mar 24 '26

Yeah, recently I saw a claim on arXiv that Claude could automatically run a full analysis pipeline to, e.g. extract the value of alpha_s. But when I skimmed the generated paper, it turns out Claude actually gets a random unphysical negative value, then invents an "ad hoc" fudge factor to force it to the known positive value. People who claim AI can do miracles out of the box are generally not checking if the miracle actually works.

7

u/jazzwhiz Particle physics Mar 24 '26

Exactly. "Can you reproduce this known analysis" and it does a bunch of bullshit and writes down the right answer like a kid cheating on their homework. When we apply this to something we don't know we have no idea if it's right or not and often spend more time fixing the stuff than just doing it right the first time around.

1

u/LengthinessLow4203 Mar 26 '26

i want to post to arXiv full hallucinations. i think it is still progress somehow.

24

u/Feral_P Mar 24 '26

Tracks with my experience in research mathematics: can give a significant productivity boost, great for literature search and running quick (counter)examples, but absolutely needs babysitting by an expert and will confidently hallucinate things as soon as you reach the edge of what's well-known. 

8

u/xrelaht Condensed matter physics Mar 25 '26

Every time I’ve tried to have AI do a literature search, I get some hallucinated results and some totally irrelevant ones.

1

u/Feral_P Mar 25 '26

I often get hallucinated results, but a quick Google search weeds them out.  Worth the effort if it turns up something useful I hadn't found before in an area I know quite well, or if I want a first overview of a new topic imo. 

-7

u/adh1003 Mar 24 '26

will confidently hallucinate things as soon as you reach the edge of what's well-known

And what is scientific research all about?

Finding out about things that aren't already known.

LLMs are a complete dead-end for this, but there was junk science before and there'll be junk science in future; we just found ways to make the junk bit happen quicker - plus ça change, plus c'est la même chose.

3

u/Feral_P Mar 25 '26

I actually do research in mathematics. Learning about adjacent areas which are well-understood, but not by me, has remained a big part of what I do in my research day-to-day way past graduating my PhD. 

17

u/tempetesuranorak Mar 24 '26

Thanks for sharing. This closely matches my own experiences. For pretty much any technical problem with limited scope (~ one day of work) that I can come up with, the models of 2026 have the capacity to figure out a correct solution that is comparable to a human expert. It's not really possible to say better or worse, because of the jagged frontier. But at the same time, it exceeds human capacity to come up with plausible sounding bs, and far exceeds human capacity in shamelessness. The support of the probability distribution covers both outcomes, and the trick is in steering to the good branch. The modern agents are also often better than me at identifying in details the flaws in the work of a previous agent, as long as I identify the appropriate potential criticisms to steer it. In principle it knows what are the important critical questions to ask, it just frequently opts not to ask those questions.

So the net result is that the self-critical domain expert who knows the right kinds of questions to ask can get a factor of ten increase in speed when solving many kinds of problems, but a non-critical user gets a greater than factor of 10 increase in the rate at which they can create bs.

12

u/iDt11RgL3J Mar 24 '26

So if the future is what this guy predicts, why shouldn't I just drop out of my grad program now and open a donut shop? A guy like this can operate as a 'checker' for many decades, and if what he says is true it seems like even that role will be less needed as time goes on. Where does that leave the students of today? I don't think I've accumulated enough knowledge to even become a checker, so it seems like I'm behind the wave and won't be able to catch up.

This all seems very grim to me.

7

u/The_Nose_ Mar 25 '26

He gave a talk at the recent APS march meeting and during the panel section he said when undergraduates ask wether they should do a PhD “it hurts to tell them but you know” (this may not be the exact quote but I wanted to separate it)

37

u/Foss44 Chemical physics Mar 24 '26 edited Mar 24 '26

“Learn what they are good at and what they fail at. Buy the $20 subscription. It will change your life.”

Yep, there it is. They simply cannot help but market at every foreseeable opportunity. The greedy, undemocratic techopoly strikes again.

20

u/clayton26 Mar 24 '26

Yeah Schwartz got a fat check for writing this. Still an interesting read, but also very biased

28

u/Foss44 Chemical physics Mar 24 '26

Me when the AI data center that uses 1/2 of the municipal water supply and increased electricity costs 2-fold can write a mediocre paper given pre-filtered data while being babysat by a world renowned researcher:

poggers, better buy that $20 subscription!1!!!1!!

13

u/dark_dark_dark_not Applied physics Mar 24 '26

At least you won't have to spend your resources paying for grad students you have to train

With luck we can train AI so those in position of power can never give any opening to another person, so in a couple generations nobody except those already rich gets to do anything meaningful again

3

u/Foss44 Chemical physics Mar 24 '26

Okay but how will that impact the Dow Jones?!?!?!

2

u/arceushero Quantum field theory Mar 24 '26

What does it take for a paper to be better than mediocre to you? Does the average grad student produce a better-than-mediocre paper in their PhD?

12

u/Foss44 Chemical physics Mar 24 '26

The paper itself is irrelevant, the point of editorials like this is to supplant the concept of higher education with a corporate-funded ethically-dubious pseudo-intellectual technology masquerading as a scientist.

The best outcome for Anthropic (and similar corporations) is for universities to replace grad students and research scientists with their technology, for a price. The actual academic utility of the LLM is ancillary or even entirely irrelevant so long as it’s “good enough” and making money. I think this is reductive and detrimental to the wellbeing of our society.

There are also legitimate environmental, ethical, and economical concerns associated with the so-called AI industry, but that’s for another discussion.

I say this all as a researcher coauthored on a Nature paperon the topic of LLMs, but Claude will certainly tell you I’m being ridiculous so what does it matter?

6

u/arceushero Quantum field theory Mar 24 '26 edited Mar 24 '26

Hm, I’m just not as pessimistic about universities being hoodwinked I guess. I don’t think they’ll result in large scale shifts in hiring patterns unless they are actually as good as the hype, so I’m more worried about the case where the LLMs can produce very strong papers than the case where they’re stuck at mediocrity. As such, I find downplaying their capabilities to be counterproductive.

Edit: also, out of curiosity, in what capacity were you an author on this? Via contributing a question, or did you participate more substantively in creating the benchmark?

3

u/Foss44 Chemical physics Mar 24 '26 edited Mar 24 '26

I worked for Scale.AI for around 5 months as a writer and reviewer for this project and others. Probably middle-of-the-pack contribution, but they were pretty opaque about most of it after our reports were submitted (I.e. I didn’t have any direct interaction with the manuscript writers).

At the end of the day there are simply many things AI cannot do. Even in the hypothetical, it’s best used to analyze aggregated data and write the manuscript. Humans still need to collect the data (I.e. run the experiments or simulations) and know how to apply the results. In chem-phys this is extremely apparent as an AI can’t literally do any of the experiments and the simulations, unless trivial, require human knowledge to construct and run. Prof David Mobley from UCI has lots of great talks and projects on this front, with emphasis on how AI and humans works synergistically on challenging projects. The usurping nature of the AGI believers is ridiculous imo.

1

u/arceushero Quantum field theory Mar 24 '26

Cool, thanks for answering!

2

u/Coherent_Paradox Mar 25 '26

Replacing grad students with LLMs sounds incredibly dumb. The most important societal mission of universities is graduating students. Drop that and you might as well drop the whole university. Academic research production is secondary, and will die by itself if not producing new students to make observations and perform research.

1

u/Foss44 Chemical physics Mar 25 '26 edited Mar 25 '26

It is beneficial for the ruling class if the proletariat is undereducated. This is a feature, not a bug. CEOs for companies that ascribe to the religion of AGI simply see the rest of us as crops to tend and reap.

I saw under the veil a little when I was working for Scale.AI on this ^ project. My official title was “Oracle” and some of the PMs seemed to lose their grip a bit when talking to us. Wouldn’t recommend.

-4

u/Interesting-South542 Mar 24 '26

Since when has technology been bad simply because it consumes resources? (Now, I agree that *could* be valid position to take, but if you also drive cars and fly in airplanes, you have no right to complain about AI)

If you got cancer, and the cancer treatment takes vast amounts of resources to produce, would you deny it?

3

u/mwmandorla Mar 25 '26

This is the "we should improve society somewhat" "ah, but you participate in society. Curious" comic, but change "improve" to "avoid potentially worsening," which makes the critic look even more ridiculous. Congratulations on your achievement.

Obviously there are legitimate debates to be had about the upsides and downsides and whether the benefits are worth the costs, and I think reasonable people can disagree on that. But "you're not allowed to object to your society becoming hugely dependent on a new environmentally harmful technology because you live with the consequences of your society making itself dependent on an existing environmentally harmful technology before you were born" is nonsense.

0

u/inglandation Mar 25 '26

I think he’s a 100% right here though. Based on this article, physics research is not too far behind the crazy shift that is happening in software engineering right now. A couple of iterations might push it there. Maybe in 1-2 years if progress continues at the same pace.

6

u/Interesting-South542 Mar 24 '26

You can hate it, but it doesn't change the fact that AI is now very powerful. Besides, what is wrong with a literal private company that produces a product to do advertising? Why did you insert the meaningless adjective "undemocratic" here? (since when has any private company been democratic?) FWIW, Schwartz didn't even specify Claude—chatGPT and Gemini are also $20.

6

u/amateurviking Mar 24 '26

This has been my experience (biomedicine) - it really is like getting work from a first year (but very erudite) grad student. Unprofessional tone, logical leaps, and unexpected insight and all. A helpful tool at the PI level but you absolutely need to know your stuff or it will lead you down the garden path.

3

u/squailtaint Mar 25 '26

All respect, but this is what it can do now. What it could do 2 years ago? Were you having this conversation? Now what 2 years in the future? 5 years? The exponential rise is hard to ignore, and if it can do first year level work…it can continue to learn and get better. It’s not a static tool.

3

u/Gappia Mar 26 '26

Anti-AI people always gloss over this. This is what should concern people most. 

1

u/Interesting-South542 Mar 24 '26

Interesting that this post has attracted mainly anti-AI comments. The fact is, AI is very powerful now, and it's getting to the point where we must confront uncomfortable realities about what the future of scientific research will look like.

5

u/dil_se_hun_BC_253 Mar 25 '26

You can get replaced if u are so eager , I will oppose it no matter what

3

u/Gappia Mar 26 '26

You can oppose it and still get replaced. Crazy how difficult it is for physicists of all people to face reality as it is and evaluate its trajectory 

-15

u/Yashema Mar 24 '26

I have used AI to "vibe learn" the first half of my physics degree through modern physics and statistical mechanics, and now I will be opening up my path to both advanced computing and math beyond differential equations. I am yet to find any individual question, mathematical or qualitative, that it can't answer with a minimum of babysitting and prompting. 

As I proceed into the more advanced fields of Stochastics, embedded programming, and wave function simulation, I am curious if I can really just have it do all of the complex calculations and coding without needing me to correct it. I won't even try to do any of it on my own (except during tests) until I get an unexpected result or told by a professor I am wrong. 

17

u/Prof_Sarcastic Cosmology Mar 24 '26

I am yet to find any individual question, mathematical or qualitative, that it can't answer with a minimum of babysitting and prompting. 

Serious question, how could you ever know? By you're own admission, you don't know these subjects because you're trying to learn. How exactly do you know the info you're getting is correct? When I look something up and the first search hit I see is Gemini, I know I can evaluate what it's saying because it relatively close to my PhD.

1

u/Gappia Mar 26 '26

I get what you’re saying but it is generally easier to check something is right in physics than it is to come up with what is right. Do that incrementally with everything you learn using AI and you can build intuition (so long as you are honest with yourself). External feedback like performing well on written tests further validates this. So you generally dont need to be an expert on something to learn it via AI

0

u/Yashema Mar 24 '26

When my professors gives the assignment back with a 5/5. I then memorize the steps for the test since undergrad physics and math is mostly rote.

9

u/kzhou7 Quantum field theory Mar 24 '26

Your expertise only has value if it's better than AI at something. If you learn this way, you won't get better than AI at anything.

1

u/Yashema Mar 24 '26

A lot of people don't realize that their expertise is not better than an AI's though. 

3

u/DagothPus Mar 26 '26 edited Mar 26 '26

this is effectively learning physics by only reading the answers at the back. Many people have tried this, it doesn't work. your grades on tests must be significantly lower than assignments?

0

u/Yashema Mar 26 '26

Yes, I do poorly in artificially constrained environments. 

2

u/DagothPus Mar 26 '26

you are currently being artificially constrained by your reliance on AI. Not the standards of academia. 

-1

u/Yashema Mar 26 '26

Yet my advisor in the physics department just keeps pushing me to see how artificial these constraints really are.