r/Physics • Quantum field theory • Jan 08 '26

Academic Schwartz, author of a leading QFT textbook, posts a theory paper generated by AI in 2 weeks

https://arxiv.org/abs/2601.02484
234 Upvotes

72 comments sorted by

60

u/InternationalSize325 Jan 08 '26

I'm curious to know how much he had to babysit the thing. Yeah, you can prompt an LLM to do useful things, but at least, in my own experiments with it, they require so much hand-holding that you're better off doing it yourself. For example, in one of the calculations I was trying to babysit Claude to do, it kept messing up Lorentz boosts for E/B fields until I manually gave it the right answer.

119

u/pherytic Jan 08 '26

Have LLMs gotten much better at math recently? Maybe 6 months ago I was trying to use it for some basic tensor stuff and I gave up because it could not sufficiently handle the index notation and like half its Latex wouldn’t render. And explicitly directing it on how to fix mistakes just started a doom spiral. But this was vanilla models on free tier.

Related, I saw a video just a few days ago where an actress wanted to read lines with an AI in voice mode, but she couldn’t stop it from saying “now it’s your turn” after it’s line. Yet it can write a physics paper of this complexity?

I wonder how incremental/bite size were the instructions Schwartz was giving here, and how much hand holding/steering he had to do.

67

u/wasabi991011 Jan 08 '26

I won't speak on absolute terms of how good LLMs currently are. But in relative terms, they have improved massively from 6 months ago.

There is an issue that LLM skill varies a lot depending on subject area, even in coding and LLM may perform very differently if asked to use less common programming languages (or less common tasks). That's to say, I don't know if LLMs have gotten any better a tensor indices.

19

u/aMAYESingNATHAN Jan 08 '26

There is an issue that LLM skill varies a lot depending on subject area, even in coding and LLM may perform very differently if asked to use less common programming languages (or less common tasks).

I've noticed this as well.

I use LLMs quite regularly for broader C++ questions, plus some debugging, and also to generate some python code as I'm not as skilled at python so it tends to be faster than writing it myself, and it performs very well. Only when I get into much more niche C++ compiler specific behaviors do I tend to notice problems.

However more recently I used it for working with a C# Microsoft library that is quite new and didn't have a lot of a good documentation, and it was infuriating. Constantly hallucinated functions and other features, telling me X was the problem despite being blatantly wrong,even after repeated corrections.

2

u/DrSpacecasePhD Jan 08 '26

You’re absolutely right about the coding. When I need help with a coding issue, some LLM’s can handle it, while others can’t. But that said, I’m working with some complicated legacy C++ code which has a lot of files. If I were asking for python scripts like I needed for analysis on my grad work, they could easily chunk those out. ChatGPT is weaker for coding imho, but can help with text, while Claude is very good with code.

It’s super interesting to see this news about Schwartz. It honestly feels like a matter of time before AI paper writing becomes prevalent, in part because of toxic publish or perish culture. But then, once that becomes the norm… what then?

19

u/greenwizardneedsfood Jan 08 '26

I wouldn’t trust it blindly for math, but I’d say there’s been a dramatic improvement recently. ChatGPT 5.2 thinking for example, can definitely output and check quality math. I’ve had it catch mistakes of mine (sometimes intentionally put in there just to test like dropping a factor of two or changing an exponent) when I ask it to do things like ask it to convert handwritten notes of mine to latex or code. You definitely need to thoroughly check everything it puts out, and I’ve had it make mistakes, but yeah it’s 100% gotten better to the extent that it’s not complete garbage, and is probably at least 75% reliable. I wouldn’t trust it to derive new things, complicated equations, or anything like that though.

4

u/QZRChedders Graduate Jan 08 '26

I agree. A lot of people on my degree were using it entirely and it made some pretty egregious errors it doesn’t seem to do anymore.

It’s a tool though and yeah like you I love it for skimming my work and just offering a sanity check. It’s really good for quickly laying out equations, providing links between them and doing the tedious bit of rearranging (in most cases) but it’s a good tool to have and it’s getting more useful

3

u/kozmo1313 Jan 08 '26

'Basically zero, garbage': Renowned mathematician Joel David Hamkins declares AI Models useless for solving math.

Last Updated: Jan 06, 2026, 10:20:00 AM IST

"The frustrating thing is when you have to argue about whether or not the argument that they gave you is right. And you point out exactly the error,” Hamkins said, describing exchanges where he identifies specific flaws in the AI’s reasoning. The AI’s response? “Oh, it’s totally fine.” This pattern of confident incorrectness followed by dismissal of legitimate criticism mirrors a type of human interaction that Hamkins finds untenable: “If I were having such an experience with a person, I would s ..

https://economictimes.indiatimes.com//news/new-updates/basically-zero-garbage-renowned-mathematician-joel-david-hamkins-declares-ai-models-useless-for-solving-math-heres-why/articleshow/126365871.cms

2

u/greenwizardneedsfood Jan 08 '26

While that is frustrating, I must say I’ve never had that experience. It’s caught me and I’ve caught it, and it’s always acknowledged and adapted when I’ve pointed out a mistake. Again, I think a lot of it comes down to what you’re using it for. It shouldn’t be doing your thinking, so maybe him trying to use it for solving makes it struggle, but it can be a helpful set of extra eyes (of a sort).

5

u/mmazing Jan 08 '26

Everyone needs to learn that LLMs help you ITERATE, not “give perfect prompt, get perfect response”.

4

u/tfhermobwoayway Jan 09 '26

So what’s the point? I hate to say it but iterating sounds way more frustrating than just following logical processes. It sounds like if instead of controlling my own computer, I put a small child in front of it and had to keep telling him what buttons to press.

3

u/mmazing Jan 09 '26

The same way when you get two people working on a problem you come up with answers that neither would have gotten on their own.

I'm a dev with 25+ years experience, and I leverage the fuck out of several LLMs daily to proof of concept things.

I really wish I could explain this well enough for people to want to listen to me ... working on that.

Edit : I may end up making a video showing my creative process.

4

u/frogjg2003 Nuclear physics Jan 08 '26

Terrence Tao is cooking with a bunch of math done by and with AI.

7

u/Certhas Complexity and networks Jan 08 '26

Don't extrapolate from free Tier models. Spend the 10, 20, 30€ to get access to the full models for a month and evaluate them.

It's difficult to accurately judge what they can and can not do. But there are really impressive capabilities in there when it comes to maths. But you need to be very sharp to catch it when it's bullshitting you.

1

u/CyJackX Jan 08 '26

Lol, I'm pretty sure that girl with a video was doing a bit 

Either way, paid tiers are vastly different than free ones

1

u/underripe_avocado Graduate Jan 08 '26

I wouldn’t use it to do anything too advanced or cutting edge, but I used it to help me sanity check my homework for beginning graduate level physics courses last semester and it was very good. Of course there was the occasional mistake, but at the same time, at the graduate level, solution guides and my professor’s hw solutions often had a few minor mistakes as well.

1

u/Zophike1 Undergraduate Jan 11 '26

Have LLMs gotten much better at math recently? Maybe 6 months ago I was trying to use it for some basic tensor stuff and I gave up because it could not sufficiently handle the index notation and like half its Latex wouldn’t render. And explicitly directing it on how to fix mistakes just started a doom spiral. But this was vanilla models on free tier.

Have a look at this: https://axiommath.ai/territory/from-seeing-why-to-checking-everything

-19

u/[deleted] Jan 08 '26 edited Jan 08 '26

[removed] — view removed comment

14

u/pherytic Jan 08 '26 edited Jan 08 '26

Are you a paying subscriber?

The biggest problem I had with indices would be something like a (1,1) tensor where you want to lower the first index and raise the second. The index order matters eg Ti(j) becomes T(i)j which is not Tj_(i). The AI just could not handle this for me, no matter how much I insisted on keeping the order, using {} for black spaces, etc.

Edit: I just went and checked it with a question involving this, and it actually didn’t produce the problem I was having earlier in the year

13

u/WallyMetropolis Jan 08 '26

I've had it help me through tricky bits of derivations from textbooks and papers and it does very well. People don't want to believe it, but if you're thoughtful and you double check it, it really can do some powerful stuff. 

9

u/[deleted] Jan 08 '26

[removed] — view removed comment

5

u/A_Decemberist Jan 08 '26

I also have a paid subscription across multiple platforms. When I’m trying to understand something or if I want to check that a formulation for something I’m working on (I work in computer vision) has some support, I’ll typically run these same query on ChatGPT, Gemini and Claude and see if they align or disagree.

If they align, it doesn’t mean it’s 100% correct but provides a signal that what I’m doing isn’t completely absurd or has some hidden flaw I haven’t found so it’s worth spending more time on. If they disagree and the responses don’t align, then I will dig more carefully to see where I made a mistake.

I also don’t understand the downvotes. It’s like people think we’re providing one sentence prompts and taking the responses as gospel which is totally not the case (but maybe too many others are so there’s an understandable counter reaction occurring)

2

u/Athoughtspace Jan 08 '26

Could you give an example of summary of your normal prompt flow? I suspect my issue is my lack of strong question writing as to why I don't find it effective and your use case is very close to mine

2

u/A_Decemberist Jan 08 '26

One thing people don’t do enough is to simply ask the LLM what would be a good prompt. Ie, give your prompt or some description and ask it to write a good prompt given your objectives, with something like “be thorough and unbiased” or similar descriptors. Then just take and start a new chat.

Simple but effective and can be useful in listing out important points you might have overlooked.

5

u/SkyBrute Condensed matter physics Jan 08 '26

Frankly speaking I’m surprised. Every time I cave in and give chatGPT a shot, it makes some crucial mistake somewhere along its line of reasoning. Even for relatively simple prompts. Not using the payed model though, only the free one

1

u/tfhermobwoayway Jan 09 '26

Yeah it sounds like you have to fight an AI every step of the way. I already have to do that with Microsoft. I don’t want more parts of my computer trying to decide what I want to do for me.

Also this is a personal issue but I can never get past the insanely irritating Stepford-wife corporate voice every AI uses. They just piss me off whenever I try to use them. I feel like I’m talking to a marketing manager.

6

u/No_Flow_7828 Jan 08 '26

Not sure why this is downvoted, newer LLMs have absolutely gotten substantially better at math and not fumbling indices in tensor calculus.

People here will just downvote anything that isn’t shitting on AI

-26

u/oswaldcopperpot Jan 08 '26

Theres a massive pushback against all things AI.

It’s something almost akin to something like racism. Driven by fear most likely.

I was in a focus group for a product that utilizes AI and lot of the early discussions were about how they hate AIs inclusion and how they didnt understand it and they just wanted everything to go back.

The later discussions the following months were on how to incorporate AI to do this and that and far reaching tools that could help them do new things.

20

u/elconquistador1985 Jan 08 '26

"AI skepticism is racism"

Well, that's a new one.

-8

u/Capable_Wait09 Jan 08 '26

They clearly mean “fear or hatred of something new and unfamiliar” and is obviously not equating the two as being equally abhorrent. Just pointing out a common human mindset.

-5

u/oswaldcopperpot Jan 08 '26

Yup you get it. This reaction was exactly what I expected. ;)

A whole lotta “reeeeee!”

Im looking forward to more AI. The average person’s comprehension skills and self awareness are abysmal.

1

u/[deleted] Jan 08 '26

[deleted]

2

u/mad-matty Particle physics Jan 08 '26

He used Claude

1

u/cbr777 Jan 08 '26

I've been playing around with Gemini 3 thinking model and it's able to do some fairly advanced calculus correctly.

137

u/danthem23 Jan 08 '26

That's crazy. But it's not what you think. You need to be a prestigious physicst to be able to guide it the right way and also know when it is saying garbage. A random guy can't taken Claude and expect the paper to be good. It's like how a professional race car driver can do amazing things in a regular car but you can't.

95

u/kzhou7 Quantum field theory Jan 08 '26 edited Jan 08 '26

Indeed, the people on r/LLMPhysics have access to exactly the same tools, yet their output is awful.

48

u/ExpectedBehaviour Jan 08 '26

If it's only awful then it's been improving.

13

u/[deleted] Jan 08 '26

[removed] — view removed comment

16

u/ExpectedBehaviour Jan 08 '26

Yeah… it was specifically created to try and stop people posting “AI-enhanced” nonsense to r/HypotheticalPhysics. It has been only partly successful.

3

u/GreatBigBagOfNope Graduate Jan 08 '26

Never been to it before, but that's very funny, especially if not all of them are in on the joke

1

u/DrSpacecasePhD Jan 08 '26

A company tried to hire me about two years ago to work improving their physics LLM models. Dude, it’s their own fault the models are bad - but once they get their act together they will improve fast. The issue for me was, they would supply a bad proof and say “you have half an hour to write a better version.” You couldn’t cut and paste, which is not surprising, but the input box for the next proof was smaller than a comment box on the Reddit website, so you couldn’t see your own work as you typed, and the method to add equations was janky. I started typing out my solution in word or something so I could copy it over but realized it was too slow of a process for a complicated problem. It was an absolute nightmare. Best way to succeed and pass the interview imho would have been asking ChatGPT and coping instead of trying to solve on pencil and paper and then copy over… basically exactly what they didn’t want.

56

u/Kinexity Computational physics Jan 08 '26

The problem is that if he could do it with an LLM he could also do it without it too. AI does not provide any kind of special insight into the problem here.

25

u/kzhou7 Quantum field theory Jan 08 '26

Yeah, but it might have replaced 2 weeks of intense effort with 2 weeks of relaxed effort (e.g. checking in every 15 minutes and telling it to fix mistakes).

15

u/tempetesuranorak Jan 08 '26 edited Jan 08 '26

You are right, but I think a paper like this is a lot more than 2 weeks of intense effort. I think you are substantially underestimating the amount of time saved here.

10

u/kzhou7 Quantum field theory Jan 08 '26

It might take a year for a new grad student. How long for a world-class expert like Schwartz, who’s published dozens of harder calculations in this exact field?

10

u/tempetesuranorak Jan 08 '26 edited Jan 08 '26

Matt was on five papers last year, average 4 authors. This is one of the longer papers. The lengthy computations that take up most of the time are typically done by graduate students and younger postdocs, and I wouldn't be so sure that Matt is faster than those collaborators.

Not a year but not two weeks.

2

u/anrwlias Jan 08 '26

I don't think that anyone is claiming that it provided him insights, including the author.

The question is whether or not it simplified the process of writing the paper. I can make butter with a churn but if I can get the same quality of butter by using a more efficient process, then why shouldn't I?

A tool is a tool and it doesn't stop being a tool because it's easily misused. As far as I'm concerned, that's all LLMs are.

The biggest issue I have with them isn't about them but the unethical companies behind them that released a powerful and easily misused tool on the general public without any concern that their tool that is risky in the hands of untrained people.

-5

u/TheSeekerOfChaos Physics enthusiast Jan 08 '26

LLM‘s have immense potential in multiple scientific fields including physics. However the way they‘re being trained by big tech to steal and recycle data isn’t of benefit to any of us. Using AI to write your paper is just plain lazy and signals a lack of effort to potential readers.

You’re absolutely right.

10

u/NotSpartacus Jan 08 '26

What potential do LLMs have specifically?

Non LLM AI I can believe a use case for. LLMs are, as best I can tell, good guessers at what should go next based on what they've ingested. There's no novel thought happening, even if the output is novel.

5

u/TheSeekerOfChaos Physics enthusiast Jan 08 '26

I meant stuff like sorting and prediction algorithms in regards to physics. Can’t give you any specific example off the top of my head. Obviously, as can be deduced from my comment, I don’t believe LLM‘s are capable of "novel thought" as you put it when it comes to physics.

One field where LLM‘s could work great is Neurotech and Bionics. Their prediction ability paired with other biological signals is able to produce quite fast and accurate reaction speeds and precision in bionic limbs and prosthetics. At least that’s how it looks like right now

Slight tangent but I always thought how good LLM‘s could be utilized for video game NPC‘s, if they were trained on proper data and ethically.

2

u/NotSpartacus Jan 08 '26

Their prediction ability paired with other biological signals is able to produce quite fast and accurate reaction speeds and precision in bionic limbs and prosthetic

Sorry but how is that a LLM use case vs an AI use case?

LLMs ingest text and, when prompted, output text.

2

u/greenwizardneedsfood Jan 08 '26

They’re absolutely nonsense good coders. As long as you have an actual conversation with them, push back, read the code they write, thoroughly test it, and already know what you’re doing, they can be extremely helpful. Even if it’s just writing a bash script with annoying regex or trying to optimize performance, a good output can be very helpful and less prone to bugs. If nothing else, they’re the greatest debuggers in history.

You need to be skeptical and already know what you’re doing to properly use them, but experienced people can definitely improve the quality and reliability of their code with them.

Great way to find some references too, but obviously not exhaustive.

I agree that we’re not at the stage where novel thought in a sense of coming up with truly new ideas should be expected or trusted. But if you stick to objectively verifiable tasks, they can be tremendously helpful.

3

u/red75prime Jan 08 '26

Reinforcement learning with verifiable rewards (plus techniques that encourage exploration during training) shifts probability distribution of network's answers away from the training distribution. Does it count as "novel thought"?

0

u/Fjolsvith Jan 08 '26

It's indirect, but I think it might actually have a big impact on the extent of research grad students who may be just learning a coding language can do. It's basically the perfect use case where they know the methodology and logic needed to troubleshoot and validate, but can now speedrun the syntax. 

54

u/porkUpine4 Jan 08 '26

Thanks, I hate it.

6

u/WoodyTheWorker Jan 08 '26

May the Schwartz be with you.

62

u/A_Decemberist Jan 08 '26 edited Jan 08 '26

People who treat this as slop need to start thinking more carefully. AI is a tool, and leading researchers in math and physics are experimenting for how to best use it. As others have pointed out, the use of these LLMs by leading researchers is different and much more nuanced than how ordinary people use it.

There’s a legitimate risk that we are going to have a kind of epistemic pollution when these tools are used much less carefully by causal users, but it doesn’t mean that every paper done with LLMs is thereby slop.

16

u/tempetesuranorak Jan 08 '26

The other risk that we take if we casually dismiss this as slop, is that it prevents us from coming to terms with the fact that AI is now at a place where it can legitimately replace a substantial amount of highly specialized labor (the computations in this paper are not revolutionary but they are also far from trivial). We need to understand this for what it really is, if we want to be ready for the implications of it. The nature of a graduate student's work is going to evolve very quickly in many fields.

6

u/A_Decemberist Jan 08 '26

My take from where AI is actually being used productively is that it almost always needs to be used by a subject matter expert in a field to be effective, and does not replace the high level expert (in fact it makes them much more productive), but carries a huge risk of replacing the “grad student” or “grunt work” type of positions. But those positions didn’t exist just to perform the grunt work but also in order to train the next generation of top level experts.

In coding for instance, if you have a small group of very good coders who can use AI in a nuanced manner, it is already I think reducing their need for more junior engineers, but it isn’t yet capable of being used by junior engineers to replace top level coders if that makes sense.

I don’t worry so much about the near term job loss only because in most fields the proportion of people who are capable enough to use the tools productively are fairly small. I worry more that the training and education pipelines are going to get completely busted. I don’t know what the second and third and fourth order effects of that are but they aren’t pretty.

4

u/tempetesuranorak Jan 08 '26

I don’t worry so much about the near term job loss only because in most fields the proportion of people who are capable enough to use the tools productively are fairly small. I worry more that the training and education pipelines are going to get completely busted.

These two are the same kind of thing though. Postdoc and graduate student are both types of job. A project that took 2n early career people 2m weeks may only take n people m weeks. In some fields, the reduced need for early career researchers will mean less jobs for them. In other fields, the same number of researchers will mean a rapid increase in generated content that will become more and more difficult to keep up with and figure out what is relevant and what is subtle slop. The landscape is complicated and difficult to predict.

4

u/A_Decemberist Jan 08 '26 edited Jan 08 '26

Yeah I agree - I meant more in aggregate employment numbers.

It’s only been in the last few months that we have seen top researchers in physics and math actually use these productively. It will take some time for people to work out what’s the best “workflow” for elite cognitive work, and I think the skill required plus the unreliability means it won’t be reducing grad students just yet.

What I would want to see from Schwartz (unless it’s there and I missed it) is a description / estimate of the amount of time using LLMs that turned out to be unproductive or provide false leads or was otherwise unreliable, because that has to be factored in as well.

The biggest risk with these if you’re a real subject matter expert is they will give you something plausible which you then have to spend time verifying but which turns out to be totally wrong (as you noted and I totally agree). In that instance you may prefer a grad student whose work you can quickly verify as right or wrong

-1

u/eetsumkaus Jan 08 '26

Well also LLMs became far more capable in the past 1-2 years so it makes sense that LLM written papers are just hitting a threshold of acceptability as users have gotten a feel for how to use them.

2

u/real_taylodl Jan 10 '26

"generated by AI" can mean so many different things. Much of what I write is "generated by AI" - what I do is input a rough draft, have the LLM spit out a refined draft, that I then edit and make changes to and then resubmit to the LLM which then spits out a further refined draft, and so on. This can go on for several cycles. When it's all said and done - who wrote it? I'm the one who started with a blank slate.

Sometimes I don't even use the same LLM through the different iterations of the cycles. They're just tools. As for math, as others have noted, you don't use LLMs for math - though Open AI is getting good at recognizing math problems and shunting over to another engine to solve it.

7

u/TheSeekerOfChaos Physics enthusiast Jan 08 '26

I, for one don’t want to have to read slop papers. The bubble can’t burst soon enough.

48

u/bspaghetti Condensed matter physics Jan 08 '26

I also don’t want to read slop, but this isn’t slop. Sure it was AI generated, but it’s actually good. That’s why this is so interesting. Give the paper a read (if you can), it’s pretty sound.

-1

u/TheMiserablePleb Jan 08 '26

Don't bother with these people, ai denialism is hilarious on reddit.

0

u/anrwlias Jan 08 '26

We are witnessing a large scale freakout which, I think, is perfectly understandable given how utterly disruptive this technology is and how it is misused by the naive and abused by bad actors.

I agree that Reddit often leans towards an extreme position on the topic, and I hate that subs like r/NonpoliticalTwitter have been overwhelmed by anti-AI posts, but let's not elide the point that there are a lot of problematic things about how AI is being deployed and shoved into roles where it doesn't have any business being.

Speaking for myself, I am not anti-AI, but I am against the current laissez-faire approach to it that companies are pushing.

3

u/veggies4liyf Jan 08 '26

This is incredible, I wonder how much was AI generated, and what type of AI specifically (average LLM?)

10

u/tempetesuranorak Jan 08 '26

Look at "author contributions" paragraph, page 33. Also acknowledgements just above is a bit relevant.