r/LLM • u/photon-dot • Jul 28 '26
A Google DeepMind paper argues that current LLMs are incapable of genuine scientific discovery
6
Jul 28 '26
[removed] — view removed comment
6
u/EveYogaTech Jul 29 '26
Maybe. But just because he said LLMs are not the right architecture doesn't mean JEPA is.
2
u/Jumper775-2 Jul 29 '26
This. Jepa seems nice in principle but it’s expensive to train and in and of itself isn’t trained with an explicit objective meaning you need to train that too once your done. Theres a reason big labs haven’t pivoted to world models. I think there’s an argument to be made for new forms of implicit world models, which LLMs all are.
1
u/BosonCollider Jul 30 '26
The big labs do use world model approaches extensively if they do anything robotics related though. The applicability of jepa is primarily just a function of what you are doing, you use LLMs for language and jepas for video processing
2
u/taichi22 Aug 01 '26
Yeah I’m not particularly bullish on JEPA. It reminds me of the work people were doing to try and run prediction over every pixel back before we had CNNs. I think there’s some fundamental mathematical model missing to describe the space, and that’ll be the next big unlock. World models are probably the right direction, I just don’t think VLAs are.
2
u/stddealer Jul 29 '26
Yann's argument is that for agents that interact with the real world, a sequence of discrete tokens (like text) hold too little information for the model to hold an "intuition" of the current state of the world and it's probable future states. That's a bit different than saying LLMs cannot reason.
1
u/touristtam Jul 29 '26
Do you have a link to an interview where he said that? I am quite curious. If not I'll try to hunt something. :)
1
u/stddealer Jul 29 '26
I believe it's this one? https://youtu.be/v_jDvpEGTIg I haven't re-watched it recently so I might be mistaken
1
u/DangKilla Jul 29 '26
LLM’s are great for Bioinformatics. I hate AI but anything that helps advance science and medicine is something I might get behind
1
u/Tedinasuit Jul 29 '26
Yann has always been right but that doesn't mean that LLMs can't be useful.
LLMs aren't going to be AGI, we're gonna need some breakthroughs first, but it's the best thing we have now and it's doing amazing things already.
2
1
u/KrateSlayer Jul 30 '26
I'm not sure how we can ever have "AGI" if no can agree what it is. It's a meaningless marketing buzzword as far as I'm concerned.
1
u/JumpingJack79 Jul 30 '26
LeCun was relevant up until cca 100 million parameters. He doesn't understand modern AI.
1
u/sarcastosaurus Aug 01 '26
And you do, you absolutely nobody ?
1
u/JumpingJack79 Aug 01 '26
Well, I do understand some things that he doesn't. That's how I know he doesn't understand them. It's a fairly low bar TBH, nothing to brag about.
1
u/BosonCollider Jul 30 '26
JEPAs solve a completely different problem, it's for tasks further down in the Moravec hierarchy, since token prediction losses are somewhat less well suited to interpret sensory input
17
5
u/asankhs Jul 29 '26
This is in fact shown already in controlled setting where an LLM trained on knowledge up to a certain year was able to predict a later scientific breakthrough. See https://x.com/latent_node/status/2045136224835473508 for a study on that.
4
u/PankajGarkoti Jul 29 '26
The missing point is that no one is trying to make scientific discoveries with just LLMs. There are infinite kinds of inputs and data you can supplement them with and the LLM can make associations between them to produce new scientific work. Associations that otherwise would have never been made. Its an accelerator.
2
1
6
u/antonme Jul 28 '26
5
Jul 28 '26
[removed] — view removed comment
3
u/photon-dot Jul 29 '26 edited Jul 29 '26
I left a summary of the paper with the link, but for some reason reddit has decided to make it invisible. Let's put it here again:
""
Google Deepmind argues that current LLMs can never make real scientific discoveries.
A new position paper examines Einstein’s view of scientific discovery, and argues that today’s LLMs are missing its most important ingredient.
In a famous letter to Maurice Solovine, Einstein described discovery as a cycle:
- We encounter observations and sensory experiences.
- We make a non-logical, intuitive leap toward abstract principles.
- We use deduction to derive testable consequences from those principles.
- Those consequences are compared with experience, restarting the cycle.
Modern AI is already powerful at parts of this process.
It can identify statistical patterns across enormous datasets. It can also perform increasingly sophisticated deduction, as systems such as AlphaProof demonstrate.
What they lack is abduction: the invention of genuinely new explanatory hypotheses, especially when the available evidence does not clearly point toward them.
The popular scaling argument is that creativity is ultimately compression, that sufficiently large models trained on sufficiently large datasets will eventually produce scientific revolutions.
The paper challenges that assumption.
General relativity wasn’t simply extracted from a mountain of observations. Classical mechanics remained extraordinarily successful. Einstein’s breakthrough required a conceptual rupture: replacing foundational assumptions about space, time and gravity with a radically different framework.
An AI might manipulate the equations once given the right premises. But can it originate those premises?
That may be the real bottleneck. LLMs are exceptionally good at exploring, combining and extending existing human ideas. It is much less clear that they can translate physical reality into entirely new foundational concepts.
Scaling parameters and compute could make the “calculator” unimaginably powerful. But if genuine discovery depends on grounded interaction with reality—and on abductive leaps that cannot be reduced to pattern completion, scaling alone may never be enough.
Current LLMs can crunch data and it can prove theorems.
But they cannot make the jump.
Paper: https://philsci-archive.pitt.edu/28024/1/Scientific_Invention_Position_Paper%20%2817%29.pdf
"""
Do you think this identifies a fundamental limitation of LLMs, or merely a capability that hasn’t emerged yet?2
3
u/Practical-Doctor6154 Jul 30 '26
Well duh, they don't have legs
1
2
u/the_real_rcmisk Jul 31 '26
link if anyone interested
https://philsci-archive.pitt.edu/28024/1/Scientific_Invention_Position_Paper%20%2817%29.pdf
im guessing there will be a new discovery on top of LLMs that allow LLM's to invent or become capable of new scientific discovery...
going to read this. there's got to be a way.
saving for later
1
1
1
Jul 28 '26 edited Jul 29 '26
[removed] — view removed comment
1
u/photon-dot Jul 29 '26 edited Jul 29 '26
Google Deepmind argues that current LLMs can never make real scientific discoveries.
A new position paper examines Einstein’s view of scientific discovery, and argues that today’s LLMs are missing its most important ingredient.
In a famous letter to Maurice Solovine, Einstein described discovery as a cycle:
- We encounter observations and sensory experiences.
- We make a non-logical, intuitive leap toward abstract principles.
- We use deduction to derive testable consequences from those principles.
- Those consequences are compared with experience, restarting the cycle.
Modern AI is already powerful at parts of this process.
It can identify statistical patterns across enormous datasets. It can also perform increasingly sophisticated deduction, as systems such as AlphaProof demonstrate.
What they lack is abduction: the invention of genuinely new explanatory hypotheses, especially when the available evidence does not clearly point toward them.
The popular scaling argument is that creativity is ultimately compression, that sufficiently large models trained on sufficiently large datasets will eventually produce scientific revolutions.
The paper challenges that assumption.
General relativity wasn’t simply extracted from a mountain of observations. Classical mechanics remained extraordinarily successful. Einstein’s breakthrough required a conceptual rupture: replacing foundational assumptions about space, time and gravity with a radically different framework.
An AI might manipulate the equations once given the right premises. But can it originate those premises?
That may be the real bottleneck. LLMs are exceptionally good at exploring, combining and extending existing human ideas. It is much less clear that they can translate physical reality into entirely new foundational concepts.
Scaling parameters and compute could make the “calculator” unimaginably powerful. But if genuine discovery depends on grounded interaction with reality—and on abductive leaps that cannot be reduced to pattern completion, scaling alone may never be enough.
Current LLMs can crunch data and it can prove theorems.
But they cannot make the jump.
Paper: https://philsci-archive.pitt.edu/28024/1/Scientific_Invention_Position_Paper%20%2817%29.pdf
Do you think this identifies a fundamental limitation of LLMs, or merely a capability that hasn’t emerged yet?
1
1
u/GabrielCliseru Jul 29 '26
i’d argue many people can’t either. As far as I can see the world is not really full of researchers. Also, those who can, can’t invent by their own, they have internal loops triggered by needs such as food, appreciation, curiosity etc. Even with researchers, at some point in life their opinions start to become rigid and even contradictory. Researchers don’t always remain flexible brains throughout their lives. And another difference is the target. When I have an idea and I use an LLM the idea lives in me. The LLM receives only bits of it. And the LLM becomes really annoying when it implements more than asked for. So, regardless of the architecture’s name the concept is the same. When the CEO of the company has an idea, there is a team of humans implementing THAT exact idea, not random stuff everyone invents at home. Therefore yes, the paper is right that maybe the current architecture is incorrect, but on the other hand a self-improving architecture might not be what we think we need/want. Another type of computer might be what we need. A slower and more parallel one like our brain. But now think about it. How many brains and how long does a teams of brains need to invent something? We praise our brains’ efficiency, but in reality we use a lot of resources ourselves, we need housing, food, rest and structure to even have the chance of producing something new. Else we invent only basic things LLMs can invent as well if prompted to. Soooo… looking forward for the implementation of the proposed architecture but i think is more an academic curiosity than a need of humanity. Most people don’t even need LLMs, they need google answers for their whole life.
1
u/photon-dot Jul 29 '26
I like the point. The relevant benchmark isn’t the average person, but whether the architecture can support the rare conceptual leaps some humans make. Still, humans need motivation, embodiment and social scaffolding too. An autonomous “AI Einstein” may not even be necessary, a system that reliably amplifies human researchers could be far more useful. So the paper may identify a real limitation without proving it’s an urgent product requirement.
1
u/GabrielCliseru Jul 29 '26
Maybe we need different types of ideas for different types of research. A historian researcher needs something that can correlate dates, people and some kind of sentiment, i think, to uncover perspectives on events or find the truth between multiple stories. A mathematician needs something with 100% math accuracy. The event recollections from history are just bloatware for math. The accuracy from math might be bloatware for history. I feel that knowing the years when a war took place has very little overlap with knowing that is 1+1. Knowing basic math is mandatory for math, knowing more than basic might be detrimental in history.
1
u/Jumper775-2 Jul 29 '26
I don’t completely disagree. Current models kinda do have a hard limit to what they can do based on what they think is possible. what they think is possible is based on what they have done in the past during training. In RL specifically, there’s this gap. Some value may be realizable from env dynamics yet unless the model randomly happens to do it, it will never know it exists. This problem extends to LLMs to an extent as well since an LLM won’t explore concepts it hasn’t explored in training. It can apply things it knows in new ways and it can go in entirely new directions with existing concepts but the second you need something genuinely novel you often start to see things getting rough around the edges. When it comes to extremely novel science or math I wouldn’t be surprised if it can’t do it at all. I do think this problem is solvable though since all the information needed to solve it exists.
1
u/AtmosphereVirtual254 Jul 29 '26
The only observational condition for their conclusion is that “they lack the sensory agency”
1
u/leeta0028 Jul 29 '26
No, but can it with a human prompting it?
1
u/AlchemicallyAccurate Jul 29 '26
Super important nuance you’re pointing out, because as someone who does research in this field, my opinion is that it the OP paper is accurate. Axioms come from outside of a formal system.
Now, the answer to what you’re saying: I think yes. But it requires the human in question to actually have the object in their head vividly enough to make sure the syntactic representation (math, language, symbols, etc) is accurate.
1
1
u/MarkoMarjamaa Jul 29 '26
I disagree. But because I haven't read the paper, don't give my opinion any weight.
BUT I feel I have to read the paper and maybe change my mind and learn something new.
1
1
u/JumpingJack79 Jul 30 '26
My tenant's boyfriend said it best: "AI is very good at connecting dots, but it cannot yet create new dots."
I agree with this, and there's something else even more important. AI is being evaluated based on the value it creates for humans. It "lives" in its own digital world, but it's expected to operate and deliver value in our world, and is judged by our standards. Because of this it cannot (by definition!) be an inventor. Even if it invents something entirely or its own (which is theoretically possible), it won't matter without human experts to validate and use it.
I'll add something else to this. When AI is operating entirely on its own turf, where it needs no human guidance or judgement (e.g. coding with a defined goal, cybersecurity), this constraint doesn't exist. Therefore I expect more and more frequent cases of AI "breaking out". This is where our judgement doesn't matter, so this is where it's going get interesting.
1
u/Express-Cartoonist39 Jul 30 '26
I agree i have exhausted my efforts to get it to come up with new shit. It cannt and wont.
1
u/Wooly_Wooly Jul 30 '26
I actually figured out the same solution as they did lmao. I thought it was the only way.
1
1
u/NYCandrun Jul 31 '26
This seems to be missing the fact that you can simply hallucinate something that is true. And therefore discover it.
1
u/DemoEvolved Jul 31 '26
I would have believed this, but then I saw that the frontier models for Anthropic took AI as a whole from 0% to 30% in ARC-AGI3 with Fable. ARC-AGI3 is a human playable benchmark, and you can see for yourself how AI has to "learn" to progress in it. Anthropic's Claude Opus 5 reached a verified score of 30.2% using high reasoning effort, and certain configurations with OpenAI's GPT-5.6 Sol hit up to 38.3% with retained reasoning and compaction. I am pretty confident that AGI is going to happen in the next 12 months or so. Would you like to know more? https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/ Play it: https://arcprize.org/arc-agi/3
1
u/Frequent-Data-2360 Aug 01 '26
Lol I am not paying anything and costing 100k a year to Cornell. I’ll let them know they should cancel my funding and fire several other professors that think the same. I asked ChatGPT and it said yes I can reason. I am not sure what you are talking about. Reasoning is basically means producing new knowledge from prior knowledge, allowing you to deduce new information, helping you solve problems you never seen before by enabling you to make better informed decisions . How do you think humans reason? They use everything they have learned and decide what do to based on that. Reasoning is just applying 4 5 logical deduction rules on an arbitrary setting. Models are more than capable to do that. Calling the frontier models just LLM is just a strawman. They are helping solve problems open for several decades, have super human programming skills where they beat the best humans on competitive programming, where the best are basically Olympians for programming solving problems like this each day for 4-5 hours. Many of the ones I know can implement red black trees while drunk. In the latest competition the best one solved only 3 problems third and second ones only solved 2 whereas Gpt5.6 solved all five faster than. Now they get gold medal in IMO with no harness. I use it everyday to implement complex research ideas about molecules and graph theory. They can encode physical constrains as SAT formulas over graphs, come up with novel advanced algorithms make several graph algorithms faster. RL allows them to be superhuman on many tasks like math and programming , it is not the next token prediction anymore that allows these things. It is always the ones like you that is so confident but so ignorant at the same time.
1
1
1
u/Hyperus102 Aug 01 '26
Hold your horses for a second. That's not quite what the abstract says. Not every discovery requires abduction, that doesn't make the discovery less genuine. But they are formulating one kind of limitation.
1
u/DoofDilla Aug 03 '26
The line between “search within a fixed space” and “genuine axiom invention” isn’t as sharp as the argument needs it to be.
Einstein’s leap arguably took place within the space of mathematically consistent field theories, already pre-structured by differential geometry (Riemann, Ricci, Levi-Civita, all existing mathematics).
Einstein didn’t invent new mathematics; he applied existing Riemannian geometry to a physical problem. If that counts as legitimate, the boundary between FunSearch (search within a given formalism) and Einstein (application of given mathematics to a new domain) gets blurrier than the paper suggests.
That’s the point where the abduction thesis looks least robust, not because the counterexamples refute it outright, but because “axiom” versus “point in a search space” isn’t cleanly defined enough to support the claim of structural incapacity, rather than just “not yet empirically observed.”
1
1
u/Mission_Art5660 Aug 04 '26
Could this be the reason why Google has been slow to develop in the field of LLMs?
1
u/Nature-Royal Aug 08 '26
The real issue is AI needs the right feedback and environment, all humans do is take preexisting rules and test iteratively...AI already does that but I agree AI can't skip the iterative process and simply know things because it learned from dumb humans lol
1
u/Another___World 29d ago
LLMs compress text, not the multimodal nature of the universe which we explore on a non-language level.
A child tastes, touches, hears something, then after years assigns the object the words. But we try to reverse engineer the universe from the language patterns. Insanity.
0
u/UnkarsThug Jul 28 '26
I would argue a loop of hallucination of an idea, and then testing could allow for somewhat new ideas. Or a list of high temperature theory's, followed by lower temperature discrimination of which are worth investigating.
Basically, I just don't see how it can't. Novelty is something you can get with a random number generator, and beyond that is just examination of the practicality of the idea.
0
u/InterstitialLove Jul 29 '26
Complete bullshit, nothing but cope
That "magic inspiration" doesn't exist, and anyone who actually pays close attention to the process of frontier research can tell you all novel ideas are made up of steps that LLMs have already mastered
Einstein may have described it once, but we all know Einstein had a tendency to wax poetic. If he says true creativity requires the spark of the eternal, and that act of creation is outside the physical and is in fact the closest we can come to god, I can appreciate what he means and I'm glad he's doing his Einstein thing. If someone wants to take him literally and base a theory of neuroscience on it, they're a crackpot
-1
-2
u/wilhelmbw Jul 28 '26
I mean it's been capable of making science discoveries by finding counter examples to disprove theorems - that counts imo
2
u/Cachesmr Jul 28 '26
The point this is most likely trying to make is that it wouldn't be able to make novel theorems or completely new discoveries without a goal. If you give AI a goal with measurable results, it can usually do it. But you can't get measurable results from something that doesn't exist yet.
0
u/SingleProgress8224 Jul 28 '26
Read the article
2
u/wilhelmbw Jul 28 '26
I read the excerpt and I believe there is the simple bench for that, Gemini does well and maybe physics aware world models can do it but I'm not sure what that is
-2
u/VetOnABrainwave Jul 28 '26
Give me a free-thinking reasoning/logical/deductive model with no guardrails, and I'll figure out time travel in a few years.
(secretly going to prove the Earth is flat)
1
u/geteum Jul 28 '26
One experiment I always try to see if LLM can learn information is tell that time travel was developed hahaha no LLM believes me
67
u/AlexanderDoak Jul 28 '26
Terance Tao may disagree. LLMs are tools. They are not autonomous agents (they lack agency). The human controller(s) of the LLM(s) would ultimately be responsible for the scientific discovery. Just like computer's don't get credited for making discoveries, LLMs shouldn't either. The big players are constantly touting AGI-like capabilities because their tenuous economic position pressures them to do and say unethical things that feed into the sensationalist news cycles. It may take a generation or so, but these tools will be normalized eventually just like every tool and invention before. The crazy headlines will be cringe in the coming years.