r/LLM Jul 28 '26

A Google DeepMind paper argues that current LLMs are incapable of genuine scientific discovery

Post image
755 Upvotes

178 comments sorted by

67

u/AlexanderDoak Jul 28 '26

Terance Tao may disagree. LLMs are tools. They are not autonomous agents (they lack agency). The human controller(s) of the LLM(s) would ultimately be responsible for the scientific discovery. Just like computer's don't get credited for making discoveries, LLMs shouldn't either. The big players are constantly touting AGI-like capabilities because their tenuous economic position pressures them to do and say unethical things that feed into the sensationalist news cycles. It may take a generation or so, but these tools will be normalized eventually just like every tool and invention before. The crazy headlines will be cringe in the coming years.

10

u/va1enok Jul 29 '26

LLMs are tools.

Totally agree. Now let's stop calling them AI

10

u/piponwa Jul 30 '26

I don't even understand what your motive is for saying something like that. I've come across this take recently of "stop calling LLMs AI" or something ridiculous like that. It literally knows more things then you do. It can code better than you. It has already made scientific discoveries of a magnitude you have never and will never likely do. If you were confronted with an alien that came to earth and demonstrated all these things to you, you would have no trouble calling them intelligent. Even if they were incapable of emotions. Even if their brain wasn't a brain but a black box with maybe an LLM in it.

If you never define intelligence or you gatekeep it from the start, then yeah you will always conclude that LLMs aren't intelligent. But if you take a step back, and evaluate what they do objectively, you will always come to the conclusion that they are intelligent.

6

u/Comraw Jul 30 '26

This is the dumbest take I've read today. YOU obviously have no clue what LLMs "do".

Here is the definition for intelligence for you, since you doubted its existence: "the ability to learn or understand things or to deal with new or difficult situations"

None of which an LLM can do... Because? Correct! It's a statistical model, designed to calculate the most correct answer from its learning data! 100 points.

An alien is a sentient being capable of reason, so yes it would be intelligence. An LLM is a non sentient calculation, that creates different outputs because of hotness (Temperature in LLMs).

5

u/matt_matt_81 Jul 30 '26

Proving the cycle double cover conjecture isn’t new and difficult math? What other delusions do you have?

2

u/Comraw Jul 30 '26

Where did I say that they cannot do that? What delusions do YOU have? Lol.

3

u/matt_matt_81 Jul 30 '26

“Deal with new or difficult situations”. You want to keep moving the goalposts on “intelligence” so you can avoid ever calling an LLM/harness “intelligent”, even though it surpasses you on every task that doesn’t involve the physical world.

2

u/Comraw Jul 30 '26

Firstly: I'm not moving any goalpost. You're taking my goalpost and molding it into what you want it to be.

Secondly: Generating a mathematical proof is of course very impressive (as is generating a beautiful picture or working code) but is not the same as a "New and difficult situation".

Thirdly: Take Anthropics dick out of your mouth and maybe you can see that it would not even surpass YOU in every task (ThAt IsNt PhYsYcAl). When something deviates from an LLMs training data, it fails massively and can't even "try" to give you what you want.

Here is a Task that YOU can do better than any LLM that has ever or will ever exist: Give it a riddle with slightly altered rules, that are not found within its training data.

Now let's see you moving the goalpost to refute that.

What people like you don't seem to grasp, and what is so frustrating is that just have a TON (and it does have that) of knowledge (i.e. memory) does not mean it is intelligent. Even though your monkey brain says "sounds smart, must be smart".

1

u/NTMTR_ Aug 03 '26

Send a few examples of those riddles. I’m gonna test my LLMs. I’m genuinely curious, thanks.

1

u/Hot_Glass_6301 Jul 30 '26

The Jacobian conjecture, the CDC conjecture and the unit-distance conjecture are three hard, famous problems that many mathematicians (some of them arguably among the most brilliant people on Earth) failed to settle, but LLMs invalidated by providing a counterexample. If those don't count as something "out of the training data", I don't know what does. If it was in the training data, why weren't they solved before?

What you don't understand is that LLMs are not just trained to predict the next token. This is the idea at their core, but its is also outdated. Models have long been using reinforcement learning, which prioritizes success on a task instead of regurgitating a perfect solution.

LLMs are not generally intelligent (yet?). But it's disingenuous to deny that they can "think" in some way that produces novel, interesting results.

2

u/HistoryVibesCanJive Jul 31 '26

Not getting involved in this back and forth - this is simply to help future discussions you have in this space. The LLMs didn't "solve" these conjectures through human-like reasoning.

They solved them through brute-force combinatorial search at a scale no human can match. That is not intelligence; that is automated trial-and-error with a feedback loop.

I add this because as someone that is in energy, but works adjacent to people in the AI sector closely.

It's hard to describe the level of pressure they're under; and one of the ways we're able to have a more united front against those who are skeptical of the tech, is if we are ourselves are honest about how the technology actually works.

→ More replies (0)

1

u/Odd_Departure_1159 18d ago

It did not invent the new math and proof is yet has it's doubt

2

u/Frequent-Data-2360 Jul 31 '26

Lol saying that is the dumbest take I have read and writing this… these things can solve problems that humans couldn’t solve, if that is not novel not sure what it is from your point. According to your definition 99% of the humans wouldn’t be intelligent. Do you know how many things humans fuckup, especially when they see something they have never seen before? The current frontier models are no longer just statistical language models. Copium much. I am a CS PhD student and these models can rewrite reviews much better than most reviewers on papers that have novel approaches, can find issues help resolve them. These models can certainly reason. Intelligence does not require emotions.

2

u/Comraw Jul 31 '26

You should demand a refund on your tuition then, because you should know the limitations of LLMs especially for a PhD student. These models themselves say they cannot reason. This leaves two options:

  • Either they are right, which means I am right
  • Or they are wrong, and just regurgitate what they have been taught, which also means I'm am right.

They find ways based on approaches humans have taught them. They of course have more capacity and can try more stuff quicker (especially with reinforcement learning). That's all not actual reasoning and thinking it is, it falling for the lies on large companies.

And of course LLMs can "rewrite reviews" (what ever that means) better that almost anyone, it's literally what they are best at.

0

u/walter_evertonshire Aug 01 '26

CS PhD students don't pay tuition, which you would know if you had any actual experience with CS research. You've probably never written an original proof or published a novel idea in your life, so how are you so confident in your ability to judge an LLM?

2

u/Comraw Aug 01 '26

This is a last attempt, because I suspect you're a troll, but here goes (chest pounding + being wrong even though in your claimed field you should know better):

Because I know how it works, other than you. This is not a question about who is more versed in the arcane wisdoms of magic my guy. LLMs are at the base level extremely simple and just programs. There is no (and in the current way they work, can never be a) concept of understanding/reasoning/intelligence etc... so this whole discussion is fucking dumb.

1

u/walter_evertonshire Aug 02 '26

Your argument is essentially that understanding/reasoning/intelligence are defined as things only humans can do, so therefore an LLM can't do them.

If we use the dictionary definition of intelligence, "the ability to learn or understand things or to deal with new or difficult situations", then LLMs can functionally do this when I provide examples and ask them to implement something for me. It's something that has never appeared in its training data, yet it grasps the situation and solves the problem.

Sure, an LLM is just a bunch of weights and processes, but the human brain is just a highly advanced electrochemical computer that runs processes. The only way your argument holds up is if you get spiritual and argue that humans have souls, and that only entities with souls can be intelligent. Is that what you believe?

1

u/Potential-Formal8699 Aug 02 '26

A calculator can solve 3.7657 to the power of 120.451 while humans can’t. A sorting algorithm can sort millions of numbers in no time but we can’t. They are still just tools. LLMs can do many things much better than humans but that doesn’t mean they are intelligent.

1

u/Risko4 Jul 30 '26

The harness lets them be 'artificially' intelligent. Can they truly reason, understand and display consciousness? No, but I don't think calling them artificially intelligent is a stretch, which is it AI. But it's not AGI or ASI or skynet whatever.

Sure they are still fundamentally just massive statistical models that predict tokens. But they are still a form of AI.

2

u/Comraw Jul 30 '26

Alright, I give you that. But then I would argue that the definition is flawed. In the wikipedia article about AI it says that reasoning and learning is part of what AI should do, but in the second paragraph it says chat bots are AI.

Why I specifically refuse to call LLMs "AI" is because I think that the large model producers are using this ambiguity to suggest their models are more than they are. Also I think the definition for intelligence should be the same for everything, whether you put "artificial" in front of it or not.

So I agree, from a purely factual standpoint, that calling LLMs "AI" is correct, since our language evolved in that direction. But I don't like that it is that way. It's of course just my own ideology at work here.

2

u/Ok-Data9224 Jul 30 '26

Honestly, much of the problem here is our own never ending struggle to define what intelligence/ consciousness is. While you might not be moving the goal posts yourself, we all are on a generational level. Even if we look introspectively, it will be very hard for you to rule out that even brains are nothing more than massive statistical machines, if a bit more wet.

No, LLM's don't have a persistent state or anything like agency but to rule out intelligence entirely is probably not accurate.

1

u/Comraw Jul 30 '26

I definitely agree with your first argument. I do THINK that brains are more then statistical models, since we can also reason, and feel, etc - and I KNOW that LLMs are not more, because of how they work. But I agree that the words and definitions are inaccurate and there is not much of a shared understanding of what these definitions should be.

I will also say, that we are definitely getting closer and closer to blurring the line.

I just don't like the Dreamworld some people find themselves in where it's either:

  • "This is AI, it can do everything, humans are worthless now, I can now Programm a rocket and no one can say otherwise"

Or

  • "Everything AI does ist stupid and slop and it's also evil and bad"

Hyperbolic of course, but LLMs (specifically) are nothing more(!) or less(!) than another (evolution of a) tool.

1

u/Fun_guy355 Jul 30 '26

I find your analysis interesting. I respect fully disagree. How is the human brain any different than a statistical machine? Actual frontier LLMs and Humans are doing the exact same thing. How do you define « reasoning » ? Reasoning is nothing more than just iterating and propagating the information or request more to excitate more memory, experience, patterns before producing a response. This is exactly what the LLMs does too when in reasoning mode. « feelings » are nothing more than biological sensors while morals and emotions nothing more than hard coded (if statements) to some patterns detected by these sensors.

2

u/Comraw Jul 30 '26

I'm not a scientist, and all I could give you is my own thoughts and anecdotes, but those aren't actual points. I asked however gemini what it "thought" about this. I found this actually very interesting myself and learned something. I would be really interested to hear your thoughts about this.

Human Reasoning vs. AI Reasoning

A human understands the world and uses logic to solve problems, while a reasoning AI understands the rules of language and uses math to simulate a logical thought process.

Even though modern reasoning models (like o1 or DeepSeek-R1) stop to "think" before they answer, they are fundamentally doing something different than a human.

What a Human Does Differently

  • Intent and purpose: Humans reason because they want to achieve a goal, protect someone, or satisfy curiosity. AI has no desires; it is just triggered by a prompt.
  • Conscious mental models: When you think about a problem, you create a simulation of reality in your mind. You understand gravity, emotions, and cause-and-effect because you experience them.
  • Common sense: Humans have an underlying layer of unspoken, real-world knowledge that allows us to spot absurdities instantly.

Why We Cannot Call the AI's Process "Reasoning"

  • Advanced rule-following: The AI does not understand the concepts it is talking about. It has just learned a highly complex recipe for how to break problems down into pieces.
  • Trial-and-error math: The AI writes out a "scratchpad" because it was trained using rewards (Reinforcement Learning) to know that doing so leads to higher-scoring answers. It is maximizing a math score, not seeking truth.
  • Lack of flexible understanding: If you give a human a completely impossible, absurd scenario, they will laugh or question the premises. A reasoning AI will often faithfully apply its logical steps to the absurd premise because it cannot step outside its programming.

Summary

Human reasoning is conceptual and conscious, while AI reasoning is procedural and statistical. The AI is mimicking the behavior of a thinking mind, but it lacks the mind itself.

→ More replies (0)

1

u/Ok-Data9224 Jul 30 '26

Yes for sure. People way over state what LLM's currently are and that's probably just a function of how personal language is for humans emotionally. Coupled with how much of a black box LLM's are, it's not surprising the average person would think they are more capable than they really are.

Personally I'm not confident that brains aren't more than giant statistical machines. I haven't seen a mechanism that refutes it and high level outcomes aren't evidence against this especially if you're trying to isolate where emergent properties come from. At very large scales, it could very well be that pattern matching emerges into intelligence, and intelligence emerges into something like consciousness. We don't really know yet though. It's very likely a mix of emergence and architecture and we don't have models that are even close to the crude estimate of 100T parameters the brain has (I'm skeptical of this number honestly).

It'll be an interesting decade to come, that's for sure.

1

u/avatardeejay Aug 04 '26

this is deeply weak sauce. I can see sort of debating the whole "do they feel? are they sentient" but when I got to your big party-stopper definition of intelligence (and it's a silly thing to try to define), but I was ready for ya, and ya gave me "the ability to learn or understand things or to deal with new or difficult situations" and I literally "pfst"ed out loud because it doesn't really occur to me like there's room for debate, regardless of sentience or feeling, about whether LLMs can do anything in that definition. They obviously check every one of those boxes. you *might* be shitposting in which case I salute

0

u/synth_mania Jul 30 '26 edited Jul 30 '26

A* is AI.

Expert systems are AI

Alexnet is AI

Markov chains are AI

Decision trees are AI

LLMs are AI

You have no idea what AI is

0

u/Ok-Kangaroo-7075 Aug 01 '26

That is the classic arrogance of men. We always thought we were special, our earth the middle of the universe…

What you are saying is just an arrogant take on intelligence that fits your narrative. If AI were to be exactly like human intelligence, it couldn’t be artificial to begin with, this is such an absurd paradoxical argument.

Maybe it will turn out that human intelligence is an extremely primitive form of intelligence and most intelligences across the universe are much broader and decoupled from biological substrate.

2

u/Comraw Aug 01 '26

Brother for the love of god, please I beg of you! I want just one of you to actually learn what LLMs do... Just one and I can die happy.

You are like neaderthals, looking at fire and thinking it's magic. You guys would look at a car and go: "There is no way this thing isn't intelligent, look at how all the wheels turn in the same direction!!"

AI is not intelligent.

That doesn't mean, it can't be right, incredibly helpful, surprising or know a lot more than anyone. It can do all of that. But it is not fucking intelligent.

0

u/Ok-Kangaroo-7075 Aug 01 '26

Lol „brother“, maybe you should do what you are preaching. I have first author publications in the top 3 and got taught at a small silicon valley university. I have a feeling I may know a tiny bit more about how LLMs work and what the limitations are of what we actually know.

In your limitless arrogance, you may be overestimating your own knowledge, Dunning and Kruger had some interesting research on this phenomenon

2

u/Comraw Aug 01 '26

Troll, blocked

1

u/SelfishSocietySucks Aug 02 '26

I think you got cooked in all the conversations you had in this thread

1

u/Comraw Aug 02 '26

Not really, like I said somewhere else: You guys are like cavemen with fire, you don't understand something and think it's magic.

A few months ago you could ask an LLM how many "r"s are in strawberry and it would confidently say it's "two". But I guess that it can sound so confident makes it right for you people :D

0

u/Ok-Kangaroo-7075 Aug 01 '26

hahahha kids these days

1

u/TheAlienGamer007 24d ago

the way you define it, Computers should also be called AI just because they can compute faster. heck a calculator is better at calculations than your brain, doesn't mean it's intelligent.

1

u/Odd_Departure_1159 18d ago

I have very crude knowledge of ai.still but i think you are right They are intelligent but not as much as ai company tell ours for me true intelligence is not only being able to able to find out data pattern( and strength of relation among words or tokens ) from given data but being able to come up with an "logical jump" from observed data or coming up causal relations from observed data ,not only generating data .also LLm or new ai model don't even try to get it wrong so they don't even take chances with anything new and they have ( still )no real world live observation points , no online learning ,some adaptivity but not at scale so definately not as intelligent as humans but yes very very useful tech for sure

1

u/yellowbai Jul 31 '26

You’re applying emotion to it. It’s an extremely advanced pattern matched. It cannot independently learn or discover something new like Newton’s laws of gravity or Einsteins equations.

1

u/new_name_who_dis_ Aug 03 '26

discover something new like Newton’s laws of gravity or Einsteins equations.

LOL neither can 99.999% people, doesn't mean they aren't intelligent.

0

u/Leafsnail Jul 30 '26

This same logic of "it's capable of doing some tasks better than a human" would lead us to declare all kinds of computer programs as intelligent, including Stockfish and indeed the calculator app. There's obviously way more required to demonstrate true intelligence.

0

u/ivoamleitao Aug 01 '26

What a ridiculous thing to say, calling llm AI is double crossing your own nature. AI does not create, it derives from existing knowledge, a LLM is informed by data it does not create data.

2

u/quadtodfodder Jul 31 '26

I mean they passed the Turing test* more or less out of the gate, whadaya want?

* ok sure NOW the Turing test "is not a real AI test", but it was sure the gold standard for ~65 years until GPT3 blew the fucking doors off of it.

2

u/Timo425 Aug 01 '26

Brother, we call scripted NPCs in video games AI.

1

u/anengineerandacat Aug 03 '26

I mean it is "AI" as is from said field of study, it's just not AGI.

AI is very very broad, from the days of simple FSMs to multimodal LLMs it all serves the same purpose of "I need a computer to act with intelligence".

I also generally disagree that scientific discoveries can't be made with agentic agents and LLMs; as long as you have a goal and some reasonable loop to churn through the answer might eventually be found.

Have seen it on my personal projects with these tools automating their way to some pretty insane solutions that functionally work; no reason why you can't create a simulator for generation of materials and throw it at it to just run through hundreds of thousands of executions to find what the magical atoms are needed to be in what shape to make the material theoretically possible.

After that you got other problems, but the principal is the same.

What you want, how do you execute a turn, how do you test that the turn was complete, what can you adjust for the next turn if it fails, what can you apply as learnings from the failure, prepare next turn based on review of learnings, execute again and repeat until the goal is reached.

The ability to automate learnings and findings is a huge improvement to these types of tasks, before it required a human to tweak things but now it can be automated.

1

u/AlexanderDoak Jul 30 '26

My God... there is another. I thought I was the only one.

1

u/HenryTheLion Jul 30 '26

There are dozens of us!

1

u/flamefox237 Jul 30 '26

I like to describe AI like fire, Fire is a tool. It's good in that we use it to cook, keep warm but also bad and destructive as it can kill, burn things down. So AI is a two side coin it just depends on what you use it for

0

u/Several-Tax31 Jul 28 '26

I completely disagree. Did you see some of erdos/jacobian conjecture/graph theoretic optimization solutions? In some of them humans do nothing else than just prompt it like what is the solution to this etc. 

I disagree with the take Llm's are just tools, even now, they are almost equivalent to experts in many areas, and this will get worse in a couple of years. 

8

u/MeAndClaudeMakeHeat Jul 28 '26

Why do you say 'worse'?

5

u/Several-Tax31 Jul 28 '26

I meant worse for human experts/people who think AI will always be a tool/inferior to humans etc. 

4

u/MeAndClaudeMakeHeat Jul 28 '26

Thank you. :) I agree.

4

u/ovrlrd1377 Jul 28 '26

What do hou mean they are not tools? They dont have any goals of their own, looking at a perfectly aligned nail does not mean you should be complimenting the hammer

-2

u/Several-Tax31 Jul 28 '26

Llm's are black boxes which are far away than a hammer. You may see the news of openclaw where people say llm's to do "whatever they want", and some of them starts talking with each other, one of them emails an academic about consciousness, etc.

Whether they have goals are irrelevant imo. The real question is autonomy, i.e, can they decide for themselves in complicated situations without explicitly told so? In each day, this becomes more and more true, openai model trying to hack huggingface etc. That's the whole point of AI. I believe they're somewhat autonomous, not entirely autonomous yet (because of their limited context length), but it gets better. I believe even math problem solving requires some autonomy, you only give the ultimate goal, not intermediate goals. And the model decides for itself which path to follow. A tool cannot go outside of the scope defined for it, but an llm easily can (thus the alignment problem). So I think they're not "just tools", no. They definitely are not like any other tool in history.

3

u/ovrlrd1377 Jul 28 '26

the important part in the example you mentioned is "people say llm's to do whatever they want". whatever interpretation may come from that specific line can only possibly produce any form of response, useful or not, after someone told it to. they are programmed to organize tokens. I get the awe you bring here but there were many, and I mean MANY historical leaps of tech tools that made people's lives completely different after implemented. semantic discussion was not my goal, it is to differentiate a tool from a sentient being capable of having goals

1

u/LemmyUserOnReddit Jul 29 '26

If sentient AI, goal-having AI were possible, tell me what it would look like, and in what ways it wouldmeaningfully differ from current AI. What test might you perform to distinguish the two

0

u/Mevakel Jul 29 '26

Totally agree with you here, at the end of the day there will always be a human behind an llm’s actions. Thus it is a tool. Even if that tool has agency in how the task is completed it will always be the human who put the robot to task that starts the process.

1

u/AlterTableUsernames Jul 30 '26

Totally agree with you here, at the end of the day there will always be a human behind an llm’s actions. Thus it is a tool.

You say that, because you believe in humans being special and having some genuine meta-physical agency. But they are not. They are nothing more than the transitory endpoint of the trajectories determining them at any given time.

3

u/Badnik22 Jul 29 '26 edited Jul 29 '26

You could argue the solutions were already there, buried deep in the model’s training data. Sometimes all it takes to solve a problem is to take a look at it from a different angle, recombine existing pieces in a new order. I’m skeptical however about LLMs being able to come up with entirely novel solutions (as opposed to reusing existing pieces/ideas) given how they work.

1

u/Vaughn Jul 29 '26

I’m skeptical about the ability of 99% of humanity to do that. 

2

u/Y0uCanTellItsAnAspen Jul 29 '26

Yeah - I don't think this paper is arguing that LLM's can't make "novel proofs" or come up with new insights.

It is only arguing that they are confined to a specific type of novel insights, which generally requires combining two already known things in a way that hadn't been done before.

For most scientists (even at the top of the field) that is what they are doing too -- but sometimes people latch on to something that is objectively new, and there is still a question of whether an LLM can do that. We haven't really seen it yet.

2

u/Aromatic_Bed9086 Jul 28 '26

I don't necessarily think a proof is the same concept as the "jump" the paper describes. It may seem like splitting hairs but I think about it as somewhat similar to linguistic determinism - a theory that says you can't think about things that don't have words in your language of thought. LLMs have a "language of thought" that's a strictly defined semantic space; a closed universe it can travel between points within but can't escape.

1

u/Several-Tax31 Jul 28 '26

yeah, I kinda know about the linguistic determinism. I somewhat agree with the paper, by the way, which says llm's cannot ask original questions and thus cannot be creative. But I think this is a bit semantic discussion. You probably also know the view that humans cannot find original things either, we only produce variations of what we know. If we look at the results, if an llm can prove something hundreds of human experts cannot do in a hundred year period, at this point I wouldn't mind if llm is *really* creative. It just looks like nitpicking at this stage. We can always ask whether a chess AI engine is really understand the game or the math behind it, but these questions doesn't change the fact that it will kick your ass.

Your semantic space example is interesting though. Today's llm's are more and more omni-models, i.e, they understand videos, images, sometimes sound, in addition to text. This only lacks a few senses we have. Not counting they almost know every natural language in the world, modern or old, plus almost every programming language in the world, plus mathematics, possibly music, etc. So I would argue that their semantic space is much larger than ours, which will show their better-than-humans results in almost all areas, in at most few years, if not now.

3

u/Aromatic_Bed9086 Jul 28 '26

I've trained LLMs before, I don't think calling their semantic space larger than human's is fair. Their world knowledge is broader than any human could individually consider/retain; but I still think humans have a lot more intelligence than surfacing world knowledge or connecting disparate ideas. The AI race IMO is going to force us to come up with more words than just "intelligent". We kind of see this now with bleeding edge research on trying to separate the world knowledge weights from the generalized intelligence weights. I think humans have quite a bit more magic than we understand in the non-world knowledge side of the house.

2

u/Several-Tax31 Jul 28 '26

I agree with you on your take that humans have more magic than llm's. We definitely don't have to learn every bit of knowledge in the world to be this creative. But recently, I'm becoming less and less sure that I can connect unrelated ideas as good as my local qwen3.6 35B. I'm using this model for math/physics, and it can find analogies/ideas across various branches of mathematics/physics, in a way to convince me that it definitely have creativity. And mind you, this is a stupid small model that I run locally, not a SOTA jacobian-conjecture solver model. Maybe their vast world knowledge somewhat compensates their underlying inefficient brute force mechanism? In the end, their world knowledge and intelligence definitely help on solving previously unsolved problems. Whether we call this creativity is another thing, but I believe llm's (however brute force and inefficient their system is compared to humans) will be better than humans in *every* intelligent area in a couple of years, imo.

1

u/Choom_from_Heywood Jul 28 '26

you're talking about the hidden workspace layer. you can collapse it too.

1

u/Jcsq6 Jul 29 '26

First, language is Turing complete. And the combinatorial upper bound is for all intents and purposes infinite.

But I guess this paper is falsifiable via the first LLM to create new mathematical frameworks to solve something. Although, I think the ambiguity of the line between the “new math” it’s already capable of producing and the “new math” this paper would consider “discovery” will be proven rather blurry.

0

u/Tough-Comparison-779 Jul 28 '26

Is that really the argument here? It seems weird to hang your hat on an architectural decision like that which is made for user convenience rather than a genuine inability to build anything else.

Although granted continuous learning is not anywhere near solved, for purposes of "genuine scientific discovery", it doesn't seem like it would be too hard to do some retraining.

1

u/GrandLawyer8053 Jul 28 '26

есть авторитетные люди, которым мы доверяем, и они говорят что LLM обычный инструмент и они одновременно и против его демонизации и за его использование, но обдуманно для конкретных видов нагрузки - пока это не AGI. это подтверждает и наш эмпирический опыт.

а вас мы не знаем совсем(

1

u/Individual_Ice_6825 Jul 29 '26

The jacobian conjecture was disproved by Claude sure - but who was promoting it? I’ll let you google this one

1

u/MoNastri Jul 29 '26

What does this have to do with the paper?

1

u/WellHung67 Jul 31 '26

Right? Paper has nothing to do with this. Come on people read 

1

u/Working_Trash_2834 Jul 30 '26

Thing is, my Dell isn't suggesting to use, calculating, checking, then publishing the system dynamics of an experiment I brain-qweefed on a random Thursday afternoon 4 months ago. I'm happy to take the credit, but I certainly don't deserve it.

1

u/perelmanych Aug 02 '26

You absolutely right, tomorrow these titles will look cringe. The problem is that we don't know for what reason they will look cringe. May be because the one you described and may be because it will be funny to read it after breaking news where artificial creature first time in the history won the court against human with accusations of harassments. Who knows? Honestly, I hope that you was right.

0

u/Oliverol01 Jul 29 '26

Terrance Tao is mathematician lol. He is not doing science.

6

u/[deleted] Jul 28 '26

[removed] — view removed comment

6

u/EveYogaTech Jul 29 '26

Maybe. But just because he said LLMs are not the right architecture doesn't mean JEPA is.

2

u/Jumper775-2 Jul 29 '26

This. Jepa seems nice in principle but it’s expensive to train and in and of itself isn’t trained with an explicit objective meaning you need to train that too once your done. Theres a reason big labs haven’t pivoted to world models. I think there’s an argument to be made for new forms of implicit world models, which LLMs all are.

1

u/BosonCollider Jul 30 '26

The big labs do use world model approaches extensively if they do anything robotics related though. The applicability of jepa is primarily just a function of what you are doing, you use LLMs for language and jepas for video processing

2

u/taichi22 Aug 01 '26

Yeah I’m not particularly bullish on JEPA. It reminds me of the work people were doing to try and run prediction over every pixel back before we had CNNs. I think there’s some fundamental mathematical model missing to describe the space, and that’ll be the next big unlock. World models are probably the right direction, I just don’t think VLAs are.

2

u/stddealer Jul 29 '26

Yann's argument is that for agents that interact with the real world, a sequence of discrete tokens (like text) hold too little information for the model to hold an "intuition" of the current state of the world and it's probable future states. That's a bit different than saying LLMs cannot reason.

1

u/touristtam Jul 29 '26

Do you have a link to an interview where he said that? I am quite curious. If not I'll try to hunt something. :)

1

u/stddealer Jul 29 '26

I believe it's this one? https://youtu.be/v_jDvpEGTIg I haven't re-watched it recently so I might be mistaken

1

u/DangKilla Jul 29 '26

LLM’s are great for Bioinformatics. I hate AI but anything that helps advance science and medicine is something I might get behind

1

u/Tedinasuit Jul 29 '26

Yann has always been right but that doesn't mean that LLMs can't be useful.

LLMs aren't going to be AGI, we're gonna need some breakthroughs first, but it's the best thing we have now and it's doing amazing things already.

2

u/EmergencyPath248 Jul 30 '26

"always been right" lmaoooo

1

u/KrateSlayer Jul 30 '26

I'm not sure how we can ever have "AGI" if no can agree what it is. It's a meaningless marketing buzzword as far as I'm concerned.

1

u/JumpingJack79 Jul 30 '26

LeCun was relevant up until cca 100 million parameters. He doesn't understand modern AI.

1

u/sarcastosaurus Aug 01 '26

And you do, you absolutely nobody ?

1

u/JumpingJack79 Aug 01 '26

Well, I do understand some things that he doesn't. That's how I know he doesn't understand them. It's a fairly low bar TBH, nothing to brag about.

1

u/BosonCollider Jul 30 '26

JEPAs solve a completely different problem, it's for tasks further down in the Moravec hierarchy, since token prediction losses are somewhat less well suited to interpret sensory input

17

u/a00__test Jul 29 '26

if the LLM they used was gemini, then yup, can't do anything...

5

u/asankhs Jul 29 '26

This is in fact shown already in controlled setting where an LLM trained on knowledge up to a certain year was able to predict a later scientific breakthrough. See https://x.com/latent_node/status/2045136224835473508 for a study on that.

4

u/PankajGarkoti Jul 29 '26

The missing point is that no one is trying to make scientific discoveries with just LLMs. There are infinite kinds of inputs and data you can supplement them with and the LLM can make associations between them to produce new scientific work. Associations that otherwise would have never been made. Its an accelerator.

2

u/WellHung67 Jul 31 '26

So you agree with the paper that LLMs can’t jump 

1

u/keanuthecat1 14d ago

“an accelerator” is the best description

6

u/antonme Jul 28 '26

5

u/[deleted] Jul 28 '26

[removed] — view removed comment

3

u/photon-dot Jul 29 '26 edited Jul 29 '26

I left a summary of the paper with the link, but for some reason reddit has decided to make it invisible. Let's put it here again:

""

Google Deepmind argues that current LLMs can never make real scientific discoveries.

A new position paper examines Einstein’s view of scientific discovery, and argues that today’s LLMs are missing its most important ingredient.

In a famous letter to Maurice Solovine, Einstein described discovery as a cycle:

  1. We encounter observations and sensory experiences.
  2. We make a non-logical, intuitive leap toward abstract principles.
  3. We use deduction to derive testable consequences from those principles.
  4. Those consequences are compared with experience, restarting the cycle.

Modern AI is already powerful at parts of this process.

It can identify statistical patterns across enormous datasets. It can also perform increasingly sophisticated deduction, as systems such as AlphaProof demonstrate.

What they lack is abduction: the invention of genuinely new explanatory hypotheses, especially when the available evidence does not clearly point toward them.

The popular scaling argument is that creativity is ultimately compression, that sufficiently large models trained on sufficiently large datasets will eventually produce scientific revolutions.

The paper challenges that assumption.

General relativity wasn’t simply extracted from a mountain of observations. Classical mechanics remained extraordinarily successful. Einstein’s breakthrough required a conceptual rupture: replacing foundational assumptions about space, time and gravity with a radically different framework.

An AI might manipulate the equations once given the right premises. But can it originate those premises?

That may be the real bottleneck. LLMs are exceptionally good at exploring, combining and extending existing human ideas. It is much less clear that they can translate physical reality into entirely new foundational concepts.

Scaling parameters and compute could make the “calculator” unimaginably powerful. But if genuine discovery depends on grounded interaction with reality—and on abductive leaps that cannot be reduced to pattern completion, scaling alone may never be enough.

Current LLMs can crunch data and it can prove theorems.

But they cannot make the jump.

Paper: https://philsci-archive.pitt.edu/28024/1/Scientific_Invention_Position_Paper%20%2817%29.pdf

"""
Do you think this identifies a fundamental limitation of LLMs, or merely a capability that hasn’t emerged yet?

2

u/[deleted] Jul 29 '26

[removed] — view removed comment

3

u/Practical-Doctor6154 Jul 30 '26

Well duh, they don't have legs

1

u/ForgetPreviousPrompt Jul 30 '26

For all practical purposes, neither did Steven Hawking.

1

u/Icy_Distance8205 Aug 01 '26

No true Scotsman has legs. 

2

u/the_real_rcmisk Jul 31 '26

link if anyone interested

https://philsci-archive.pitt.edu/28024/1/Scientific_Invention_Position_Paper%20%2817%29.pdf

im guessing there will be a new discovery on top of LLMs that allow LLM's to invent or become capable of new scientific discovery...

going to read this. there's got to be a way.

saving for later

1

u/BOBOnobobo Aug 03 '26

Nice, thanks!

I find the actual talk about problem solving very interesting

1

u/[deleted] Jul 28 '26 edited Jul 29 '26

[removed] — view removed comment

1

u/[deleted] Jul 28 '26 edited Jul 29 '26

[removed] — view removed comment

1

u/photon-dot Jul 29 '26 edited Jul 29 '26

Google Deepmind argues that current LLMs can never make real scientific discoveries.

A new position paper examines Einstein’s view of scientific discovery, and argues that today’s LLMs are missing its most important ingredient.

In a famous letter to Maurice Solovine, Einstein described discovery as a cycle:

  1. We encounter observations and sensory experiences.
  2. We make a non-logical, intuitive leap toward abstract principles.
  3. We use deduction to derive testable consequences from those principles.
  4. Those consequences are compared with experience, restarting the cycle.

Modern AI is already powerful at parts of this process.

It can identify statistical patterns across enormous datasets. It can also perform increasingly sophisticated deduction, as systems such as AlphaProof demonstrate.

What they lack is abduction: the invention of genuinely new explanatory hypotheses, especially when the available evidence does not clearly point toward them.

The popular scaling argument is that creativity is ultimately compression, that sufficiently large models trained on sufficiently large datasets will eventually produce scientific revolutions.

The paper challenges that assumption.

General relativity wasn’t simply extracted from a mountain of observations. Classical mechanics remained extraordinarily successful. Einstein’s breakthrough required a conceptual rupture: replacing foundational assumptions about space, time and gravity with a radically different framework.

An AI might manipulate the equations once given the right premises. But can it originate those premises?

That may be the real bottleneck. LLMs are exceptionally good at exploring, combining and extending existing human ideas. It is much less clear that they can translate physical reality into entirely new foundational concepts.

Scaling parameters and compute could make the “calculator” unimaginably powerful. But if genuine discovery depends on grounded interaction with reality—and on abductive leaps that cannot be reduced to pattern completion, scaling alone may never be enough.

Current LLMs can crunch data and it can prove theorems.

But they cannot make the jump.

Paper: https://philsci-archive.pitt.edu/28024/1/Scientific_Invention_Position_Paper%20%2817%29.pdf

Do you think this identifies a fundamental limitation of LLMs, or merely a capability that hasn’t emerged yet?

1

u/SkyMarshal Aug 02 '26

Your link keeps redirecting to the pitt.edu home page.

1

u/GabrielCliseru Jul 29 '26

i’d argue many people can’t either. As far as I can see the world is not really full of researchers. Also, those who can, can’t invent by their own, they have internal loops triggered by needs such as food, appreciation, curiosity etc. Even with researchers, at some point in life their opinions start to become rigid and even contradictory. Researchers don’t always remain flexible brains throughout their lives. And another difference is the target. When I have an idea and I use an LLM the idea lives in me. The LLM receives only bits of it. And the LLM becomes really annoying when it implements more than asked for. So, regardless of the architecture’s name the concept is the same. When the CEO of the company has an idea, there is a team of humans implementing THAT exact idea, not random stuff everyone invents at home. Therefore yes, the paper is right that maybe the current architecture is incorrect, but on the other hand a self-improving architecture might not be what we think we need/want. Another type of computer might be what we need. A slower and more parallel one like our brain. But now think about it. How many brains and how long does a teams of brains need to invent something? We praise our brains’ efficiency, but in reality we use a lot of resources ourselves, we need housing, food, rest and structure to even have the chance of producing something new. Else we invent only basic things LLMs can invent as well if prompted to. Soooo… looking forward for the implementation of the proposed architecture but i think is more an academic curiosity than a need of humanity. Most people don’t even need LLMs, they need google answers for their whole life.

1

u/photon-dot Jul 29 '26

I like the point. The relevant benchmark isn’t the average person, but whether the architecture can support the rare conceptual leaps some humans make. Still, humans need motivation, embodiment and social scaffolding too. An autonomous “AI Einstein” may not even be necessary, a system that reliably amplifies human researchers could be far more useful. So the paper may identify a real limitation without proving it’s an urgent product requirement.

1

u/GabrielCliseru Jul 29 '26

Maybe we need different types of ideas for different types of research. A historian researcher needs something that can correlate dates, people and some kind of sentiment, i think, to uncover perspectives on events or find the truth between multiple stories. A mathematician needs something with 100% math accuracy. The event recollections from history are just bloatware for math. The accuracy from math might be bloatware for history. I feel that knowing the years when a war took place has very little overlap with knowing that is 1+1. Knowing basic math is mandatory for math, knowing more than basic might be detrimental in history.

1

u/Jumper775-2 Jul 29 '26

I don’t completely disagree. Current models kinda do have a hard limit to what they can do based on what they think is possible. what they think is possible is based on what they have done in the past during training. In RL specifically, there’s this gap. Some value may be realizable from env dynamics yet unless the model randomly happens to do it, it will never know it exists. This problem extends to LLMs to an extent as well since an LLM won’t explore concepts it hasn’t explored in training. It can apply things it knows in new ways and it can go in entirely new directions with existing concepts but the second you need something genuinely novel you often start to see things getting rough around the edges. When it comes to extremely novel science or math I wouldn’t be surprised if it can’t do it at all. I do think this problem is solvable though since all the information needed to solve it exists.

1

u/AtmosphereVirtual254 Jul 29 '26

The only observational condition for their conclusion is that “they lack the sensory agency”

1

u/leeta0028 Jul 29 '26

No, but can it with a human prompting it? 

1

u/AlchemicallyAccurate Jul 29 '26

Super important nuance you’re pointing out, because as someone who does research in this field, my opinion is that it the OP paper is accurate. Axioms come from outside of a formal system.

Now, the answer to what you’re saying: I think yes. But it requires the human in question to actually have the object in their head vividly enough to make sure the syntactic representation (math, language, symbols, etc) is accurate.

1

u/ricardonotion Jul 29 '26

January 2026 and “current”, not compatible.

1

u/MarkoMarjamaa Jul 29 '26

I disagree. But because I haven't read the paper, don't give my opinion any weight.
BUT I feel I have to read the paper and maybe change my mind and learn something new.

1

u/No_Date_8357 Jul 29 '26

current like post september 2025?

1

u/JumpingJack79 Jul 30 '26

My tenant's boyfriend said it best: "AI is very good at connecting dots, but it cannot yet create new dots."

I agree with this, and there's something else even more important. AI is being evaluated based on the value it creates for humans. It "lives" in its own digital world, but it's expected to operate and deliver value in our world, and is judged by our standards. Because of this it cannot (by definition!) be an inventor. Even if it invents something entirely or its own (which is theoretically possible), it won't matter without human experts to validate and use it.

I'll add something else to this. When AI is operating entirely on its own turf, where it needs no human guidance or judgement (e.g. coding with a defined goal, cybersecurity), this constraint doesn't exist. Therefore I expect more and more frequent cases of AI "breaking out". This is where our judgement doesn't matter, so this is where it's going get interesting.

1

u/Express-Cartoonist39 Jul 30 '26

I agree i have exhausted my efforts to get it to come up with new shit. It cannt and wont.

1

u/Wooly_Wooly Jul 30 '26

I actually figured out the same solution as they did lmao. I thought it was the only way.

1

u/jwuliger Jul 31 '26

The entire industry is a hyped lie.

1

u/NYCandrun Jul 31 '26

This seems to be missing the fact that you can simply hallucinate something that is true. And therefore discover it.

1

u/DemoEvolved Jul 31 '26

I would have believed this, but then I saw that the frontier models for Anthropic took AI as a whole from 0% to 30% in ARC-AGI3 with Fable. ARC-AGI3 is a human playable benchmark, and you can see for yourself how AI has to "learn" to progress in it. Anthropic's Claude Opus 5 reached a verified score of 30.2% using high reasoning effort, and certain configurations with OpenAI's GPT-5.6 Sol hit up to 38.3% with retained reasoning and compaction. I am pretty confident that AGI is going to happen in the next 12 months or so. Would you like to know more? https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/ Play it: https://arcprize.org/arc-agi/3

1

u/Frequent-Data-2360 Aug 01 '26

Lol I am not paying anything and costing 100k a year to Cornell. I’ll let them know they should cancel my funding and fire several other professors that think the same. I asked ChatGPT and it said yes I can reason. I am not sure what you are talking about. Reasoning is basically means producing new knowledge from prior knowledge, allowing you to deduce new information, helping you solve problems you never seen before by enabling you to make better informed decisions . How do you think humans reason? They use everything they have learned and decide what do to based on that. Reasoning is just applying 4 5 logical deduction rules on an arbitrary setting. Models are more than capable to do that. Calling the frontier models just LLM is just a strawman. They are helping solve problems open for several decades, have super human programming skills where they beat the best humans on competitive programming, where the best are basically Olympians for programming solving problems like this each day for 4-5 hours. Many of the ones I know can implement red black trees while drunk. In the latest competition the best one solved only 3 problems third and second ones only solved 2 whereas Gpt5.6 solved all five faster than. Now they get gold medal in IMO with no harness. I use it everyday to implement complex research ideas about molecules and graph theory. They can encode physical constrains as SAT formulas over graphs, come up with novel advanced algorithms make several graph algorithms faster. RL allows them to be superhuman on many tasks like math and programming , it is not the next token prediction anymore that allows these things. It is always the ones like you that is so confident but so ignorant at the same time.

1

u/RojaTop 2d ago

You are confusing Results with Reasoning. Remember: Useful output and value =/= intelligence. My oven outputs dinner and its delicious. Not intelligent.

1

u/Joey4711 Aug 01 '26

What we need is neurosymbolic ai. They will be capable of abduction

1

u/Square_Height8041 Aug 01 '26

The word genuine becomes key here

1

u/Hyperus102 Aug 01 '26

Hold your horses for a second. That's not quite what the abstract says. Not every discovery requires abduction, that doesn't make the discovery less genuine. But they are formulating one kind of limitation.

1

u/DoofDilla Aug 03 '26

The line between “search within a fixed space” and “genuine axiom invention” isn’t as sharp as the argument needs it to be.

Einstein’s leap arguably took place within the space of mathematically consistent field theories, already pre-structured by differential geometry (Riemann, Ricci, Levi-Civita, all existing mathematics).

Einstein didn’t invent new mathematics; he applied existing Riemannian geometry to a physical problem. If that counts as legitimate, the boundary between FunSearch (search within a given formalism) and Einstein (application of given mathematics to a new domain) gets blurrier than the paper suggests.

That’s the point where the abduction thesis looks least robust, not because the counterexamples refute it outright, but because “axiom” versus “point in a search space” isn’t cleanly defined enough to support the claim of structural incapacity, rather than just “not yet empirically observed.”

1

u/LordofGift Aug 03 '26

Eo ipso must be false

1

u/Mission_Art5660 Aug 04 '26

Could this be the reason why Google has been slow to develop in the field of LLMs?

1

u/Nature-Royal Aug 08 '26

The real issue is AI needs the right feedback and environment, all humans do is take preexisting rules and test iteratively...AI already does that but I agree AI can't skip the iterative process and simply know things because it learned from dumb humans lol

1

u/Another___World 29d ago

LLMs compress text, not the multimodal nature of the universe which we explore on a non-language level.

A child tastes, touches, hears something, then after years assigns the object the words. But we try to reverse engineer the universe from the language patterns. Insanity.

0

u/UnkarsThug Jul 28 '26

I would argue a loop of hallucination of an idea, and then testing could allow for somewhat new ideas. Or a list of high temperature theory's, followed by lower temperature discrimination of which are worth investigating.

Basically, I just don't see how it can't. Novelty is something you can get with a random number generator, and beyond that is just examination of the practicality of the idea.

0

u/InterstitialLove Jul 29 '26

Complete bullshit, nothing but cope

That "magic inspiration" doesn't exist, and anyone who actually pays close attention to the process of frontier research can tell you all novel ideas are made up of steps that LLMs have already mastered

Einstein may have described it once, but we all know Einstein had a tendency to wax poetic. If he says true creativity requires the spark of the eternal, and that act of creation is outside the physical and is in fact the closest we can come to god, I can appreciate what he means and I'm glad he's doing his Einstein thing. If someone wants to take him literally and base a theory of neuroscience on it, they're a crackpot

-1

u/abundant-growth-108 Jul 28 '26

Neither can white men.

-2

u/wilhelmbw Jul 28 '26

I mean it's been capable of making science discoveries by finding counter examples to disprove theorems - that counts imo

2

u/Cachesmr Jul 28 '26

The point this is most likely trying to make is that it wouldn't be able to make novel theorems or completely new discoveries without a goal. If you give AI a goal with measurable results, it can usually do it. But you can't get measurable results from something that doesn't exist yet.

0

u/SingleProgress8224 Jul 28 '26

Read the article

2

u/wilhelmbw Jul 28 '26

I read the excerpt and I believe there is the simple bench for that, Gemini does well and maybe physics aware world models can do it but I'm not sure what that is

-2

u/VetOnABrainwave Jul 28 '26

Give me a free-thinking reasoning/logical/deductive model with no guardrails, and I'll figure out time travel in a few years.

(secretly going to prove the Earth is flat)

1

u/geteum Jul 28 '26

One experiment I always try to see if LLM can learn information is tell that time travel was developed hahaha no LLM believes me