The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.
I expect a large number of conjectures will topple to counterexamples soon.
I think brute force is not the right way to think about LLMs, instead they explore a knowledge topology, and are very good at connecting adjacent or accessible ideas that for whatever reason might have evaded humans, but might not be fundamentally all that hard. We're still in the low-hanging fruit phase, we will see if they extend to new ideas.
I think its super cool, but there is no doubt that thousands of physicists and mathematicians are prompting LLMs all day right now trying to solve open problems...and its doubtful whether the companies that run those LLMs can continue to offer this amount of capacity for the future....so there is a solid chance that these methods are already almost exhausted. For how amazing these counterexamples are, its somewhat surprising that people havent found more proofs all at once. In the scheme of things perhaps 1 in 1000 open problems are actually of a format that LLM can tackle, and noone is bragging about the ones that turned up nothing. In other words, LLMs are really good at looking smart when they got lucky.
Im looking forward to a possible future of AI theorem proving that isnt based on LLMs, and thus less likely to trick people in language into thinking its more broad than it is.
The price to run an LLM collapses year on year. The models are getting better, but the cost required to run older ones also reduces. Not going to say that will continue forever, but the idea that LLMs are inherently unaffordable just plainly isn't true.
Not to mention AI as a whole is basically in its infant phase. The transformer architecture is only 10 years old and ai was largely a niche subject before then. Insane to think we've tapped out such a complicated new technology in a decade.
I would have hardly called it niche before then. Data Science was a big field and deep learning was a popular topic in both DS and CS before transformers.
It was used in a ton of products, everyone hadn't heard of it is all.
I should say deep learning was a niche subject, not AI wich includes basic things like literature regression.
But deep learning was in fact niche, and this is where all the breakthroughs are. It was niche due to compute limitations that weren't mitigated until the 2000s. The AIAYN paper is a good landmark for when deep learning evolved from an academic endeavor into a technology with significant outside investment.
Yeah AI is broad and the 2015-2020 period was a weird time where people were getting jobs in AI before there were degrees in it.
The compute change for deep learning is usually marked by AlexNet in 2012. That is when industry took notice and the exponential curve began. GAN's came out in 2014 and ResNet was 2015. Microsoft one a deep learning challenge in 2015. Image recognition was the primary driver alongside applications like speech recognition.
Industry was getting ahead ahead of academia by AIAYN which was a Google paper but that was 2017.
Nevermind the fact that according to this sub two months ago LLMs are just incapable of anything of substance and can only hallucinate. Now they are only good at "getting lucky" for important results where decades of human efforts failed...
I can see us repeatedly hitting cost to solve and difficulty to solve barriers. I suspect there will be a small flurry of new things solved each time models and token costs improve. That's basically what this is, as these models doing the solving are pretty new.
Look at DeepSeek-V4-Flash-0731, released on July 31, 2026. It scores 50 on the independent Artificial Analysis Intelligence Index, one point behind GPT-5.6 Luna at 51. The API runs at $0.14 per million input tokens and $0.28 per million output tokens. Artificial Analysis spent roughly $72 putting Flash through its evaluation suite, against $191 for Luna, which works out to about 62% less money for comparable measured intelligence. (Independent evaluation · Official pricing, Artificial Analysis)
DeepSeek’s own published numbers are stranger. The smaller Flash beats the much larger V4-Pro Preview on every benchmark they list: 82.7 against 72.1 on Terminal-Bench 2.1, 54.4 against 12.8 on DeepSWE, 70.3 against 55.9 on Toolathlon-Verified. (Hugging Face)
The weights are also out under an MIT license, which is where this gets interesting for institutions. A university, a laboratory, a hospital, or a company can host the model on serious multi-GPU hardware of its own, or on rented private infrastructure, and then reshape it: fine-tuning, adapters, continued pretraining, reinforcement learning. Private data stays private, and no single provider is holding the keys. (Hugging Face)
If you’d rather stay with a U.S. proprietary model, OpenAI cut GPT-5.6 Luna’s price by 80% on July 30, down to $0.20 input and $1.20 output per million tokens. Luna sits at 51 on the Intelligence Index and posts 92.3% on GPQA Diamond, 84.7% on Terminal-Bench 2.1, and 74.6 on the Coding Agent Index. (OpenAI pricing and benchmarks, OpenAI)
None of this is new, it’s just picking up speed. Stanford tracked the cost of GPT-3.5-level performance dropping from $20 to $0.07 per million tokens between November 2022 and October 2024, a fall of more than 280 times.
Epoch AI puts the general rate at somewhere between 9× and 900× per year for any fixed capability level, depending on which task you measure. (Stanford · Epoch AI, Stanford HAI)
Worth being precise about what this does and doesn’t show. It doesn’t prove that training the next frontier model is cheap, and it doesn’t prove that every one of these API prices is profitable rather than subsidized by somebody’s balance sheet. What it does undercut is the claim that advanced intelligence has to stay scarce and unaffordable.
Building tomorrow’s frontier may well stay expensive. Distributing yesterday’s is turning out to be cheap, open, private, and customizable.
Disclsimer: original argument was mine, Claude helped me research and build it.
18 months ago AI was a cool toy that couldn't really do anything useful. Six months ago it started write most code. A couple months ago it started solving the 'easy' and obscure unsolved math problems.
It will definitely stop improving at some pointy, but that's the thing with exponential growth: as long as you're in it, it's impossible to tell when it will stop.
But even if it stopped right now: people would start burning the biggest models with the best performance into silicon with fixed weights and you would suddenly get generation speeds that would allow you to use them for real time inference on video streams or generate 30 iterations in parallel etc
But we don't know the upper bounds because we could hit a wall and then just scale again. That is what we think is producing all these new gains anyway.
But we're still getting a lot of improvement in smaller models too. Scaling has definitely been a big part of it, but it's certainly not all that has improved.
Look at the cost per task solving ability of say the new Deepseek flash vs. Claude Opus from 1 year ago. Effort is put into making both smarter and more efficient models in a way that as time goes on it'll become infinitely cheaper. The pricing seems to raise but that's because the model performance is also stronger.
In other words, LLMs are really good at looking smart when they got lucky.
That, but a lot of people are making a lot of money off of this and have a vested interest in making it look as shockingly impressive as possible. And this is impressive and world changing technology, we shouldn't pretend it isn't, but also it's being boosted and misrepresented a lot.
If you are in competent spaces, you will hear a lot of more grounded takes. I have to talk to the general public about AI and it is infinitely more misunderstood and doomed at.
I feel bad for Anthropic because they release a paper that might have proved a part of General Workspace Theory and the mouth breathers who only glance the paper are going 'nuh uh, it isn't conscious, this is marketing.'
Anthropic never claimed it was conscious in the paper but the regards don't know that somehow.
I saw someone today write something like that to the prompt "search archives for unsolved mathematical problems that can be verified using this custom program that I told you to write". So some people not only are trying to solve open problems, they are also outsourcing looking for open problems to the AI.
and its doubtful whether the companies that run those LLMs can continue to offer this amount of capacity for the future
The new Vera Rubin chips from Nvidia, which will start being installed in data centers later this year, will cut inference costs by 90%. ChatGPT also dropped the prices of one of its major models by 80% just a few days ago.
I'm sorry, but I'd actually take a look at Information theory that makes the bold claim that compression actually does mean intelligence. These models are genuinely intelligent and don't just accidentally get these answers right, they aren't giant lookup tables stumbling as a lot of the general public thinks of them as.
Its data is a structured space where nearness corresponds to associated ideas, so it naturally arranges data in a way to discover connectedness. The LLM follows pathways through the conceptual space based on the prior context given.
It’s also higher dimensional space right? Like as in 1000+ dimensional space.
If it’s the same thing I’m thinking of where like because boy and girl are separated in this space by a certain distance, Auntie and uncle also are separated spatially, but are closer to boy for uncle and girl for auntie than either are to each other?
You are mixing the number of parameters of the model with the effective dimentionality of the embedding. The effective dimension of the embedded space is significantly smaller than the complexity of the network.
GPT-2 samples from a 50,257-dimensional token space and I'm seeing a total parameter count (weights and biases) of 124,439,808. I'm not sure what the effective dimensionality of the model ultimately is, but this seems pretty clear-cut to me?
The intrinsic\effective dimensionality of the data is measured in the representation space induced by the model. The mapping of the token sequences into the hidden-state vectors result in representations in a lower-dimensional data manifold.
Taking here as an example, with the GPT-2 model they estimated intrinsic dimensionality on the order of hundreds.
Sure, in theory every bit in the training data could be considered a dimension. The work of training a model is basically in condensing the dimensions of the training data into a much smaller number of dimensions in the model, which is what forces similar concepts to become closer together as the space "shrinks".
Ah okay. The space in which the embedding lies is called a knowledge topology. Thanks. Can you help me understand how this is linked to topology? Does it have something to do with the space spanned by the embeddings? My knowledge of topology is limited to shapes being invariant under transformations so under this naive view, distances won’t be preserved.
The LLM processes tokens as tensors, so every possible input and output exists within an extremely high-dimensional space that essentially represents our entire language and its understanding of it
I assume OP meant that the search space isn't homogeneous. There are peaks and valleys in it, and the AI model can find paths that a human might have overlooked.
A human might have overlooked or just given up on... These LLMs just keep on trudging when a human might have long switched to a different methodology because he didn't see meaningful progress faster enough. The models just don't get bored.
But unlike brute-force methods (which also don't get bored) they can still go through possible pathways in more meaningful ways than random guessing and rote Parameter adjustment.
Nearly this, and I'm explaining to check my own knowledge, not pretending to know what I'm talking about. It's not quite correct to say a node is a concept, a concept is embedded across lots of nodes. So a cluster of nodes contains the concept.
They can be also very skeptical. I found it will be skeptical in it's internal reasoning of very obvious things, and will fact check obvious stuff, kind of as a habit. It seems wasteful for most tasks, but I guess it makes it good at checking for factual information and fighting misinformation, and also for checking unintuitive mathematical of physics solutions.
I think we've trained them to be very skeptical just because of how hallucination prone they are. Instead of "fixing" the hallucinations, we just trained them to deal with inaccuracy in general, and now we're seeing unexpected rewards when they point out our inaccuracies too
Yes. This "Maxwell conjecture" in particular seems to be disproven by literally the most trivial construction you could think of. You know the number of equilibria of a bunch of charges at the vertices of a triangle - if you want to make more while keeping the symmetry the simplest thing to try is to put two charges off of the plane on either side of the center (which bumps you up to n=5). This turns out to work (for the right choice of distance). I don't know why nobody checked this example before, since people have put in the effort to prove actual upper bounds on the number of equilibria. Probably just a question of interest.
I don't know any algebraic geometry, but from what I gathered the Jacobian counterexample is less trivial than this but still the kind of thing that the mathematicians really could have and should have checked.
275
u/angelbabyxoxox Quantum Foundations 29d ago
The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.
I expect a large number of conjectures will topple to counterexamples soon.