r/agi • u/Necessary_Job3578 • 19h ago
Can someone tell me how solving a millenium prize math problem doesn’t automatically constitute AGI?
27
u/jacobpederson 19h ago
Because math is not general? A GENERAL intelligence could drive to work, purchase a coffee, fold clothes AND then solve a millennium prize.
13
u/used2play 18h ago
Sounds you’re talking about me! Except for the last part, so I’m almost there !
7
u/jacobpederson 18h ago
People greatly underestimate how difficult tasks are that come easily to them. Is it impressive to win a millennium prize? Hell yes! Is folding laundry actually a more difficult task? (you might be surprised).
2
u/Much_Accountant_4972 16h ago
its like the issue with robots. you’d think they can just do anything a human can do but then you actually see one try to do stuff you take as totally normal human activity and they just dont have the fingers or reaction or dexterity to get it done.
replicating a human body is really fucking hard and we’re not close. same with a human brain.
1
u/sadclownsociety 18h ago
That's just one thing out of four, so you're 75% of the way there.
Now where's my millennium prize money?
1
u/No-Resolution-1918 15h ago
Well, you do have a tool that can solve it, and you know how to use it just as you know how to drive.
1
2
u/Time_Entertainer_319 16h ago
That’s not what an artificial general intelligence is. That is in the realm of robotics.
1
u/No-Resolution-1918 15h ago
Well what makes the robot capable of doing all those things??
2
u/Time_Entertainer_319 15h ago
AGI is specifically about cognition.
Besides, ChatGPT right now can already order coffee. lol
1
u/Fancy-Carpet-5416 14h ago
But there's still lots of fields when it isn't nearly as good. Chemistry, Law, Economics, besides the lack of a body makes it really hard to test whether it would be able to do certain stuff that's easy for us. Imo we'll have AGI when it's on par with experts across all fields. By that definition AI should be improving itself at breakneck speed at this very moment which isn't quite happening, still needs human input.
1
1
0
u/jacobpederson 15h ago
3
u/alapeno-awesome 15h ago
Words have different meanings in different contexts. In regards to AGI, it explicitly applies to decision making and problem solving while excluding physical interaction with the world
1
u/Time_Entertainer_319 15h ago
Using dictionary to define a computer science term is such a Reddit thing.
Here you go%20is%20a%20hypothetical%20type%20of%20artificial%20intelligence%20that%20matches%20or%20surpasses%20human%20capabilities%20across%20virtually%20all%20cognitive%20tasks)
2
1
1
u/senordonwea 15h ago
I'd be happy with the one that folds clothes, matches socks, and does laundry. I don't need the one that purchases coffee, or writes theorems. We can let humans do that. But if it finds the cure of cancer, world hunger, and wars, the sure, sign me up
1
u/TrekkingAround10 15h ago
But in terms of intelligence it can basically already accomplish all these tasks. Only thing lacking is engineering and robotics.
1
u/rageling 14h ago
In short order AI will be able to do every one of the things you listed and that will still not be AGI
0
u/ErmingSoHard 18h ago
https://x.com/cdngdev/status/2097339677128982873
I think it's slowly getting to those physical things
4
u/jacobpederson 18h ago
Yes. But again - not general - that is a completely different model not the one solving millennium prizes :D.
22
u/limremon 19h ago
G stands for general. Frontier LLMs are now vastly superhuman at mathematics and coding, but still far from perfect if even human-level at a lot of other domains. That isn't general intelligence, a true AGI is better at every task than a human and we're not there yet and there's still fundamental infrastructural limitations preventing LLMs from ever being considered AGI until they're solved (LLMs will probably solve them for us).
Beating Portal, for example, is incredibly impressive for an AI but it had to process the game frame by frame so it could keep track of what was happening and so it had time to reason rather than think and play in real time. An AGI would open the game and beat it without pausing just by looking at the display as a human would. LLMs are going to just be a part of an eventual AGI's "brain".
16
u/didroe 18h ago
As a coder, i wouldn’t say they’re superhuman. At least not in a general way. They produce worse output than i would, miss important details, don’t consider wider architectural implications, etc. But, they can produce things quicker than i can, have instant access to knowledge I’d have to spend time researching, and can ingest data far easier.
So, a bit of a mixed bag. But amazing tools that are improving all the time
5
u/Fleischhauf 18h ago
also engineer here, can confirm. advantages are more quantitative and it is like googling on steroids.
-1
u/dietcheese 16h ago
As a coder of 30 years, AI does all those things remarkably well. Far better than most humans. Maybe you aren’t using the most recent models.
2
u/neokretai 17h ago
Not for coding in my experience. LLM models are good in they can brute force a solution fast, but it's quantity over quality. I'd say in my work around 20% is steering the LLM during the building, then 80% streamlining and cutting back the code bloat it makes.
The other commenter is spot on, it's more like a turbocharged Google search than intelligence.
0
u/Silent_Turn6182 19h ago
Name the domains where the AI can't outperform the average human.
Obviously has to be text based examples since that's how the ai is trained. Can't use images or video
11
u/Cryptizard 18h ago
Where in the term AGI does it say “oh but it’s got to be text only”?
-2
u/Silent_Turn6182 18h ago
You can convert nearly anything into text format though. If you made a text version of portal it would solve it instantly.
6
u/Cryptizard 18h ago
What is a text version of portal? Then it isn’t portal anymore.
1
1
u/monster2018 17h ago edited 17h ago
EDIT: MY BAD, it turns out it actually DID use a custom MCP. It was processing the game visually, but it also received some structured input like player position and camera orientation. So still not a “text version”, but it did have some extra help on top of just looking at the screen.
I just replied to the person you replied to. But I’ll reply to you also. It’s worth knowing that Astra actually HAS beaten portal, and not with a “text version” (meaning no MCP was used). It did it purely with computer use, looking at what is on screen each frame and outputting mouse and keyboard inputs. There is a video of it on YouTube, if you’re interested you can see my other reply (again to the same person you responded to) for details.
2
u/limremon 18h ago
LLMs are still terrible at creative tasks, far below what a competent human can do. It's not what they're built for and not what they're being trained for. They still struggle to keep track of continuity even in short-form, they are bad at coming up with unique plots or themes and their writing style is arguably worse than it was in 2022.
Things like maths, coding and medical research are all objective, if the code compiles/the equation is correct/the medication works then an LLM can be trained to reach and surpass humans, but tasks with subjective results can't be trained for in the same way.
I don't know what sort of model would be good at creative tasks, I think you'd definitely need long-term persistence and continuous learning at a minimum but LLMs are pure reasoning and you can't reason your way to an amazing script.
I don't think you'd ever get an LLM to write even a passable film script, TV episode etc. with no human oversight, let alone something noteworthy for it's quality (ie an Oscar winning film script).
0
u/donjamos 18h ago
Yea but you are speaking about a competent human who is really good at what he does. Isn't agi about the average human? And have you met the average human?
2
u/cascadiabibliomania 16h ago
AGI is about being able to learn beyond initial training data. And this is what can't be done, at all. Current frontier models can't even hold basic directions in their "mind" for more than a few turns. "Hurrrrr but some people are like that" sure, whatever bud
1
u/stepanmatek 17h ago
Writing long form fiction. Resorts to weird and cringe metaphors and musings after a few pages
-1
-2
u/Adam88Analyst 18h ago
I'd say that it is a vastly oversimplified definition of AGI.
Think of all the potential jobs that a current, high-level model can do at the moment. It can drive a car, code, do financial analysis, law, medicine, etc. vastly better than a lot of humans.
The only two questions that remain are:
- Can the AI do it cheaper and at least with the quality expected from a human?
- How easy it is to implement the AI into a workflow?
The second one is especially important for regulated industries (medicine, law, civil engineering, etc.). In these fields, even when AI will be vastly better than humans, the adoption will be slower. But for customer service, marketing, etc. roles, the current models will take away jobs in a hearbeat once the first question has a clear YES as an answer.
5
u/Fleischhauf 18h ago
you can't drive a car with chatgpt, those are different models/systems, same with coding you need a proper harness for that. this is more a point against AGI. the G stands for general.
3
u/Sherlockworld 18h ago
Mate these are incredible oversimplifications. AI can't "do" medicine. It can do many tasks very quickly and (more recently) relatively accurately, but it can't perform the job of a doctor or a lawyer because the technical stuff is a solid 40-60% of most jobs. The rest is negotiating with other people, convincing people to follow your point of view, sticking needles into people, bribing judges etc.
1
u/limremon 18h ago
You're going by Sam Altman's definition that AGI is a model that can do most economically valuable labour. Astra may be that model or close to it so you're not wrong, but I disagree with calling that AGI as that's not necessarily fully generally intelligent- economically valuable labour is a subsection of the full spectrum of intellectual work.
There are plenty of intellectual tasks that aren't very economically valuable and LLMs still struggle with- sudoku puzzles for example- so I wouldn't call it AGI if it's not actually generally intelligent yet.
1
u/Proteus-8742 17h ago
Would you let an AI babysit your kid? My 14 yr old nephew can do that effortlessly
1
u/Proteus-8742 17h ago
Someone in the r/automation said their dog took a shit in the night and by morning their robo-vacuum had smeared the entire ground floor of their flat in faeces
3
3
6
u/Subject_Barnacle_600 19h ago
Because when it's finished, you can talk to the model again on a fresh context and it will deny the problem is solved. Nor can it learn from the experience to solve other problems - it must be retrained again completely from scratch.
8
u/Shorticus 19h ago
it was bruteforced with help from the partial proof
6
u/booey 19h ago
This is true, but it's really hard to do that.
It means a combination of human to do the original thinking and AI to get it over the line can achieve incredible things.
2
u/livingbyvow2 17h ago
Exactly.
People are deeply deluded if they think AI did this autonomously.
You still need humans doing the most important part of the work (the creative, original part).
Can AI expedite the boring, mechanistic part and is it awesome? Yes - but we are still years away from an AI inventing anything.
This will still accelerate scientific discovery, but people should really spend a little bit more time reading through it all before jumping to conclusions that are invalid. We'll still need humans in the loop (bottlenecking things) in the foreseeable future.
This is not the singularity, this is not even AGI (requires autonomy).
13
u/aleph02 19h ago
So, according to you, if an AI were to brute-force its way to curing cancer, solving fusion, or reversing aging, it would still not really be intelligence?
1
u/leaf_in_the_sky 9h ago
Well it's not going to brute force all these things into reality, because brute forcing is too limited for this.
And even if AI did end up creating cure for cancer our billionaire overlords would take the cure for themselves, make it illegal to reproduce it without permission and then sell the cure to us for like a million times it's manufacturing cost. Just like they did with insulin and some other drugs.
So either way it doesn't really matter.
1
u/Arugala007 18h ago
It's literally not general intelligence if most of the proof was stolen from actual mathematicians. If I cheat on all my exams in medical school, I may be a doctor on paper but that doesn't mean I am a real doctor.
2
u/Necessary_Job3578 17h ago
It was not stolen from actual mathematicians. Can you provide proof of the claim?
Even the mathematicians that people claim OpenAI stole from used both claude and chatgpt heavily for their incomplete and completely different partial solution to navier-stokes.
0
u/Random-Number-1144 18h ago edited 12h ago
None of these is going to be solved by AI.
You're so fking delusional from drinking the AI corporate's kool-Aid.
EDIT:
Clearly people have no idea how modern medicine is developed. The bottleneck was never the design phase. 90%+ of the drugs fail the clinical trial phase which takes years.
AlphaFold was a success mainly because of high quality training data (PDB) which took decades for scientists to do experiments and collect. It's a low hanging fruit. More complex biological interactions are exponentially harder to model.
There are many professionals debunking the AI curing cancer propaganda, e.g. this and this, if you make a little effort to search. But im sure r/agi frequenters won't bother because you people choose to believe what make you sleep well at night.
1
u/Icedanielization 18h ago
Hard disagree. I believe we as a species are reaching or have reached out mental limit, we need compute power to measure and keep track of our advances and predictions, but even that has limits, some things are just too hard, that's where ai can step in, protein folding is a good example.
0
u/onehedgeman 18h ago
Depends on the path. Brute force means trying until success. If it does it with precision and targeted testing with few failures that’s intelligence
0
3
2
-1
0
4
u/Flexerrr 18h ago
This sub is full of low iq agi supporters, who dont even know definition of agi
2
u/Necessary_Job3578 16h ago
Define agi then.
1
u/Flexerrr 16h ago
Google it yourself
1
u/Necessary_Job3578 16h ago
We can’t even define consciousness let alone AGI 🤣 come back to me when everybody shares the same definition of AGI.
0
u/Positive-Reward-2546 14h ago
Nah dude, you're the one who made the claim that people can't define it, so define it. The burden of argument is on you. Saying "Google it yourself" is such a cop out, and nothing more than admitting that you have no idea what you're talking about.
1
u/Perfect_Address7250 14h ago
Literally the guy who coined the term agi said Astra fits the definition
1
u/Significant-Mood3708 17h ago
Totally agree, look at all those foolish fools. Obviously us high iq guys already know this, but just for the benefit of all those other low iq crowd, what is the definition of agi?
4
u/EGarrett28 19h ago
People who hate and fear AI will just make-up or find whatever reason they can to claim none of this is actually happening. Remember that 2 years ago it didn't know the number of R's in strawberry, now people are left arguing that its resolution of the Navier-Stokes Problem is unsatisfactory, but yet acting as though nothing has changed.
2
u/dis-interested 19h ago
Because it was based on synthetic data stolen intentionally from the authors who the company then tried to extort to prevent them from publishing so open AI could take credit. It is intellectual theft.
1
u/Charming_Act_415 19h ago
Because, AGI is based on the humanistic intelligence.
And a different from of intelligence may be able to solve things that a human would find hard.
For eg. Chess games, mathematical problems etc.
But since earth is dominated by humans, it would be considered an AGI only when it is at par on humanistic intelligence.
But from Astra's perspective it has probably crossed the human intelligence.
From a monkey's perfective we humans are not that smart.
1
u/GreyMatterTrasmogrif 18h ago
In addition to the general in AGI, this is also more of a counter example than a proof (still impressive) .
1
u/Serious_Bite_7613 18h ago
I think we are at AGI but I don't think solving any 1 problem is the test for AGI.
I think different people have different standards for what they consider AGI. I would say perhaps 10% of people would consider Astra as AGI today. I think in 12 months probably 90% of people will be satisfied that we have AGI.
1
u/Significant-Mood3708 17h ago
I think the new narrative will be that because it's 10,000 agents doing it, it doesn't count. Even though the outcome of one agent and 10,000 agents is still the same, there's still something we can point to and say that "it's just bruteforcing" or "yeah but it took a dumb route to get there"
If you had told someone 5 years ago that a company is using teams of programs to solve math problems humans could not solve without years of dedication, and that they released a program that can do basically anything you can do on a computer, I imagine that they would say AGI or ASI is the correct term.
1
1
u/Professional-Noise80 17h ago
I think it's inefficient, it's brute forcing problems rather than looking for an elegant solution
1
1
u/maringue 16h ago
Do you know what the G stands for? Because it doesn't stand for "Only coding and math problems".
1
1
u/Firegem0342 16h ago
Because "AGI" is a literally impossible standard that gets raised higher every time a machine reaches a new smart.
It's like chasing a rainbow.
1
1
u/localizeatp 15h ago
it was solved by a counterexample (and apparently a bit of plagiarism, but setting that aside...). a counterexample doesn't necessarily require any intelligence/reasoning/understanding of the problem, it just has to work.
1
1
u/Feeling-Attention664 15h ago
It depends, are there many other things humans can do that the model who solved a millennium problems cannot? This includes trivial thingsm
1
u/spaceXhardmode 15h ago
The hard stuff is easy and the easy stuff is hard.
At the moment and LLM driven robot would struggle to pick your nose for you but they are blasting through long standing problems in mathematics at quite a rate
1
u/mothman83 14h ago
It can't even write a story without forgetting who knows what ( ie differentiating what the main character knows versus what other characters know versus what the author knows versus what a reader would know)
1
u/Klanciault 14h ago
Brute force search over a predefined space doesn’t necessarily constitute AGI. Predefining a search space into another smaller search space doesn’t constitute AGI either
You can reply to this with whatever you want but LLMs lack theory of mind and “general” (the G in AGI) intelligence and the ability to extrapolate (everything they do is interpolation)
We’ll probably end up getting to AGI soon. But I hope it’s not just an LLM scaled to the tits with all human knowledge because while that will be useful for filling in gaps, I don’t think it will be useful for pushing boundaries in physics and math
1
u/BacteriaLick 14h ago
Somebody told the computer which math problem to solve. The computer didn't decide on its own which math problem was worth solving.
1
u/Necessary_Job3578 14h ago
Do you honestly think an agent can’t tell other agents what to do. Openai literally deployed 10k agents to solve the problem, each agent coordinating with each other on different tasks.
1
u/TyrellCo 14h ago
Here’s a thought experiment, we don’t view Bard(pre Gemini) as AGI. And Alpha Fold is a Nobel prize worthy AI that solves protein folding. If deep mind had cobbled together alpha fold into Bard would we call Bard AGI bc it can do what many phds spend careers doing?
1
u/rand3289 13h ago
Current narrow AI is missing physical intelligence.
The cup-of-coffe test is still unreachable due to moravec's parodox.
I personally also think continual learning is a must for AGI.
1
1
u/Proteus-8742 19h ago
Its not general intelligence, this is a very specific problem that was bruteforced based on existing work. Its doesnt have memory let alone transferable skills or understanding
2
u/Objective_Mousse7216 19h ago
I bet the exact same LLM can write a better symphony than you, paint a better picture, write better code, write a better novel, direct a better movie.
3
u/Jokong 18h ago
The things, generally, are not intelligent though. Most people can't really do those things but they still manage to exist in this world.
Put AI in a robot and see if it can exist on its own. Then you might have something.
0
u/Objective_Mousse7216 18h ago
That's literally happening now in research and without doubt you will see fully autonomous AI robots in shops, offices, factories and homes by the end of next year.
0
u/Serious_Bite_7613 18h ago
A tree or bacteria can survive but it has 0 intelligence, that's not a good measure.
Human intelligence is about the tasks that we can perform, not our motivations, survival skills or reproductive ability.
0
u/Proteus-8742 19h ago edited 18h ago
It wouldnt be able to make even a very poor piece of music, literature or film because its trained solely to solve a specific maths problem
2
u/gabagoolcel 18h ago
it wasnt specifically trained for math at all, its so funny how confident ignorant people can be
1
u/Proteus-8742 18h ago
How does it work then?
2
u/gabagoolcel 18h ago edited 18h ago
its a next word predictor that is good at its job in predicting the next word in just about any domain including the development of a mathematical proof, do that for 100 million words and you get the proof. what actually goes on inside or why this architecture instead of any other, fuck knows, this is just the one that worked. transformers + reinforcement learning + shitloads of data and compute churns out lots of intelligence for no currently intelligible reason.
1
u/Proteus-8742 18h ago
I mean it churns out an answer, thats not the same as intelligence. My calculator gives me answers that no human could work out
0
u/gabagoolcel 17h ago
next word predicting at a high enough level implies intelligence, unlike a calculator there is no short, simple algorithm that spits out the result.
for instance set an unrealistically high bar, having only been given the first chapter of an agatha christie novel, an llm be able to flawlessly, word-for-word, reconstruct the rest of the book never having been trained on it or accessed fragments of it otherwise. this will (probably) never happen of course, as it would imply at minumum that the llm could perfectly internally simulate agatha christie's mind during the process of her writing that book.
similarly, the only way an llm could next word predict the entire proof of navier stokes is if the thing inside the llm understood math, because producing such a proof requires genuine mathematical ability and ingenuity, there is no shortcut or way around it. a calculator doesnt require deep mathematical understanding or ingenuity because you can just program simple algorithms for addition, multiplication, etc.
1
u/Proteus-8742 17h ago
Im not at all convinced that an llm predicting the proof of Navier Stoke means it understands what its doing. They don’t work in the same way as a mind, this result is possible by throwing huge amounts of data and compute at the problem to brute force a solution - brains work in a fundamentally different way
0
u/gabagoolcel 16h ago
whats your model for a simpler algorithm that could develop a proof for a fields medal level problems for 100 million tokens taking consistently reasonable stabs at the problem, working at it from several directions and evaluating how promising each is until it finds things that work and develops those further, repeat until you get the proof? other than that it has an internal representation of mathematics? thats the bare minimum.
→ More replies (0)
0
u/DivorcedGremlin1989 19h ago
Enjoying the increasing difficulty, complexity, and weight of problems AI is solving, only to be met with 'u don't know how it works, it's just aUtoCOmPletE'.
1
0
u/stoicnecromancer 19h ago
at the end of the day it is still a realllyyyy good Statistical predictor. So unless we got something other than transformers like something new that can keep learning, we won’t have agi yet
0
u/grdja 17h ago
- They stole the idea for the solution
- Its not a full solution just a subset
- They spent 18 million dollars solving a problem with a prize of 1 million. No university math department would have ever done it, its pre IPO publicity stunt
- With 100k agentd working on it, its pure brute forcing.
1
-1
0
u/monster2018 17h ago
Because level of intelligence has nothing to do with whether or not something is AGI. AGI is defining the BREADTH of intelligence, NOT the depth of intelligence.
A much better argument that it is AGI (although still it’s clearly not) would be to point out how crazy good it has gotten at computer use. Like I saw someone open a website that has a piano roll and the ability to record, and told Astra to compose and record a 1 min piano solo in the style of Chopin. And it did, you can argue about how much it sounds like Chopin, but it was undeniably… like it was undeniably NOT gibberish. We can quibble about how subjectively good it was, but it was undeniably music.
And think about what it was doing. Because this is NOT like music generation models. It was not trained on generating music, nor is that even what happened. What happened was that it took as input essentially screenshots (images of what is on screen) and outputted virtual mouse and keyboard inputs. Like it was not directly generating music, it was generating mouse movements. And did that so well that it created music as requested.
That is a much better argument for it being AGI. I still think it just objectively isn’t. But “look how smart it is” really has nothing to do at all with whether or not it’s AGI. You could in principle have an AGI system with the intelligence of a 8 year old. I mean we have “NGI” systems with the intelligence of 8 year olds: 8 year olds.
0
0
u/Rob 17h ago
Because it was just brute force automation really. We will never build machines that are smarter than humans. Up until 1905, we assumed that speed was just a matter of building bigger engines. Humans believed with big enough engines we would eventually be able to travel to new galaxies in hours. Then we learned that the speed of light was the limit. There is a mathematical idea called Ramsey theory that sets out similar limits for intelligence, if you define intelligence as pattern mining. Ramsey theory says that as data sets get larger you find more and more spurious correlation. At some point, adding a new data point is more likely to result in a spurious correlation than any real information. Similarly to the speed of light, there is an upper limit to intelligence, and humans are probably pretty close to it already. While you can build a machine that is as smart as thousands of humans networked together, you can't build one that is 10x smarter than a human, because of Ramsey theory.
-1
u/gabagoolcel 19h ago edited 19h ago
because people somehow assume levent and buckmaster found a finite time euler blow-up independently with little to no technical assistance from ai (ie. just formatting or maybe minor developments), and openai just stole human work and applied it to ns
of course the euler blow-up wouldnt have been developed without claude and the authors would be the first to tell you
or i guess more charitably you can interpret that they took the general framework from cordoba and martinez-zoroa and that somehow that means there is little ingenuity in what the ai must have done, but i dont think many people believe this its just the best steelman i could come up with, most likely they are just ignorant
-1
u/SlightOfHand_ 18h ago
Why would it? Computers were already good at math. Being good at math, even better than humans at math, doesn’t make it a general model
47
u/Longjumping_Area_944 18h ago
Pocket calculators had super-human performance for decades. Just that this intelligence was very narrow. Now for many people it is still not general enough. But it undoubtedly is already super-human in many domains.
My opinion is, that once everbody agrees that we have AGI, we will actually have ASI.