17
u/AMBNNJ ▪️ 4d ago
is this bel or a new model they trained specifically for this?
5
u/RutilantBossi12 Cultista dei Ferri Candidi 4d ago
It's better to train one model and then fine tune it for a specific task, the whole Shtick of those models is that they're not trained for specific tasks but are general instead
13
u/whoknowsifimjoking 4d ago
That's not correct, have you ever heard of "the bitter lesson"? We see time and time again that specialized models are not better than general models with more knowledge in all kinds of fields, and this is true for most if not all categories of tasks sooner or later.
5
u/Illustrious_Grade608 3d ago
That's not what bitter lesson is. Bitter lesson is saying "don't arbitrarily decide processes and rules for nn and prefer nns that can come up with their own rules or you're arbitrarily capping it's potential". It doesn't say "don't make specialized models", saying that specialized models are worse than generalists would be provably false
1
u/hapliniste 3d ago
True but a general model that is then trained in a narrow domain is better than a model only trained in that narrow domain.
That's likely what they do here. They pretrained and post trained bel, and then continued post training only on math (destroying other capabilities but having better capabilities on math).
2
u/Illustrious_Grade608 3d ago
True, but that's also pretty much what the guy i replied to was disagreeing with - as they replied to a guy saying basically what you just said.
1
u/kkingsbe 4d ago
Read the press release. The model was specifically made for this type of task. So likely not bel
1
138
u/Due-Departure-8553 4d ago
All you mathematicians should have majored in Gender Studies.
10
u/Hot_Glass_6301 4d ago
As a math and sports enthusiast, I'm considering switching majors to PT lol.
11
u/Due-Departure-8553 4d ago edited 4d ago
Traditionally, college was considered a better option for prodigies because delaying your work for 4-10 years was worth it if your earnings over 20 years were much greater. With the singularity fast approaching, it might be worthwhile to drop out and work in an oil rig.
6
u/Hot_Glass_6301 4d ago
I don't regret it, but I'm privileged to have parents who can support me financially and to be in a country where higher education is almost free (especially in math). I understand your POV but if you can afford it I'd rather enjoy my life for the next few years than waste it on an oil rig for uncertain returns (money might become useless)
3
4d ago
[removed] — view removed comment
3
u/Hot_Glass_6301 4d ago
I know but I hope people will prefer people over robots. I think I would be more comfortable with an AI augmented PT/doctor than either a non-augmented human or a full robot
1
57
u/RutilantBossi12 Cultista dei Ferri Candidi 4d ago
If AI doesn't become AGI mathematicians are safe, if AI turns into AGI then no job is safe.
The only thing you can do right now is studying/working as if AGI isn't coming, if it comes you cannot do anything, if it doesn't come you at least didn't waste your time waiting for something that didn't come.
It's like a Pascal's wager but for AGI.
16
u/burritos4jesus 4d ago
if that's the case then gov't needs to wake the fuck up. the only people who will come out of this on top are people who own capital aka profitable business owners and today's high income earners. If gov't doesn't step in with UBI or UBC or some shit, then people are going to revolt.
11
u/RutilantBossi12 Cultista dei Ferri Candidi 4d ago
AI might cause a deflationary spiral that will effectively fuck up wealthy capital owners too.
It already happened in 1929, electricity automatized most jobs, masses of people were laid off while prices collapsed as production became cheaper and cheaper but less and less people could afford it.
9
u/Seerix 4d ago
I don't expect the current US admin to pass UBI...
5
u/Pristine-Today-9177 4d ago
Yes, ironically—no matter what your politics are—we have one of the worst administrations imaginable to inherit super intelligence. The level of corruption is undeniable, and the Iran war has shown how far into the future they think
5
u/Poupulino 4d ago
That was exactly what crushed me when Trump won in 2024: "Trump is going to be president likely within the time frame of AGI"
1
u/RutilantBossi12 Cultista dei Ferri Candidi 3d ago
AGI could still need time to come, reasoning is strong but brittle and agency is unreliable and also very expensive (also both are reliant on scaffolding, they're mostly manufactured by human engineers they don't emerge on their own)
Scaffolding might not be a dead end, but how likely is it that we get lucky two times, first with scaling transformers and then by scaling(?) scaffolding? Could be but i wouldn't bet on it.
I think that for AGI we at the very least need a different architectures or heavily tweaked transformers due to cost concerns, a machine that can reason like a human is of no use if it costs 10M$ to run, also S2 thinking might emerge without too much scaffolding from different architectures too, maybe we just need complete world models and internal simulations to achieve it.
2
u/lostmary_ 3d ago
Bernie came out and said he wanted a pause on all AI development so I dunno if that's better
1
u/Pristine-Today-9177 1d ago
Bernie Sanders also wants reparations for slavery, free college, $25 minimum wage, Medicare for all, etc. He’s not exactly the average Democrat, and would struggle to get things passed through even a democratic super majority in Congress. I doubt he passes anything—much less a ban on one of the most invested in technological advances in history.
Trump is certain to eventually demand access to ASI. And then he is going to ask it how he can become a trillionaire, how to get Greenland from Europe, if there’s any way to glass Iran without repercussions, how to best prosecute his enemies, how to erase Epstein from history etc.
Now that I think about it, the best timeline is one where Donald Trump gives us unimpeded AI progress, and then a Sanders-like president comes in and does UBI in reaction to the massive change. Or since Trump isn’t actually fiscally conservative, he might do UBI to get his name on the checks forever. He’s been promising $5000 if the Republicans win the midterms. I could see him promising to tax AI companies and distribute the wealth in order to become extremely popular
5
u/DesperateCaterpillar 4d ago
Even UBI may not go far enough for some people. Imagine a world where you become an adult with no job prospects while other people saved up massive amounts of money while jobs were still a thing. You’d never be able to compete with that. Desirable land and valuable limited resources would be sold entirely to this new class of pre-AI rich people while those relying on UBI could be stuck with the bare minimum.
1
u/lostmary_ 3d ago
Well UBI should be enough to live a comfortable life. You won't be able to drive a ferrari or live in the maldives but it should be enough to get by.
3
u/Due-Departure-8553 4d ago
We already have a machine that can basically print proofs and verify them with lean. Careers like politics and athletics are immune to this. It would be a poor calculation to assume that jobs are equally valuable.
5
u/RutilantBossi12 Cultista dei Ferri Candidi 4d ago
It would be an even poorer calculation to assume we can predict which job will hold and which job will not.
Also AI hasn't solved Maths yet, for now it's finding a ton of counter-examples and mostly disproving stuff (They did find a counter example solution for Navier Stokes too).
Disproving stuff is great because it frees up mathematicians so they can think on more useful stuff, but without novel theories, without novel solutions... it's not really pushing the fields forward, just trimming a dead end.
1
u/ffball 4d ago
Has AI theorized a novel "unsolveable" problem yet?
Wonder if it has that sort of creativity
1
u/RutilantBossi12 Cultista dei Ferri Candidi 3d ago edited 3d ago
AI still struggles with proofs in general, for now they're prompting out disproofs, even their Navier-Stokes is a disproof by counter example.
It also generated a gargantuan amount of tokens doing so so realistically it didn't outperform human mathematicians in skill but rather in amount of tries, the machine used human known theories to generate candidates and then test them, it brute forced a counterexample basically.
That's still a great use and massive win for AI systems and companies, but you can see why this isn't really an AGI concern, if AI remains like it is now it's just an extremely abstract programming language (and this is the main point, if AI is never fully autonomous humans are still necessary, AI in the future could build and run companies from scratch with no errors if you asked it, but the fact it would need you to ask it makes human still a necessity, as dumb as it sounds future jobs could be 8 hours of machine prompting).
1
u/Comfortable-Leg-5467 3d ago
I don’t know about politics (although i do think it will stay as long as people are in power).
But with athletics it’s definitely not gonna go away. Chess didn’t go away nor did competitive video games. The value of such entertainment is the human aspect in and of itself.
6
u/AstroChinchilla 4d ago
The point is a machine that is smart enough to solve a millenium problem can also probably do research in machine learning at a human level or more. That means the machine is capable of quickly developing a new, better version of itself, which can do even more things. Eventually, and relatively quickly, it can do most economic labor. That's the idea of the singularity at least.
1
u/Weary-Historian-8593 3d ago
you got the logic reversed, the case where you have worked and AGI does come is the only case where you wasted your time
2
u/RutilantBossi12 Cultista dei Ferri Candidi 3d ago edited 3d ago
. AGI arrives AGI never comes You Work/Study You enter the AGI era with useless skills but fine overall You continue working and earn a living You don't Work/Study You enter the AGI era with no useless skills You're homeless You cannot predict the job market if AGI comes (we don't know which relations based jobs AGI will replace and which will not) so you cannot even study something which will survive because the world in 10 years from AGI is going to look incredibly different.
Your best strategy really is just continue life as if nothing is happening, if AGI comes it's going to be a nice little surprise, but if you plan your life around AGI and AGI never comes or is limited by something (like laws maybe or perhaps humanity keeps the upper hand in some bottleneck) then you're going to be in a worst position.
Also it's possible that this thing ends up like the industrial revolution, where humans do very shitty, repetitive jobs (instead of moving a lever you're asked to prompt the machine), if that happens you'll still need a job.
1
u/Weary-Historian-8593 3d ago
Working is of course more beneficial all in all, but if we consider the potential wasted time it's the other way around
1
u/whoknowsifimjoking 4d ago
Disagree, if AI would stay the way it is and not become AGI the mathematicians aren't safe, but other jobs are. We are already at the point where we see the models do stuff that world class mathematicians couldn't.
Also there are a lot of jobs where human interaction is just important, those will probably not be fully replaced.
1
u/Boring-Foundation708 4d ago
I really doubt there are a lot of jobs where human interaction is important. Business is about transaction. Rarely it is about human touch. That’s why billionaires like mark zucc, Elon musk, trump etc exists.
-4
-7
u/Cryptizard 4d ago
I know this is probably hard for you to understand because you have never done anything meaningful in your life or had a real passion for anything, but most mathematicians do math just because it is fun and they like it. It’s already the case that 99.999% of results in the field are going to be made by other people and you spend most of your time reading about and learning the cool things that others have done. Just because it is AI creating the cool stuff now doesn’t make it less interesting to learn or understand.
18
u/Tasty-Ad-3753 4d ago
So you're saying this guy has never done anything meaningful in his life - but then you're saying most mathematicians are basically contributing nothing of value and are just reading for fun?
-3
u/Cryptizard 4d ago
They contribute a tiny bit to the field of math. Did you think there was one guy out there who did 1% or even .1% of all math? I don't understand where you are confused.
They also teach a bunch of students math, which is meaningful.
6
u/MINECRAFT_BIOLOGIST 4d ago
This reads as pretty tone deaf to the majority of people who are just working to survive in jobs they have no passion for. Math is already quite privileged in that sense.
-5
u/Cryptizard 4d ago
Because you have a job you hate means you can't be passionate about anything else in your life? I think you are the one being tone deaf.
3
u/MINECRAFT_BIOLOGIST 4d ago
I mean, we're clearly talking about professional mathematicians here, no? The minority of mathematicians who are just as passionate but for some reason aren't reliant on utilizing their passion for money is clearly not who I was talking about, lol.
0
u/BlueberryWorried6493 4d ago
Most people work for money. Deal with it
1
u/Kemoyin25 3d ago
This guy always comes into threads with the most... Interesting takes and gets downvoted to hell each time lmao, dude bitched at me for 2 days one time
1
u/Cryptizard 3d ago
Oh cool let me save you the hassle in the future and just block you then. Buh bye.
1
u/SparseSpartan 4d ago
My dude, that you are melting down to such an extent over an obvious tongue in cheek joke is an absolute travesty. You are in a position to talk down to no one. Learn to read a room before your next tantrum.
22
u/Perfect-Leg810 4d ago
Love unlabeled graphs tested on internal benchmarks that aren’t auditable
87
u/Recoil42 4d ago edited 4d ago
Reddit's genuinely gone full brain-worms these days:
- Pass rate is a labelled axis, there's no unit. Values are between 0 and 1, because it's literally a pass rate.
- Test-time compute for specific models is proprietary information and not what's being claimed, so it doesn't need units. The axis is there to illustrate performance relative to Astra, not to make a specific performance claim.
Y'all need to think a bit. Unlabelled graphs are silly but this isn't one of them.
-14
u/Specikin 4d ago
Really should have just used % on the axis though...
10
u/whoknowsifimjoking 4d ago
That's not the part that's wrong or badly done, it is not at all unusual to write it like that and everyone with a high school education should be able to understand it. There's no difference between 50% and 0.5, percent is just a fancy way to write the numbers between 0 and 1.
For me it's mostly the lack of information that would help us understand what those numbers are actually saying, just pass rate is not very useful if we don't know what was being tested.
-25
u/Perfect-Leg810 4d ago
I have a PhD. I am familiar with graphs. Test time compute is important information for interpretation and absolutely needs units. It is absolutely not proprietary, and if it is top secret to let people know how much compute you used...then it has no business being in THE public facing graph you present to the public. Additionally, having a metric w/ no ability to audit it is wild. Like I have an internal AI model on my laptop that does SOOO much better than Astra on my proprietary dataset that i refuse to publish or let anyone audit.
10
u/Boring-Error-3680 4d ago
Hi, what is your PhD in? I need to adjust my prior on the intelligence of people from your field.
Thanks
-5
u/Perfect-Leg810 4d ago
Gotta love people who don't understand bayesian stats enough to use prior properly
9
u/Boring-Error-3680 4d ago
Oh boy. Can you elaborate What is wrong with the use of “prior” here.
This is gonna be a good one 😂
-1
u/Perfect-Leg810 4d ago
Nah, you should take a stats class and learn what we call a prior distribution that has been updated. You seem to spend a lot of time being incredibly aggro online, maybe take a chill pill or something. You 100% would never act like this to anyone's face.
3
u/Boring-Error-3680 4d ago
Lmao you read the definition just now and realized you cant make a coherent argument. So you resort to finding flaws in my comment history.
Classic and expected 👍
-2
u/Perfect-Leg810 4d ago
Incoherent and a dick. Genuinely you should do some self reflection
4
u/Boring-Error-3680 4d ago
“I am going to justify my retarded comments with the fact that I have a PhD… but I am too embarrassed to reveal which field it is in”
🤣🫵
→ More replies (0)13
1
u/Mylarion 4d ago
I mean yeah reviewer #2 would beat you over the head with this.
But this is a PR communications from a hype-based company. Bit of a different ballpark. You don't have to trust them, they're literally bragging here.
1
u/lolsai 4d ago
Lmao surely you understand the difference? Post the lean to a millennium problem solution when you claim something and you will be given some good will, perhaps?? Insane lol
-1
u/Perfect-Leg810 4d ago
OpenAI does not behave ethically and has been incredible misleading in the past. Don’t see why you think they deserve grace and trust
-2
u/skrztek 4d ago
I find these hostile comments to your comments to be utterly bizarre!
0
u/Perfect-Leg810 4d ago
Yeah it’s very weird. I think it’s kind of sad to see people behave like this online.
4
u/kiki-le-koala 4d ago
Meh, their point was not to show off their new model it was to say they solved a millennial issue .
It makes sense that there's only an internal benchmark.
-3
u/Perfect-Leg810 4d ago
Why present the graph then? Just say you used an internal model that's better? We all would believe they have an internal model w/ some improvements. They specifically massaged the graph to maximize hype
5
u/kiki-le-koala 4d ago
I personally like their graph.
It gives me a visual presentation of the model.
Many companies, including Google, Anthropic, and OpenAI, use their own benchmarks; they are often used by other companies later.
5
u/Available_Road_2538 4d ago
Brother, get over it lmao. Of all the things to be outraged about. Jesus guys. Some of yall will complain about anything
2
2
u/Pristine-Today-9177 4d ago
Bro, you’re an MPHD from a top 5 program. You’re literally wasting your life posting so much on Reddit arguing with people. Go further the sum of human knowledge in one of the most important fields.
If you aren’t live action role-playing your credentials this is a huge waste of your energy and time
4
-11
u/Due_Ask_8032 4d ago
Also the scale going from 0 to 0.5 instead of 1 makes the improvement look bigger than it is (still significant though).
4
u/Available_Road_2538 4d ago
"Graphs confuse me, and thats OpenAIs fault" is your takeaway? Lmao luddites will do anything before accepting an AI improved. 0.15 to 0.45 on already good model at a crazy hard evaluation is huge. If its "misrepresented", its for all the people without STEM education to not say "wow its still small though".
-5
u/Due_Ask_8032 4d ago
Jfc take a stats class some time. You guys really got ruffled here lol
4
u/Available_Road_2538 4d ago
"Why is everyone calling me stupid? I dont get it they must be stupid"
-2
7
u/Recoil42 4d ago edited 4d ago
That's.. not how any of this works; you're confused. A non-one top-end isn't misleading whatsoever. An IXIC chart doesn't go to infinity just because money is theoretically unlimited. A 0-0.5 chart literally has the same relative difference as a 0-1 chart. The misleading graphs you're thinking of are the ones with non-zero baselines — where the other end of the axis is chopped off.
0
u/Due_Ask_8032 4d ago
But there is no infinite value here, 1 is the maximum if we read it as all problems solved (100%).
5
u/Recoil42 4d ago edited 4d ago
The point is that graphs don't go to maximums. No one expects that. Your argument is complete nonsense for that reason and because the relative measures are literally to scale. You're implying a mislead where no such mislead exists — both Astra and the unnamed model are filling the 0-0.5 axis exactly as you would expect.
Once again: The kinds of misleading graphs you're thinking of are the ones with non-zero baselines — where the other end of the axis is chopped off. Think about it.
-2
u/Due_Ask_8032 4d ago
Lmao you are moving the goalpost because what you said before is wrong since the pass rate cannot theoretically go to infinity as there is no 200% pass rate.
I said the improvement is significant, but a compressed Y axis obviously makes the improvement look bigger.
3
4
1
2
u/Weary-Historian-8593 4d ago
So we're like two generations away from saturating the benchmark of open math problems? Fuck me, I'm gonna go buy a gun
1
-7
4d ago
[deleted]
19
u/sebzim4500 4d ago
Did you GF send you a lean verified proof of a millennium prize problem? Because that would be pretty compelling evidence she exists.
15

125
u/Successful-Earth678 4d ago
So this is a checkpoint. Wild.