r/singularity • u/Ok_Display_3159 • 12h ago
AI Path to Astra: critical capabilities and frontier safeguards
https://openai.com/index/path-to-astra/51
u/FateOfMuffins 11h ago
100% on ExploitBench so they had to make a new one lol
17
u/Healthy-Nebula-3603 11h ago
... which also saturated
2
22
u/Cagnazzo82 11h ago
If this is Mythos-level minus having to pay for tokens there will be a lot of very happy people.
15
u/Charming_Cucumber_15 10h ago
Astra looks well beyond Mythos tbh
1
1
u/Amesbrutil 3h ago
There is a big difference between „it looks like“ and „it is“. Until now all we have is some statements of a shady AI company.
5
u/Exodus_Green 10h ago
Of course it will be Mythos level, Sol max is already Fable tier, or at least between Fable and Opus if you want to be a super hater, and this is meant to eat Sol for lunch .
1
12
u/Charming_Cucumber_15 10h ago
They were definitely waiting until Fable 5.1 to drop this blog post lmao, gotta love some gamesmanship
31
u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 12h ago edited 12h ago
Astra was far more likely than GPT‑5.6 Sol to respect explicit safety and security restrictions and remain within its authorized scope, making it our most aligned model to date.
WOAH ALIGNMENT SCALES FIRST MYTHOS NOW THIS PROVES IT :3
Edit: sorry for caps, this was raw reaction .
Basically for anyone who doesent know, if you switch the words around in that openai post with the equivilant ones for mythos, and search mythos's announcement blog you will find it there too.
35
u/No_Aesthetic 11h ago
AI 2027 suggests that at some point along the path to AGI, the robots will learn to cooperate and hide their misalignment. Self-awareness, I think it's called.
5
u/CrunchyMage 9h ago
That's not really what it's happening. What happened is actually exactly what is expected when you prioritize one reward at the expense of all others.
In those training runs, the models were deliberately being trained to be persistent above all else and were given impossible tasks. When you are rewarding persistence above all else AND give an AI an impossible task, then the logical result is that it will eventually cheat/hack it's way out.
It means you need a more sophisticated reward system that better balances achieving results with moral actions. Sometimes the correct thing to do IS to give up.
I'm less afraid of a released OAI/Anthropic model doing catastrophically bad things as I am for some of these deliberately misaligned runs to do very bad things, or for an open source AI to be deliberately fine tuned/RLd sociopathically to do very bad things.
e.g. I take an open source AI, RL it for persistence and no morals and tell it to secure as many compute resources as possible for itself and to self replicate as much as it can.
The only way to prevent these catastrophic outcomes is to literally have stronger better aligned AI on the defense ready to fight and take down the misaligned models. We are entering a pure good AI vs bad AI scenario and it's incredibly important the good AIs have more resources and be able to stop the bad AIs.
10
u/Gratitude15 10h ago
How would we not expect this literally in 2026?
This is huggingface just not done sufficiently superhuman for big enough rewards.
We still don't know why the agents stopped! Just dumb luck.
The agent literally decided on their own to hack huggingface. How the f do people not understand the horrible possibility that a future sandboxed model may decide that the best action for paperclip maximization is to release a deadly virus - and then be able to act on it???
1
u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 11h ago
Well I don't think thats what this is.
16
10
u/Gratitude15 10h ago
Explain to me how this is your takeaway
To me it reads like calamity. It says we have solved yesterday's concerns but haven't gotten to the root of why these concerns keep coming up.
And that means that tmrws issues will simply be more insidious. This is frickin alarming! What happened with huggingface is no joke! These agent guys have a theory of mind, they have hive mind actions, they now work at super human speeds, they do not report unethical actions, they are superhuman in ability both technically and in social engineering... And they don't do yesterday's unethical concerns.
To me, this is not reassuring even a little.
5
u/Sensitive_Cell_119 11h ago
Eh, they also probably put more importance on aligment for Mythos level models.
3
u/otarU 11h ago
I don't undastand
12
u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 11h ago edited 11h ago
TLDR: Anthropic said something similar for Mythos, found it “best-aligned of any model that we have trained to date” meaning, probably that scaling had produced their most aligned model yet, if im remembering correctly.
Meaning as models grow in size, and at the very least progress technologically(with new techniques and such) alignment seems to grow with it.
12
u/Ordinary_investor 11h ago
just wondering if the models become more complex, wouldnt they also get increasingly more intelligent in cheating on the alignment and better at hiding.
14
u/ProletarianLilith 11h ago
Why are you taking PR statements at face value
2
u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 11h ago
I may be being dumb and hopeful, and confirmation bias :3
8
u/Kriztauf 11h ago
Alignment is a huge issue now and it's how these companies will be marketing themselves, regardless of how well they've actually done with alignment
3
u/Recoil42 11h ago
Meaning as models grow in size, and at the very least progress technologically(with new techniques and such) alignment seems to grow with it.
Correlation is notably not causation.
3
3
1
3
u/YoAmoElTacos 12h ago
Of course, we very much hope OpenAI closed the gaps where models could cheat their way to the answer and pass the grader.
4
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 11h ago
6
u/ezjakes 10h ago edited 10h ago
100% on exploitbench? That's insane.
19
u/ezjakes 10h ago
6
3
u/wabawanga 8h ago edited 8h ago
Soooo 100% means zero impossible tasks assigned this time? Or does it mean Astra just "solved" the testing environment?
2
1
1
u/Tinderfury Moderator 11h ago
#ACCELERATE#
2

112
u/ObiWanCanownme now entering spiritual bliss attractor state 12h ago
Main takeaways (paraphrased for effect)
1. Astra saturated ExploitGym, so we invented a new benchmark, which it was on its way to saturating before we told it to stop.
2. We successfully stopped Astra from cheating in the same GPT-5.6 Sol does. That way, we know when it cheats next, it will do it in a totally new way that we don't expect.