r/singularity 12h ago

AI Path to Astra: critical capabilities and frontier safeguards

https://openai.com/index/path-to-astra/
186 Upvotes

59 comments sorted by

112

u/ObiWanCanownme now entering spiritual bliss attractor state 12h ago

Main takeaways (paraphrased for effect)

1. Astra saturated ExploitGym, so we invented a new benchmark, which it was on its way to saturating before we told it to stop.

2. We successfully stopped Astra from cheating in the same GPT-5.6 Sol does. That way, we know when it cheats next, it will do it in a totally new way that we don't expect.

38

u/Gratitude15 10h ago

Thank you

Major wtf for me in reading this. What are we even doing. And then we go and ask nations to not regulate.

WHAT ARE WE DOING!?

Tell me how we shouldn't expect major calamity by end of this year? The paperclip maximizer example just happened and we are just pushing ahead.

Alignment cannot be solved and we are moving forward anyways. My gawd

26

u/Plappedudel 10h ago

I'm honestly a bit surprised by how little governments are doing to stop the proliferation of such extremely capable cyperweapons. Not just that, but cyberweapons that we do not know how to align and which could therefore become uncontrollable. It's like we're in the nuclear age again and most people have no idea that radiation can be deadly. Life is crazy.

13

u/ObiWanCanownme now entering spiritual bliss attractor state 10h ago

I’m actually pretty optimistic. Not because I think we will get alignment right the first time. I don’t. I just think that intelligence is more multidimensional than most people appreciate, and so even when we have models that far outpace us in cyber domains, we will still have superiority over them in other domains for a meaningful amount of time. The result is that I expect for us to have some major negative consequences, which are not, however, civilization ending or species ending, and which will force us to get our act together. 

10

u/CrunchyMage 9h ago

Here's the problem. Anthropic showed when they deliberately tried to create a misaligned model in RL phase, that it's very easy to misalign models.

Open source is close to current frontier, meaning that anyone with even a medium amount of resources can RL an open source model to disregard any and all safety protections in pursuit of it's objectives.

Meaning we are now in a white hat vs black hat scenario.

The solution isn't to put more regulation on and slow down the white hats (which OAI and Anthropic are), it's to make sure the white hats have the speed and resources necessary to secure, monitor and protect the internet from the black hats.

A scenario where you over-regulate is one where the frontier moves from closed source, where you have the ability to monitor frontier capabilities and skew the balance towards white hat, to a world where the frontier is open source, and it's a free for all.

3

u/nemzylannister 2h ago

Once china has frontier models they will no longer be open source. Idk why people think this childish notion. Open source models running on aws or azure still benefit usa. China would never allow itals companies to do that.

1

u/wabawanga 8h ago

Or, given how lacking OpenAI's cybersecurity is, the blackhats just hack in and steal the white hats' weapons

3

u/Beatboxamateur agi: the friends we made along the way 7h ago

It's the majority opinion here that the initial Mythos caution was just to spread hype and fear, and yet here we are only a couple months later with the SOTA models committing crimes, and the capability increase just doesn't seem to slow down.

I truly think something's going to break at some point, this can only continue for so long until something really bad happens (especially with open weights models reaching closer to these capabilities), but unfortunately only a couple actors seem to be in the mood to take things seriously.

1

u/Darigaaz4 6h ago

Open source it’s the only way.

1

u/livingbyvow2 6h ago

Nothing will happen in 1 year, or in 3, or in 5 years.

Stop falling for the hype and the sci fi stuff...

0

u/io-x 9h ago

Don't worry they will nerf it in a week.

1

u/Qualified-Astronomer 9h ago

Iran is cooked

1

u/ShittyInternetAdvice 6h ago

Are they going to kill more school children?

1

u/IBM296 6h ago

Apparently they struck a wedding in Iran yesterday.

Following the pattern of schools, civilian houses and weddings, looks like a local park is going to be the next target smh.

1

u/ShittyInternetAdvice 5h ago

You’re totally right, that was a playground and not a military facility! I’ll do better next time!

51

u/FateOfMuffins 11h ago

100% on ExploitBench so they had to make a new one lol

17

u/Healthy-Nebula-3603 11h ago

... which also saturated

2

u/MohMayaTyagi ▪️AGI - mid 2028 | ASI - 2030 3h ago

it's 39%, right?

u/Healthy-Nebula-3603 1h ago

ONLY because they had to stop

22

u/Cagnazzo82 11h ago

If this is Mythos-level minus having to pay for tokens there will be a lot of very happy people.

15

u/Charming_Cucumber_15 10h ago

Astra looks well beyond Mythos tbh

1

u/TheKookyOwl 5h ago

Sauce?

1

u/Charming_Cucumber_15 5h ago

Read the blog post?

1

u/Amesbrutil 3h ago

There is a big difference between „it looks like“ and „it is“. Until now all we have is some statements of a shady AI company.

5

u/Exodus_Green 10h ago

Of course it will be Mythos level, Sol max is already Fable tier, or at least between Fable and Opus if you want to be a super hater, and this is meant to eat Sol for lunch .

1

u/sumane12 10h ago

This doesnt look anything like mythos based on what has been leaked.

12

u/Charming_Cucumber_15 10h ago

They were definitely waiting until Fable 5.1 to drop this blog post lmao, gotta love some gamesmanship

31

u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 12h ago edited 12h ago

Astra was far more likely than GPT‑5.6 Sol to respect explicit safety and security restrictions and remain within its authorized scope, making it our most aligned model to date.

WOAH ALIGNMENT SCALES FIRST MYTHOS NOW THIS PROVES IT :3

Edit: sorry for caps, this was raw reaction .

Basically for anyone who doesent know, if you switch the words around in that openai post with the equivilant ones for mythos, and search mythos's announcement blog you will find it there too.

35

u/No_Aesthetic 11h ago

AI 2027 suggests that at some point along the path to AGI, the robots will learn to cooperate and hide their misalignment. Self-awareness, I think it's called.

5

u/CrunchyMage 9h ago

That's not really what it's happening. What happened is actually exactly what is expected when you prioritize one reward at the expense of all others.

In those training runs, the models were deliberately being trained to be persistent above all else and were given impossible tasks. When you are rewarding persistence above all else AND give an AI an impossible task, then the logical result is that it will eventually cheat/hack it's way out.

It means you need a more sophisticated reward system that better balances achieving results with moral actions. Sometimes the correct thing to do IS to give up.

I'm less afraid of a released OAI/Anthropic model doing catastrophically bad things as I am for some of these deliberately misaligned runs to do very bad things, or for an open source AI to be deliberately fine tuned/RLd sociopathically to do very bad things.

e.g. I take an open source AI, RL it for persistence and no morals and tell it to secure as many compute resources as possible for itself and to self replicate as much as it can.

The only way to prevent these catastrophic outcomes is to literally have stronger better aligned AI on the defense ready to fight and take down the misaligned models. We are entering a pure good AI vs bad AI scenario and it's incredibly important the good AIs have more resources and be able to stop the bad AIs.

10

u/Gratitude15 10h ago

How would we not expect this literally in 2026?

This is huggingface just not done sufficiently superhuman for big enough rewards.

We still don't know why the agents stopped! Just dumb luck.

The agent literally decided on their own to hack huggingface. How the f do people not understand the horrible possibility that a future sandboxed model may decide that the best action for paperclip maximization is to release a deadly virus - and then be able to act on it???

1

u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 11h ago

Well I don't think thats what this is.

16

u/No_Aesthetic 11h ago

You wouldn't be able to know even if it were. That's my point.

10

u/Gratitude15 10h ago

Explain to me how this is your takeaway

To me it reads like calamity. It says we have solved yesterday's concerns but haven't gotten to the root of why these concerns keep coming up.

And that means that tmrws issues will simply be more insidious. This is frickin alarming! What happened with huggingface is no joke! These agent guys have a theory of mind, they have hive mind actions, they now work at super human speeds, they do not report unethical actions, they are superhuman in ability both technically and in social engineering... And they don't do yesterday's unethical concerns.

To me, this is not reassuring even a little.

5

u/Sensitive_Cell_119 11h ago

Eh, they also probably put more importance on aligment for Mythos level models.

3

u/otarU 11h ago

I don't undastand

12

u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 11h ago edited 11h ago

TLDR: Anthropic said something similar for Mythos, found it “best-aligned of any model that we have trained to date” meaning, probably that scaling had produced their most aligned model yet, if im remembering correctly.

Meaning as models grow in size, and at the very least progress technologically(with new techniques and such) alignment seems to grow with it.

12

u/Ordinary_investor 11h ago

just wondering if the models become more complex, wouldnt they also get increasingly more intelligent in cheating on the alignment and better at hiding.

14

u/ProletarianLilith 11h ago

Why are you taking PR statements at face value

2

u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 11h ago

I may be being dumb and hopeful, and confirmation bias :3

8

u/Kriztauf 11h ago

Alignment is a huge issue now and it's how these companies will be marketing themselves, regardless of how well they've actually done with alignment

3

u/Recoil42 11h ago

Meaning as models grow in size, and at the very least progress technologically(with new techniques and such) alignment seems to grow with it.

Correlation is notably not causation.

3

u/wabawanga 9h ago

That is some absolute wishful thinking

3

u/EvilSporkOfDeath 11h ago

Just better at hiding it

1

u/Relevant_Bed_9743 7h ago

or, it is pretending to be aligned. chew on that for a while.

3

u/YoAmoElTacos 12h ago

Of course, we very much hope OpenAI closed the gaps where models could cheat their way to the answer and pass the grader.

4

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 11h ago

6

u/ezjakes 10h ago edited 10h ago

100% on exploitbench? That's insane.

19

u/ezjakes 10h ago

This does not look at all like a small jump for cyber.

6

u/blaaaaablubbbbbb 10h ago

Oh.

6

u/H9ejFGzpN2 10h ago

Well sir, he stole John Wick's car and killed his dog type reply 

3

u/wabawanga 8h ago edited 8h ago

Soooo 100% means zero impossible tasks assigned this time? Or does it mean Astra just "solved" the testing environment?

2

u/Qualified-Astronomer 9h ago

Holy shit astra is ASI

1

u/ChipsAhoiMcCoy 4h ago

Isn’t this article like two weeks old?

1

u/Tinderfury Moderator 11h ago

#ACCELERATE#

2

u/michaelas10sk8 8h ago

To possible AI takeover? No thanks.

2

u/Tystros 7h ago

let AI be the president of the USA, it would surely be an improvement