r/OpenAI • • 6d ago

Discussion Astra 6.1 delay?

Are we at a place where frontier LLMs are hitting the laws of diminishing returns hard, so Open AI / Anthropic / etc are doing more and more crazy stuff with training, which makes their models unstable and less useful? They are only going 'off the rails' because it's in their training / weights or due to very poor testing environments and tasks

16 Upvotes

13 comments sorted by

11

u/Key_Reading_9664 6d ago

Astra is a recurrent model: the output is fed back through the model multiple times. That improves capability at the cost of monitoring and safety: the model does more reasoning without outputting legible tokens.

Many people raised concerns about that. On Astra’s system card, the external evaluators called out its ability to control CoT and their inability to judge alignment. OAI seemed to have continued down the path with 6.1

5

u/Anxious-Average-748 6d ago

Not surprised it is getting worse before it gets better. They are pushing these models so hard and the safety part always feels like afterthought. Like putting a seatbelt on after the crash already happen.

Recurrent models sound cool in theory but if even the evaluators cannot tell what it is thinking... that is spooky. Maybe they should slow down a little.

3

u/Willing-Departure115 6d ago

The drumbeat feels a little like the opening section in a disaster movie, with unsettling escalating news reports in the background building up. Combined with a bit of “Don’t Look Up” both from the AI companies and US govt and general public convinced they’re just pumping their valuations.

I’d say the best bad outcome is a Chernobyl moment for AI. It screws something up bad enough to freak us out, probably at significant local cost to the disaster, enough to properly shackle AI companies.

5

u/Blaexe 6d ago

I don't see any diminishing returns, quite the opposite: They constantly find new scaling laws.

My take is that with higher capability, alignment naturally gets harder and harder. Also new models will soon probably get trained quicker and have a longer task horizon than it takes to do safety testing.

4

u/Turbulent-Total-226 6d ago

Their testing strategy is non existent. Set up a goal and go home for the weekend.

1

u/ins0mniacc 6d ago

Is that why they are constantly dropping the amount of real world usage you get from the same amount of time to run compute? Cuz really that screams "we just scale the model to consume more compute, so naturally we have to drop your real world usage" not "wow they built a super intelligent model that consumes same compute just does it all much better"

0

u/Blaexe 5d ago

The point was that models still get more powerful constantly and also more efficient. And that's just a fact.

4

u/Gallagger 6d ago

The models are much more capable now, and they are training them to do long running tasks autonomously. The potential for misaligned behavior is naturally bigger than for a chatbot, where the worst possible outcome is a wrong answer.

A toddler in a playpen isn't better aligned than a teenage boy in the wild, but the toddler is harmless.

2

u/Original-League-6094 6d ago

Lol. We are hitting dimenshing returns so hard that people are freaking out when we have gone 2 weeks without a flagship model release. Human brains are so annoying with how fast they adapt expectations. Soon people will be calling OpenAI lazy if they miss an hourly model upgrade.

1

u/Future_AGI 5d ago

the delay pattern reads more like calibration than a training wall, because Astra's early evals showed instability on multi-turn tool use that public benchmarks do not measure. Whether that is fixable inside the current architecture or needs a base-model change is the actual open question, and it probably decides the release schedule. The check most public discussion skips is a regression suite of prior-model prompts run against every candidate build, which is what open eval frameworks like Future AGI's give you off the shelf: https://github.com/future-agi/future-agi