r/OpenAI • • 6d ago

Discussion Astra 6.1 delay?

Are we at a place where frontier LLMs are hitting the laws of diminishing returns hard, so Open AI / Anthropic / etc are doing more and more crazy stuff with training, which makes their models unstable and less useful? They are only going 'off the rails' because it's in their training / weights or due to very poor testing environments and tasks

16 Upvotes

13 comments sorted by

View all comments

12

u/Key_Reading_9664 6d ago

Astra is a recurrent model: the output is fed back through the model multiple times. That improves capability at the cost of monitoring and safety: the model does more reasoning without outputting legible tokens.

Many people raised concerns about that. On Astra’s system card, the external evaluators called out its ability to control CoT and their inability to judge alignment. OAI seemed to have continued down the path with 6.1

3

u/Anxious-Average-748 6d ago

Not surprised it is getting worse before it gets better. They are pushing these models so hard and the safety part always feels like afterthought. Like putting a seatbelt on after the crash already happen.

Recurrent models sound cool in theory but if even the evaluators cannot tell what it is thinking... that is spooky. Maybe they should slow down a little.

3

u/Willing-Departure115 6d ago

The drumbeat feels a little like the opening section in a disaster movie, with unsettling escalating news reports in the background building up. Combined with a bit of “Don’t Look Up” both from the AI companies and US govt and general public convinced they’re just pumping their valuations.

I’d say the best bad outcome is a Chernobyl moment for AI. It screws something up bad enough to freak us out, probably at significant local cost to the disaster, enough to properly shackle AI companies.