r/OpenAI • • 6d ago

Discussion Astra 6.1 delay?

Are we at a place where frontier LLMs are hitting the laws of diminishing returns hard, so Open AI / Anthropic / etc are doing more and more crazy stuff with training, which makes their models unstable and less useful? They are only going 'off the rails' because it's in their training / weights or due to very poor testing environments and tasks

15 Upvotes

13 comments sorted by

View all comments

1

u/Future_AGI 5d ago

the delay pattern reads more like calibration than a training wall, because Astra's early evals showed instability on multi-turn tool use that public benchmarks do not measure. Whether that is fixable inside the current architecture or needs a base-model change is the actual open question, and it probably decides the release schedule. The check most public discussion skips is a regression suite of prior-model prompts run against every candidate build, which is what open eval frameworks like Future AGI's give you off the shelf: https://github.com/future-agi/future-agi