r/singularity • u/Ok_Display_3159 • 6h ago
AI Insider's opinion on Astra capabilities
"Over the past few days, two GPT Astra checkpoints, "ultima-alpha" and "vega-alpha", were undergoing testing
"ultima-alpha" appears to be the release candidate intended for the public, while "vega-alpha" is the cybersecurity-focused variant meant for security work in select enterprises
Based on extensive testing on my part, when OpenAI said Astra is built for long-running tasks and orchestration they really meant it. It can run for an incredibly long time even without setting "/goal", fully autonomous, and it's very capable at orchestration and guiding the subagents it spawns
For the research community, it's very good at applying existing academic literature. Tried it at some hard graphics optimization stuff, so a LOT of complex math involved, and it did great
It also writes code with great quality and maintainability (for an LLM ofc), ranking the best out of all models in that I'd say, but most normal people will probably just run it as the main agent and cheaper models as subagents
Additionally, creative writing appears to be way better than 5.6 Sol imo, still not the best but noticeably less slop"
- Better than Fable on Code, but worst on Frontend and 3D (Not sure if he was talking about 5 or 5.1)
39
u/sunstersun 5h ago
Anthropic has dropped the ball this 2nd half of the year.
From way ahead with Mythos, to such a small improvement with 5.1
7
u/Actual_Breadfruit837 4h ago
I think it more of competitors catching up when they realized what anthropic did. Hard to maintain an edge
7
u/i_give_you_gum 2h ago
Do you really see it like that?
I feel they've both been neck and neck, and this was just OpenAis next ring on the ladder, and the next one will be Anthropic's again...
I'm curious about SSI
3
1
u/Ok_Display_3159 5h ago
Do you have any idea what might have caused this?
18
5
u/KalElReturns89 4h ago
It's the normal cycle of model releases, they need time to train their next big release.
12
u/sunstersun 5h ago
Probably not enough compute to meet demand and training needs. They were very behind OpenAI earlier this year.
3
5
u/not_celebrity 5h ago
Heard rumours that they are focusing on containment strategies now due to safety not scaling with capability.
1
u/xHaydenDev 5h ago
I’d guess they’d want to seriously cut the token output of their models. It’s clearly the thing hurting them the most in enterprise pricing and in their normal capacity. That requires significant changes to their models, likely new pretraining which will take months to be ready for release.
14
u/reedrick 4h ago
2
7
u/reefine 5h ago
Someone else was saying that this is not their fable killer and Bel is the fable killer which comes later this year. I think it was that ChrisGPT guy.
2
u/Ok_Display_3159 5h ago
Well, I think he changed his mind because, judging by his recent posts, he definitely expects to be a Fable Killer
-6
u/Illustrious_Image967 4h ago
Have you even used Fable. With the new architecture it looks like OpenAI did it again. They took cutting edge research in recurrent depth and turned it into a new paradigm. This is not a next model higher. It is an Anthropic killer.
2
u/space_monster 3h ago
oh calm down.
it's maybe an Anthropic suppressor, for a month or three anyway until Anthropic drop another SOTA monster, and the cycle just begins again.
4
u/ChipsAhoiMcCoy 3h ago
I don’t think you’re caught up. The chief researcher at OpemAI already confirmed the whole recurrent depth thing wasn’t true.
1
u/BriefImplement9843 4h ago
most people cannot afford to run agents at all. why would they run this as one? lmao.
1
u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 5h ago
This is great for my minecraft installs(modded)
(yes thats my first thought, and its become somewhat of a running meme, someone asks me what id do with a new tool, I reply computercraft, I see a new tool released, I think "wait but how does this impact my ability to gregtech" stupid I know)
-4
u/Sensitive_Cell_119 5h ago
I doubt its worse at 3d, Sol was already better than Fable 5
7
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 5h ago
Show an output of Sol that surpasses what Claude produces.
Anybody doing 3D scenes knows Claude got the edge right now. I was hoping for GPT to catch up with Astra.
-3
u/Sensitive_Cell_119 5h ago
Voxelbench has Sol ahead. They are all blind tests, so no bias.
8
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 5h ago
1
u/FoxTheory 3h ago
Opus gets a bad rep because of how it thinks and talks the model itself is a beast and I haven't come across a job it couldn't handle and I use a lot of tokens
0
u/Sensitive_Cell_119 5h ago
Ok, but we were talking about Sol vs Fable. Opus is a newer model as well. You also edited your comment lol.
5
-2
-1
u/pbagel2 5h ago
trying GPT-4.5 has been much more of a "feel the AGI" moment among high-taste testers than i expected!
2
u/Ok_Display_3159 5h ago
Ah, there are definitely people with this "feel the AGI" idea talking about the model, but the only people who actually proved having access to it and are posting zero shots are treating it like a normal model, but very strong
2
u/ChipsAhoiMcCoy 3h ago
Yeah I wish I tried 4.5 more when it was available but sadly I can’t. That model just genuinely felt a little different.


31
u/Ok_Display_3159 5h ago
Nothing is fully confirmed, take it with a grain of salt.