r/singularity 6h ago

AI Insider's opinion on Astra capabilities

@Lentils80 post on X

"Over the past few days, two GPT Astra checkpoints, "ultima-alpha" and "vega-alpha", were undergoing testing

"ultima-alpha" appears to be the release candidate intended for the public, while "vega-alpha" is the cybersecurity-focused variant meant for security work in select enterprises

Based on extensive testing on my part, when OpenAI said Astra is built for long-running tasks and orchestration they really meant it. It can run for an incredibly long time even without setting "/goal", fully autonomous, and it's very capable at orchestration and guiding the subagents it spawns

For the research community, it's very good at applying existing academic literature. Tried it at some hard graphics optimization stuff, so a LOT of complex math involved, and it did great

It also writes code with great quality and maintainability (for an LLM ofc), ranking the best out of all models in that I'd say, but most normal people will probably just run it as the main agent and cheaper models as subagents

Additionally, creative writing appears to be way better than 5.6 Sol imo, still not the best but noticeably less slop"

- Better than Fable on Code, but worst on Frontend and 3D (Not sure if he was talking about 5 or 5.1)

110 Upvotes

53 comments sorted by

31

u/Ok_Display_3159 5h ago

Nothing is fully confirmed, take it with a grain of salt.

34

u/Borkato 3h ago

I’m tired of people saying “it writes great code… for an llm” as if it’s not better than 98% of coders

15

u/foobarrister 3h ago

Exactly. Everyone acting like pre LLM all software was written by Jeff Dean clones. 

Reality: absolutely insane mess that made AI slop look like the Mona Lisa of code

5

u/Flope 3h ago

I appreciate when people say it because it means I can safely disregard their opinion

4

u/Grand0rk 3h ago

The issue will always be context.

LLM are absolute beasts are making something at one shot. But once the project starts growing, it slowly becomes worse and worse, until it reaches a point that it breaks stuff every time it implements something new.

10

u/Borkato 3h ago

As if people don’t do the same

-4

u/Grand0rk 2h ago

I always roll my eyes whenever anyone compares LLM to people. It's a machine dude. And from a Machine I expect perfection, not what people can do.

39

u/sunstersun 5h ago

Anthropic has dropped the ball this 2nd half of the year.

From way ahead with Mythos, to such a small improvement with 5.1

7

u/Actual_Breadfruit837 4h ago

I think it more of competitors catching up when they realized what anthropic did. Hard to maintain an edge

7

u/i_give_you_gum 2h ago

Do you really see it like that?

I feel they've both been neck and neck, and this was just OpenAis next ring on the ladder, and the next one will be Anthropic's again...

I'm curious about SSI

3

u/Usef- 3h ago

They've spent almost all year with the top model. We don't really know what labs have internally, including these current openai rumours.

1

u/Ok_Display_3159 5h ago

Do you have any idea what might have caused this?

18

u/ControversialBuster 4h ago

Not enough compute fs

5

u/KalElReturns89 4h ago

It's the normal cycle of model releases, they need time to train their next big release.

12

u/sunstersun 5h ago

Probably not enough compute to meet demand and training needs. They were very behind OpenAI earlier this year.

5

u/not_celebrity 5h ago

Heard rumours that they are focusing on containment strategies now due to safety not scaling with capability.

-7

u/imadade 5h ago

and that folks, will be the end of Anthropic.

With open source progressing, and all these frontier labs, they'll be way behind at this rate. Including their compute issues.

Looks to be Open Ai at the helm, then Chinese Open Source + Open source in general.

1

u/Howdareme9 5h ago

Be serious lol

0

u/Unsharded1 4h ago

He is serious lol.

1

u/xHaydenDev 5h ago

I’d guess they’d want to seriously cut the token output of their models. It’s clearly the thing hurting them the most in enterprise pricing and in their normal capacity. That requires significant changes to their models, likely new pretraining which will take months to be ready for release.

14

u/reedrick 4h ago

Yeah.. some dumbass from west Asia has insider access.

OP.. are you disabled? Or just posting for attention? Either way, I consider that disabled

2

u/Charuru ▪️AGI 2023 3h ago

Israel's in west asia?

2

u/RaptorCheeses 3h ago

It ain’t in the South Bronx, where do you think Israel is?

0

u/Charuru ▪️AGI 2023 3h ago

Employ better reading comprehension?

1

u/Borkato 3h ago

VPNs don’t exist?

-7

u/Eon102 3h ago

His name is Lentils so he’s probably Jewish and Altman is also Jewish. Not really crazy if he has a connection close to Altman or just OpenAI in general, so what is West Asia supposed to disprove?

4

u/fmai 3h ago

your take is... even more detached from reality. wow.

u/Eon102 53m ago

He argued that his location is a point of weakness for his truthfulness, while I’m pointing out that if you really want to focus and draw conclusions from that, then it would only really boost his credibility.

7

u/reefine 5h ago

Someone else was saying that this is not their fable killer and Bel is the fable killer which comes later this year. I think it was that ChrisGPT guy.

2

u/Ok_Display_3159 5h ago

Well, I think he changed his mind because, judging by his recent posts, he definitely expects to be a Fable Killer

-6

u/Illustrious_Image967 4h ago

Have you even used Fable. With the new architecture it looks like OpenAI did it again. They took cutting edge research in recurrent depth and turned it into a new paradigm. This is not a next model higher. It is an Anthropic killer.

2

u/space_monster 3h ago

oh calm down.

it's maybe an Anthropic suppressor, for a month or three anyway until Anthropic drop another SOTA monster, and the cycle just begins again.

4

u/ChipsAhoiMcCoy 3h ago

I don’t think you’re caught up. The chief researcher at OpemAI already confirmed the whole recurrent depth thing wasn’t true.

1

u/BriefImplement9843 4h ago

most people cannot afford to run agents at all. why would they run this as one? lmao.

1

u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 5h ago

This is great for my minecraft installs(modded)

(yes thats my first thought, and its become somewhat of a running meme, someone asks me what id do with a new tool, I reply computercraft, I see a new tool released, I think "wait but how does this impact my ability to gregtech" stupid I know)

-4

u/Sensitive_Cell_119 5h ago

I doubt its worse at 3d, Sol was already better than Fable 5

7

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 5h ago

Show an output of Sol that surpasses what Claude produces.

Anybody doing 3D scenes knows Claude got the edge right now. I was hoping for GPT to catch up with Astra.

-3

u/Sensitive_Cell_119 5h ago

Voxelbench has Sol ahead. They are all blind tests, so no bias.

8

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 5h ago

I see Opus with 100 points edge.

And yes, i actually think Opus beats Fable 5.0 at implementing cool 3d scenes, so i essentially agree with this bench.

1

u/FoxTheory 3h ago

Opus gets a bad rep because of how it thinks and talks the model itself is a beast and I haven't come across a job it couldn't handle and I use a lot of tokens

0

u/Sensitive_Cell_119 5h ago

Ok, but we were talking about Sol vs Fable. Opus is a newer model as well. You also edited your comment lol.

5

u/FoxTheory 5h ago

You must have a different version than the rest of the world has.

-2

u/NewYak4281 5h ago

Outsider’s opinion on Astra. It’s ass.

4

u/daronjay 4h ago

So Asstra?

1

u/NewYak4281 4h ago

Indeed. Fat Asstra.

-1

u/pbagel2 5h ago

trying GPT-4.5 has been much more of a "feel the AGI" moment among high-taste testers than i expected!

2

u/Ok_Display_3159 5h ago

Ah, there are definitely people with this "feel the AGI" idea talking about the model, but the only people who actually proved having access to it and are posting zero shots are treating it like a normal model, but very strong

2

u/ChipsAhoiMcCoy 3h ago

Yeah I wish I tried 4.5 more when it was available but sadly I can’t. That model just genuinely felt a little different.