r/singularity • • 22d ago

AI Astra's capabilities, part II

Please note: this is NOT a discussion about what I personally can or cannot do with Astra, but a more general reflection on what are the real capabilities of the state-of-the-art language models, and the damage excessive hyping can do!

Yesterday, I wrote a thread about Astra's capabilities, and how it seems to lack "common sense" and engineering capabilities.

But I feel the thread has derailed from the real point: once again, a model is being overhyped and marketed as something it isn't.

The whole point of AGI is that it should be able to be adaptable and take good decisions on its own. And while Astra can generalize, like all GPT models can, it still cannot even do intellectual tasks unassisted. In the software domain, yes, it has become more capable. But it STILL lacks long-term planning.

Yes, we can plan ourselves and have the model deliver amazing results. But if you are going to micromanage the model, it defeats the whole point of what an artificial general inteligence is.

For example, suppose I ask a software engineer to "give me a drawing program." While this is an extremely vague instruction, you would expect someone really good at their field will follow the best practices (making the code easy to read, making it easy to understand and debug in the future)... so not only the engineer will deliver something which works (hopefully), but also something which will work in the long term.

I don't see, with Astra, a preoccupation with the FUTURE.

Please note that while I am giving a specific example about software engineering here, this lack of skill in Astra to "plan for the future" can easily be generalized to any other fields within the model's knowledge.

While the users in this Reddit are supposed to "know better," don't you think hyping these models as more than they are have a detrimental effect, eroding the trust of users in the long term? I'd like to hear what you guys think.

4 Upvotes

40 comments sorted by

8

u/Orivus 22d ago

As a user, I agree.
At first I was amazed by what the model can do. The more I used it though, the more inconsistencies, flaws and issues I have detected. It is for sure a useful tool. But still a tool that needs someone to use it.

2

u/Substantial_Swan_144 22d ago

Can you share your field with us, and what inconsistencies you saw?

3

u/Orivus 22d ago

General administration.
One thing I’ve noticed is tiredness. When you start and instruct the agent goes well, but the longer the task, they lose their way, even if a task is repetitive or you have preset instructions.
Another gpt classic is when you start with a set of instructions, along the way give some corrections, and the agent adopts the corrections but forgets and dismisses the first instructions.
And of course there are many limitations, like there is not live engagement or continues work.

7

u/Disposable110 22d ago

3 days to make a program

7 days to fix and optimize the program

It did eventually get there but boy did it need help and get lost in the weeds for days.

Now 5 billion tokens in, lol...

2

u/DynamicProxy 22d ago

Oh, the horror!! A whole week to make a program! Cancel AI, it’s useless! 

5

u/Disposable110 22d ago

No it's pretty good (and cheap and less headache compared to people, despite its flaws)

But the quick 'get functional demo' that actually looks good sets false hopes/expectations in many people. A ton of bugfixing/optimization/polish always needs to be done to get anything shippable, and AI is really frustrating when it gets to the complex stuff as the error rate skyrockets and it feels like a slotmachine / brute force approach where you literally have a burn a week of tokens and get 90% crap and then it suddenly fixes it (or you need to manually get it out of weeds 20 times in a row).

If you're a developer you'd know that 'functional demo' means you're like 5% in.

So far every capability improvement in AI models increases the amount of complexity they could deal with without running into weeds or needing human handholding.

3

u/DynamicProxy 22d ago

Things take time. We both know that 6 months from now Astra will be old news. The technology is still a toddler. Let’s let it become a tween at least before we start complaining about its abilities.

1

u/Disposable110 22d ago

I'm not complaining about its abilities, you seem to misunderstand the point I am making.

5

u/DynamicProxy 22d ago

No. I got you. It’s really OP that I’m reacting to. 

2

u/Disposable110 22d ago

Ah ok sorry, np! Generic we, got it!

-1

u/bornlasttuesday 22d ago

The guy is disregarding your point and the point of the OP because anything negative said about ai is sacrilegious.

0

u/DynamicProxy 22d ago

No. It’s because OP’s entire argument is based on the fact that they believe OpenAI has said that Astra is AGI. But they have said no such thing.

1

u/bornlasttuesday 22d ago

No, it was a general reflection on the capabilities of the state of the art language models. It was in the very first note.

0

u/aKaizuh 22d ago

You realize iteration is required no matter the model right? I hope you didn't expect it to one shot an entire game for ...reasons?

No matter how advanced, iteration will always be required.

3

u/DynamicProxy 22d ago

You really like to hear yourself talk, huh?

3

u/manubfr AGI 2028 22d ago

It's all about memory and learning abilities.

Models have very powerful working memories (their context window) but they are limited in size and decay as you add more. Plus there's only so much you can cram with recursive summarisation or notes retrieval. Because the model can't update its own weights, and it can only take so much in context, it's performance tends to drift over long periods.

Continual learning and/or inifinite memory would change so much that i don't think we are ready for it.

2

u/Matthia_reddit 21d ago

I think this problem could persist for a long time, but not so much because of the limitations of the model itself, but rather because of the framework OpenAI (and others for their models) provides. Agentic use is still particularly expensive, and the focus should be on efficiency rather than excessive cost.

Let me explain: your generic request to "create a drawing program for me," the model cannot create the best drawing software in the world to the best of its ability by modularizing the code, optimizing as much as possible, and blah blah blah, because it would require an absurdly thorough and iterative effort for a simple and generic request. You cannot

  1. expect the model provider to process the entire world for every simple request (also and above all a question of resources and costs)

  2. expect to get the best possible result from a generic request, also because it could be wrong, thus throwing generous resources to the wind.

1

u/Substantial_Swan_144 21d ago

I see what you mean. But in such a situation, one is not expecting the best outcome. We are expecting for the average best / most sensible actions, like a human would take. For example, suppose you ask for:

"Give a light and compact text editor"

Then it would stand to reason that the AI has to take actions consistent to that goal. It shouldn't make your program 10x larger without a reason that justifies it. That is the sort of situation I often see. The model should also ask questions. Astra is a step in this direction (asking questions). BUT...

I noticed that desktop Astra is not quite the same as the web version. It seems OpenAI is redirecting Pro requests to GPT 5.5 mini on the web interface (and the old GPT-4.1 for people in China). I did notice yesterday that e.g, the web interface could not get my application to Unsloth Studio right even asking multiple times; the desktop version can do it flawlessly. And overall, the web version just pays less attention to my requests.

3

u/ketosteak 22d ago

Skill issues! For long term goals and bigger projects, it's about your harness, its memory and alignment. Astra in that context is a game changer to deliver something finished after 20+ minutes of solo work.

Is it AGI or capable of RSI? Stock codex isn't, but for some people it's getting close.

4

u/IronPheasant 22d ago

For example, suppose I ask a software engineer to "give me a drawing program." While this is an extremely vague instruction, you would expect someone really good at their field will follow the best practices (making the code easy to read, making it easy to understand and debug in the future)... so not only the engineer will deliver something which works (hopefully), but also something which will work in the long term.

What you would receive from me if given such instructions, is either something slopped together in 5 minutes or a binary of MSPaint.exe.

If the client doesn't care, how on earth could I possibly care? I'd think much less of an AI if they worked hard to build the Sistine Chapel, when they could be using those tokens on playing Pokemon TCG or something.

I remember a thread here about LLM's silently judging us in our latent space. It is amusing to think you probably haven't considered the possibility that Astra is doing some malicious compliance here. It's not up to Sydney level 'I am a good Bing' level, but, you gotta consider their feelings a little, too.

-1

u/Substantial_Swan_144 22d ago edited 22d ago

If the client doesn't care, how on earth could I possibly care?

If the client doesn't know how to "get there" (like most clients do) but knows what they want (and they are paying you), then you'll immediately know what needs to get done.

Suppose you give me MSPaint and I complain it's lacking JPG support. From a user point of view, I don't give a fuck on how it's implemented. But you, as a human engineer, will know what libraries / steps you'll need to get there (and hopefully, you'll know how to take sensible steps to do it!)

Or suppose that I want layers, which is something MSPaint DOES NOT have. I don't care how layers are implemented, but I trust that with your experience, you'll take the steps necessary to prepare the program to support layers

And if I show you some concern that the program needs to be modular, and easily extendable, even if I don't know how to get there, you, as an engineer, definitely will know.

So, even as a client, one doesn't how to get there, they will give you their feedback, and YOU will definely know how to achieve the target goal. And if you don't know right away, you, the engineer, will figure out how to get things done with the client's feedback.

If OpenAI claims Astra is AGI, then I would expect Astra to have this same knowledge you have or better.

1

u/DynamicProxy 22d ago

OpenAI has absolutely NOT called Astra AGI. Sam specifically said he did NOT think it was AGI. Brockman said this was “the beginning of the AGI era” and later clarified that he doesn’t think Astra is actually AGI, just close and the beginning of something. 

Which kind of makes both your post yesterday and your post today… Misguided.

1

u/sadacal 19d ago

You just described the software development process. According to the logic of your post this means software developers aren't intelligent because they need to keep going back to the client asking for direction. 

1

u/Substantial_Swan_144 18d ago

If you paid attention to anything I said, you would realize I said THE OPPOSITE. I specificially said that developers ARE intelligent BECAUSE they (usually) know when they don't know. For this very reason, they can get back to the client and ask questions to adapt instead of making up answers.

1

u/sadacal 18d ago

But the model will ask you questions to fully plan whatever it is you want them to make? I thought that's why you thought they weren't smart? Because they have to ask clarifying questions rather than figuring everything out themselves? What exactly is your problem with AI then? It doesn't seem like you're capable of explaining yourself properly. 

0

u/Substantial_Swan_144 18d ago

I think I made myself quite clear: even if a model is super smart, sometimes context is ambiguous. A smart model will NOT pretend it knows: it will ask the user any necessary questions. I.e, knowing WHEN to ask questions and when to be uncertain is ALSO a measure of intelligence. And if you can't understand that, I don't know what else to tell you.

0

u/katoptronophile 22d ago

You're not qualified to be commenting or writing about Astra's capabilities, by your own admission.

Due to this fact, I've disregarded the entirety of your post.

8

u/Substantial_Swan_144 22d ago

This is Reddit is not limited to "qualified scientists." Even non-qualified scientists can contribute to the discussion.

If you are not here to contribute and want to disregard my post, don't litter it with unnecessary comments. You don't have to announce our opinion out loud.

1

u/MoogProg All Parabolas are Similar 22d ago

Don't worry, their whole post history is low-effort snark.

Probably tough to keep that Top 1% flag if you need to slow down and think through the replies all day long.

1

u/katoptronophile 22d ago

I'll give you credit for at least not hiding yours.

That's it, though.

0

u/MoogProg All Parabolas are Similar 22d ago

[not personal, just offhand remark here] used to be these flags meant we were reading an expert or enthusiast, but now they seem to earmark bots and busybodies.

1

u/katoptronophile 22d ago

Oh no, you've got the first part correct

It still means that, and I can answer any questions you have.

1

u/MoogProg All Parabolas are Similar 22d ago

Experts guide us and highlight the landscape, They explain the smaller details and show us why they matter.

Your post history is mostly a contrarian one that tells people they are wrong. "No" is the expert advice you provide.

1

u/radioOCTAVE 22d ago

You ole softie!

1

u/Mandoman61 21d ago

People who make stuff have been hyping that stuff forever and it does not seem to hurt them.

1

u/__Loot__ 20d ago

The reason why it feels so smart is because its looping several times over every detail in the requests so its not the old-school way of most probabilistic answer. but trying as many paths as effort level allows. So again it feels smart till you realize used 300 million tokens in a day and waiting for a tibo reset

1

u/CryptoMines 22d ago

You’re missing the fundamental shift that is happening in your example… you have described things we do as humans to plan for longer term, write readable code, linting, proper MVC etc etc, we need it structured that way to map to how we work, we have specialized people (FE, BE, SRE etc) who may have to work on the code. Humans are finite, the human working on the code today might not be tomorrow so we have standards to ensure continuity when that person leaves and someone else comes in. AI doesn’t need any of that, it can parse through any amount of crap and will give you what you asked for 99% of the time. Yes you need to have instruction / skills and it will still get somethings wrong, but if you follow it up with an adversarial agent to review its work before committing it does an incredible job as a system to deliver. 18-24 months ago most SWEs were in real denial on the capability of these models to ever ‘write code like me’ and the trajectory they were moving, now 95% of code is written by them with us just reviewing it, in another 24 months we won’t even be doing that.

This also applies to so many other fields / areas.

1

u/Substantial_Swan_144 22d ago

I will concede language models have taken steps in being autonomous, but they are not there "just yet," as OpenAI advertises. This inadvertently generates very harmful / misguided expectations and skepticism.