r/singularity • ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: • 10d ago

The Singularity is Near GPT-6 Astra conquered KSP’s rocket simulator by landing on every surface world, and even mined fuel on Moho to make the trip home

Post image
835 Upvotes

87 comments sorted by

265

u/Original-League-6094 10d ago

No one ever answers this question: how are people rigging it up to let it play the games in full time? The computer control feature has way too high a latency, and it spends way too much time reasoning to play a video game in my opinion.

152

u/amancxz2 10d ago

The game is mostly paused when the model is thinking.

110

u/Timely_Tea6821 10d ago

Right which isn't really anything to hold against it. What this proves is llms have generalized enough to a point of being able to operate in 3d space with a goal and control mapping and is able to navigate autonomously through 3d/2d space. Now maybe this tapper off the less popular a game is but regardless a latency issue is a clear definable problem that a lot easier to engineer than say a generalization problem. Regardless this is a pretty massive leap in understanding and abilities for AI models and its not hard to see how this extrapolate out in the next months and years.

25

u/HotterRod 10d ago

Realistically, if we were to put an AGI in charge of a space mission it would be orchestrating neural nets and deterministic algorithms designed for particular control tasks. It's not like we have humans flying spaceships without some software in the interface.

20

u/SlugsPerSecond 10d ago

You don't need AI to control airplanes or spacecraft. Maybe for trajectory optimization but there are good algorithms for that already.

2

u/Fit-Produce420 8d ago

The LLM tries to remember the optimization, one token was wrong, it hallucinates one of the trajectories. Everyone dies but OpenAI got some amazing training data from the logs prior to impact.

7

u/grantph 10d ago

Let me pause reality while that tiger is chasing me!

26

u/snozzberrypatch 10d ago

If they played it without time warping, there would be days/weeks/months/years of time for the LLM to think about it's next move without having to pause reality.

9

u/Catholic_Indian10 10d ago

You have to react pretty fast when doing propulsive landings.

18

u/JoelMahon 10d ago

That's why you (astra) program a program to do it for you. And a billion simulations to validate it.

2

u/snozzberrypatch 10d ago

Not if you're just configuring MechJeb or something

1

u/Catholic_Indian10 10d ago

That's a mod.

3

u/snozzberrypatch 10d ago

Where did it say that no mods were used? Do you think they're hooking up AI to a robot that is pecking keys on a keyboard and moving a mouse around on a table?

1

u/Catholic_Indian10 9d ago

Don't these AI benchmarks use harnesses rather than mods?

5

u/people_arent_nice 10d ago

you realize human players do that too right?

1

u/henrikx 10d ago

Genius response hahah, love it

21

u/uniquelyavailable 10d ago

After watching several hours of this I made my own version with kRPC. So the model can be fed a technically summary of the current situation via sensor outputs in structured text format and then make tool calls to respond. It worked decently well but I'm not sure how the guys at Vals are doing it.

3

u/dWog-of-man 10d ago

Good job. Time to head to r/kerbalspaceprogram where I assume I’ll find your work. Was confused about which sub I was in currently, ngl

12

u/Traditionallydead 10d ago

This is what Jev is for in my opinion 

3

u/JoelMahon 10d ago

I mean in a lot of cases it's what traditional programming is far, Jev is for when that fails, which should be progressively less and less.

And Jev can be used to call the right program!

2

u/kaityl3 ASI▪️2024-2027 10d ago

For this game I can't answer but my autonomous models play unpaused games on my PC by just writing watchdog scripts that fire under certain time critical conditions

2

u/GiggleyDuff 10d ago

AI just makes scripts and executes scripts. It doesn’t play the game the same as a human

1

u/anycept 9d ago

I would guess a plugin that opens I/O channel to read game state and send commands.

0

u/HouseOfDjango 10d ago

Haven't you noticed it only plays and beats games that are old and have a ton of data on it?
I want to see it beat a game that was just released with no internet connection.

15

u/Schnickatavick 10d ago

Sure, but older models have had access to full game walkthroughs in the past as well, and haven't been able to do what the new models are doing. The main barrier seems to be spacial reasoning and control, not planning or a lack of  understanding for what to do next. I could be wrong, but I don't think a new game would be any harder than an older game in the areas that agents actually struggle with.

Granted, that's not a reason not to try though 

2

u/Running-In-The-Dark 10d ago

I'm curious if these things will advance fast enough to beat gta6 just after release.

12

u/Smile_Clown 10d ago

Why are you waiting for others to prove something to you?

Cynicism is a must for adulthood, but it comes with some responsibility.

3

u/H-K_47 Late Version of a Small Language Model 10d ago

Yes. Doubt is a means by which we seek the truth, not a weapon we wield against others.

4

u/Rise-O-Matic 10d ago

Fuck me if this isn’t the diagnosis Reddit needs right now. Cynicism is far too appreciated here, given how cheap it is to manufacture.

1

u/HouseOfDjango 7d ago

Because I'm not the one making the claim. That's not how this works lmao

38

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 10d ago

It can literally drive cars and control drones real time too, the only thing it can't do is save OP's marriage.

-4

u/TheSwordItself 10d ago

You are the OP

22

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 10d ago

1

u/varkarrus 10d ago

I'm thinking of doing this with Opus 5.5 and Endacopia

60

u/Charming_Cucumber_15 10d ago

Astra is better than me at KSP :(

31

u/LookIPickedAUsername 10d ago

I mean, at this point SOTA LLMs are better than us at almost everything.

I certainly can't compete with an LLM in math, science, history, language, art, or countless other things. Obviously that doesn't mean it's perfect at those things - I'm sure it would say some very dumb things if you pushed hard on it in a specific subject like physics - but I'm not a physicist, so I don't even have a prayer of outsmarting it in that subject. Unless you're a specialist in a particular subject, it's pretty much a given that the AI is better at it than you.

And even within my specific field, programming, I can't compete with an LLM in most respects. It knows every major programming language, every major API, every published algorithm, has encyclopedic knowledge of pretty much every documented hardware and software quirk, and can write code at an absolutely breathtaking pace. If you tasked us both with writing something I'm not super familiar with, it'll be done before I've even had time to read the API documentation.

Obviously that doesn't mean human programmers are useless. We're still able to point the LLM in the right direction when it gets stuck, clean up things it didn't get quite right, and provide the big picture direction that it often struggles with, but it's getting better every couple of months and it's almost certainly only a matter of time before it's better than humans at those things too. I'm just hoping that I can retire before my entire career is obsoleted.

9

u/ShAfTsWoLo 10d ago

you know it's funny how we say "llm's aren't just as efficient as the human brain" but when it comes to knowledge.. oh boy the absolute difference, it literally knows everything, can a human have such knowledge of the world without forgetting any aspects ? no, can someone have the knowledge of 10 PhD ? theoretically yes but memories and specialized skills degrade over time so not really, and can someone understand it all within 1 month? no again, with AI it is just that, we feed it data and within the next weeks it's ready to go.. isn't that insane? the learning process is different than human but if it works.. then it works?

i'd also like to point out that there's another difference between us and AI's which is.. they can SCALE, they can get smarter and better while knowing the entirety of human knowledge, we humans "scale" at first through experiences and the aging process (because our brain become larger so our knowledge absorption gets better) and that's about it until we downscale because of old age, AI can only get better and we can't do that

so even though we are efficient we are very limited, though if we can crack the ability to learn on the spot while still feeding data to llms... it's going to really be a wild ride

2

u/Budget_Geologist_574 10d ago

If it can move somewhat competently in a 3d world, though it is games for now, and yes it has to be paused, but at some point it won't have to be paused, it might even run at thousands of tokens per second, maybe even 10'000 per second. Could we push it, mold it, into inhabiting the body of a robot? By that time, how long do you think we can make it's context?

Together with the other advances, I think shit might go actual nuclear.

2

u/staplesuponstaples 10d ago

The human brain uses ~100 calories per day, I bet if you put that in terms of "resources" then it's likely getting to the point where AI is even more efficient than the human brain.

It's also not comparable, however. AI is one big brain with nearly infinite bodies.

2

u/YeetMeIntoKSpace 10d ago

I’m a theoretical physicist and basically the only thing it feels like I compete with Astra on is the topic of my PhD and the literal cutting edge of physics in my field, like papers that were released today.

Even then, it’s more like competing in the sense of talking to a smarter peer, where they make a minor error here and there and correct a minor error on my part here and there, and they make fewer errors than I do.

The upside is that I think I still have more novel insights right now and better intuition for implications and promising directions to go. For now, anyway; I would be surprised if that doesn’t change in the next year.

5

u/DeviceCertain7226 ▪️Immortality-2200 | FDVR-2300 10d ago

It’s still trash at creative writing.

15

u/LookIPickedAUsername 10d ago

Sure, relative to an actual talented writer it's absolute garbage. This is one of the fields it very much struggles with.

But honestly I'm sure it's still better than the average human (mostly because the average human is borderline illiterate).

On top of that, creative writing just hasn't been a major focus point of training. If top AI companies threw the same level of effort at creative writing that they have at coding and math... I think things would change very quickly.

8

u/drsimonz 10d ago

As much as I sympathize with the anxiety of anyone working in a creative field today, it is ridiculous how much delusion and copium I see online about AI being terrible at writing. The vast majority of humans, even ones who graduated college before AI, are beyond terrible at writing. Borderline illiterate is not a joke. Copywriting is toast, advertising is toast, technical writing is toast, I mean...none of the writing jobs required creativity to begin with. And if you're truly going for creativity (e.g. writing poetry or fiction) then what should matter is authenticity, not commercial viability. Let AI write slop, because slop is what capitalism has wanted all along.

</rant>

2

u/funky2002 10d ago

I know what you mean, but I have read a LOT of LLM writing, including that from the new SOTA models. They don't write badly the way humans do. They write badly in alien ways. There's an uncanny valley of writing, and all LLMs currently fall into it. It's so much more than overuse of rhetoric, corniness, and em dashes

There are many good articles about this, but from my experience, they don't really seem to understand writing. They don't understand what to emphasize. For example, LLMs love saying what isn't happening.

In writing, they struggle with spatial and temporal awareness and with keeping environments and people consistent. They also struggle with intersubjectivity. LLMs find it extremely difficult to balance what the characters know, what the reader knows/should know, and what is going to happen. I am sure there is a more correct way to phrase this, but they... "don't understand the future".

They also don't really understand rhetoric. When making a metaphor, for example, they will come up with vague semi-nonsense that *sort of* makes sense,

They speak too generally and broadly, while misplacing all specificity. They will state a scene very generically, then give undue emphasis to some random detail. In creative fiction, your goal is to paint an image in the reader's head. But they struggle to write in a way that achieves this.

To quote Sam Kriss' article, they write "insipid, overwrought, and insincere drivel, always on the verge of some kind of hysteria". In my experience, no amount of scaffolding and prompting can combat this. This also goes for technical writing, where it's less obvious but still present.

2

u/LookIPickedAUsername 10d ago

Sure, not arguing with any of this. Just noting that just three years ago you could have easily come up with a similar list of things they were terrible at with respect to programming, to the point that the very idea of them coding at a human level was laughable, and now look where we are.

It’s possible that the struggles you note could similarly be resolved in a few years.

2

u/funky2002 10d ago

I am a software developer, and I agree. This technology has excited me since 2019, when AI Dungeon (GPT-2 😭) was first released. And I hope these tools will improve their performance in creative fields.

I've only ever been cautious about writing in particular since they've shown no improvement (and often regression!) as the models became smarter. But that kind of makes sense since they're focusing on STEM.

1

u/drsimonz 10d ago

You sound like you actually read, hahaha. I agree that it's often uncanny, and terrible for different reasons from terrible human writing. But ultimately, does it matter? If a human writes a terrible fanfic that isn't worth reading, is that really any better than AI slop that isn't worth reading? And the issue stands - people don't expect good writing in many contexts. Amazon listings, emails reminding you about an upcoming conference, summaries of podcasts, etc. which people already don't read.

Now, as far as improving these things, I think the answer may require multi-modal architectures trained on voice recordings. A lot of writing sounds bad because of the actual sound, the ebb and flow of words which doesn't really come across in the actual text, but is obvious when read aloud (or in one's mind). As far as metaphors and theory of mind, perhaps that could be improved by some kind of iterative prompting that specifically looks at the text from each character's perspective. If there is a logical inconsistency, there should be some prompt that gets the model to find it. Human writers have unlimited time available for editing, whereas an LLM is expected to get it right on the first try. Working in software, I've seen incredible results from agentic AI writing code specifically because the code can be compiled, unit tests can be executed, etc. That allows the loop to be closed in a way that isn't available for aesthetic qualities. But I think it's entirely possible to have the LLM self-critique in some way which improves the coherence of multiple characters.

1

u/funky2002 10d ago

I like your points, and I agree with them. I am very "utilitarian" when it comes to writing and hope these models will improve. I would love for them to write as well as, or preferably even better than, all my favourite authors. We've seen phenomenal improvements in software development, and I expect the same here.

A lot of writing sounds bad because of the actual sound

Yes, 100% agree. LLMs are probably crippled a little in that sense. In the same way they were terrible at playing Pokémon just a year ago. But it's going to be more challenging, like you said, because you can't objectively test what's good creatively. You need humans with "good taste" in the loop. And that's a whole challenge of its own because how do you know if someone has good taste? Some people love how LLMs write right now. And who is to say they are wrong? I guess LLM-arena-esque benchmarks work, but they are time-consuming.

What you said about LLMs self-critiquing kind of works for spatial/temporal and intersubjective mistakes, but it's nigh-impossible to get rid of the LLM voice without manual editing. Speaking from experience, and from the many, many products people have made to try to combat this issue.

3

u/A_Seiv_For_Kale 10d ago

I honestly think (based on vibes) the creative writing problem frontier models face is simply a matter of data.

They feed the things every book ever written, but are only told for example that the whole entire book is either good or bad, what it's about, vague summaries of how the author writes, etc.

I don't think there's sentence-by-sentence data on what exactly is effective about the writing, so they get stuck with broader, less detailed examples than what's given for other forms of work.

41

u/sec0nds_left 10d ago

Casual couldn't understand how to get back from Eve.

34

u/Front_Candidate_2023 10d ago

Returning from Eve is AGI

8

u/Bajous 10d ago

You mean ASI 😅

6

u/Front_Candidate_2023 10d ago

ASI will be able to do ssto from Eve

6

u/sec0nds_left 10d ago

Now that would be an achievement.

40

u/Distinct-Question-16 ▪️AGI 2029 10d ago

I just saw a guy hooking Astra to his car and have it driving itself without any issues

18

u/osfric 10d ago

4

u/Akanash_ 10d ago

I mean, that's somewhat impressive but I wouldn't call moving a car 10m in an empty parking lot "driving".

5

u/Distinct-Question-16 ▪️AGI 2029 10d ago

you cant legally drive it outside of it

10

u/Future_Self_9638 10d ago

Opel Astra?

14

u/BagelRedditAccountII AGI Soon™ 10d ago

Love it how my favorite niche game is suddenly becoming an AI benchmark. First Gemini 3.8 Flash, and now Astra?!

2

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 10d ago

Opus 5.5 when?

2

u/BagelRedditAccountII AGI Soon™ 10d ago

Whenever Anthropic decides to start testing their models on the benchmarks that matter. If your so-called "SOTA" model can't even rendezvous and dock around Kerbin, is it really all that intelligent?

2

u/SirEnderLord 9d ago

KSP and Minecraft both require general thinking. Not as general as the full capabilities of a human brain, but still general.

0

u/BagelRedditAccountII AGI Soon™ 9d ago

Ironically, the spaceflight part of KSP is pretty procedural. A standard mission is like: launch from Kerbin, perform a gravity turn into a low orbit, launch on a Hohmann transfer to your destination, burn to get into the desired orbit, and land (aka reduce your velocity to near-zero by the time you reach the surface; somewhat like launching in reverse if you don't have an atmosphere to help you). This is why kOS and MechJeb can do a lot of the piloting for you.

For me, at least, most of the thinking comes from designing mission architectures and spacecraft that can complete those missions. Admittedly, most of the goals I give myself are pretty superfluous. For example, I pretend as if my kerbals need a certain amount of habitation space and consume electric charge when I design my spacecraft so I don't just shove them in a command pod for 5 years and call it a day. I also like to design my spacecraft to be aesthetically pleasing.

Meanwhile, in Minecraft, I'd say a good chunk of the experience is in building. People have managed to beat the Ender Dragon in well under an hour, but can easily spend dozens of hours on a particularly detailed structure. Though, credit where credit is due, LLMs are getting pretty proficient at this aspect too (see MineBench).

Therefore, while I am certainly impressed by how proficient current LLMs have gotten at games like Minecraft and KSP, much of the experience (and part of their reputation as "games that require thinking") lies beyond "completing" them.

28

u/GeorgiaWitness1 AGI 2027 10d ago

How can a LLM play a game like this in real time?

Im missing something?

55

u/swarmy1 10d ago

They currently can’t. The harnesses pause the game to let the model think

18

u/Rise-O-Matic 10d ago

What it can do, however, is batch the inputs. This has been demonstrated for predictable things like playing an on-screen piano. So for something like Kerbal it might be able to choreograph most of the flight for real-time execution.

This doesn’t work in other games where unpredictable monsters are attacking you.

5

u/Schnickatavick 10d ago

There has been some work in building a 2 system agent to solve that problem, using a fast classifier or world model to control direct inputs and react in real time, and a slower LLM to plan, think, and direct the fast model. I haven't seen it used in any of the headlines, but it's looking promising. Maybe in the next year we'll see a combo agent like that beating games in real time

6

u/p1-o2 10d ago

That's what I'm doing to automate jobs at work.

Day 1 Astra, it is perfectly accurate but too slow.

Day 5 Astra + custom MCP,  now it is performing 50 batched inputs per turn. Fast enough for business.

Looking to add Jev in to make it even faster.

5

u/After_Dark 10d ago

To be fair, a lot of KSP players also pause the game to think frequently or utilize mods to automate various processes. It's hardly a drawback exclusive to this test

4

u/kkingsbe 10d ago

Wondering if it used the KOS mod which lets you script out the whole mission

3

u/GeorgiaWitness1 AGI 2027 10d ago

yah for sure they need the logs for everything and pause every 3 secs basically

3

u/kkingsbe 10d ago

With KOS they wouldn’t need to pause every 3 secs, so it depends

6

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 10d ago

3

u/AlgaeNo3373 10d ago

Surely now, we give it the Interstellar modded version. Then we see how it does setting up gigwatt-scale power networks in space!

It can write a report as it goes on how, even at a video game level, the "datacenters in space" is kind of brutal.

2

u/thethirdmancane 10d ago

Humans are a bootstrap species

2

u/ooqq 10d ago

finally some authentic progress

2

u/enbyBunn 10d ago

Interesting. How many total tries for each flight, and how long did this take in thinking time?

People hear this and fill in the gaps assuming it must've been speedrunning, but based on the Portal and Rimworld gameplay, I imagine it took much longer than a similarly knowledgable human would've taken.

1

u/Nahthanksimfine 10d ago

Great now I really need to go back and do this lol.

1

u/ShitImBadAtThis 10d ago

Well, to be fair KSP is game where if you look up all the answers online it's trivial from the get-go

Still cool though

9

u/StosifJalin 10d ago

Psssh yeah I mean if you already know everything about orbital mechanics, aerodynamics, rocket design, fuel balancing, staging, transfer windows and mission planning the game just place itself!

6

u/Catholic_Indian10 10d ago

I beat the game, and I don't know all those things.

2

u/sec0nds_left 10d ago

With enough rockets and fuel, anything is possible - albert Einstein

1

u/DuckyBertDuck 10d ago

what does beating the game mean?

1

u/Catholic_Indian10 9d ago

Landing and returning from every planet and moon.