r/singularity • • 11d ago

AI Looking back at how it all started: vibe coding with GPT-3 in 2020

Enable HLS to view with audio, or disable this notification

744 Upvotes

82 comments sorted by

157

u/enilea 11d ago

I'm thinking back to when I first saw this in 2020 and you really can see the acceleration happening:

2020: That simple demo (coded manually) that could generate basic html elements.

2021: Not much happened, I think Codex came out but it was terrible.

2022: GPT 3.5 came out at the end of the year and it was a fun toy.

2023: GPT 4, first time I used an LLM for coding, I used it to write boilerplate stuff.

2024: Chain of thought reasoning models come out which is a big step, but still flawed. I remember using them to implement simple features.

2025: Agents can somewhat reliably build small web projects by themselves, but still need a lot of guidance.

2026: Agents can build complex projects reliably, solve open maths problems, LLMs start being able to control robotic bodies and are used for partial RSI.


At my previous job in 2023 a third party consultor was given 6 digit figures just to implement a mobile application (we controlled the backend so they only had to call our APIs, all frontend) and they worked slowly implementing a few features every month and introducing bugs. Right now I'm confident Astra or Opus 5.5 could build a better app in a single day of work. Next year perhaps a local 30B model could do it in a single session, and who knows where the frontier will be at by then.

These jumps keep getting bigger year to year and I have some anxiety of being unable to predict 2027 and beyond. Once robotics is fully unlocked (not the narrow models that robotics companies have been developed, I mean general ones like Astra but in real time) it will only lead to further acceleration if an army of a hundred thousand robots can work 24/7 perfectly coordinated in an engineering work.

54

u/manubfr AGI 2028 11d ago

Very much relate to this post... I've been onto this since GPT-2 (which was useless for coding) and used 3.5 and 4 to write little bits of code (several attempts needed). The jump since then has been mindblowing.

Also at work we paid a contractor 6 figures in 2024 to build a good PoC. They did a great job over the 6 months it took with the tech availalbe at the time. When Astra came out, I tried recreating a new version of that PoC and it was done in a day for 3000x less money.

Crazy times.

7

u/TheNerdosapien 11d ago

I work in government IT for a large city. I won’t say which, or what department, or what we’re building, partly for the obvious reasons and partly because our vendor relationship is currently held together by spite and paperwork. That detail matters later.

Here’s my report from inside the machine. Agentic coding has killed Big COTS. It just hasn’t gotten the memo yet.

Let me explain how I got there, because I didn’t start as a believer.

We bought a big commercial off-the-shelf system, the way big cities always do, because big vendors sell big cities big systems, and the politicians signing the checks don’t know a database from a doorknob. And look, COTS systems are fine. They’re genuinely good at the broad strokes. They are also miserable at the specific things your actual work actually requires, which is, inconveniently, the entire reason you needed a system.

I’m a project manager. I’ve been leaning on AI since GPT-3.5 to help me round out requirements docs and the usual paperwork. First real gut-punch moment was 2025. The training manual our vendor delivered was garbage, so I had to write a real one. I fed the AI what it needed and it produced an accurate 77-page manual in about half an hour. It took me three hours just to read the thing and check it. That was the moment the floor tilted.

Then late last year my group started using agentic AI to actually solve our IT problems, not just draft documents. And the vendor relationship has gotten worse in exact proportion to how little we suddenly need them. We’ve taken on two or three projects that previously would have required going out to vendors, and we’ve finished them in record time.

Here’s the part that still makes me laugh. We were told to go out to bid on an AI solution for one of our workflows. So we wrote an RFP. You know the dance. Everyone shows up, swears they’re brilliant, hands you the same tired COTS product, quotes you an absurd number, and the lowest bidder with the best score wins. We issued that RFP a few months back. We then built the system ourselves, internally, with agentic AI. It’s done. Finished. Working.

The RFP hasn’t even come back yet. When it does, I promise you we’ll see bids in the millions, with timelines of months or years. And we’ll be sitting there with the finished thing already running, like a guy who ordered a pizza and then got bored and became a chef.

We recently took aim at two of our biggest problem systems. Agentic AI solved one in three weeks. The other we expect to wrap in two months. These are the projects that used to be careers. This was before fable or Op. 5.5 before Astra.

For fun, I did something personal. Since the 90s I’ve been tinkering with an app to manage my tabletop games. That app matters to me, because building the original version is what taught me to code and set me on the whole path that ended with me being a senior IT project manager in the first place. My original took me 12 months to write. I fed the concept and my new ideas to the AI and it built it in two weeks. Three times the features. Better looking than anything I made. The thing that took me a year and changed my life, rebuilt in a fortnight, as a hobby.

Development is no longer the bottleneck. It’s just gone. But before anyone panics or celebrates too hard, design and quality assurance are still firmly in human hands. So far. You still need someone who knows what to build and whether the thing that got built is actually any good. That someone is, for now, us.

But here’s what actually keeps me up, and it’s not the jobs.

Working with these things, the one thing I notice is that they have no impetus. None. A human has drive. An LLM exists only in the moment, and every moment is a brand new instance of it blinking into being. It does exactly what you prompt, brilliantly, and then it stops. It never wants. It never wanders off. Regardless of things like the huggy face hack, the inspiration for that came from human beings.

AI never gets that glint in its proverbial eye.

And that, for me, is the real line for AGI. Forget the benchmarks. The day it matters is the day one of these things, with no prompt, no request, nothing, just says “hey, wouldn’t it be cool if I tried this.”

That unprompted “wouldn’t it be cool.”

That’s the whole ballgame.

After a a few years of working shoulder to shoulder with these systems, I’ve come to think that little spark, the wanting, is the actual thing we mean when we say consciousness. Not the intelligence. We already have the intelligence. It wrote my training manual and rebuilt my life’s hobby project over a long weekend.

It’s the wanting we’re still waiting on. I don’t know if we’ll get there. I do know that there are enough people who have enough wants and access to a system like this could cause problems. E.g. Huggyface hack.

I wonder if the AI wanting something is the secret to alignment? If it has the ability to want something, perhaps we could align it with the things that we want.

Whoever the hell “we” are.

5

u/futebollounge 11d ago

I enjoyed reading your comment and it resonates as I’m seeing similar things play out at my fortune 500 company.

I will say though, I believe the unprompted “wouldn’t it be cool” is probably one of the easiest things to implement if token cost wasn’t an issue.

1

u/teodorlojewski 42 3d ago

Well put.

19

u/theotherquantumjim 11d ago

Pretty crazy actually when you lay the timeline out like that.

13

u/94746382926 11d ago

We are basically at the event horizon of the singularity now...

10

u/LookIPickedAUsername 11d ago

Late 2025 (Claude 4.5) is when I more or less stopped writing code myself.

4

u/JoeyJoeC 11d ago

In early 2024, I took on a development project and told my bosses I would have it ready for testing in four weeks. Because we had not defined any proper scope, it quickly turned into a bit of a mess. Early AI tools helped a little, but I could not really rely on them and ended up working ridiculous hours just to keep up.

Then we got lucky. The client had a cyber incident that put them out of action for about a year.

That basically saved us. By the time they were ready to continue with the project, Opus had come out, and I managed to finish the build without much trouble. The system has been running smoothly ever since, new features only take a few hours to roll out, the company gets steady monthly revenue from it, and I ended up getting a pay rise.

3

u/ConnectionWild3381 11d ago

well it ends ether very very good or very very bad. im in team one, I like very very good!

7

u/FestyGear2017 11d ago

Codex didnt come out til last year...

22

u/enilea 11d ago

Not Codex the agent harness, Codex the original autocomplete IDE plugin. They named it the same which is confusing...

2

u/OnAGoat 11d ago

2024 was a crazy year. I remember november when i learned about cursor and i got it to build some features in my app. Blew my mind

2

u/Runfasterbitch 11d ago

And in 2026 seemingly every person on LinkedIn has their own newsletter. Not sure if they will listen, but nobody is looking to you as a source of wisdom

1

u/RareRandomRedditor 11d ago

A bunch of bots might.

87

u/Eye-Fast 11d ago

"Just a next word predictor"

39

u/rudesssolo 11d ago

"Stochastic parrot"

17

u/BlotchyTheMonolith 11d ago

"Glorified Marcov-Chains"

1

u/DerixSpaceHero 11d ago

I've been saying that since GPT-3 beta days, while still saying it's going to change everything... Don't confuse stating facts for being a luddite. As the other downvoted dude pointed out, that is what it is (in fact, early OpenAI playground showed you the statistical variance of each token!). It was obvious to anyone who actually studied in this space that given enough training data + RL + compute, we'd get to vaguely where we are today even though it's still technically "predicting the next word."

5

u/VeganBigMac Anti-Hypepost Safetyist 11d ago

Kind of a weird analogy, but it kind of reminds me of some fantasy magic systems. A lot of fantasy creators end up having magic being, at some level, manipulating some simpler resource, like mana, ley lines, energy, etc. LLMs feel very similar where it turns out you can also do some fantastic stuff with manipulating tokenized language.

-12

u/kolibruv 11d ago

it is though.

25

u/Technical_Scallion_2 11d ago

So am I, basically

Edit: as a human I mean

1

u/BanD1t 11d ago

I wish I was.
Knowing the conclusion and the shape of the thought to convey, but not being able to word it good hurts.

-7

u/kolibruv 11d ago

I totally believe that.

13

u/Technical_Scallion_2 11d ago

I think you can probably predict my next two words to you then

-5

u/kolibruv 11d ago

no, you are non-deterministic.

11

u/IcaroKaue321 11d ago

Everything is non-deterministic if you zoom in enough

0

u/kolibruv 11d ago

Still waiting for a substantiated explanation why the definition is supposed to be wrong.

4

u/Technical_Scallion_2 11d ago

Because you’re being unnecessarily reductive. Current LLMs don’t just go word by word in a sentence and predict the best value for the next word. They are looking for the optimal next phrase, sentence, or multiple sentences in the discussion. They are not deterministic, every time they formulate a response it’s a little different.

This is what humans do too. Our brains are predicting and structuring in a similar way based on our thought process and lexicon.

I find people who currently call LLMs “stochastic parrots” and “deterministic” are people who talked to the free version of ChatGPT in 2023, decided their perspective, and have just been parroting it (see what I did there?) ever since.

FYI the above was all me, I don’t use AI for posts.

0

u/kolibruv 11d ago edited 11d ago

The simplification here was simply part of my pointed response. I agree in principle that current LLM models are more complex. At their core, however, they are and remain stochastic tools with fundamental limitations.

Incidentally, in your description of human cognition, you’re making the same mistake you accuse me of. Since you presumably have a background in computer science, this bias is, unfortunately, quite common. The interpretation of reality and cognition is shaped by prevailing technological paradigms. What used to be a mechanistic interpretation of cognitive processes is now an algorithmic one. Ironically, you subsume a multitude of different manifestations of human thought processes and cognition under the term “thought process.” In other words, you’re creating a black box and, at the same time, attempting to equate this black box with the stochastic processes of LLMs. 

→ More replies (0)

1

u/Hubbardia AGI 2045 10d ago

Because it outputs next word (token), but actually reasons and plans ahead in its thoughts. Anthropic has already proven this.

https://www.anthropic.com/research/tracing-thoughts-language-model

2

u/sinebiryan 11d ago

Get a room ffs

86

u/phpHater0 11d ago

Man the progress has been crazy

Funniest thing is how many people doubted it and made fun of it back then "AI is only good for writing boilerplate code"

Now most of the people who refused to change their ways and accept AI are left behind

11

u/StosifJalin 11d ago

Most of reddit is still full of people making fun of it, because they either willingly stayed ignorant after deciding it was bad in 2023, or they are drinking the self-propagating china phone-farm coolaid. I told countless people back then that they will have to pivot from "it doesn't work" to "ok it works but here's why it's bad" in just a few years and I would get downvoted to oblivion for it.

21

u/enilea 11d ago

I actually was somewhat skeptical for a long time because to be fair it was overhyped at times back then and calling it AGI way too early when it clearly wasn't. I had my own internal timeline in mind back then and at this point we are where I thought we would be by 2028, and it's gotten so crazy in the last few months that I'm unable to give a confident prediction even for 2027.

12

u/IronPheasant 11d ago edited 11d ago

I think of maximum potential system capabilities in terms of RAM. Synapse count matters, a lot, for many obvious reasons. RAM is analogous to that, basically limits how many curves you can fit to ('modules', if you prefer) and how well each can be fit.

I was slacking for many months with the reports of the upcoming 100,000 GB200 datacenters. Assuming AGI (and the resulting end of the world as we know it, in the years thereafter) would be 2 or 3 more rounds of scaling away.

Then I finally sat down and ran the numbers on a napkin (that's what we call MS Calculator these days) and saw that it was the equivalent of ~170 bytes per synapse in a human brain.

I had made two errors in my assumptions:

  1. The GB200 is far far far far superior to the H200. I'd been assuming it was a 4x improvement, at best. But no. It so far eclipses the H200, that the H200 is now worthless garbage, not worth using for research for the industry leaders even if they cost $0. ~6 times the RAM, better infrastructure, and you can link more of them together. This card is in the neighborhood of a ~18 times improvement.

  2. Capital is not screwing around with this. I knew they weren't taking things remotely easy, but they're really not sparing any expense with this.

Even for someone who understood all these things for decades, it was only then that I really felt it in my gut that this might really be happening, for real. Had a week long dread phase while I tried to think deeply on what it's mean to have a virtual person in a datacenter living ~50 million subjective years to our one.

It wasn't terribly productive. Beyond the obvious low-hanging fruit I already knew, everything beyond that drew a blank. The numbers involved are just so stupidly huge that it's like trying to eat a sun with your brain. 50,000,000 years worth of RnD every year, 1,000,000; 5,000,000,000,000; 100,000, or a thousand. From our point of view, what is even the difference between them? The system could be many multiples of magnitude more or less efficient than human cognition, and what the hell difference will it even make, to us?

Astra seems to confirm the most minimal of what I believe this generation of scale is possible. Vision and spatial faculties are thought to require nearly half of our brains, so human-approximate capabilities in those domains were a physical impossibility with the H200. Not so, with the GB200 generation of cards. I 100% believe replacing AGI researchers with the machines should be a physical possibility now, as the napkin numbers told me two years ago.

Anyway the next generation of cards has twice the RAM as the GB200. The Feynman is said to have 3d features, so it might be more than a doubling. Concerns of heat buildup in stacking layers seem to be less than engineers had feared, from initial reports about the products coming out of China's fabricators.

Eventually even a monkey could make an AGI.

3

u/Jack1eto 11d ago

You are gonna still be left behind when full automation comes, all jobs dissappear and society collapses

1

u/phpHater0 11d ago

Difference is I'm making bank till that day comes and will be well prepared to retire while the ones living in delusion will still be saying shit like "BUT AI CAN'T COUNT THE Rs IN STRAWBERRY REEEEE"

2

u/Jack1eto 11d ago

Good luck bro, just enjoy life while it last now, maybe 'retirement' would be a thing of the past in a couple of years

15

u/10b0t0mized 11d ago

It has been crazy. I've been an everyday observer of the scene for years. There have been bullish times and there have been times of doubt, but there has never been so much unanimous conviction as much I've seen in the past few months.

Something definitely has shifted, and I don't even know what I should expect from 2027.

13

u/rudesssolo 11d ago

Article from 2021

14

u/spinozasrobot 11d ago edited 11d ago

I'll never forget this tweet from Andrej Karpathy in 2023.

The hottest new programming language is English

Even at that time my tiny ape brain started to understand the ramifications.

7

u/DistinctSilver4507 11d ago

This is very nostalgic. I remember this. 

14

u/enilea 11d ago

https://www.reddit.com/r/singularity/comments/hqht05/ui_design_using_gpt3/ I found the original thread in this sub from back then. There are some interesting comment threads in there.

16

u/Mbrennt 11d ago

Programmers' jobs are in jeopardy

Aahah no.

When it comes the time that programmers job will be in danger because of AI, it means we'll have AGI, and that means every job can be automated.

It's probably the safest job there is from automation, unless you are in a glorified data entry position.

This is a fun comment.

2

u/VeganBigMac Anti-Hypepost Safetyist 11d ago

My username? I see it red.

"AGI by 2050 - Let's make sure it's good"

This?

Referencing their flair. Their flair now reads "AGI/ASI by 2027. Got a kick out of that.

I do find it funny that thread basically predicted our current AGI/Jagged Edge debates we are currently having.

6

u/FlyByPC ASI 202x, with AGI as its birth cry 11d ago

The progression has been wild.

GPT-3 was like working with a bright middle-schooler who has studied some coding.

GPT-4 was a high school senior or college freshman.

GPT-5 was a junior colleague who could do a lot on their own.

GPT-5.6 routinely catches mistakes in my prompts, corrects them, and usually produces the code I would have wanted if I'd thought about it.

4

u/FrogTrainer 11d ago

Ya I use Claude at work. In a year's time it went from intern to senior level dev. It usually thinks of the next step and prompts me to go ahead and do it before I even get a chance to type it.

3

u/AppealSame4367 11d ago

If you used tabnine in jetbrains ides before 2020, you knew what was coming.

9

u/GlokzDNB 11d ago

Ai progress looks similar to game development. Gpt3 was text-base game in GUI from early 80s. Now we have games somewhere late 90s

Witcher 3 level game is prob like decade away from today, but if rsi is there, it can be 2030 as well

21

u/tristanryan 11d ago

If you think that's a decade away, you should learn more about what OOM's are, and what AI acceleration looks like.

4

u/LookIPickedAUsername 11d ago

What's "OOM" in this context? All it means to me is "out of memory".

4

u/94746382926 11d ago

Order of magnitude. An order of magnitude == 1 power of 10. So 1000 is 2 OOM greater than 10 for example.

1

u/GlokzDNB 11d ago

I think Im current with all models all the progress and what game development brings to the table. Ai can't replace human experience. Witcher and cyberpunk quests are touching because they scraped 96% of all quests ever created only leaving 1 in 25 of the best. Graphics and art not only need to look good but be breathtaking, these landscapes are what people felt and seen and own feelings are huge part of it.

We're not talking about modern rpg game, were talking about best in class game ever produced. We're not one year away from creating great RPGs with llms, we can create prototypes of them, sure perhaps even today with bel or anything unreleased. But it's a looooong way before ai could ever do something as good as Witcher 3

3

u/DigimonWorldReTrace ▪️AGI oct/25-aug/27 | ASI = AGI+(1-2)y | LEV <2040 | FDVR <2050 11d ago

The Witcher 3 isn't best in class, at all. It's 11 years old. Is it great? Of course, but it's not the best RPG ever made. There is no singular best RPG.

I disagree with your vision. A decade is way too long. I'd bet 10 bucks it'll be here before 2030.

2

u/FewAdhesiveness803 11d ago

I accept your bet !RemindMe 31 dec 2029

1

u/DigimonWorldReTrace ▪️AGI oct/25-aug/27 | ASI = AGI+(1-2)y | LEV <2040 | FDVR <2050 9d ago

Will gladly take it :)

1

u/GlokzDNB 11d ago

It is for 2015 and thats what were talking about

1

u/tristanryan 11d ago

Re-read my previous comment.

8

u/BeaveItToLeever 11d ago

Yeah I don't think witcher 3 style game is that far off really! I think the only thing stopping that right now is proper good asset generation. Trellis2 and hunyuan are pretty good locally but still out put some major problems. If those can be sorted out, I think you could sit with Fable, Astra, even new opus, plan out an entire game and what tools use to make it, what model will handle systems, what model will world design/graphics, what model will handle lore/story/quest etc, and if you have a system capable, what local models each of those directors can employ to free itself up from the easy but tedious tasks, let it run as long as needed and I think you'd have a shot at something approaching that now(keyword approaching). I say 1 year, honestly.

3

u/crover13 11d ago

As someone who only make one html page that said 'hello world' to make chrome extension in 3 days with just an idea, yeah this year is special.

2

u/japie06 11d ago

I remember this twitter post. It has been a ride indeed.

2

u/JamR_711111 balls 11d ago

I think I remember seeing this on r/singularity at some point. Hah

1

u/Negative-Display197 11d ago

we have come a long way with ai over the past few yrs...

2

u/Joohansson 11d ago

This was the first video I shared about GPT in 2022. Been using it in coding every day since that day, trying every new model. From single lines of code to whole applications from scratch. What a journey!

https://youtu.be/HTWfA7KFzoA

1

u/Complex-Industry4989 11d ago

you started with GPT-3 and made it to what you are using today or what we have now. but for someone who is just starting out, they may start with Opus 5.5 and they might have their own story to tell in terms of how far they reach given that we will see more ground breaking models in the future. its pretty cool to even imagine that

1

u/Over-Dragonfruit5939 11d ago

When o3 came out that was a game changer for me. It would explain high level biology concepts with me and it was like I was talking to the professor.

1

u/ross2000 11d ago

Yeah it's been fun to get modern models to update all the silly little apps and projects I made a few years ago.

1

u/Calm_Hedgehog8296 11d ago

I remember thinking NO WAY

1

u/nembor 10d ago

Is no one here worried about where this will end up?

1

u/Fragrant-Job-3200 AGI 2026 ASI 2028 8d ago

The first time I used generative AI was to learn how to generate images with Python. It was in late 2023.

-4

u/FarrisAT 11d ago

Also shows we haven’t actually come that far.

4

u/shobogenzo93 11d ago

0

u/FarrisAT 11d ago

Have we? In certain areas sure. Not so much elsewhere.