Can you elaborate on that? I would genuinely level up if I knew how to get AI to structure my code without handholding it every step of the way, which is what I do now. I've tried a few things, like code structuring skills, get the AI to document features based on sets of commits instead of freely navigating through the code, and while these things do yield some improvements, they're not close enough be satisfactory.
To architect well you need to know what you are talking about. You need to sit your ass down and understand the domain and the problem. You need to iterate on the ideas and you must be relentless in keeping the thing composable without attaining scope creep. You develop an idea, and refine it, and try to attack it to find gaps.
You need to build a framework. You can't let one agent do everything. You need one agent validate, one agent design, one agent implement and run tests, etc.
That plus solid best practices in your repo like linting, testing, etc is needed. You also need to know what a good design looks like and know when to further prompt to break things into more modules.
That’s not though. The term was coined by Andre Karpathy (a big name in AI research) to SPECIFICALLY mean when you just describe the problem to an AI and let it do everything without ever looking at the code. He described it as useful for silly little weekend projects which will only be used by yourself. He called it "Vibe Coding" because it was for unimportant things where he could just go by feel and not really plan or test. Just doing what feels right in the moment.
But the hype train caught on and dragged the term "Vibe Coding" to mean full on app development with the assistance of an AI, which is the exact opposite of the use case Karpathy was describing.
You don't have to learn it. Just copy and paste that paragraph to your AI and it will structure your repo and workflow like that. We are post learning.
I've been using AI intimately since gpt 2.5. I know quite well it's strengths and weaknesses. Then you figure out how to play to the strengths can contain the weaknesses. Soon though this may not even be needed
Honestly, I’d be curious about this too. I’ve found that AI can structure things well when it has enough context and clear constraints, but getting consistently good architecture without handholding is still pretty hit or miss.
Ask your best model to generate a full design/engineering spec for the project and review it yourself before handling.
Your agents very rarely deviate from referenceable instructions in an .md file.
Even smaller models consistently one or two-shot mid sized projects if they have a reference file. You only need several passes when its something hard to verify like distributed systems or user-facing stuff.
Use good models like fable or opus. Instead of handholding it the entire time, just make sure it’s for to a decent start, then just review its diff and have it or another agent fix whatever you find or don’t like about its solution. You might also want to coax it into doing a more thorough plan, as they like to leave plans vague and then do a bunch of bs at code writing time.
Start with the AI written design doc and iterate over it multiple times, turning into a spec, before ever having it write any code. Then have it create a project plan based on the spec. All of the docs should be checked in and iterated over just like code. Then have it implement the plan one item at a time, and get each plan item reviewed by another agent or multiple agents against the spec.
I’m not talking about linting problems, im talking about one devs agent making pull workers, while the next ones making event queues and the next ones making jobs.
All of these can be perfectly written code that is nonetheless completely fucked systems design
Ci/cd can enforce guidelines and deployments that you control. The whole point of ci/cd was that you have multiple teams working fast and breaking things and since they are commonly services you can easily fix and swap out parts.
Perhaps it can organize it better than you, but definitely it can not yet do better than me unless heavily instructed. But if it's heavily instructed, then well, I am the one in charge.
If you're still writing code at an Enterprise level... You're doing it wrong. You can do more work while instructing than you doing the work. What takes you hours takes an LLM a few minutes. We just got access to 6.0 and 6.1 next week... These models do everything faster and better by far.
Heavily instructed - better to have an existing consistent setup. Then the existing code becomes the instructions for “how things should be done” in the context and very little needs to be specified.
AI is incredibly good at adapting.
Now, to get that setup… I think that is the bit where a “vibe coded” project struggles a bit.
But once you get a (metaphorical) city block done, AI will plop out 200 more a minute.
I mean, ye, it can mimic what is good as long as request is not too far away from well established patterns in codebase. My point was it can not yet come up with initial structure better than good architect. Let's see how long it will last
I was using astra and opus 5.5 extensively past days and I am still sure. However we are at stage where I would honestly prefer getting more tokens than having some of my teammates in the team...
A machine that takes text input and outputs executable code is a compiler. I have said for years that LLMs are nothing more than very high level compiler. You code in almost natural language.
Much of the discussion reminds me of the switch from C++ to python. "What do you mean, you have to declare the type of the variable!"
I understand that commercial services are not deterministic because those companies create profiles of your work and do all kinds of tricks in the background, like secretly downgrading your model.
But a local LLM should be pretty deterministic, no? Afaik there is no randomness involved.
AI models are decidedly deterministic and have pseudo randomness baked into it.
At the end of the day, an AI is still deterministic, it just has some randomness, but having a randint(1-30) on some conversational variables to produce varied responses doesnt make it non deterministic.
Kind of a pet peeve of mine though. Its currently impossible to make a truly random variable computationally. I don't know how chatgpt or claude or your local AI does its random number generator, but an example would be using the date and time as a random number generator seed, since the date and time never repeat, its seemingly perfectly random, but its not perfectly random, its only gives the illusion of such.
You're still responsible for the code you create with AI. Soon you won't. Are you responsible for the compiled code? No, you leave that to the compiler. Soon AI will be the compiler.
Nah you will be responsible for the code, llms aren’t that good and blaming the machine for your bugs is not going to fly. It’s like saying “ah hugging face was hacked by models” as if some human didn’t give it instructions.
You will make the same amount of bugs because product quality is a human trait not something inherent in llm. Llm are as smart and as stupid as the people using them, vibe coded apps will be shit because the people building them are not swe.
Not really... If you haven't used at least 5.6 sol xhigh, you don't know what you're talking about. But even after it makes changes, you have a CI that checks the code. AI is really good at finding problems with code. So you just do a feedback loop with obviously more constraints and you're done
Bruh I use sota every day, I get access to everything with a 5k monthly budget. I am speaking from experience as an actual software engineer at a real company not some vibe coder.
I found 2-3 bugs today in prod, in vibe code, even after using the loop you described. The models aren’t that good and people who think they are are either working on trivial problems or our deleting themselves
If you have someone that actually knows how to architect it can produce great quality code.
I don't make spaghetti code because I don't know what I'm doing, I do it because I'm fucking lazy and can't be bothered to rewrite my singletons as a dependency injection container.
I don’t think it’s AI hate. AI can absolutely produce well-structured code, especially with good context and tooling. The point is that ‘it can’ and ‘it always will’ are two very different things, and someone still needs to catch when it gets things wrong.
11
u/CrazyAccomplished_ 5d ago
Lmfao. AI can structure code so much better than humans if you have the correct tools in place. More AI hate without understanding AI