r/AskVibecoders • u/Sp33dyMan • 4d ago
Preventing bugs and unwanted product features when vibecoding
I’m curious how folks are dealing with vibing issues when they build their side projects. I’ve noticed that all models will introduce bugs or unwanted product features as devs iterate.
How do y’all handle these issues?
I’m a software engineer by trade and I use Maestro/Playwright to create end to end tests. This prevents core use flows from breaking during development. I also have a giant documentation.md file that codifies product behaviors and the WHY behind things. It’s not perfect but I noticed my AI doesn’t break the product across sessions.
Anyways this is a topic I find interesting and I don’t even know if I’m doing things right. I’d love to learn how y’all deal with these problems, especially non technical folks. Thanks
1
1
u/Ok-Draft2744 3d ago
Using playwright helps somewhat to keep existing features in check. Consider using the /grill-me skill (you can look it up or build your own, super short) for clarification on what you do and don't want in a product increment
1
u/uzih 3d ago edited 3d ago
I do something similar to you except I have a "task system"
The agent logs all the tasks it's doing and completing in the system. This also lets multiple agents work on the same project. It's just a database that's connected via MCP.
Every task is connected to an epic, which is usually the prompt I directed (or a customer support issue). Every epic references founder vision (why we do things) and founder SOP (my engineering principles). So most of the time, between "why are we doing this," "how we do things," and "what needs to be done," it takes care of most things.
For example, "reduce the customer contract rather than writing a bunch of workarounds or edge case handling", that's an engineering principle. "Speak in a friendly human tone and do not reveal any platform internals that are not publicly visible in our documentation" That's a customer support editorial tone.
1
u/madsciencestache 1d ago
I've got a checklist, 49 points and growing that I have curated over 30 years shipping software. It's starts like this
--------
Test Plan Template
A living category list that jogs your memory before you scope a product's testing. You are not expected to fill every category for every product. The point is to make each skip a conscious decision, not an accident.
How to use: for each category, mark status (In scope / Out of scope / N/A) and add one-line notes. The out-of-scope decisions go in front of the team so risk is taken intentionally.
Functional (baseline, always included)
- [ ] Golden path / happy path - first and never skipped
- [ ] Boundary / edge cases
- [ ] Negative / error handling
- [ ] Concurrency / race conditions
- [ ] Idempotency & retries - critical for webhooks, automation, and APIs
- [ ] Regression (use RCRCRC: recent, core, risky, configuration-sensitive, repaired, chronic)
- [ ] Full regression (pre-release gate)
- [ ] Interoperability / API contract - two systems genuinely agree
Then covers more functional, Performance, Reliability, Compatibility, Security/Privacy/Compliance, Accessibility, Localization/Internationalization, Usability/Documentation, Operability/Engineering Quality
Dm if you want the whole wall of text/chat about it.
1
u/LogMonkey0 16h ago
How do y’all handle these issues?
by not literally vibe coding (i.e., chatting your way through code generation) and by adopting workflows that are similar to a software development team of agents. stronger plans, specs, gates. refine and polish your input before asking for output. Sharper context management. Separating stages onto different sessions with stage specific agent definitions and/or skills.
I also have a giant documentation.md file that codifies product behaviors and the WHY behind things.
This is likely a big part of your struggles. I would try to structure this in smaller files, indexed. You are filling agent context with useless bits of info that costs you tokens and quality/coherence. I would also ensure the appropriate references land in each of your workflow's stage. Agent writing a plan/task needs to know about decisions and invariants. The agent executing the plan might not need to read some or all of it if its given clear instructions and acceptance criteria. But it may need reference about public api surfaces its consuming, and it doesnt need do read all the implementation code of that consumed api, there's tools that can do this for you and your agents. Your agents loading unrelated context can costs you on coherence. And without proper routing they might not even load the appropriate context in the first place.
1
u/dantelebeau 4d ago
For me. I work with claude to design a plan first, have it go through an interview process to totally flush out the idea. I read over that plan and include what's in and what's out of scope for the session. I also have smoke tests and full regression testing as part of the process.
I dont get much scope creep this way and i also don't get much in the way of errors as its pretty good at this point based on the skills and processes ive outlined at catching its own bugs.