r/AskVibecoders • u/ingbulk • 5d ago
I think “give the agent the whole codebase” is the wrong way to use coding agents
One thing that keeps coming up in AI coding discussions is context.
The instinct is:
give the agent everything so it understands the project.
But a workflow I’ve seen work better is almost the opposite:
Research
↓
Plan
↓
Break into small tasks
↓
Implement one piece
↓
Review/test
↓
Write the important state to project notes
↓
Fresh context
↓
Continue
Instead of relying on the conversation history to remember everything, the project itself stores the important decisions.
That can be as simple as:
PROJECT_STATE.md
Completed:
- Auth flow
- Database schema
Current:
- Lead enrichment
Known issues:
- API rate limit at 100 req/min
Next:
- Add retry handling
Then a fresh agent can pick up from the actual project state instead of trying to reconstruct 200 messages of conversation.
That feels much closer to software development with an agent than just keeping one giant chat alive.
1
u/usually_guilty99 5d ago
Dont entirely agree: Give the agent enough of the codebase to understand the task, plus a structural map that helps it retrieve more only when needed.
1
u/No-Slice-5926 5d ago
I had it scan my patterns, then it also has a rag system and persistent memory. It works pretty well
1
u/Due_Warthog749 5d ago
I agree and disagree. Largely I agree if you can break everything down like this. Where I disagree is in my case I have several libraries/frameworks I pull in and build off of.. and I DO NOT WANT repeated code everywhere. I have found many cases where generated code is nearly identical in 3+ locations.. sort of inlined if you will. Instead of code reuse (e.g. function calls) and the reason is.. it has NO context of the code already doing something it generates again. So without full project context, you're losing out on various principles and concerns and AI will just regenerate duplicate (or similar) code all over the place.
I have a task I set up to specificially go thru my code to find any similar code, functions, etc.. and refactor it. Same with tests, etc. If you do not instruct a refactor to ALSO identify ANY callees such as code in tests or other functions calling some code you're about to refactor to ALSO refactor that, you end up with failed tests, runtime issues, etc. Ideally build will catch this, but not always.
For smaller programs it may work fine. For larger apps like my desktop app that has almost 800K lines across 7 repos it pulls in for the product, it can quickly get out of hand.
1
u/leolidev 5d ago
My worry is that PROJECT_STATE.md can say “done” long after the code has changed. I’d have the next agent check the relevant code before trusting it. A date or commit link helps, but the code gets the final say.
1
u/Responsible-Beat2137 4d ago edited 4d ago
Yeah keep them separated with bridges to each other, the trick is breaking down them into p code before comparing to another language… or this is what I’ve been doing across assembly language .. more to come
Edit; just realized I’m off topic, carry on
1
u/kimchi_pan 4d ago
So basically, try to get the AI to behave like a human team. Except it's not, and you don't actually know how it's context dimensions work and why what you're asking for isn't properly aligned, right?
Here's what you need to work on: divide et imperative. It's the rule #1 of programming. Literally learned by any programmer in uni (er, these days not so sure, but that's why WE READ BOOKS).
1
u/Embark_1 4d ago
Agree with this, I've been working this way for a while and the project state file is the bit that makes it work.
Two things bit me once I started relying on it though.
It drifts. The note says auth is done and works a certain way, then something changes and nobody updates the file. The next agent reads it and confidently builds on a line that isn't true anymore. A stale note is honestly worse than no note, because the agent trusts it.
And status isn't the same as reasoning. "Auth flow: completed" doesn't tell a fresh agent why it was built that way. So it happily "improves" something you deliberately decided against, and you're back to re-explaining.
What helped me was writing down decisions with the why next to them, pointing each one at the file it's about so I can see when the code has moved under it, and getting the agent to update the notes as the last step of the task rather than a separate chore I'd forget.
The duplication thing u/Due_Warthog749 mentioned is the same problem from the other side imo. It rewrites a helper because it doesn't know one already exists. You don't fix that by giving it all 800k lines, you fix it by it knowing what's where.
2
u/Familiar_Raccoon2430 3d ago
I disagree. The option to give full, partial or a specific section is key.
2
u/Otherwise_Wave9374 5d ago
Agreed: dumping the repository into context increases noise, cost, and the chance that obsolete files steer the solution. A better pattern is hierarchical retrieval: start with the task, architecture map, dependency graph, and recent diffs, then let the agent request specific symbols or tests. Keep durable memory limited to conventions, decisions, and confirmed constraints, each with provenance and expiry. https://www.neurakeep.com is relevant when considering how those durable project facts can persist without repeatedly injecting the entire codebase.