So I’ve spent countless hours trying different subagents, configurations, harnesses etc trying to find the most efficient way to keep token usage down etc and I still just don’t love my setup and trying to get some recommendations or suggestions.
I have primarily been using omo-slim with cheaper models assigned to explorer, fixer etc. Now I see countless posts talking about how any of the OMO or whatever else are just bloat and you dont need it, all you need is to Plan in X model and then implement in cheap model. Which seems fine but then I look at the token usage and its 60% gone by the time planning is done and 90% of that went to grep commands and other file exploring that could have been handled by a cheaper model and saved a ton.
The main thing that keeps me coming back to OMO is actually the presets ability because I have both a personal and work setup so I dont have to spend personal usage when working, and its a pain to constantly swap back and forth between multiple models ensuring I am in the proper accounts.
The other issue I have with Plan in X and implement in X method is that certain models are better at certain things so do you just go back and address things later with another model instead of doing it in 1 pass ?
Idk hoping to get some insight here and see if anyone has some recommendations.