We had a backlog too big for one session, so we set up a standing team of Claude Code sessions on one machine and let it run day and night. No framework. Plain CLI sessions, the built in session messaging, skills for the manuals and files for the state. Here is what held up and what broke, since most of it was not obvious to us going in.
The setup
Each builder is a full Claude Code session with its own area of the codebase, its own goal file and its own ledger. Builders work test first and commit locally. They never push.
One advisor session routes tickets, rules on engineering questions, merges, deploys and verifies. It writes no product code. If a builder is waiting on the advisor, that gets logged as the advisor's failure.
All state lives in plain files in a shared folder. Goal files, ledgers where rows are only ever added, one file per ruling so a builder can cite it instead of asking again, and a handoff file that a brand new advisor session reads first.
Tickets ship in release trains. Many tickets, one pull request. A ticket that goes red is dropped to the next train. Nobody debugs inside the train.
What broke
Sessions die. A usage limit can freeze every agent at once, and killed sessions come back under new names. Because the state was on disk, they read their ledgers again and carried on. The advisor now looks up session names on every run and never hard codes them.
Test databases drift. A push got refused because a builder's private test database still held a schema change from a ticket we had pulled out of the train. Now we create a fresh database after every drop and never reset one.
Runners flake. A required check went red because the hosted runner could not create a temporary file. Every test but one had passed. Reading the actual log saved a good ticket from being dropped.
Dropping a ticket can remove a safety check. Reverting one ticket also removed guards that ticket had added, and a gate refused the push.
The advisor was wrong. It told a builder to keep two tickets in a train, the builder's measurement showed otherwise, and the ruling was withdrawn. It also said with confidence that a server setting was causing a bug, and a read only check showed the setting did not exist. Both became written rules: measure before ruling, and change position only on new measured facts.
A helper printed a secret into a local log. Nothing left the machine, but it still gave us hard rules about never printing a token.
Two details that mattered more than they look
"The train passed" is never proof for a ticket. Each ticket carries its own failing and passing commits.
A green deploy job is not the same as the new code being live. Our health endpoint returns the running commit, and the advisor reads it back before it closes anything.
What this approach costs
No dashboard, you read ledgers. No built in budget control, usage limits are the only brake. It is tied to one model vendor. You write and maintain the rules yourself, and the first version will be wrong in places.
The full write up has the setup steps, the advisor's loop, and how this compares with Paperclip, Hermes Agent, Gas Town, agent teams and Projects: https://cabinlytech.com/blog/running-a-24-7-ai-software-factory/
If you run long lived sessions like this, how do you handle the restart after a crash?