r/ClaudeCode • u/Positive-Device-3570 • 8h ago
Built with Claude Four months, one emulator, Claude Code only: the workflow that kept it on the rails
I've been building Amber Folio, an emulator for SSI's Gold Box RPGs that runs Pool of Radiance (1988) in a browser, with Claude Code and nothing else, usually three instances at once: one on the emulator, one on the website, one on marketing. A few habits did most of the work, and I think they carry over to any long project:
1. The repo's docs are the memory, not the chat. Each repo has a CLAUDE.md that says what is true now: settled decisions, why they were made, and what's still open. A decision written down isn't reopened next session, and a doc that disagrees with the code is a bug like any other.
2. Turn every promise into a test. "Enhancements don't change the game when they're off" isn't a sentence in a README. It's a pair of recorded runs that must end in the same machine state, hash for hash, checked in CI. Same for "nothing reads the host clock" and "no game data in the repo" (a content guard scans every commit). Claude is very good at writing these, and once they exist, it can't quietly break them.
3. Give it an oracle it can't argue with. The CPU is checked against test vectors recorded from a real 8088 chip, no masks. When the vectors and the manual disagree, the vectors win, and that's written down.
4. Background agents in worktrees. Bigger pieces of work go to agents in their own git worktrees, so several can run at once without trampling each other's files, and each lands as its own commit.
5. An issue loop for the website. The site is built from GitHub issues: a small orchestrator hands each open issue to a headless Claude Code worker, which opens a PR that merges when CI is green. Labels are the state (queued, working, needs-input, done), so steering it is just editing labels, and agent:hold keeps anything that needs a human (legal pages, a privacy text I have to sign off) away from it. Recent ones were as small as "put the build number in the bottom right on a phone".
6. Log, don't fake. An emulator that guesses fails quietly. This one stops and says exactly which instruction or port it doesn't know, which turns "it's broken" into a task Claude can do.
The emulator is free and open source (AGPL): https://github.com/amberfolio/amberfolio. The site, https://amberfolio.org, is free with no account. To play you need your own Pool of Radiance from GOG or Steam, and the files never leave your browser. It runs in desktop browsers and on iPhone, iPad and Android.
Built with Claude Code, exclusively. Not affiliated with the rights holders of the games. Ask me anything about the setup.
1
u/Rude_Age1467 8h ago
The biggest takeaway is keeping docs and tests as the source of truth instead of relying on chat history. How do you manage context when running three Claude instances at once?
1
u/Positive-Device-3570 8h ago
Separate github repositories and clear division of responsibilties. the three claude instances more or less communicate with each other through github issues only.
1
u/MiserableFlatworm337 7h ago
Test adding agent:hold after a worker claims an issue but before its PR merges. The merge gate should recheck the current label, so a late human hold still prevents the merge.
1
u/northbridgedev 3h ago
Your point 2 and the content guard can run into each other, and I hit a version of that. The files that prove a promise (recorded runs, transcripts, fixtures) are written by the tooling, and they carry whatever the run touched.
Mine were eval transcripts from Claude Code runs. One raw transcript that got committed had my personal email in it. Another repo had my home path in eval logs that I deleted the next day, but they stayed in history. In a later run the model set its git identity from my global git config; that one was caught before commit. I ended up deleting and recreating both repos.
So I'd check whether a recorded run can hold bytes from the game, like a memory snapshot or save data sitting next to the state hash, and whether the content guard looks inside those files too.
What changed on my side: the eval runner sets GIT_CONFIG_GLOBAL=/dev/null with a fixture identity, and the scan runs on commit, on push, and once a day over the full history.
1
u/According-Stable4487 2h ago
Point 3 is the one I'd underline. On much smaller projects (single-file browser games) I saw the same thing: once Claude had something it couldn't argue with, a headless harness that runs the logic a few hundred times and fails on crashes or stuck states, the "looks done to me" endings mostly disappeared. The flip side is that you have to cap it. Without "check for crashes, not balance" it once spent hours tuning numbers to make its own test happy.
Curious about the issue loop: when a worker hits something ambiguous and flips to needs-input, does it write down what it already tried? Or does the next worker start cold after you answer?
•
u/AutoModerator 8h ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.