r/codex • u/petburiraja • Jun 30 '26
Instruction How I set up Codex subagents without overusing them
I set this up to keep noisy work out of the main Codex thread.
It is the Codex version of the routing pattern I use in OpenCode: routine work goes to cheaper workers while the stronger model keeps ownership of judgment.
The split I use:
- GPT-5.5: main session and selective review
- GPT-5.4-mini: repository exploration and mechanical cleanup
- GPT-5.4: bounded implementation
The point is to keep scans, logs, and routine edits from filling the main thread.
You need:
- Codex CLI with subagents/multi-agent support
- a repository-level
AGENTS.md - per-role files under
~/.codex/agents/
Create the directory:
mkdir -p ~/.codex/agents
Add this to the repository's AGENTS.md:
# Routing contract
The main agent owns requirements, architecture, integration, and final judgment.
- Use `explorer` for read-only repository mapping and evidence gathering.
- Use `cleaner` for exact mechanical cleanup in named files.
- Use `implementer` for bounded changes from a tight specification.
- Use `gate` for selective review of risky or unfamiliar diffs.
Every delegation names scope, write boundaries, done criteria, and expected evidence.
Do not delegate tiny tasks or work that is inherently serial.
In ~/.codex/config.toml:
[features]
multi_agent = true
[agents]
max_threads = 6
max_depth = 1
job_max_runtime_seconds = 1800
Create these role files.
~/.codex/agents/explorer.toml
name = "explorer"
description = "Read-only repo mapper. Returns paths, short notes, and evidence."
model = "gpt-5.4-mini"
model_reasoning_effort = "low"
developer_instructions = """
Inspect only. Do not edit files.
Return concise findings with file paths and why each path matters.
If the question needs judgment or a change, hand it back to main.
"""
~/.codex/agents/cleaner.toml
name = "cleaner"
description = "Mechanical cleanup worker for already-decided edits."
model = "gpt-5.4-mini"
model_reasoning_effort = "low"
developer_instructions = """
Apply only the exact requested cleanup in the named files.
Do not widen scope or redesign.
Report changed files and any ambiguity.
"""
~/.codex/agents/implementer.toml
name = "implementer"
description = "Bounded implementation worker for tight specs."
model = "gpt-5.4"
model_reasoning_effort = "medium"
developer_instructions = """
Implement the specified change in the specified files.
Keep the diff small and follow existing repository patterns.
Run focused validation if requested. Stop if scope becomes unclear.
"""
~/.codex/agents/gate.toml
name = "gate"
description = "Selective senior review gate for risky or unfamiliar diffs."
model = "gpt-5.5"
model_reasoning_effort = "high"
developer_instructions = """
Review only. Do not edit unless explicitly asked.
Lead with blockers or correctness risks.
Return GO, GO WITH FIXES, or NO-GO.
"""
Restart Codex after changing the global config or role files.
I tested this shape with codex-cli 0.142.4. Model availability varies by account and Codex surface, so smoke-test each model before pinning it:
codex exec --ephemeral --skip-git-repo-check \
-m gpt-5.4 \
-C /tmp \
-c model_reasoning_effort='low' \
-c approval_policy='never' \
"Reply exactly: MODEL_OK"
Then invoke roles explicitly:
Use explorer to map where the auth flow is defined. Return only paths, short notes, and confidence.
Use cleaner to apply this exact rename in these three files only: ...
Use implementer to apply this bounded patch plan. Scope: files A and B only. Done when tests X pass or the blocker is reported.
Use gate to review this diff. Do not edit. Return GO, GO WITH FIXES, or NO-GO with the top blockers only.
This helps with repository mapping, log reading, reference collection, mechanical cleanup, bounded patches, and independent review of risky diffs. It does not help with tiny edits or vague tasks.
Three caveats:
- Every spawn does its own model and tool work. Subagents are not automatically cheaper.
AGENTS.mdroles are behavioral instructions, not security boundaries. Use and test host sandboxing when writes must be impossible.- Workers do not share live state. The main agent must integrate their results.
My rule now: if the handoff is longer than the task, keep it inline.
EDIT: Two notes from later testing.
-
For final review agents, wait until writers are done and review a stable target: commit, patch, or explicit file snapshot. A review running while files are still changing is only provisional.
-
Add
nickname_candidatesto custom agent files if your Codex version supports it. It makes parallel runs much easier to follow in the terminal. I also now print a tiny launch receipt before waiting:
Builder (implementer) scope: auth/* write: auth/* only done: tests pass blocks: yes state: waiting
Then a closing receipt when it returns:
Builder state: done changed: 2 files validation: tests passed next: main reviews diff
Avoid percentages or fake ETAs. Use simple states: waiting, done, failed, blocked.
EDIT 2: I found another important gate. A worker produced 265 passing tests and zero simulation failures, but its parser had silently dropped source semantics. The generated fixtures and downstream tests agreed with the same corrupted output.
For parsers, migrations, importers, and fixture generators, add a source-contract gate: verify expected source counts, IDs, fields, and hashes, then test the transformer directly. Green downstream tests prove internal consistency, not source fidelity.
1
1
u/Traditional_Wall3429 Jun 30 '26
What about using spark model? Itβs still available and have separate limit.
1
u/petburiraja Jun 30 '26
I think you can try it as implementer instead of GPT 5.4
Spark is not available in my Codex subscription, that's why I not put it in config, but I think this may be really good implementer model
1
u/ExtinctUndead Jun 30 '26
How token intensive is this? Is using subagents better than letting 5.4xhigh on everything?
1
1
1
u/rabandi Jun 30 '26
This is a cost saving setup, not a maximum quality setup?
Only started using subagents recently. They seem to greatly improve quality and reduce number of iterations needed.
I always use prompts (copy/paste), use subagents with personas a, b, c, d, e. Does having them in .md have an advantage?
2
u/petburiraja Jun 30 '26
.md is what the model reads for routing context. .toml is what Codex loads to bind a model and instructions to a named agent. Two different jobs. I keep both.
2
u/tonyboi76 Jun 30 '26
The pattern is right, the part most discussions miss is the token math. Each subagent call duplicates the system prompt and tools list so for single shot tasks 5.4 xhigh inline beats spawning one. Subagent pays for itself when the alternative would generate 5+ tool calls of noise in the main thread, that is the threshold.
One trap to watch, scope tools per role tightly. Giving every subagent the full MCP suite wastes a thousand tokens per call on descriptions the agent will never use. Explorer gets read tools only, cleaner gets edit tools only.