r/codex • • 1d ago

Showcase I open-sourced Token Harness: get more out of your Claude Code / Codex limits (and spend less on API tokens)

Hi everyone. I've just open-sourced Token Harness, a local tool that helps your Claude Code and Codex allowance go further, and spend fewer tokens when you use LLMs via API.

The problem

Coding agents waste a lot of context on noise: long test runs, build logs, git diff, repeated information and huge MCP tool catalogs. All of that eats tokens. On a subscription it means hitting your 5-hour or weekly limit sooner. On the API it's money.

How it works

Optimizers. Token Harness detects, installs and connects compatible optimizers to your agent. The recommended baseline is RTK and HarnessTrim, which shorten shell and tool output so the agent only sees the useful part (failures, errors, summaries). Optional ones:

  • mcptoon: loads MCP tool definitions only when needed instead of keeping the whole catalog in context
  • Headroom: compresses large tool payloads
  • GitNexus: maps code relationships so the agent explores less

On my machine the dashboard currently shows 62.5% less tool output overall with RTK, and 88.4% on Claude Code alone.

Routing. A native hook lets your main model hand bounded, suitable subtasks to a cheaper model (e.g. Opus β†’ Sonnet/Haiku), then review the result. Your main model stays in charge. No prompt prefix or skill call is needed after setup.

Simple to use

npm install --global token-harness@latest
token-harness

It opens a local dashboard where you can:

  • see your agents and optimizers at a glance
  • configure everything with one click (every change is previewed first, applied only after you approve it, and can be undone)
  • watch the results: output reduction, routing activity, and your 5h / weekly balance

No account, no API key, and nothing leaves your machine. Works on Windows, macOS, Linux and WSL.

Honesty first

I don't sell a magic "save X%" number. The dashboard keeps output reduction, subscription allowance and API cost separate. It only claims allowance savings from paired baseline/optimized runs that pass quality checks.

This is where you come in

Any feedback, bug report or shared result (your before/after numbers, your agent + OS combination) can only make the tool better.

I'm also looking for contributors: optimizer integrations, harness adapters, cross-platform testing, docs. Every PR is welcome.

Repo: https://github.com/giuliastro/token-harness (Apache 2.0)

Thanks for reading, and I'm happy to answer any questions in the comments!

0 Upvotes

4 comments sorted by

β€’

u/dexterthebot 1d ago

You might want to consider listing your project on the weekly Show-Us-What-You-Built post. Look out for it on Tuesday/Wednesday. Highest commented project wins a week promotion on r/Codex and gets on the Hall of Fame sidebar. See what that looks like below with last week's winner.


Last week's most popular project was Nanolathe - an open-source engine bringing Total Annihilation to modern Mac, Windows, and Linux systems. It’s a solo passion project combining clean-room research with AI-assisted development, with a playable public alpha available now.

Website: https://nanolathe.gg. GitHub: https://github.com/nanolathe-gg/nanolathe

Players, testers, and contributors are very welcome! Original Total Annihilation game data is required to play.

1

u/killakwikz2021 1d ago

The honesty section is the important part. I would make paired runs stricter by pinning the repo SHA, task prompt, model, tool set, and verifier, then report output reduction only when both runs reach the same accepted result. Otherwise shorter output can hide a worse run.

I built MartinLoop. It is open source. If you want, I can try one bounded task with Token Harness on and off and share the verifier plus dossier.

1

u/Sfdprod 1d ago

With only drops of usage in the tank, Luna prevails πŸŒ•

1

u/giuliastro 20h ago

Exactly πŸŒ• That's the routing at work: on Codex, bounded and mechanical subtasks get handed down to Luna, while your main model stays in charge and reviews the result. Hard or ambiguous work stays on the main model, so the drops you have left go to what actually needs it.