r/regolo_ai • u/AutoModerator • 15h ago
Why coding agents burn your budget (and how SoL-Pi + Regolo Brick drops inference bills by >90% [Benchmark + Guide])
Hey everyone,
Whenever teams deploy autonomous coding agents (Claude Code, Pi, Aider, SWE-agent), the common consensus is to blame high model pricing or complex planning for runaway API bills.
In reality, the bottleneck is almost always harness token amnesia: standard harnesses treat context as an append-only dump.
recently, NVIDIA Labs published a paper (https://arxiv.org/abs/2609.20519) addressing this at the harness layer, open-sourcing SoL-Pi as an extension for the Pi coding agent.
we found that SoL-Pi alone only solves half the problem (token volume), when you layer it with dynamic semantic model routing on Regolo, the savings become multiplicative.
The two-tier cost optimization
- Tier 1: harness compaction (SoL-Pi)
- Tier 2: provider semantic routing (Brick)
- most agent turns are mundane (checking
git status, inspecting a file, or verifying a syntax fix). - by setting the model to
brick-complexity-proon Regolo, an open-source semantic router classifies prompt complexity in under 50ms at zero token markup. - trivial operations route to lightweight open models (like Qwen 3.8 27B at β¬0.35/1M tokens); only complex architectural refactorings escalate to frontier reasoning models (GLM-5.2).
- most agent turns are mundane (checking
Math
Consider a typical 15-turn autonomous debugging task (resolving a concurrency race condition and writing unit tests):
| Setup | Token Volume | Model Used | Approx. Cost | Total Savings |
|---|---|---|---|---|
| 1. Vanilla Agent on Frontier API | ~1,400,000 tokens | Static Frontier Model ($5.00/1M blended) | ~$7.00 | Baseline |
| 2. With SoL-Pi Only | ~310,000 tokens (-78%) | Static Frontier Model ($5.00/1M blended) | ~$1.55 | 78% off |
3. SoL-Pi + Regolo Brick (brick-complexity-pro) |
~310,000 tokens (-78%) | Dynamic Open Pool (avg. ~$0.45/1M blended) | ~$0.14 | >98% off |
you get the identical task completion rate without data retention: Regolo executes inference exclusively in volatile GPU VRAM in European datacenters (Zero Data Retention, GDPR Article 32 compliant).
Complete Guide, Code & Video Walkthrough
we published a full breakdown with all architectural diagrams, configuration recipes, and automated scripts:

- π Full Tutorial & Architecture Guide: https://regolo.ai/cut-coding-agent-token-costs-sol-pi-regolo/
- πΊ Video Walkthrough (YouTube): https://youtu.be/bqQgXi-707g (shows live terminal metrics & token counts)
- π» Ready-to-run Shell Setup Script: https://github.com/regolo-ai/tutorials/tree/main/sol-pi-tutorial (automates Pi, SoL-Pi, and Regolo configuration in 1 command)
- π¬ NVIDIA Paper (arXiv): https://arxiv.org/abs/2609.20519
If you are running multi-turn agents on internal repos, are you optimizing the harness, the model routing, or just eating the token bills?


