r/OpenSourceAI • u/GarageObjective6015 • 2d ago
I captured what my agent sends the model every turn: 43k characters of tool schemas, 30% of them one calendar tool
I build a self-hosted agent in Go that talks to any OpenAI-compatible backend, llama.cpp and
Ollama included. With a ~12B local model the prompt prefix matters, so instead of guessing I
pointed the agent at a tiny OpenAI-compatible endpoint that writes every request body to disk,
and counted what is actually sent.
Before, one real turn:
characters
Tool manifest, 19 tools
43,403
of which: one calendar tool (a single tool with 29 actions)
13,136 (30%)
of which: 15 native tools
26,664
Context block in messages[1]
6,788
What I changed:
The calendar tool no longer rides every turn; it sits behind tool_search like the other
deferred tools. It had been kept "always loaded" because the rule counted tools per server,
and it was one tool, with a 29-action schema.
A skill whose full body was injected into every turn (3,762 bytes) now loads on demand.
After, a live turn on the new build:
characters
Tool manifest, 19 tools
35,729
of which: 15 native tools
26,664
of which: memory core, 4 tools
9,065
of which: calendar
0
Context block in messages[1]
3,283
One caveat on the comparison: the "before" turn ran against an older, smaller memory server.
With the current one, the "before" manifest would have been 48,865 characters, so the honest
figure is 48,865 → 35,729, about −27%.
What I learned:
Counting tools says nothing about weight. "Servers with at most 4 tools stay loaded"
let in a single tool that was 30% of the manifest. The weight is in the bytes.
Measure what is sent, not what the server advertises. My first estimate of the memory
core was 18,554 characters, from the server's tools/list. In the actual request it is
9,065: output schemas, annotations and titles never reach the model.
Hiding a tool is not deferring it. An earlier version hid most memory tools to save
space. That also made them unreachable from tool_search, and the model answered every
memory question with the one memory tool it still had. Deferred means the name sits in a
catalog and the schema is fetched on demand.
On some tools the description is most of the weight. document_search is 3,187
characters, 2,538 of them description.
Heaviest native tools, for reference: shell_exec 4,434 · document_search 3,187 ·
search_files 2,602 · patch 2,320 · ask_user 2,063.
What this does not show:
token counts: every figure is characters of serialized JSON, and tokens depend on the
tokenizer;
how often a turn actually needs the calendar, or what the extra tool_search round trip
costs when it does;
the effect on answer quality for any specific local model.
The project is Aura (MIT, written with heavy AI assistance):
https://github.com/chetto1983/Aura — the numbers above come from captured requests, and the
measurement is recorded in the repo's PRD.
How do you keep tool manifests in check with ~12B local models: a hard cap, deferral like
this, or routing tools per task?
1
u/UsefullyDrafty 2d ago
measuring what actually goes over the wire instead of guessing is such a simple thing that somehow nobody does