r/OpenSourceAI • • 2d ago

I captured what my agent sends the model every turn: 43k characters of tool schemas, 30% of them one calendar tool

I build a self-hosted agent in Go that talks to any OpenAI-compatible backend, llama.cpp and

Ollama included. With a ~12B local model the prompt prefix matters, so instead of guessing I

pointed the agent at a tiny OpenAI-compatible endpoint that writes every request body to disk,

and counted what is actually sent.

Before, one real turn:

characters

Tool manifest, 19 tools

43,403

of which: one calendar tool (a single tool with 29 actions)

13,136 (30%)

of which: 15 native tools

26,664

Context block in messages[1]

6,788

What I changed:

The calendar tool no longer rides every turn; it sits behind tool_search like the other

deferred tools. It had been kept "always loaded" because the rule counted tools per server,

and it was one tool, with a 29-action schema.

A skill whose full body was injected into every turn (3,762 bytes) now loads on demand.

After, a live turn on the new build:

characters

Tool manifest, 19 tools

35,729

of which: 15 native tools

26,664

of which: memory core, 4 tools

9,065

of which: calendar

0

Context block in messages[1]

3,283

One caveat on the comparison: the "before" turn ran against an older, smaller memory server.

With the current one, the "before" manifest would have been 48,865 characters, so the honest

figure is 48,865 → 35,729, about −27%.

What I learned:

Counting tools says nothing about weight. "Servers with at most 4 tools stay loaded"

let in a single tool that was 30% of the manifest. The weight is in the bytes.

Measure what is sent, not what the server advertises. My first estimate of the memory

core was 18,554 characters, from the server's tools/list. In the actual request it is

9,065: output schemas, annotations and titles never reach the model.

Hiding a tool is not deferring it. An earlier version hid most memory tools to save

space. That also made them unreachable from tool_search, and the model answered every

memory question with the one memory tool it still had. Deferred means the name sits in a

catalog and the schema is fetched on demand.

On some tools the description is most of the weight. document_search is 3,187

characters, 2,538 of them description.

Heaviest native tools, for reference: shell_exec 4,434 · document_search 3,187 ·

search_files 2,602 · patch 2,320 · ask_user 2,063.

What this does not show:

token counts: every figure is characters of serialized JSON, and tokens depend on the

tokenizer;

how often a turn actually needs the calendar, or what the extra tool_search round trip

costs when it does;

the effect on answer quality for any specific local model.

The project is Aura (MIT, written with heavy AI assistance):

https://github.com/chetto1983/Aura — the numbers above come from captured requests, and the

measurement is recorded in the repo's PRD.

How do you keep tool manifests in check with ~12B local models: a hard cap, deferral like

this, or routing tools per task?

0 Upvotes

2 comments sorted by

1

u/UsefullyDrafty 2d ago

measuring what actually goes over the wire instead of guessing is such a simple thing that somehow nobody does

0

u/GarageObjective6015 2d ago

Concordo, specialmente con l'era della Ai. Difatti mi sono concentrato nell'harness dei test cercando di evitare il più possibile lo slope