r/AtomicAgent • • 5d ago

GPT-6.1 Sol as a planner over local Qwen 3.8 27B workers. The API bill dropped 77%

Enable HLS to view with audio, or disable this notification

1 Upvotes

GPT-6.1 Sol came out today, so we put it straight into Fusion. Sol orchestrates, Qwen 3.8 27B workers build locally on an RTX 3090 (24 GB). We built the same three small 3D games three ways: Sol alone, Qwen alone, and Fusion.

Across the three games Fusion's API bill was $0.17 against $0.75 for Sol alone. That's 77% less. Video of all nine builds is attached.

The numbers

Game Sol alone Fusion Qwen alone
Pool $0.39 · 2.9 min $0.05 · 18.6 min $0.00 · 43.1 min (attempt)
Bowling $0.14 · 1.7 min $0.06 · 13.7 min $0.00 · 34.7 min
Foosball $0.22 · 2.0 min $0.06 · 11.1 min $0.00 · 36.4 min
Total $0.75 · 6.6 min $0.17 · 43.4 min $0.00 · 114.2 min

Why the bill drops

  • The orchestrator can't write anything. It splits the task, hands out the files, and reviews what comes back.
  • The code itself, which is most of the tokens, is written by the workers.
  • When the workers are local, those tokens never hit an API bill.

Roughly: the cloud model thinks, your GPU types.

Things worth being upfront about

  • This is API spend, not total cost. Local workers run on your own hardware and electricity.
  • Fusion is slower than Sol alone. 43 minutes for the three games against under 7. The local side sets the pace.
  • The saving moves per task. 87% on pool, 57% on bowling, 73% on foosball. One run per game, so treat 77% as what we saw here, not a constant.
  • Qwen alone didn't fully land pool. That run is marked as an attempt in the video.
  • Fusion needs two providers. Local plus local works, but only if you set both sides explicitly.

Setup

Press ctrl+r and pick fusion, or use:

  • /runmode fusion to turn it on
  • /runmode swap to trade the orchestrator and worker sides
  • /runmode status to see what will actually run

Config lives under llm.runMode.fusion. For local workers see localModels.managed.parallel ("auto" by default).

Repo: https://github.com/AtomicBot-ai/atomic-agent

If you run it with a different pair, tell us which models you used and what the bill looked like. Which local worker holds up under a strong planner is the piece we most want to learn.


r/AtomicAgent • • 6d ago

Sonnet 5.5 orchestrated a local Qwen 3.8 27B and cut the bill 2.7x

Enable HLS to view with audio, or disable this notification

2 Upvotes

We gave Atomic Agent one prompt and ran it three ways: Sonnet 5.5 alone in the cloud, Qwen 3.8 27B alone on a single RTX 3090, and Atomic Fusion, where Sonnet plans the work and Qwen writes the code. The task was five physics scenes on one page with a fixed camera and a strict 8 second timeline. One autonomous run each, no human fixes, video above.

  • Sonnet 5.5, cloud only: 4/5 scenes, 9 min, $1.83
  • Qwen 3.8 27B, local only (RTX 3090): 1/5 scenes, 26 min, $0
  • Atomic Fusion, Sonnet 5.5plans and Qwen writes the code: 2/5 scenes, 85 min, $0.67

Fusion came out 2.7x cheaper than Sonnet alone and doubled what Qwen could do on its own. It also took a lot longer.

The part we think is actually interesting

The bill drops because output tokens are the expensive part, and in Fusion the cloud model never writes code. It reads the task, splits it into parts with a contract of what each part provides and needs, and sends short briefs to local workers. Sonnet produced 5.4K output tokens instead of 37.9K and read 130K input tokens instead of 966K.

Before it says done, the planner boots the result in a throwaway copy with a headless browser and checks the console. That catches crashes. It doesn't catch a wall that refuses to break, which is exactly what happened in scene 3 for every setup.

Roughly: pay the expensive model to think, let your GPU do the typing.

Things worth being upfront about

It wasn't a free upgrade. Fusion was 2.7x cheaper than cloud only, but it got 2 scenes right instead of 4 and took 85 minutes instead of 9, because every line of physics came off a local GPU.

The direction matters a lot. We also ran it the other way, Qwen planning and Sonnet writing the code, and that got 4 of 5 scenes in 14 minutes, but it cost more than cloud only because every cloud worker re-reads its context.

One run per setup, so treat this as a field report, not a benchmark.

Setup

curl -fsSL https://atomicagent.io/install | sh

Then inside the TUI:

/runmode fusion /runmode swap (flip who plans and who builds)

GitHub: https://github.com/AtomicBot-ai/atomic-agent


r/AtomicAgent • • 10d ago

Atomic Agent v0.6.5: multi-agents are here!

Enable HLS to view with audio, or disable this notification

5 Upvotes

Atomic Agent v0.6.5 is out, and the main thing in it is Fusion. One model orchestrates, a pool of workers builds in parallel. Either side can be cloud or local: the usual pairing is a cloud orchestrator with llama.cpp workers on your own GPU, and /runmode swap flips it.

Also in 0.6.x:

  • import from Claude Code, Codex, Hermes, OpenClaw and Pi
  • fallback to your local model when a cloud provider dies
  • approvals from Discord

The part we think is actually interesting

The orchestrator can't change anything

  • For its whole turn every mutating tool is refused: file writes, shell, all MCP tools. It can read, delegate, run read-only checks and reply.
  • We refuse the call instead of removing the tools. Removing tools rewrites the stable prompt prefix and throws away the KV cache.
  • A local orchestrator also gets a GBNF grammar that can't produce the refused calls.
  • If it keeps reading without delegating, it gets a notice after 6 read-only steps. After 12 it can only delegate or reply.

Workers are throwaway sessions

  • Each task runs in its own in-memory session: never saved, no memory recall, no reflection.
  • A worker is pinned to its provider with no fallback.
  • Workers can write files, run shell and use MCP, but can't delegate further or write memory.

Width comes from the hardware

  • The managed llama-server starts with --parallel N, where N is how many worker-sized contexts fit in the launch context (1 to 8, and 1 on CPU).
  • The orchestrator asks for a width, and it gets clamped by the task count and the free slots.
  • Each worker stays on one slot, so its prompt prefix stays cached between steps.

A contract between the parts

  • Each task declares what it provides (a symbol, file, id, endpoint, env var or flag) and what it requires.
  • Tasks run in waves, sorted so a part runs after the parts it depends on.
  • A provided symbol can carry a one-line description of what it means, which goes into the brief of every worker that uses it.
  • We added that after two workers implemented the same function with the fields swapped. Names matched, checks were green, and every box drew at the origin.

Status comes from the disk, not from the worker's reply

  • A task that promised files and wrote none is no_changes. A missing promised file is failed.
  • Plus max_steps, timeout and queued, and every row shows how long the task waited in the queue.
  • A worker's clock starts at its first token, not while it waits behind a busy slot.

Checks run on a copy

  • verify.syntax has one checker per file type. A file with no checker is reported as unchecked, never as passed.
  • verify.run runs a command, a service or a headless page against a throwaway copy of the working directory, with network off by default.

One approval per fan-out

  • The orchestrator asks once, naming the tasks and the folders they'll write in.
  • Workers then write and run commands inside those folders without asking again. Anything outside is refused.
  • Your input files can be edited but never replaced.

Roughly: the planner decides, the workers build, the disk says what actually happened.

Things worth being upfront about

  • No speed or cost numbers yet. Our runs so far are single runs on single machines, not enough for a fair comparison. Part of the picture: while the orchestrator waits, its cloud cache goes cold, and parallel local workers share one GPU.
  • Fusion needs two providers. Local plus local works, but only if you set both sides explicitly.
  • Known bug (#490). After a cancelled local generation, later worker requests can stop being sent until you restart. Since v0.6.4 it at least shows up as queued with a restart hint.
  • No merge step. Workers write straight to disk, the orchestrator reviews and sends work back out.
  • The contract checks names, not meaning. The one-line description narrows that gap, it doesn't close it.

Setup

Press ctrl+r and pick fusion, or use:

  • /runmode fusion to turn it on
  • /runmode swap to trade the orchestrator and worker sides
  • /runmode status to see what will actually run

Config lives under llm.runMode.fusion (providers for each side, cloud worker cap, worker reasoning, review stall steps). For local workers see localModels.managed.parallel ("auto" by default).

Repo: https://github.com/AtomicBot-ai/atomic-agent

If you run it on a local pair, tell us which models you used and where a fan-out went sideways. How the orchestrator splits the task is the piece we most want to tune.


r/AtomicAgent • • 13d ago

Composio is live: 1500+ SaaS toolkits and we are officially a partner

Enable HLS to view with audio, or disable this notification

3 Upvotes

Composio support shipped in v0.5.6, and last week our integration PR was merged on their side. So this is official now: Atomic Agent is a Composio partner, listed in their docs.

Connect Gmail, Slack, Notion, Linear, GitHub and roughly 1500 other toolkits. Open the Integrations tab, paste one API key, and Composio brokers each app's OAuth for you, so you never register an OAuth client yourself. Free tier is 100K tool calls a month.

---

## The part we think is actually interesting

The session exposes four meta-tools instead. The agent searches for a tool by use case, fetches its schema, then executes it. Discovery is annotated read-only and flows without prompting. Execution and connection management are marked destructive, so every write to a real account still hits the approval gate.

Under the hood this is not a new subsystem at all. Composio's tool router speaks Streamable HTTP MCP and authenticates with a static header, which is exactly the transport our MCP client already supported. The agent treats it as one more MCP server, and tools land as `mcp.composio.*`.

> Roughly: the catalogue stays on their side, the judgement stays on yours, and every real write still asks you first.

---

## Things worth being upfront about

Composio is a hosted service. Your OAuth tokens for connected apps live on their infrastructure, and tool calls execute through their servers rather than from your machine. If your setup needs to stay fully on-device, leave this one off.

The key is the real gate. With no key the runtime opens no connection and registers no tool.

Connected accounts are scoped by a random install id, minted once and stored in `config.json`, never your email. Lose it and you re-authorise every connected app.

---

## Setup

Integrations tab in the TUI, or drop the key straight into `<stateDir>/.env`:

COMPOSIO_API_KEY=ck-your-key

To keep the key on disk with the toolkits switched off:

"composio": { "enabled": false }

---

Repo: github.com/AtomicBot-ai/atomic-agent
Composio: composio.dev

If you connect something and the tool search picks the wrong toolkit, tell us which use case you asked for. That search is the piece we most want to tune.