Skip to content

Fusion run mode

Fusion is one of Atomic Agent’s three run modes, next to Local and Cloud. Instead of one model doing everything, Fusion gives the job to two: an orchestrator that plans, writes a brief for each part and reviews what comes back, and workers that execute those parts, several at once.

Either seat can be a local model or a cloud model. That makes Fusion a way to pair a model on your own machine with a cloud model and get results closer to what the cloud model reaches on its own, while choosing which side does the thinking and which side does the building.

How Fusion works

In Fusion, the orchestrator never builds anything itself. It reads your project, splits the job into self-contained parts and sends them out through a tool called fusion.delegate. Each worker takes one part, writes its files straight to disk and reports back. The orchestrator then checks the result and sends weak parts out again.

Orchestrator

Holds the conversation with you. Reads enough to choose an approach, breaks the job into parts, writes a brief for each, reviews the replies and runs checks. It cannot write files or run commands during a Fusion turn.

Workers

Each worker runs one part in a fresh, throwaway session. It sees its brief and your original request, not the chat history or memory. It cannot ask you questions, request approval or delegate further.

A few details worth knowing:

  • There is no merge step. Workers write their files directly into the working directory. The orchestrator reads what is on disk, verifies it and re-delegates anything that falls short.
  • Workers get your request, not just the brief. Each worker receives the orchestrator’s brief plus your original message (up to 16,000 characters), so a short brief does not leave it guessing.
  • The orchestrator is locked out of changes for the whole turn. Read-only tools, fusion.delegate, reply and finish are allowed. Anything that writes, runs a command, or comes from an MCP server is refused. This is by design: the workers build, the orchestrator judges.
  • Parts can depend on each other. A fusion.delegate call can carry a contract (owners, provides, requires, checks). Tasks that need something another task provides run in a later wave, and name mismatches are caught before workers start.

Choosing a direction

Fusion runs in two directions, and which side plans and which side builds changes the result. Cloud plans with local workers is the default pairing. Local plans with cloud workers suits jobs with a lot of output, such as large code files, because the heavy writing happens on the cloud side.

This is what /runmode fusion sets up by default: a cloud model as orchestrator and your local model as the workers.

Fits when the job splits into many small, self-contained parts and you want the bulk of the execution on your own hardware.

Watch out for jobs that need a lot of generated output per part. Every line of code a worker writes comes from your local model, and parallel local workers share one GPU. For large code tasks, the reverse direction is usually the better fit.

Two cloud models or two local models also work, as long as the two seats are two different providers. Pointing both seats at the same provider is refused.

Turning Fusion on

The quickest way is from the TUI: run /runmode fusion. You need two configured providers, usually one cloud provider with an API key and one local model that is already downloaded. Atomic Agent pins both seats, saves the mode to your config and starts the local model server if it is down.

From the TUI

/runmode open the "where it runs" switch (same as ctrl+r)
/runmode fusion switch to Fusion
/runmode swap trade the two seats (orchestrator becomes worker and back)
/runmode workers 3 set the default number of workers (1-8)
/runmode status show what the mode resolves to right now
/runmode cloud leave Fusion for Cloud
/runmode local leave Fusion for Local

The first time a session switches into Fusion, the chat shows a short introduction naming the two models it resolved. After that, /runmode status prints the orchestrator, the worker provider and how many workers can run at once.

From config.json

The mode lives under llm.runMode in <stateDir>/config.json (by default ~/.atomic-agent/config.json). These keys cannot be set with atomic-agent config set, so either use the TUI or edit the file by hand.

{
"llm": {
// "providers": [ ... ] keep your existing providers as they are
"activeTextProvider": "openrouter",
"runMode": {
"mode": "fusion",
"fusion": {
"orchestratorProvider": "openrouter",
"workerProvider": "local-llama",
"workers": 2
}
}
}
}

The provider ids must match entries in your llm.providers list. local-llama is the id Atomic Agent gives its managed local model.

All Fusion settings

Key (under llm.runMode.fusion)DefaultRangeWhat it does
orchestratorProviderthe active cloud provider, else the first cloud providera provider idWhich provider plans and reviews.
workerProviderthe first local provider that is not the orchestratora provider idWhich provider runs the workers.
workers2 (a local worker seat runs 1 at a time until you set this)1 to 8Default fan-out width when the orchestrator does not name one.
cloudWorkers41 to 32Hard cap on parallel workers when the workers are in the cloud.
workerMaxSteps601 to 1000Step budget for one worker.
workerTimeoutMs2,700,000 (45 min)1,000 to 86,400,000Time budget for one worker.
workerReasoningprovider defaultlow, medium, highReasoning effort sent with worker requests.
workerMaxOutputTokensmodel maximum1 to 1,000,000Output cap per worker step.
reviewStallSteps60 to 1000 (0 turns it off)How long the orchestrator may only read before it is pushed to delegate or reply.
orchestratorModel, workerModelunsetany labelDisplay labels only. They do not change which model serves.

How many workers run

The orchestrator decides the width of each fan-out itself. Your workers setting only fills in when it names no number. What really limits parallel workers is physical: the number of tasks, the local server’s request slots for a local worker seat, and the cloudWorkers cap for a cloud worker seat.

  • Local workers. Every worker needs a request slot on the local model server, and all slots share one GPU and one context pool. If the orchestrator names no width and you have not set workers yourself, local workers run one at a time. Extra tasks queue and the result tells the orchestrator how many slots there were.
  • Cloud workers. There are no server slots to size. A request above cloudWorkers is clamped and the result says so.
  • Changing the count. /runmode workers N saves the default width and also sets the local server’s slot count (localModels.managed.parallel) to the same number. A server that is already running keeps its old slot count until you restart it (Manage › LLM › Local, then s).

One fusion.delegate call holds at most 16 tasks, and each brief is capped at 32,000 characters.

Verification and re-delegation

The orchestrator reviews every part before it accepts the work. It has two read-only check tools: verify.syntax checks the syntax of the files a part produced, and verify.run runs a command, a local service or a web page in a throwaway copy of the working directory. Anything weak goes back out as a new, more specific brief.

What each check covers:

  • verify.syntax picks a checker per file type (JavaScript and JSON, TypeScript through the project’s tsc, Python, shell, HTML inline scripts, CSS brace balance). A file with no checker is reported as unchecked, never as passing.
  • verify.run has three kinds: command (a test runner, compiler or script), service (start it, wait for a port, send requests) and page (open a local HTML file or URL in a headless browser, collect uncaught errors, console errors and missing selectors). Nothing it writes reaches your real workspace.
  • Re-delegation. A task that comes back failed, cancelled, needs_orchestrator or simply not good enough goes out again in another fusion.delegate call that says what was wrong and what good looks like.
  • Stalled reviews. If the orchestrator only reads for 6 steps in a row (the reviewStallSteps default), it is told to delegate or reply. At twice that, it can only call fusion.delegate, reply or finish. The threshold is halved when your message reads as a repair request.
  • Stalled workers. A worker that has written no file by half its budget is handed back early with what it found, so the orchestrator can split the task and try again.

What you see while it runs

A Fusion turn shows the plan, one approval for the fan-out, and a live row per worker. Each row names the task, the model running it, the tool in use and the elapsed time, beside the orchestrator’s time estimate when it gave one. A summary line then reports how many tasks came back ok.

The approval. When approvals are on, Atomic Agent asks you once per fan-out, before any worker starts. The prompt lists the task titles and the directories the workers may write to and run commands in. Your answer covers the rest of the turn, so a review that re-delegates five times does not ask five times. A later fan-out that reaches a new directory asks again.

Anything a worker cannot do alone (because it needs a person, or a wider write scope than you approved) comes back to the orchestrator marked needs_orchestrator instead of stalling.

Limits and gotchas

  • Pin a real model on the cloud side. Routers such as openrouter/auto can pick small, cheap models that ignore Fusion’s instructions. Choose a specific model for the cloud provider (/model or /llm).
  • The local seat needs real hardware. A 27B-class local model is practical from a GPU with around 24 GB of memory. On smaller cards the local seat can be slow enough that a single run stretches into hours.
  • Your data is not confined to your machine. Whichever seat is in the cloud sees your request and the file contents it works with.
  • Both seats must be different providers. Two cloud models or two local models are fine, but not the same provider twice.
  • The orchestrator cannot call MCP tools during a Fusion turn. They are treated as tools that change things and are refused.

Troubleshooting

Most Fusion problems come from which provider is active, a missing second provider, or the local server’s slot count. Start with /runmode status: it prints what the mode resolves to, how many workers can run and, when the stored and the effective mode differ, why.

What you seeWhat it meansFix
stored fusion, effective cloud in /runmode statusThe active provider is no longer the orchestrator.Run /runmode fusion again.
”Fusion needs a cloud orchestrator. No cloud provider is configured. Staying on local.”No orchestrator could be resolved.Add a cloud provider (Manage › LLM › Cloud, or /llm), or pin both seats in config.
”Fusion needs two providers, one to orchestrate and one to run the workers.”Only one provider is configured.Add a second provider in Manage › LLM.
swap needs fusion/runmode swap only works while Fusion is on.Run /runmode fusion first.
config set failed: unknown key llm.runMode...These keys are not settable from the CLI.Use /runmode in the TUI or edit config.json.
A note that work was sent but only some workers ran at a timeThe local server has fewer request slots than tasks.Set /runmode workers N and restart the local server, or let the orchestrator send fewer, larger tasks.
The orchestrator’s writes are refusedThis is Fusion working as designed.Nothing to fix. The workers do the writing.
A worker “handed back early: no file written by half the budget”The worker read but did not write.Usually handled by the orchestrator re-splitting the task. If it repeats, give the task a clearer deliverable.
Unsure whether Fusion is really activeThe mode chip alone does not prove the orchestrator has the delegate tool.Run /tools and check that fusion.delegate is listed.

FAQ

Which direction should I start with? For code-heavy jobs, try a local orchestrator with cloud workers (/runmode swap after /runmode fusion). For jobs that split into many small parts, the default cloud orchestrator with local workers is the natural fit.

Does the cloud model see my files? The cloud seat sees what it is given and what it reads: in the default direction, your conversation and the files it reviews; with cloud workers, your request (up to 16,000 characters) plus the files each worker opens.

Can I use two cloud models, or two local ones? Yes. Any two different providers work, for example a careful cloud model directing a fast one, or a large local model directing a small one. With two local providers, pin both seats.

How do I leave Fusion? Run /runmode cloud or /runmode local, or pick another mode in the switch (ctrl+r).