Orchestrator
Holds the conversation with you. Reads enough to choose an approach, breaks the job into parts, writes a brief for each, reviews the replies and runs checks. It cannot write files or run commands during a Fusion turn.
Fusion is one of Atomic Agent’s three run modes, next to Local and Cloud. Instead of one model doing everything, Fusion gives the job to two: an orchestrator that plans, writes a brief for each part and reviews what comes back, and workers that execute those parts, several at once.
Either seat can be a local model or a cloud model. That makes Fusion a way to pair a model on your own machine with a cloud model and get results closer to what the cloud model reaches on its own, while choosing which side does the thinking and which side does the building.
In Fusion, the orchestrator never builds anything itself. It reads your project, splits the job into self-contained parts and sends them out through a tool called fusion.delegate. Each worker takes one part, writes its files straight to disk and reports back. The orchestrator then checks the result and sends weak parts out again.
Orchestrator
Holds the conversation with you. Reads enough to choose an approach, breaks the job into parts, writes a brief for each, reviews the replies and runs checks. It cannot write files or run commands during a Fusion turn.
Workers
Each worker runs one part in a fresh, throwaway session. It sees its brief and your original request, not the chat history or memory. It cannot ask you questions, request approval or delegate further.
A few details worth knowing:
fusion.delegate, reply and finish are allowed. Anything that writes, runs a command, or comes from an MCP server is refused. This is by design: the workers build, the orchestrator judges.fusion.delegate call can carry a contract (owners, provides, requires, checks). Tasks that need something another task provides run in a later wave, and name mismatches are caught before workers start.Fusion runs in two directions, and which side plans and which side builds changes the result. Cloud plans with local workers is the default pairing. Local plans with cloud workers suits jobs with a lot of output, such as large code files, because the heavy writing happens on the cloud side.
This is what /runmode fusion sets up by default: a cloud model as orchestrator and your local model as the workers.
Fits when the job splits into many small, self-contained parts and you want the bulk of the execution on your own hardware.
Watch out for jobs that need a lot of generated output per part. Every line of code a worker writes comes from your local model, and parallel local workers share one GPU. For large code tasks, the reverse direction is usually the better fit.
Your local model orchestrates and cloud models execute. Switch to it with /runmode swap.
Fits when the plan is short but the output is large. A plan is a small amount of text that a local model handles well, while the long writing goes to cloud workers. In our own runs this direction reached the same level as cloud-only on a code-heavy task, where the local model alone fell well short.
Keep in mind that cloud workers see your request and the files they read, so your data does reach the cloud provider. The local orchestrator runs with a single request slot on the local server, which Atomic Agent sets automatically when the server’s slot count is left on "auto".
Two cloud models or two local models also work, as long as the two seats are two different providers. Pointing both seats at the same provider is refused.
The quickest way is from the TUI: run /runmode fusion. You need two configured providers, usually one cloud provider with an API key and one local model that is already downloaded. Atomic Agent pins both seats, saves the mode to your config and starts the local model server if it is down.
/runmode open the "where it runs" switch (same as ctrl+r)/runmode fusion switch to Fusion/runmode swap trade the two seats (orchestrator becomes worker and back)/runmode workers 3 set the default number of workers (1-8)/runmode status show what the mode resolves to right now/runmode cloud leave Fusion for Cloud/runmode local leave Fusion for LocalThe first time a session switches into Fusion, the chat shows a short introduction naming the two models it resolved. After that, /runmode status prints the orchestrator, the worker provider and how many workers can run at once.
The mode lives under llm.runMode in <stateDir>/config.json (by default ~/.atomic-agent/config.json). These keys cannot be set with atomic-agent config set, so either use the TUI or edit the file by hand.
{ "llm": { // "providers": [ ... ] keep your existing providers as they are "activeTextProvider": "openrouter", "runMode": { "mode": "fusion", "fusion": { "orchestratorProvider": "openrouter", "workerProvider": "local-llama", "workers": 2 } } }}The provider ids must match entries in your llm.providers list. local-llama is the id Atomic Agent gives its managed local model.
Key (under llm.runMode.fusion) | Default | Range | What it does |
|---|---|---|---|
orchestratorProvider | the active cloud provider, else the first cloud provider | a provider id | Which provider plans and reviews. |
workerProvider | the first local provider that is not the orchestrator | a provider id | Which provider runs the workers. |
workers | 2 (a local worker seat runs 1 at a time until you set this) | 1 to 8 | Default fan-out width when the orchestrator does not name one. |
cloudWorkers | 4 | 1 to 32 | Hard cap on parallel workers when the workers are in the cloud. |
workerMaxSteps | 60 | 1 to 1000 | Step budget for one worker. |
workerTimeoutMs | 2,700,000 (45 min) | 1,000 to 86,400,000 | Time budget for one worker. |
workerReasoning | provider default | low, medium, high | Reasoning effort sent with worker requests. |
workerMaxOutputTokens | model maximum | 1 to 1,000,000 | Output cap per worker step. |
reviewStallSteps | 6 | 0 to 1000 (0 turns it off) | How long the orchestrator may only read before it is pushed to delegate or reply. |
orchestratorModel, workerModel | unset | any label | Display labels only. They do not change which model serves. |
The orchestrator decides the width of each fan-out itself. Your workers setting only fills in when it names no number. What really limits parallel workers is physical: the number of tasks, the local server’s request slots for a local worker seat, and the cloudWorkers cap for a cloud worker seat.
workers yourself, local workers run one at a time. Extra tasks queue and the result tells the orchestrator how many slots there were.cloudWorkers is clamped and the result says so./runmode workers N saves the default width and also sets the local server’s slot count (localModels.managed.parallel) to the same number. A server that is already running keeps its old slot count until you restart it (Manage › LLM › Local, then s).One fusion.delegate call holds at most 16 tasks, and each brief is capped at 32,000 characters.
The orchestrator reviews every part before it accepts the work. It has two read-only check tools: verify.syntax checks the syntax of the files a part produced, and verify.run runs a command, a local service or a web page in a throwaway copy of the working directory. Anything weak goes back out as a new, more specific brief.
What each check covers:
verify.syntax picks a checker per file type (JavaScript and JSON, TypeScript through the project’s tsc, Python, shell, HTML inline scripts, CSS brace balance). A file with no checker is reported as unchecked, never as passing.verify.run has three kinds: command (a test runner, compiler or script), service (start it, wait for a port, send requests) and page (open a local HTML file or URL in a headless browser, collect uncaught errors, console errors and missing selectors). Nothing it writes reaches your real workspace.failed, cancelled, needs_orchestrator or simply not good enough goes out again in another fusion.delegate call that says what was wrong and what good looks like.reviewStallSteps default), it is told to delegate or reply. At twice that, it can only call fusion.delegate, reply or finish. The threshold is halved when your message reads as a repair request.A Fusion turn shows the plan, one approval for the fan-out, and a live row per worker. Each row names the task, the model running it, the tool in use and the elapsed time, beside the orchestrator’s time estimate when it gave one. A summary line then reports how many tasks came back ok.
The approval. When approvals are on, Atomic Agent asks you once per fan-out, before any worker starts. The prompt lists the task titles and the directories the workers may write to and run commands in. Your answer covers the rest of the turn, so a review that re-delegates five times does not ask five times. A later fan-out that reaches a new directory asks again.
Anything a worker cannot do alone (because it needs a person, or a wider write scope than you approved) comes back to the orchestrator marked needs_orchestrator instead of stalling.
openrouter/auto can pick small, cheap models that ignore Fusion’s instructions. Choose a specific model for the cloud provider (/model or /llm).Most Fusion problems come from which provider is active, a missing second provider, or the local server’s slot count. Start with /runmode status: it prints what the mode resolves to, how many workers can run and, when the stored and the effective mode differ, why.
| What you see | What it means | Fix |
|---|---|---|
stored fusion, effective cloud in /runmode status | The active provider is no longer the orchestrator. | Run /runmode fusion again. |
| ”Fusion needs a cloud orchestrator. No cloud provider is configured. Staying on local.” | No orchestrator could be resolved. | Add a cloud provider (Manage › LLM › Cloud, or /llm), or pin both seats in config. |
| ”Fusion needs two providers, one to orchestrate and one to run the workers.” | Only one provider is configured. | Add a second provider in Manage › LLM. |
swap needs fusion | /runmode swap only works while Fusion is on. | Run /runmode fusion first. |
config set failed: unknown key llm.runMode... | These keys are not settable from the CLI. | Use /runmode in the TUI or edit config.json. |
| A note that work was sent but only some workers ran at a time | The local server has fewer request slots than tasks. | Set /runmode workers N and restart the local server, or let the orchestrator send fewer, larger tasks. |
| The orchestrator’s writes are refused | This is Fusion working as designed. | Nothing to fix. The workers do the writing. |
| A worker “handed back early: no file written by half the budget” | The worker read but did not write. | Usually handled by the orchestrator re-splitting the task. If it repeats, give the task a clearer deliverable. |
| Unsure whether Fusion is really active | The mode chip alone does not prove the orchestrator has the delegate tool. | Run /tools and check that fusion.delegate is listed. |
Which direction should I start with?
For code-heavy jobs, try a local orchestrator with cloud workers (/runmode swap after /runmode fusion). For jobs that split into many small parts, the default cloud orchestrator with local workers is the natural fit.
Does the cloud model see my files? The cloud seat sees what it is given and what it reads: in the default direction, your conversation and the files it reviews; with cloud workers, your request (up to 16,000 characters) plus the files each worker opens.
Can I use two cloud models, or two local ones? Yes. Any two different providers work, for example a careful cloud model directing a fast one, or a large local model directing a small one. With two local providers, pin both seats.
How do I leave Fusion?
Run /runmode cloud or /runmode local, or pick another mode in the switch (ctrl+r).