Skip to content

Atomic Agent

A local-first AI agent that runs on your own machine. Private by default, no per-token fees, hackable all the way down.

Atomic Agent is an AI assistant you run yourself. It uses small, quantized language models through llama.cpp right on your laptop, so your files and conversations stay on your machine by default.

It does real work: browses the web, reads and edits files, runs shell commands, extracts text from documents, remembers things across sessions, and schedules tasks to run in the background. You drive it from a terminal UI, a CLI, an OpenAI-compatible HTTP API, or chat apps like Telegram and Discord.

Atomic Agent is built by AtomicBot, a member of NVIDIA Inception. The project stays MIT-licensed and runs on whatever hardware you already own — an NVIDIA GPU is not required.

Why it’s different

Local-first

Inference runs on your hardware via llama.cpp. No account, and your files and conversations stay local by default.

Private by default

Files, sessions, and memory live in a local SQLite store. Browser, HTTP, cloud models, MCP and shell are tools you can see and control. Anonymous usage analytics and crash reports are on by default; turn them off with "analytics": { "enabled": false } in config.json.

No per-token fees

Run as many turns as you want on a local model. The only cost is the electricity your GPU draws.

Hackable all the way down

Add skills, wire up MCP servers, swap models, tune the prompt budget. Everything is inspectable and configurable.

Start here

How it works

Every turn runs the same tight loop: build a compact prompt, ask the model for tool calls, execute them, fold the results back into durable state, and repeat until the agent replies or finishes. A byte-stable prompt prefix keeps the KV-cache warm so each step stays cheap, and memory grows externally in SQLite instead of bloating the context window.

flowchart TD
    User["User message<br/>or scheduled task"]
    TC["TurnController<br/>per-session FIFO"]
    Prompt["Prompt builder<br/>stable prefix + bounded tail"]
    Model["llama.cpp / cloud model<br/>GBNF grammar"]
    Calls["JSON tool calls<br/>1..N actions"]
    Exec["Batch executor<br/>parallel reads, approval gates"]
    Tools["Tools<br/>browser · files · shell · MCP"]
    State["Durable state<br/>sessions · memory · tasks"]
    Reflect["Reflection<br/>async, off the hot path"]
    Reply["Reply / finish"]

    User --> TC
    TC --> Prompt
    Prompt -->|KV-cache reuse| Model
    Model -->|constrained JSON| Calls
    Calls --> Exec
    Exec --> Tools
    Tools -->|compressed results| State
    State -->|next step| Prompt
    Exec -->|terminal call| Reply
    Reply -.->|end of turn| Reflect
    Reflect -.->|distill facts & notes| State

    classDef hot fill:#e1f5ff,stroke:#0369a1,color:#0c4a6e;
    classDef store fill:#e8f5e9,stroke:#2e7d32,color:#1b5e20;
    class Prompt,Model,Exec hot;
    class State,Reflect store;

The model never sees a giant chat log. Conversation history is bounded and compressed, recalled memory arrives as compact pointers, and dangerous actions (shell, file writes, HTTP) pause for approval. Read more in Architecture and Local-first.

Explore the features

Go deeper