Skip to content

Memory

Most agents “remember” by stuffing the entire chat history back into the prompt. That gets expensive fast and crowds out the model’s attention. Atomic Agent keeps its long-term memory in a local SQLite database instead. Within a session, recent conversation stays in the prompt, bounded by the context window (agent.conversationMaxPairs, default 200 pairs, trimmed as the window fills). Across sessions it does not replay transcripts: it feeds the model only the few stored facts that matter for the turn in front of it.

The result is an agent that recalls your name, your project conventions, and lessons it learned last week — without dragging a giant transcript along for the ride.

What you can do with it

Store facts about you

Pin durable facts (memory.profile.set) like your name, OS, or preferred language. They render into every relevant prompt automatically.

Save and search notes

Drop freeform notes (memory.notes.store) and recall them later by keyword search (memory.notes.recall). Full-text indexed, no setup.

Let it learn on its own

After each turn, a background reflection pass quietly extracts facts and notes from the conversation — you don’t have to ask.

A database you own

The memory database lives in <stateDir>/memory.sqlite on your machine. Recalled memory goes into the prompt, and reflection runs on the active model, so with a hosted provider that memory text is sent to the provider.

The three channels

Memory is split into three independent channels that share one SQLite file. Each answers a different question.

ChannelHoldsTool surfaceRenders as
ProfileKey/value facts about you and your environmentmemory.profile.set/remove/list/history### profile
NotesSearchable freeform observationsmemory.notes.store/recall/forget### recalled + ### memory-index
ReflectionAutomated end-of-turn extraction(runs itself)feeds the other two

On top of these, the memory-v2 fabric adds distilled lessons and advisory procedures (memory.lessons.recall, memory.procedures.recall) plus link graphs and vote curation — all enabled by default — and optional embeddings. More on those below.

Memory that grows outside the prompt

This is the core idea. The prompt the model sees is small and bounded. Memory grows on disk.

Each turn, Atomic Agent pulls only what’s relevant from SQLite and renders it into the variable tail of the prompt — never the cached stable prefix. The tail sections each have a hard token cap:

  • ### profile — pinned facts + facts whose keywords hit your message (cap: memory.profile.maxTokens, default 512)
  • ### recalled — top-K matching notes as short previews, about 160 characters each (memory.recallInjection.previewChars); the model fetches the full body with memory.notes.recall { id } (cap: memory.recallInjection.maxTokens, default 400)
  • ### memory-index — compact pointers to other notes by id (cap: memory.index.maxTokens, default 300)
  • ### lessons — distilled patterns (cap: memory.lessons.maxTokens, default 300)
  • ### procedures — advisory how-to templates (cap: memory.procedures.maxTokens, default 400)

Because these always live in the tail and never the prefix, your KV-cache survives. Memory injection costs a few hundred tokens, not a full transcript.

How memory is written and recalled

Every turn touches memory in two ways: a hot read before the model runs (refreshed after each step), and a fire-and-forget write after it replies.

graph TD
    A["Turn starts<br/>runTurn(userMessage)"] --> C["MemoryContextProvider<br/>.buildMemoryContext()"]

    subgraph Recall["Hot read (before the first step, refreshed after each step)"]
        C -->|recallHybridAsync| D["MemoryStore<br/>BM25 over memories_fts"]
        D -.optional cosine.-> E["EmbeddingStore"]
        C -->|expand| F["LinkStore.expand()<br/>BFS neighbors"]
        C -->|recall| G["LessonStore + ProcedureStore"]
        C -->|get active| H["ProfileStore<br/>superseded_by IS NULL"]
    end

    D --> I["Render variable tail:<br/>### profile / ### recalled<br/>### memory-index<br/>### lessons / ### procedures"]
    F --> I
    G --> I
    H --> I

    I --> J["Agent loop runs steps<br/>→ assistant_reply"]

    subgraph Write["Hot write (post-turn, async, never awaited)"]
        J -->|fire-and-forget| K["ReflectionRunner<br/>on the side-call slot"]
        K -->|micro-prompt| L["Extract SET / NOTE / EVOLVE"]
        L -->|SET| M["ProfileStore.set()"]
        L -->|NOTE| N["MemoryStore.storeAsync()<br/>dedup → insert/merge"]
        N -->|overflow| O["utility-weighted eviction<br/>(default)"]
        N -.fire-forget.-> E
        K -.phase 2.-> P["LinkGeneratorRunner"]
        K -.phase 7a.-> Q["VoteRunner"]
    end

    subgraph Cold["Cold path (periodic consolidator)"]
        R["ConsolidatorJob<br/>every ~6h"] -->|connected components<br/>over link graph| S["Distill cluster → Lesson + Procedure"]
        S -->|archive parents| T["memories.consolidated_into"]
        S -->|decay| Q
    end

    style C fill:#e1f5ff
    style K fill:#fff3e0
    style R fill:#f3e5f5
    style I fill:#e8f5e9

Two things to notice:

  1. Recall runs before the first step and is refreshed after each step, using your message plus recent tool results, so a note stored mid-turn can surface on the next step. Fusion worker turns do not read memory.
  2. Reflection is never awaited. It runs in the background, on a reserved side-call slot when llama-server has more than one slot (on a single-slot server it takes an idle slot), so it doesn’t disturb the main agent’s KV-cache. A new reflection aborts the previous one for the same session. Reflection runs on the active model: with a hosted provider, the conversation it reads is sent to that provider.

Working with profile facts

Profile facts are key/value metadata that render straight into the prompt. They’re the agent’s “always-on” knowledge about you.

Terminal window
# The agent calls these tools itself, but here's the shape:
memory.profile.set { key, value, pinned?, keywords? }
memory.profile.list # shows * for pinned, ~ for contextual
memory.profile.history { key } # bi-temporal version chain
memory.profile.remove { key }

There are two flavors:

  • Pinned facts (pinned: true, the default) always render. Use for identity-level truth: your name, OS, primary language.
  • Contextual facts (pinned: false + keywords) only render when your message contains one of their keywords. This keeps cold facts reachable without bloating every prompt.

Working with notes

Notes are freeform text, full-text indexed (FTS5), and recalled by relevance.

Terminal window
memory.notes.store { content, tags? } # max 4000 chars, up to 16 tags
memory.notes.recall { query? | id?, k?, scope?, tags? } # scope: "all" | "project"
memory.notes.forget { id }

The agent recalls notes two ways:

  • Automatic injection — before each turn, the top-K notes matching your message are rendered into ### recalled (default k = memory.recallInjection.k, 3).
  • Explicit recall — the model can call memory.notes.recall with its own query to pull more, or look up a specific note by id.

Storage stays bounded

Memory can’t grow forever. Atomic Agent enforces caps at multiple layers:

  • Per-reflection caps: at most memory.reflection.maxFactsPerCall facts (default 3) and memory.reflection.maxNotesPerCall notes (default 2) per turn.
  • Hard note cap: memory.notes.maxEntries rows (default 1000). On overflow, excess rows are evicted in a single SQL statement.
  • Utility-weighted eviction (memory.eviction.utilityWeighted, on by default): rows are evicted lowest vote_score first, then least recalled, least recently recalled, then least recently updated, so downvoted, never-recalled, stale notes go first. Set it to false for plain eviction by least recently updated.
  • Lessons and procedures are capped at 500 each, and ones that never helped a turn are deprecated after 30 days. The profile is capped at 500 unpinned facts.

The memory-v2 fabric

Beyond facts and notes, Atomic Agent layers a richer fabric inspired by Complementary Learning Systems — a fast “hippocampal” write path and a slow “neocortical” consolidation path.

The phases, and what each adds

Everything below is on by default except phase 1B (embeddings).

  • Phase 1A — dedup + eviction: FTS5 candidate fetch + Jaccard similarity merges near-duplicate notes (memory.dedup.enabled, threshold memory.dedup.fts5Threshold default 0.85). Utility-weighted eviction (memory.eviction.utilityWeighted).
  • Phase 1B — embeddings: a second llama daemon produces vectors for hybrid recall (memory.embeddings.enabled). Cosine search is brute force, up to bruteForceCeiling rows (default 200); above that, recall is BM25-only. The one phase that’s off by default: it needs the embedding daemon running.
  • Phase 2 — link graph: LinkGeneratorRunner writes typed edges between related notes; recall expands neighbors via BFS (memory.links.enabled, memory.links.expansionDepth default 1, memory.links.maxExpanded default 12).
  • Phase 3 — evolution: reflection refines tags on existing memories as it learns more (memory.evolution.enabled). Note content stays append-only regardless.
  • Phase 5 — lessons: the cold-path ConsolidatorJob clusters linked episodes, distills each cluster into one Lesson (a 1–3 sentence principle), and archives the parents (memory.lessons.enabled, memory.consolidation.enabled).
  • Phase 7a — voting: up/downvotes curate ranking, clamped per item (memory.voting.maxVotePerItem) and decayed each consolidator tick (memory.voting.signalDecay, default 0.95).
  • Phase 7b — procedures: the same clusters yield advisory Procedures (2–8 plain-text steps). They’re searchable but never auto-executed (memory.procedures.enabled).

Distilled lessons and procedures

Lessons and procedures both render into the prompt tail (### lessons, ### procedures) and are recalled with their own tools:

Terminal window
memory.lessons.recall { query?, id?, k? }
memory.procedures.recall { query?, id?, k? }

Adding lessons (phase 5) and procedures (phase 7b) are the only two changes the memory fabric makes to the byte-stable prompt prefix. Each one invalidates the KV-cache once at rollout, so plan a fresh session pool when you first enable them.

Query rewriting before recall (v2.5)

Raw user messages make poor search queries. “What did we decide about that?” has almost no keywords for BM25 to bite on, so recall comes back empty exactly when context would help most.

The v2.5 query rewriter sits in front of recall and fixes this. It’s enabled by default (memory.retrieve.rewriter.enabled) and gated by a heuristic (gateMode: "heuristic"), so it only fires when a message actually looks like it needs rewriting — a short or referential query gets expanded using the last few turns, while a keyword-rich one goes straight through untouched.

{
"memory": {
"retrieve": {
"rewriter": {
"enabled": true,
"gateMode": "heuristic",
"historyTurns": 3,
"timeoutMs": 10000
}
}
}
}

Two design choices keep it cheap. The rewriter runs on the reserved side-call slot it shares with reflection (or any idle slot when the server has only one), never the main agent’s slot. And it’s bounded by timeoutMs (default 10000): if the rewrite doesn’t come back in time, recall proceeds with the original message rather than stalling the turn. Like reflection, it runs on the active model.

Configuration

Memory is tuned entirely under the memory.* block in <stateDir>/config.json. The most common knobs:

{
"memory": {
"profile": { "enabled": true, "maxTokens": 512, "contextualKeywordGate": true },
"notes": { "enabled": true, "maxEntries": 1000, "recallDefaultK": 5 },
"recallInjection": { "enabled": true, "k": 3, "maxTokens": 400 },
"index": { "enabled": true, "limit": 20, "maxTokens": 300 },
"reflection": { "enabled": true, "maxFactsPerCall": 3, "maxNotesPerCall": 2 }
}
}

If memory.notes.enabled is false, the MemoryContextProvider is never built and the agent skips the note-based sections (### recalled, ### memory-index) in the prompt. Profile facts still render.

To change one key, use atomic-agent config set <key> <value>, for example atomic-agent config set memory.retrieve.rewriter.enabled false.

Inspecting memory

The whole fabric is a single SQLite file you own:

<stateDir>/memory.sqlite

It holds the profile_facts, memories (+ memories_fts), memory_embeddings, memory_links, lessons, procedures, and vote_events tables. To browse it live, open the TUI:

Terminal window
atomic-agent tui
# → Manage → Memory tab: profile, notes, lessons, procedures, links and votes
# (in notes, `f` toggles the archive filter and `g` shows a note's linked neighbours)

To read your memory outside the agent, export it to an Obsidian vault as Markdown files with [[wikilinks]]:

Terminal window
atomic-agent memory export --vault ~/Documents/MyVault
# or: OBSIDIAN_VAULT_PATH=~/Documents/MyVault atomic-agent memory export

The export covers notes, lessons and procedures. It is one-way (the database is opened read-only and nothing syncs back); re-running it updates files in place and prunes files for records that are gone. --folder <name> picks the vault subfolder.

Gotchas worth knowing

  • Recall is refreshed every step. Memory is fetched before the first step and again after each one, so a note stored mid-turn can show up on the next step.
  • Reflection runs on the side-call slot. With more than one llama-server slot it never touches the main agent’s KV-cache, and a new reflection aborts the prior one for the same session.
  • Memory tools never prompt. Writes (memory.profile.set, memory.notes.store, and so on) run without approval, one at a time; recalls are read-only and can run in parallel.
  • Dedup is strict. The existing row must be a tag superset of the new entry to merge (Jaccard ≥ threshold). Otherwise both are kept.
  • Embeddings are computed in the background. A freshly stored note may not have its vector yet on the very next recall; keyword search still finds it.
  • Migrations are idempotent and one-way. Re-running on a current-version database is a no-op; downgrades are refused.

Memory tools

Full reference for memory.profile.*, memory.notes.*, memory.lessons.recall, memory.procedures.recall.