Private agent substrate

Memory that gets better.

Antumbra plugs into the coding agent you already use. It remembers what your work teaches it, keeps track of where each lesson came from, and trains what keeps proving true into small experts that run on your own hardware.

Private previewClaude Code todayAny MCP client

An annular eclipse, seen from the antumbra: the moon sits inside the sun's disc and a ring of light stays lit.

A session with Antumbra

Your agent starts every session already knowing your work.

Three hooks connect Claude Code to Antumbra. This is one session in a payments service, replayed: what was recalled, what the agent did with it, and what went back in.

  1. Session start

    The bootstrap recalls standing conventions and the memories that matter for this repository and branch, and checks each one's git anchor before your agent reads it.

  2. Every prompt

    Each message recalls the three memories most relevant to it, scoped to the repository and branch you are on.

  3. The work

    Your agent uses what it recalled and says which memory it leaned on.

  4. Capture

    Before the turn ends, what was learned goes back in with its repository, commit and branch.

  5. Standing

    A memory that helped is reinforced. One that misled is penalized. Nothing ages out on a timer.

~/src/ledger · fix/webhook-retry
claude SessionStart · antumbra-session-start Antumbra · 9 memories for acme/ledger on fix/webhook-retry (a41c9e2) (world) Ledger amounts are i64 cents, never f64. [live] (bank) CI needs SQLX_OFFLINE=true or sqlx dials a database. [live] (bank) Refund rounding fixed on release/2026.08. [not-on-head] (world) Payout batches run from feat/batch-payouts. [orphaned] + 5 more recalled but left out; recall_memories for them. the webhook retry test hangs again after the tokio bump UserPromptSubmit · antumbra-prompt-recall (bank) 0.84 tokio 1.40 hung tests/webhook_retry.rs until the test paused the clock with tokio::time::pause(). [live] (world) 0.90 Retries back off 1s, 2s, 4s through RetryPolicy::exponential. [live] (opinion) 0.75 Prefer paused virtual time to real sleeps in async tests. This matches a hang from September: the test waits out real backoff delays. Pausing tokio's clock fixed it then, so I'll do the same here. Edit tests/webhook_retry.rs (+3 -1) Bash cargo nextest run webhook_retry 6 tests passed in 0.41s Stop · antumbra-capture store_memory (bank) tokio 1.41 needs tokio::time::pause() in tests/webhook_retry.rs too; the fix holds across minor bumps. provenance github.com/acme/ledger @ 9d1e4b7 (fix/webhook-retry) reinforce_memory (bank) tokio 1.40 hung tests/webhook_retry.rs... confidence 0.84 -> 0.88 · reinforced 3 times

What it keeps

Three kinds of memory, kept apart.

Facts, experiences and judgments earn trust differently, so Antumbra stores them in separate networks. Every memory carries a confidence, a count of how often it has helped, and the evidence behind it.

  • Facts world

    What is true about your code, your stack and your conventions.

    Ledger amounts are stored as i64 cents. Never f64.

    0.95reinforced 6×acme/ledger@a41c9e2

  • Experiences bank

    What happened when something was tried, and how it turned out.

    The tokio 1.40 bump hung tests/webhook_retry.rs until the test paused the clock.

    0.88reinforced 3×acme/ledger@3f9c2a1

  • Judgments opinion

    How you like the work done. Opinions need more confidence before they train anything.

    Prefer paused virtual time to real sleeps in async tests.

    0.75reinforced 2×stated by you

How it stays useful

A memory your agent can trust, and one that grows.

  • Recall

    Answers that clear a bar, or nothing at all.

    Dense and full-text search are fused, so exact identifiers still surface. A cross-encoder reranks the candidates and a calibrated floor turns the rest away. When nothing is relevant, your agent hears nothing_cleared_the_floor instead of five near misses.

    Try the recall bench
  • Provenance

    It knows when a memory has gone stale.

    A memory about code carries its repository, commit and branch. Before your agent reads it, the anchor is checked against HEAD and tagged live, not on HEAD, or orphaned. A merged pull request carries what its branch learned onto main.

    Watch an anchor move
  • Experts

    What keeps proving true becomes a skill.

    Reinforced, verifiable memories graduate through a consolidation gate into frozen LoRA experts over one small shared model. A router sends a task to the expert that covers it, and escalates with a reason when none does.

    Meet the population
  • Ownership

    On your hardware, under your keys.

    The store, the embedder and the reranker run where you run them. Every call carries a signed token, and tenant and compartment isolation is enforced inside the database engine. Sharing is explicit, and revoking it takes effect at once.

    See how sharing works

Measured

Numbers from the repository, with their sources.

Each figure below comes from an ADR, an experiment or the stack's own notes. Where a result is preliminary, it says so.

  • 0.885

    F1 of the relevance floor with a calibrated gte-reranker-modernbert-base, at an expected calibration error of 0.043.

    ADR-0024
  • 65–80 ms

    To rerank 32 memories on an RTX 3090 Ti. The same batch takes 3.5 to 5 seconds on a CPU.

    docker-compose.gpu.yml
  • 170 ms

    For an answer from a resident expert, 1.2 seconds cold. A latency figure only: answer quality is still being measured.

    EXP-022
  • 0.37 → 0.67

    Recall@10 on 400 real memories when each is indexed as overlapping chunks. Measured, and not yet shipped.

    ADR-0025, proposed

Works with

It plugs into the agent you already use.

Antumbra speaks MCP. It ships 27 tools; the agent profile trims them to the eleven a coding agent needs, highlighted below.

  • Claude Code tested

    Hooks for session start, every prompt, capture and compaction, in bash and PowerShell.

  • Any MCP client

    Local stdio, or streamable HTTP with a signed token on every call.

  • Terminal console

    antumbra-tui shows the population, memory, the loop and evals. --demo runs it on a seeded store.

  • Dashboard

    Recall, compartments, documents and the dependency graph in the browser, served by the MCP server itself.

  • GitHub

    A webhook re-anchors memories at merge, orphans them when a branch is deleted, and ingests documents.

  • recall_memories
  • store_memory
  • reinforce_memory
  • penalize_memory
  • recall_documents
  • ingest_document
  • route
  • answer
  • leave_handoff
  • handoffs
  • complete_handoff
  • list_memories
  • forget_memory
  • relate_memories
  • get_neighbors
  • list_documents
  • population
  • workspace_stats
  • create_compartment
  • propose_compartments
  • list_compartments
  • share_compartment
  • revoke_compartment
  • record_dependency
  • list_dependencies
  • blast_radius
  • record_merges

Start with your next session.

Stand up the store, mint a token, add three hooks. Your agent's next session opens with what the last one learned.