Recall

Answers that clear a bar, or nothing at all.

Most memory tools return the five closest things they have, relevant or not. Antumbra reranks what it finds against your question and drops whatever falls under a calibrated floor, so your agent's context holds what applies and nothing else.

The recall bench

Ask a payments service what it remembers.

Type a question, or pick one. Move the floor to see what your agent would and would not be handed.

You are in github.com/acme/ledger on fix/webhook-retry, with 37 example memories.

  1. Dense12nearest by meaning
  2. Full text9BM25, exact words
  3. Fused13reciprocal rank, k = 60
  4. Cleared1p ≥ 0.50
0.50
  • banklive

    The tokio 1.40 bump hung tests/webhook_retry.rs until the test paused the clock with tokio::time::pause().

    confidence 0.84reinforced 2×acme/ledger

    dense0.83
    bm2512.7
    p0.97
  • floor 0.50 · below it, your agent sees nothing
  • worldlive

    Run the tests with cargo nextest run; plain cargo test misses the per-test timeouts in .config/nextest.toml.

    confidence 0.85reinforced 2×acme/ledger

    dense0.49
    bm252.5
    p0.12

This bench runs a small imitation of the pipeline in your browser. A hand-built concept space stands in for the MiniLM embedding and a logistic stands in for the cross-encoder; the stages, the fusion, the floor and the scoping work the way Antumbra's do.

What happens to a question

Find widely, then keep only what applies.

  • Two searches, one list

    A 384-dimension HNSW index finds memories close in meaning; BM25 full text finds the ones that use your exact words, so SQLX_OFFLINE still surfaces. Reciprocal rank fusion merges the two.

  • A reranker reads each one

    A cross-encoder scores every candidate against the question. If the reranker is down, recall falls back to the fused order instead of failing.

  • A floor that means something

    The rerank score is calibrated per model into a probability, so the floor is the same bar whichever reranker you run. When nothing clears it, recall says nothing_cleared_the_floor.

  • Scoped, never hidden

    Pass the repository and branch, and memories from elsewhere are ranked below the ones that apply here. They stay visible, because sometimes the answer is in another repository.

  • Several questions, several searches

    A prompt that asks for more than one thing is retrieved part by part as well as whole, so the second question is not drowned out by the first.

  • Short answers first

    Recall returns a bounded prefix of each memory, so a survey does not use up the context window. Ask for the full text once you know which memory matters.

Measured

How well the floor sorts relevant from not.

  • 0.885

    F1 of the floor with gte-reranker-modernbert-base calibrated, at 0.882 accuracy. The bge-reranker-base calibration reaches 0.797.

    ADR-0024
  • 0.043

    Expected calibration error: on average, the probability the floor reads is within about four points of how often memories at that score turn out relevant.

    ADR-0024
  • 65–80 ms

    To rerank 32 memories on an RTX 3090 Ti, using about 900 MiB. The CPU takes 3.5 to 5 seconds for the same batch.

    docker-compose.gpu.yml
  • 0.21 → 0.42

    Mean reciprocal rank on 400 real memories when each is indexed as overlapping 400-character chunks. Measured, and not yet shipped.

    ADR-0025, proposed

Recall is half of it.

The other half is knowing whether a memory still holds: its anchor, its standing, and who may read it.