Recall
Answers that clear a bar, or nothing at all.
Most memory tools return the five closest things they have, relevant or not. Antumbra reranks what it finds against your question and drops whatever falls under a calibrated floor, so your agent's context holds what applies and nothing else.
The recall bench
Ask a payments service what it remembers.
Type a question, or pick one. Move the floor to see what your agent would and would not be handed.
You are in github.com/acme/ledger on fix/webhook-retry, with 37 example memories.
- Dense12nearest by meaning
- Full text9BM25, exact words
- Fused13reciprocal rank, k = 60
- Cleared1p ≥ 0.50
-
banklive
The tokio 1.40 bump hung tests/webhook_retry.rs until the test paused the clock with tokio::time::pause().
dense0.83bm2512.7p0.97 - floor 0.50 · below it, your agent sees nothing
-
worldlive
Run the tests with cargo nextest run; plain cargo test misses the per-test timeouts in .config/nextest.toml.
dense0.49bm252.5p0.12
This bench runs a small imitation of the pipeline in your browser. A hand-built concept space stands in for the MiniLM embedding and a logistic stands in for the cross-encoder; the stages, the fusion, the floor and the scoping work the way Antumbra's do.
What happens to a question
Find widely, then keep only what applies.
-
Two searches, one list
A 384-dimension HNSW index finds memories close in meaning; BM25 full text finds the ones that use your exact words, so
SQLX_OFFLINEstill surfaces. Reciprocal rank fusion merges the two. -
A reranker reads each one
A cross-encoder scores every candidate against the question. If the reranker is down, recall falls back to the fused order instead of failing.
-
A floor that means something
The rerank score is calibrated per model into a probability, so the floor is the same bar whichever reranker you run. When nothing clears it, recall says
nothing_cleared_the_floor. -
Scoped, never hidden
Pass the repository and branch, and memories from elsewhere are ranked below the ones that apply here. They stay visible, because sometimes the answer is in another repository.
-
Several questions, several searches
A prompt that asks for more than one thing is retrieved part by part as well as whole, so the second question is not drowned out by the first.
-
Short answers first
Recall returns a bounded prefix of each memory, so a survey does not use up the context window. Ask for the full text once you know which memory matters.
Measured
How well the floor sorts relevant from not.
-
0.885
F1 of the floor with gte-reranker-modernbert-base calibrated, at 0.882 accuracy. The bge-reranker-base calibration reaches 0.797.
ADR-0024 -
0.043
Expected calibration error: on average, the probability the floor reads is within about four points of how often memories at that score turn out relevant.
ADR-0024 -
65–80 ms
To rerank 32 memories on an RTX 3090 Ti, using about 900 MiB. The CPU takes 3.5 to 5 seconds for the same batch.
docker-compose.gpu.yml -
0.21 → 0.42
Mean reciprocal rank on 400 real memories when each is indexed as overlapping 400-character chunks. Measured, and not yet shipped.
ADR-0025, proposed
Recall is half of it.
The other half is knowing whether a memory still holds: its anchor, its standing, and who may read it.