Experts · Experimental

Competence you own.

A memory layer helps a model remember. Antumbra's experts make it better at your work: what keeps proving true is trained into small, frozen LoRA adapters over one shared base model, each with a learned boundary of what it covers. A router sends a task to the expert that covers it, and escalates with a reason when none does.

ExperimentalNeeds an NVIDIA GPUValidated on one RTX 3090 Ti

The router

Inside its boundary, the expert answers. Outside, it says so.

Each expert learns where its competence ends. The gate scores a task against every expert, answers from the one that clearly covers it, and composes two when a task needs both. When the best two are too close, or nothing is close enough, it hands the task back to your agent's own model with the reason.

An example population. Routing runs without a GPU and explains every escalation; answering needs one.

Pick a task
  • bun-conventionsUses bun, not npm, in JavaScript repositories.0.12
  • exact-pinsPins dependency versions exactly.0.08
  • sqlx-migrationsWrites reversible sqlx migrations.0.21
  • async-testsTests async code on paused virtual time.0.74
  • money-typesKeeps amounts in i64 cents.0.18
  • axum-routesWrites axum 0.8 routes and handlers.0.26
  • release-flowTags and ships through the release workflow.0.09
  • k8s-deployEdits the service's Kubernetes manifests.0.07
  • commit-styleWrites commit messages that explain why.0.11
Answered by async-testsinside its boundary

Similarity 0.74, next best 0.26: a clear margin, above the floor.

#[tokio::test(start_paused = true)]
async fn retries_back_off() {
    // the backoff sleeps now advance on virtual time
}

Served from a resident adapter in about 170 ms, 1.2 s cold (EXP-022, latency only).

The loop

From a verified outcome to a frozen expert.

It runs unattended: the server can consolidate in the background, and a GPU node picks up the training.

  1. Verify

    An outcome counts when something checked it: a test passed, a verifier ran, or you confirmed it.

  2. Remember

    It is stored with its provenance, and reinforced each time it helps again.

  3. Consolidate

    Memories that recur, hold their confidence and can be checked pass the gate. Volatile facts stay in the store.

  4. Train

    RAFT or GRPO fine-tunes a LoRA adapter over a 4-bit Qwen2.5-Coder-1.5B base, on one 24 GB GPU.

  5. Freeze and route

    The adapter is frozen with a learned boundary. The router uses it inside that boundary and escalates outside.

The experiment ledger

Every milestone is an experiment with a kill criterion.

These are Antumbra's own results, on one RTX 3090 Ti with Qwen2.5-Coder-1.5B. They are small-scale, and the ledger says where a result is a single run or still open.

ExperimentQuestionResultStanding
EXP-001Can a verifier teach a convention?RAFT pass rate 0.06 → 1.00 with a convention rule, 0.38 → 1.00 with a code-executing verifier.small scale
EXP-008Does GRPO learn faster than RAFT?GRPO 0.33 → 0.92 → 1.00 against RAFT 0.08 → 0.25 → 1.00.single run
EXP-009Can the base be 4-bit?Matches the f16 learning curve with about a quarter of the resident base memory.small scale
EXP-011Does a one-time correction stick?“Use bun” went 0.00 → 1.00 on held-out packages and survived a fresh process.small scale
EXP-013Can a router separate close experts?A learned router turned a 0.051 margin into p = 1.000 separation.calibration open
EXP-015Do experts compose?Two experts together produced bun add react --save-exact, which neither produces alone.small scale
EXP-017Can it tell what it does not know?In-distribution similarity 0.69–0.75, out-of-distribution 0.35–0.42: room for an abstention floor.small scale
EXP-018Can the population evolve on its own?evolve raised fitness 0.25 → 1.00.small scale
EXP-019Can it fill its own gaps?populate raised coverage 0.00 → 1.00.small scale
EXP-010Does training on new work erase old skills?Not settled either way.inconclusive
EXP-022Does memory to expert to answer run unattended?About 90 s end to end on the first pass with the 3.0 GB base download, about 10 s once cached. Answers in 1.2 s cold, 170 ms resident.quality not yet shown

Why small experts

A way off renting your competence.

  • Owned, not rented

    What an expert learned is a file of weights on your disk, trained on your verified work. No provider can change it or take it away.

  • One base, many adapters

    Adapters are small, so many stay resident over a single base model and swap in without reloading it.

  • A boundary for each

    Counterfactual search finds where an expert stops being right, so the router knows when not to use it.

  • Escalation is the safe answer

    Your frontier agent is still there. Antumbra takes the work it has proven it can do, and hands back the rest.

Start with memory. Experts follow.

Memory, recall and routing run without a GPU. When you add one, what your agent has learned starts turning into experts.