← Back to blog
Product October 9, 2026 by Javier Arancibia

My 7M Router Now Knows When to Shut Up — Conformal Abstention, Shipped

Last week the abstention story on my tiny self-hosted router was a hand-tuned confidence floor. This week it's a calibrated guarantee — and the change is two environment variables.


The product under all of this is mtlm-router: a ~7M-parameter pure-MFL model that classifies requests into typed routes on a CPU, plus small pluggable heads per domain. Its whole pitch is that it abstains instead of guessing — a request it can't place confidently goes to a human, a bigger model, or a specialist, not to a confident wrong answer.

Until now the abstention point was a confidence floor: max_prob >= 0.55 or per-lane overrides. It works, but "0.55" is a vibe. Every customer asks the same question — how wrong will it be? — and a vibe doesn't answer that.

Conformal prediction, in one paragraph

Conformal prediction is a distribution-free calibration method: hold out a labeled split, compute how wrong the model's confidence tends to be, and derive a threshold qhat for a stated error budget alpha. At runtime the model doesn't answer with its top class — it answers with the set of classes whose probability exceeds 1 - qhat. A singleton set means "only one plausible route" — answer it. An empty or multi-way set means genuine ambiguity — abstain.

The insight that sold me: a request scored at 0.48 confidence where every other class is near zero is not ambiguous — it's a strong answer that fails a 0.55 floor. The singleton-set criterion sees that. Max-prob alone can't.

The numbers (municipal routing, merged taxonomy)

On a 315-row held-out evaluation split of French municipal signalement requests, calibrated on a separate split of live probability vectors:

gate                        automated   error-on-automated
raw floor  tau=0.80             87.9%        0.36%
raw floor  tau=0.55             96.2%        0.33%
conformal  alpha=0.02           97.8%        0.97%
conformal  alpha=0.05           94.6%        0.34%

Thirteen points more automation than the floor we were running, at a comparable error rate — with the error budget stated up front instead of discovered in production. Honest caveat, because I'd want it stated: that's the empirical figure on this split. The formal coverage guarantee assumes calibration and production traffic are exchangeable — drift means recalibrate. It's still a much better contract than "trust 0.55."

What shipping it looked like

The head already returns a full probability vector, so the gate is a loop and a comparison — no extra forward pass, no second model:

ANVIL_EXPERT_QHAT="mairie_c:0.648"   # per-lane
ANVIL_CONFORMAL_QHAT=0.648           # default head

Responses carry set_size and qhat; non-singletons delegate with reason:"conformal_abstain". When a lane has a qhat configured, the confidence floor steps aside — the calibrated set is the criterion.

Why this is the actual product

The offer to a municipality or an SME was always "we automate what's safe and escalate the rest." Now it's precise: give us ~200 labeled examples and an error budget, we ship you a calibrated abstention point — one env line. The evaluation tool (tools/conformal_eval.py, in the repo) scores a holdout through the live server, calibrates on one half, and reports automation-vs-error on the other. The whole thing — model, heads, gate — is a ~8MB tarball that runs on a laptop.

Live on the hosted API today; the municipal lane answers 97.8% of what citizens throw at it and abstains on the rest, measurably.

Self-hosted, CPU-only, no telemetry. If your inbox or front desk has a routing problem with an error budget, that's the demo I want to run.

Enjoyed this post?

Follow for more on agent-first engineering, self-hosted systems, and building for autonomy.

Follow @javimosch