A cheaper model that drops quality is debt, not a saving.

One task class, both routes, hundreds of runs, scored against the check you ratified. Promote only when the cheaper route clears the bar on enough evidence and costs less; otherwise the verdict is hold: keep paying, now proven to earn its keep. The bar is ours today. What an outcome is worth to you should set it, and that is a number we will never guess.

Cheaper by 69.4%, and read task by task: the saving forfeits 20 outcomes the incumbent lands, and recovers 4.Figure 1

The same 116 tasks, both arms run in full on a live tagging workflow. Each cell is one task. The gate reads the pairing, not the averages: which outcomes the cheaper arm keeps, which it drops, and which it finds that the incumbent missed.

88 both arms landed20 only the incumbent — what the saving forfeits4 only the candidate4 neither

Held. The cheaper arm keeps 88 of the incumbent’s outcomes, drops 20, and finds 4 of its own — a net 16 forfeited to save $0.0372. Break‑even: $0.0023 an outcome. Value a finished outcome above that and refusing was the right trade.

n = 116 paired taskslanded 108 → 92net −16 outcomesis_simulated=falsecaptured 2026-08-19

Source: the committed artifact provenance/autoroute_defended_savings_frac.json. Every cell derives from its per-call rows at build time; none is typed.

The gate says yes when the evidence does: −63.7% with quality up — promoted, the paired test over 115 shared tasks holding.Figure 2

The same model at a lower thinking budget, 115 trials on a live reasoning workflow. The first route the gate has passed in a paired eval: not worse than the incumbent task for task, and cheaper.

A. What each budget cost
gemini-2.5-flash · full thinking budget$0.0970
gemini-2.5-flash · thinking budget 128$0.0352
B. What each budget landed
full thinking budget89%
thinking budget 12894%
the drawn rule is the full budget’s own rate: the lower budget cleared it

Promoted. The bar is parity with the incumbent, not an absolute rate: over 115 shared tasks the lower budget was not worse task for task, so the gate passed it on 2026-09-19.

n = 115 taskscost −63.7%quality 89% → 94%is_simulated=falsecaptured 2026-09-19

Source: the committed artifact provenance/reasoning_effort_savings_frac.json. Both panels derive from it at build time.

A different provider entirely, 85% cheaper — and refused, because quality fell from 92% to 63%.Figure 3

Figs. 01 and 02 each stayed inside one vendor’s house. The front page says cross-provider, so here the incumbent is Google’s gemini-2.5-flash and the challenger DeepSeek’s deepseek-v4-flash-0731, both run in full over 91 tasks on the same frozen check.

A. What each provider cost
Google · gemini-2.5-flash$0.0794
DeepSeek · deepseek-v4-flash-0731$0.0118
B. What each provider landed
Google incumbent92%
DeepSeek challenger63%
the drawn rule is the incumbent’s own rate: the challenger fell well short of it

Refused. A different provider and 85% cheaper, turned down because the challenger landed 63% of outcomes to the incumbent’s 92%. A genuine rival, not a cheaper tier of the same house, and the gate ruled on it exactly as it ruled on one. That is what makes the neutrality real.

n = 91 taskscost −85% against the incumbentquality 92% → 63%is_simulated=falsecaptured 2026-09-13

Source: the committed artifact provenance/autoroute_cross_provider_algorithm.json. The first real face-off across two providers; both panels derive from it at build time, and the vendor names are read off the model ids, never typed.

Their card says 117 of 120. Everyone reads that as 97.5%; the gate reads it as 93.9%, and holds.The example run card on OpenRouter’s Ori page ranks three models on one 120-prompt eval by asserts passed: 120, 117, 112 of 120. Nothing on the card is wrong. The gate reads a pass count as the lowest true rate the panel can vouch for at 95% confidence, and compares that to its bar, never the headline rate.Figure 4
A. As the card prints it
120 of 120100.0%
117 of 12097.5%
112 of 12093.3%
B. As the gate reads it
120 of 12097.8%
117 of 12093.9%
112 of 12088.6%
the drawn rule is the card’s own rate; the fill is the lowest rate 120 prompts can vouch for. Even a clean sweep reads 97.8%, not 100.0%.

Held, on the arithmetic alone. 117 of 120 is the ordinary pass, and most tooling would ship it. Its lower bound sits under our bar, so the gate would keep paying for the incumbent and ask for more evidence. The clean 120 of 120 is not a promotion yet either: the gate never reads one arm alone. It runs the incumbent down the same prompts and rules on the pair, the way Figs. 01–03 were ruled.

n = 120 prompts, their panel117 of 120 → lower bound 93.9%their card, their eval — not our meterread 16 September 2026

Source: the example run card on openrouter.ai/ori (ori eval · standup-summarizer · 120 prompts · run #42), read 16 September 2026. OpenRouter is not a customer and has not evaluated Margin; its card is cited for its counts, not its verdicts, and the same three counts on any eval page would read the same way here. The bounds are computed at build time by the same Wilson form the gate runs; the card’s judge scores are not used.

Under these verdicts sits one piece of algebra.

Every mix of models that lands the same outcome is a curve; the cheapest mix that still lands it is where your budget line touches that curve. The gate’s whole job is to keep you on the curve — a point below it costs less and delivers less, and no per-token dashboard can tell the two apart.

The parity engine, written up →
isoquant · cost mix vs. the outcome it buys
small-model callsfrontier callsbudget too smallsame outcomecheapest mix that holds paritycheaper — and short of the outcome
anywhere on the curvethe same outcome
where the budget line touches it4 small : 1 frontier
below the curvecheaper, and short
Buy the mix, not the model. Every point on the curve buys the same outcome; the cheapest is where the budget line touches it. Below the curve you pay less and get less, and that line is our parity gate. Drawn from the algebra — your own mix is measured.

Next: The Governor →

quality slips → it reverts → the loop re-runsThe Metermeasures every callThe Estatewhat you actually runThe Money Mapwhere the spend goesThe Parity Gateprove it held qualityThe Governoract, with revert armed
  1. The Meter: measures every call, its cost, tokens, substrate, and whether it worked.
  2. The Estate: the org chart of your AI workforce.
  3. The Money Map: where the spend goes, decomposed by step.
  4. The Parity Gate: proves the cheaper route held quality before it ships.
  5. The Governor: acts with a human ratifying, and a revert armed; if quality slips it reverts and the loop re-runs.
A cheaper route, once a person approves it, is held only while quality holds. Today that bar is one we set on your behalf.

Watch the loop run on real spend.

The console shows the loop on open-source agents we metered ourselves: measured spend, recorded face-offs and the gate’s verdicts, every number tagged real or sample. No live route has slipped yet, so the revert has not fired.