A recommendation you have to babysit is not automation.

Once you ratify a promotion, The Governor publishes it and arms an automatic revert; your code applies the switch, and if parity slips on live traffic the decision rolls back on its own and tells you it did. Autonomy is a dial you set, not a default we assume: the system proposes and proves, a person approves, and it undoes itself the moment quality drops.

The closed loop, written up →
The only promotion ever on our books, we took back — the take-back is the product working.Figure 1

In July the gate promoted claude-sonnet-4-6 → claude-haiku-4-5 on 6 paired tasks: both arms perfect, parity 1.00, a 66.9% saving on the table. Six easy tasks cannot separate two models, so that promotion was arithmetically inevitable — and when the same route ran the full 116-task panel, the gate refused it. The record was withdrawn; every call row stays in the artifact and its git history.

  1. 6 paired tasks
  2. parity 1.00 → promoted
  3. re-run: 116 tasks
  4. parity 0.85 < 0.97
  5. withdrawn — evidence kept

The same motion, on evidence instead of traffic. No customer routes moved here — this was a claim on our own books, promoted on the best result we had and taken back the moment a fuller panel stopped supporting it. On your estate The Governor runs that loop on live traffic: a person ratifies, your code applies the switch, the revert is armed from the moment it publishes, and quality slipping is what pulls it back — not someone noticing.

n = 6 → 116 paired tasksparity 1.00 → 0.85 against a 0.97 barstanding promotions today: 0is_simulated=false in bothpromoted 2026-07-19 · refused 2026-08-19 · withdrawn 2026-08-19

Source: the committed artifacts provenance/autoroute_savings_frac.json (which carries its own withdrawal note, MAR-515) and provenance/autoroute_defended_savings_frac.json. The original promoted=true record is in git history: we withdraw claims, never evidence.

The gate, ruling on live traffic.

That is what makes acting safe: taking the change back is automatic, not a 3am page. And the ledger of every face-off compounds, so the tenth workflow is cheaper to evaluate than the first.

quality slips → it reverts → the loop re-runsThe Metermeasures every callThe Estatewhat you actually runThe Money Mapwhere the spend goesThe Parity Gateprove it held qualityThe Governoract, with revert armed
  1. The Meter: measures every call, its cost, tokens, substrate, and whether it worked.
  2. The Estate: the org chart of your AI workforce.
  3. The Money Map: where the spend goes, decomposed by step.
  4. The Parity Gate: proves the cheaper route held quality before it ships.
  5. The Governor: acts with a human ratifying, and a revert armed; if quality slips it reverts and the loop re-runs.
A cheaper route, once a person approves it, is held only while quality holds. Today that bar is one we set on your behalf.

Watch the loop run on real spend.

The console shows the loop on open-source agents we metered ourselves: measured spend, recorded face-offs and the gate’s verdicts, every number tagged real or sample. No live route has slipped yet, so the revert has not fired.