One control loop. Five parts.

Every cost tool tells you what an agent operation spent. Only you know what the outcome was worth: a merged PR, a passed test, a resolved ticket. Price it and each operation reports its return; price nothing and it reports cost per outcome, labelled as one.

We turned down a 69.4% saving. 116 real tasks, a cheaper route at 30.6% of the incumbent’s bill, quality 93% to 79%, parity 0.85 against a 0.97 bar. The gate refused it. Published, tagged is_simulated=false.

quality slips → it reverts → the loop re-runsThe Metermeasures every callThe Estatewhat you actually runThe Money Mapwhere the spend goesThe Parity Gateprove it held qualityThe Governoract, with revert armed
  1. The Meter: measures every call, its cost, tokens, substrate, and whether it worked.
  2. The Estate: the org chart of your AI workforce.
  3. The Money Map: where the spend goes, decomposed by step.
  4. The Parity Gate: proves the cheaper route held quality before it ships.
  5. The Governor: acts with a human ratifying, and a revert armed; if quality slips it reverts and the loop re-runs.
A cheaper route, once a person approves it, is held only while quality holds. Today that bar is one we set on your behalf.

The Meter

Every call, metered: cost, tokens, substrate, and whether it worked, a verdict your own evals supply.

Cost per outcome is the honest default of a value-per-outcome model. No customer has priced an outcome yet, so every figure Margin computes is cost per outcome, and it is labelled as one. Price yours and the meter records it beside the cost, every outcome counted once and unweighted. A value you did not supply is never invented.

The Meter →
one metered call
resolve_ticketclaude-sonnet-4-5
cost
tokens
substrate
outcome
✓ worked
the four fields on every call

The Estate

The org chart of your AI workforce: team, environment, agent, operation.

The Estate →
your AI workforce
  • Support team
    • production
      • triage agent classify
      • reply agent draft
    • staging
      • eval agent score
team → environment → agent → operation

The Money Map

Where the spend actually goes, decomposed by step.

The Money Map →
where the spend goes
Support agentcost / outcome
Research crew
RAG assistant
Batch eval
share of spend, by step · live figures on the console

The Parity Gate

A cheaper route is promoted only while quality holds, and reverted when it slips.

The Parity Gate →
one face-off
incumbent
candidatecheaper
parity vs the bar you set
✓ clears the bar → promote
when it does not clear: hold, keep paying · runs live on the console

The Governor

The action half: the control loop, with a person ratifying every change.

The Governor →
act, with revert armed
gpt-5.4gpt-5.4-mini
✓ ratified revert armed
your code applies the switch; if parity slips, the decision reverts on its own and says so.
a person ratifies · autonomy is a dial

The measuring half is free and open.

pip install margin-cost reads a local corpus of invoices and outcome counts and prints cost per outcome for every vendor, with the retry tax, cache efficiency, token yield, and the coverage tiers that stop any of it claiming more precision than the evidence supports. It measures; it does not act. No account, no backend, and it never phones home, including to us.

What Margin adds: routing a call to a cheaper model, proving the cheaper model held quality first, promoting that route and reverting it when parity slips, continuously. Measuring spend is becoming table stakes; proving a change did not cost you quality is the commercial product.

Watch the loop run on real spend.

The console shows the loop on open-source agents we metered ourselves: measured spend, recorded face-offs and the gate’s verdicts, every number tagged real or sample. No live route has slipped yet, so the revert has not fired.