What actually happens, day 0 to day 60.
Onboarding is two plugs: wrap your LLM calls with a small SDK, and point an eval adapter at the results your harness already writes. The first read is cost per successful outcome, not per thousand tokens; price the outcome and the same read becomes the return on it.
Then we run one task class down two routes, hundreds of times. If the cheaper one holds quality, you approve the switch and we bank the saving with an automatic revert armed. If it does not, we tell you to keep paying, and that verdict is worth as much as the saving. Fig. 01 is the same loop run on our own estate first.
Four measured events on our own estate, in the order we captured them. No customer appears here; the targets below stay targets, these four are measured.
2026-07-19 · a promote, then withdrawn
withdrawnA promote at parity 1.00 over 6 tasks. Our own gate later withdrew it: six easy tasks cannot separate two arms.
2026-08-19 · the defend verdict
−69.4% refusedThe same route re-run over 116 paired tasks: quality 0.93 → 0.79, parity 0.85 under the 0.97 bar. Keep paying.
2026-09-03 · a metered read
$0.003722 / outcomemargin-oss-seed (aider-edit, crewai-solve, crewai-pipeline-solve) metered end to end: 332 calls, 280 of 306 outcomes landed.
2026-09-19 · the first promotion
−63.7% promotedSame model, lower thinking budget, 115 tasks: quality held (0.89 → 0.94) while cost fell by nearly two thirds — and the paired test over the same tasks let the gate say yes.
is_simulated=falsecaptured 2026-07-19 → 2026-09-19Source: the four committed artifacts in provenance/, read at build time; none is typed, and the withdrawn promote stays in the file it was recorded in.
Time-to-value targets
Both are targets we are building to, not measurements. No design partner has run this end to end yet. When one has, this page will say so, with the real numbers and where they came from.
00 · scope · the week before
It starts with one call, and one definition.
One workflow with a real, rising bill and a repetitive shape: support triage, document extraction, code review. Then the hardest step, saying what an outcome is: cost per outcome is undefined until you do, and that is a judgment about your business, not a config value. One precondition: the passing check has to point back to the run it graded. An eval suite does that for free; an acceptance in a reviewer’s queue that never references the call has no denominator, and wiring that link comes first. The same call asks what one outcome is worth to you. We never guess that number either.
Cost per outcome is the honest default of a value-per-outcome model. No customer has priced an outcome yet, so every figure Margin computes is cost per outcome, and it is labelled as one. Price yours and the meter records it beside the cost, every outcome counted once and unweighted. A value you did not supply is never invented.
01 · instrument · day 0, about 30 min
Thirty minutes of one engineer’s time.
Nothing new to grade and no harness of ours: the adapter reads the results file your harness already writes. The SDK is stdlib-only and fail-safe (an emit never raises into your request path; a row that fails to land is reported, not lost), the source is forced server-side to your key’s project, and no prompts are stored. This stage is plumbing, and we say so.
02 · first read · day 1 to 3
The number you did not have.
The first read is built to be uncomfortable, not flattering: a number you cannot argue with and did not have. Every number carries its provenance, so it is real or it does not render, and it is current to the last metered call, never a claim of always-on monitoring.
03 · first verdict · week 1 to 2
The counterfactual, which is the whole point.
We run the same task class down both routes, the incumbent and a cheaper candidate, enough times to be sure, and score both against the check you froze in Stage 0. That is the answer to how do you know?: the cheaper option was run and measured, not a before-and-after that controls for nothing. Two verdicts, and both are wins.
The cheaper route held parity and costs less: the saving, the trial count, and the parity number it cleared.
The cheaper route did not hold: it would cut cost and cost you a share of your outcomes. Keep paying.
the trust event of the relationship04 · first banked saving · week 2 to 4
A saving your own bill corroborates.
A human approves; the loop never promotes alone. If parity slips below the bar on live traffic, the route reverts on its own and says so. The reconcile against the provider’s bill is the “can I trust this number?” check a finance reviewer runs first, and where most tools have nothing to show.
05 · steady state · month 2 and on
It gets stickier, not stale.
Every face-off is permanent evidence: this task class, this model pair, this verdict, this date. The tenth workflow is cheaper to evaluate than the first, a new cheap model gets answered with a run rather than an opinion, and the frozen checks become your own quality definition, versioned.
What is built and what is not
Built and real
The Python and TypeScript SDKs, the ingest API and per-project keys, the eight eval adapters Stage 01 names, the optional value tag that turns cost per outcome into value per outcome, the four metrics with enforced provenance, the discover-propose-ratify flow, the auto-router with its parity gate and auto-revert, the invoice reconcile, and the console, which you sign up for and sign in to by email.
In scope, not yet shipped
Organizations above a project, and a connected repo that Margin runs for you: the next things we build, and neither moves to the left column until it is real. Billing stays a hand-sent invoice until that actually hurts.
Not built, and a real gap
The architectural lever class is advisory: it cannot auto-revert, so it is never acted on for you. A customer whose waste is mostly architectural gets a smaller number from us than their problem deserves, and hears that from us first.
See the loop run on real spend.
The console shows the loop on open-source agents we metered ourselves: measured spend, recorded face-offs and the gate’s verdicts, every number tagged real or sample.