Where Margin fits, and where it doesn’t.
Three tools get mistaken for what we do: a router, a spend dashboard, and the version your team could build. Each is good at something real, and each section below says so first. No competitor is named; every claim is about a category.
- A router
- Picks a model inside a lone call.
- A spend dashboard
- Sums whole invoices after the fact.
- Margin
- Holds effort and outcome side by side, which is what judging either requires.
Scope is the product. Inside a lone call, that route is a 69.4% win; on an invoice, a smaller bill. Only the scope that holds the call beside what it bought could see the quality go, and refuse.
is_simulated=falsecaptured 2026-08-19Source: the committed artifact provenance/autoroute_defended_savings_frac.json, read at build time. The diagram itself carries no values.
The one thing none of the three does: prove a cheaper configuration held the outcome, then undo it when it stops. Quality scoring is bundled into tools you may already pay for, and one of them now files the ticket on a regressing answer; none of them chooses the route, prices that choice against the bill, and holds the saving to a proven bar.
What each category does
| The concern | A routergateway · model picker | A spend dashboardFinOps · cost tracking | Margineconomic control layer |
|---|---|---|---|
| Picking the model | covers you. Chooses a model per call, usually to cut cost, and moves on. | leaves it to you. Sees which models were billed. Picks none. | covers you. Chooses per task class, and only where the gate has proved the swap holds. |
| The invoice | leaves it to you. Sees only the traffic it routes. The rest of the bill is elsewhere. | covers you. Every dollar across every vendor, with budgets and chargeback. | partly, with a caveat. Reads it. Does not replace it. |
| What the money bought | leaves it to you. Cost per token, per call. What the call returned is out of frame. | leaves it to you. Cost per vendor, per month. A merged ticket and a timeout cost the same. | covers you. Cost per outcome: the call and what it landed, side by side. |
| After the switch | leaves it to you. Assumes the cheaper model held. Nobody measures. | leaves it to you. Cannot see quality. It sees a smaller line item. | covers you. Parity measured on your tasks, held to our bar, before a saving is banked. |
| When quality slips | leaves it to you. Stays on the cheaper model until someone notices. | leaves it to you. Reports the saving. The regression is on somebody else's chart. | partly, with a caveat. Built to revert on its own and withdraw the saving. No live route has slipped yet, so it has not fired. |
| Who sells you tokens | partly, with a caveat. Many resell inference at a markup. | covers you. Nobody. It sells the view of the bill, not the inference. | covers you. Nobody. Neutral by design: we sell no inference. |
vs a router
A router picks the model. It never checks the switch was worth it.
A router chooses a model per request: a real job, and one move inside our loop. It leaves out what comes after the switch.
A router is a lever. Margin is the gate around the lever, and the revert when it stops paying.
A router is the right tool if
- Per-call model selection, or failover when a provider goes down, is all you need.
- You already trust the smaller model on your own traffic.
- Nobody will ask later whether quality survived the swap.
Choose Margin when
- The cheaper model should win only where it provably clears a parity floor, measured on your tasks.
- The swap has to undo itself the moment quality drops below the floor.
- Every banked dollar needs a committed artifact anyone can reproduce.
vs a spend dashboard
A spend dashboard totals the invoice. It can't tell you what the money bought.
FinOps tools put every dollar in one place; we read that data rather than compete with it. What a total cannot give you is the denominator: what those dollars produced, and which line to cut without losing one. A dashboard reports. It has no safe way to act.
Cost per outcome is the honest default of a value-per-outcome model. No customer has priced an outcome yet, so every figure Margin computes is cost per outcome, and it is labelled as one. Price yours and the meter records it beside the cost, every outcome counted once and unweighted. A value you did not supply is never invented.
FinOps owns the invoice. We turn the AI slice of it into a decision. Same data, one layer up.
A dashboard covers you if
- One picture of every dollar across cloud, SaaS and AI vendors is the goal.
- Budgets, chargeback and finance-grade reporting matter most.
- The question is what left the bank, never what it returned.
Add Margin when
- The call you face needs cost per outcome, not cost per vendor.
- You want the single highest-recovery change surfaced and proven before anyone touches it.
- Acting on the AI slice, safely and reversibly, is the point, not another chart.
vs building it in-house
You can build cost-per-outcome tracking. Trusting it is the hard part.
Any strong team can wire up per-call metering in a sprint; the measurement is not the moat. The act is, and the ways it goes wrong are not obvious until they have happened to you. On 6 paired tasks our own gate measured parity 1.00 and promoted a cheaper model. The same route on 116 tasks measured 0.85, under the 0.97 floor, and we withdrew the promotion. 6 easy tasks cannot separate two models, and a harness that stops at the first green never learns it shipped a regression. The take-back is on the record.
Measuring is a weekend. Trusting the measurement enough to act on it, and to undo the action, is the year.
Build it yourself if
- One workflow, one model, and a spare engineer describe your situation.
- Spend is too small to justify a tool, and building appeals more than buying.
- Someone will own the statistics, not just the plumbing, next year too.
What a home-built harness gets wrong
- A panel too small to decide. Parity on six tasks is the absence of evidence, and it reads identical to a pass.
- Call sites nobody instrumented. On a real bill, 3 of 6 sites were unmetered and carried all of the cache-write spend.
- A price table that drifts. One output rate moved 2.4× in a day; a hardcoded table takes every past number with it.
- A cold start on every task class, forever. The one you cannot build: a prior learned from traffic that is not yours.
See the comparison run on real spend.
The console shows the loop on open-source agents we metered ourselves: measured spend, recorded face-offs and the gate’s verdicts, every number tagged real or simulated.
or email subh@trymargin.io directly; it reaches the founder, not a queue.