Neutral · cross-provider · sells no tokens

Labor economics for your AI workforce.

You hired a workforce made of agents and nobody is measuring what any one of them produces. Margin prices every operation against cost‑per‑outcome: a merged PR, a passed test, a resolved ticket. It proves the fix before it ships.

No signup. The demo drops you straight into the live console.
example · a support pipeline at 240,000 runs / monththe workforce
intakeentry stageretrieveshareddraftspecialistreviewspecialistverifyspecialistsummarisespecialistjudgegradershipoutcome
the workforce
intakeentry stage
retrieveshared
draftspecialist
reviewspecialist
verifyspecialist
summarisespecialist
judgegrader
wage bill19.7¢ / outcome
wage bill19.7¢ / outcome
Agents are workers, and your bill is a wage bill. Eight operations across four vendors. Nobody can tell you what any one of them produced for what it cost.

The method, held to a real bar.

No inflated totals. These four are the whole method: the real face‑offs we have run, the trials each has to clear, the parity bar that promotes a route, and the flag every figure on this page carries.

Route face‑offs run

Real, head‑to‑head — is_simulated=false.

Trials each must clear
cleared

A real power floor, before anything counts as “proven.”

Parity bar to promote
enforced

And to auto‑revert — one line, both ways.

Every figure carries
is_simulated=false

Each reproduces from a committed artifact.

Our own estate, metered.

We meter our own work first: what a cheaper route has earned back at proven parity, never a projection, and the live console read behind it, refreshed when the page loads.

Margin has removed, at measured parity
parity = the cheaper model ran the same tasks and still landed the job
the counterfactual saving lands here the moment a cheaper route holds parity
Measured flows
Spend under management
the real metered base
Time to first proven saving
not yet
first call → first proven route
Counterfactual, not banked cash: real metered spend × the proven‑at‑parity rate. The curve is the demonstrated‑on‑benchmark track record. Parity‑gated: a flow with no proof contributes $0, never a fabricated win.
margin · console · the estateREAL · measured
Real metered calls
Blended cost / outcome
Proven auto-route saving
ProjectMetered callsOutcomesLast seen
Every row is a measured run, never a seeded one.See it live

What would this be at your scale?

Your spend at the blended rate from the sample estate above, where one flow routes and the rest hold. Not a best‑case single flow.

Reading the sample estate…

Per‑token price is the wrong number.

Tokens are not the unit of value. A merged PR is. Margin prices work against the only honest denominator: cost‑per‑outcome.

ranked by $ / token vs. $ / outcome
ranked by $ / tokenranked by $ / outcomewhat every console showswhat you actually paysmallmidfrontiermidfrontiersmall3.4 attempts × 1.9 the contextcheapest per token, dearest per outcome
ranked by $ / token
1. small
2. mid
3. frontier
ranked by $ / outcome
1. mid
2. frontier
3. small3.4 attempts × 1.9 context
So which one is cheapest? Price per token puts the small model first; price per landed outcome puts it last. Retries are invisible on the left and all on the invoice. A mechanism, not your estate.

You do not have to take our word for the gap.

Rippling ran about 2,100 scored runs per model over their own payroll and personnel data, then published it. Seven models finished within one point of each other on pass rate. The bills ranged from $621 to $4,359.

We thought token efficiency was the driver of cost, but really it’s just the cost per token. Given the constant shifts in both variables, we have to just track dollars.

Matt MacInnis, President and Chief Product Officer, Rippling. “RipplingBench: A comparison of open‑weight and frontier models in real‑world use cases”, .

Rippling is not a customer and has not evaluated Margin. They measured the problem on their own systems and reached the same conclusion we build on: at equal quality, the price you pay is a choice, and per‑token telemetry will not tell you which choice you made.

Picking a cheaper model is the easy part. Knowing where it breaks is not.

Every provider publishes what the small model costs. Nobody publishes which of your tasks survive on it. We run both on your own work and score how often the cheaper one still landed the job — that score is parity, and a route is adopted only above the floor we set. Parity is measured under your harness, on your tasks. A benchmark score is a claim about the harness it ran in, not the model alone, so a number measured in another harness does not carry to yours.

It points the other way just as often. We only tell you to keep spending after running the cheaper option and watching it lose. You validate the wage rather than cutting it.

isoquant · cost mix vs. the outcome it buys
small-model callsfrontier callsbudget too smallsame outcomecheapest mix that holds paritycheaper — and short of the outcome
anywhere on the curvethe same outcome
where the budget line touches it4 small : 1 frontier
below the curvecheaper, and short
Buy the mix, not the model. Every point on the curve buys the same outcome; the cheapest is where the budget line touches it. Below the curve you pay less and get less, and that line is our parity gate. Drawn from the algebra — your own mix is measured.
The closed loop · livenot reached

A router optimizes one stack. Margin proves the whole system, then acts.

Every router will tell you it cut your bill. Ask how it knows. One picking from a public benchmark never ran what it replaced. We run both, and say what our number is conditional on.

the system, not the stack

You budget for the boxes. You pay for the edges. One supervisor sits between every hand-off, so each worker you add is two more edges — and the tokens live on the edges, not in the boxes.

intakechat / email / voiceinput guardPII scrub · injection filterextractintent · order · plantriagesentiment · urgencysupervisorroutes every hand-offknowledgeRAG · vector storeorder lookupCRM · APIpolicy checkeligibility · rulesaccount actionrefund · cancelcomposedraft the replyoutput guardhallucination · toneverifyQA · sign-offsendresolvedescalateto a humanconversation memoryread at every hopa doomed run, stopped here — never billed downstreammargin · upstream validation gateevery hand-off round-trips the plannermargin · coordinatorthe full context, re-sent every hopmargin · context compressionthe same retrieval, paid for twicemargin · prompt cachingrework that lands no new outcomemargin · retry budget

A support pipeline, drawn to scale. Every worker is one box and two edges through the supervisor; the waste hides on the edges — re-sent context, rework that changes nothing, the round-trip through the planner. That is what Margin meters, and where it acts.

measure → recommend → act → revert
MeasureRecommendActparity clears the floor?yesbankedno, or parity slips later — auto-reverts to the baseline
1. measureevery operation, per outcome
2. recommendthe dominant bill
3. actonly past the gate
the gate
parity clears the floorbanked
parity slips, now or laterauto-reverts to the baseline
The arc back is the product. Anything can recommend a cheaper model. A recommendation you have to police is not automation, and a saving nobody re-checks is not a saving.
Usage consoles & prompt routers
  • Spend by day and by model
  • One stack, one vendor
  • The saving is asserted
vs
Margin: the closed loop
  • Cost per outcome, per operation
  • The whole chain, every vendor
  • Runs what it replaced
  • Auto-reverts if quality dips

Three lines, and the meter is running.

Point the meter at the OpenAI, Anthropic or Gemini SDK an agent already calls. Right: a real workflow from our estate, as recorded.

jobscraper-fit-scoring.ts
# TypeScript — Python: pip install margin-meter
$ npm install margin-meter

import { MarginMeter, instrumentGemini }
  from "margin-meter";

const meter = new MarginMeter();
instrumentGemini(model, meter, {
  workflowId: "jobscraper-fit-scoring",
});
// every generateContent() call is now metered
the meter’s readmeasured, not simulated
{
  "workflow": "jobscraper-fit-scoring",
  "model": "gemini-2.5-flash",
  "calls": 5,
  "outcomes": { "passed": 5, "total": 5 },
  "total_cost_usd": 0.004875,
  "cost_per_outcome_usd": 0.000975,
  "is_simulated": false
}
Five calls, five outcomes passed, captured 2026‑07‑18. Every figure reproduces from a committed artifact the provenance test guards.

The supply‑chain map at a buyer’s scale, and the honest read on our own books.

One card is a company at your scale with the bottleneck lit, a sample and labelled as one. The other is our own books, every figure measured.

a company at your scaleSample · simulated
The estate bill
The estate bill
Four agent systems priced into one estate bill. A labelled sample, not a real customer.
on our own factorynot reached
the real estate bill we run Margin against
Real calls
Outcomes
Cost / outcome

Small on purpose, and measured rather than seeded.

The economics of intelligence, written down.

The three elasticities people run together, how the parity engine proves a cheaper config holds, and where this goes next.

Read the write-up

Margin never sees your prompts.

A meter, not a proxy. We cannot read a request, and our ingest going down cannot stop your agents.

what crosses the wire to Margin
your agentyour provideryour request and its responseunbroken — Margin is not on this pathMarginusage numbers onlyon the wirenever on the wireworkflow nameprovidermodelinput tokensoutput tokenslatencypass or failcomputed costprompt textmodel outputtool argumentsyour users' dataplus a link you choose to attach, or nothing
your call
your agent → your providerdirect; Margin is not on it
on the wire, to Margin
workflow name
provider
model
input tokens
output tokens
latency
pass or fail
computed cost
never on the wire
prompt text
model output
tool arguments
your users' data
plus a link you choose to attach, or nothing
Margin sits beside the call rather than in front of it: we receive the usage block your provider already returns. The fields on the right are not redacted. There is no field for them.

margin-meter is public on PyPI and npm. Read it before you install it.

Don’t take our word for the numbers. A test fails the build if they slip.

We hold no compliance badge we haven’t earned and show no logo we can’t name. What we can show is the guardrail: five checks that run on every commit, each one failing closed, so a number that cannot prove itself never reaches this page, starting with the outcome the whole metric divides by.

how a grade earns its proof strength
Proposereads the code · probabilisticRatifyhuman sign-off · onceMeasurefrozen check · every runand every grade says how it was provenground trutha test passed, an exit codereference-freea judge, labelled as the softer tierunprovableskipped — never counted as a pass
1. Proposereads the code · probabilistic
2. Ratifyhuman sign-off · once
3. Measurefrozen check · every run
and every grade says how it was proven
ground trutha test passed, an exit code
reference-freea judge, labelled as the softer tier
unprovableskipped — never counted as a pass
Probabilistic to find it, deterministic to score it. Nothing grades your spend until a person agrees what “done” means. A run we cannot grade is skipped, never a pass — that is the denominator, so inflating it fakes everything above.
provenance check

The headline savings must reproduce from a committed artifact tagged is_simulated=false. Read a metric that is missing or simulated and the check fails closed, rather than passing on a blank.

auto-route adversarial check

A route is promoted only once parity clears a high bar we hold ourselves to and the candidate is cheaper, and it auto-reverts the moment parity slips back below that line.

regression gate

A build that cut cost by dropping its pass-rate fails, however large the saving. Quality is settled before economics, never the other way around.

pricing accuracy check

Per-token math is checked against published pricing, with cached reads billed at their real fraction, so a saving cannot be inflated by a friendly rounding choice.

outcome integrity check

An outcome has to be positively determined. An ambiguous or empty signal raises, instead of being quietly counted as a pass and fabricating the very number cost‑per‑outcome divides by.

These are floors, and a floor moves one way. Adding a stricter case is welcome; relaxing one is blocked at commit time. When a floor does have to move, we change it in the open with a human sign‑off, never silently.

Get your cost‑per‑outcome read.

You get the founder, not a sales rep. We’ll show you the one route worth flipping, with its parity proof. Nothing is gated.

One email when early access opens. No drip sequence, no newsletter.

Or write to subh@trymargin.io. A person replies, usually the same day.