What is the argument?
Per‑token price is the wrong unit.
The industry prices intelligence by the token, so that is what every dashboard measures. But a token is an input. Nobody buys tokens; they buy a resolved ticket, a merged pull request, an answer a human accepts. The first unit that measures anything you actually bought is the cost‑per‑outcome: dollars spent divided by outcomes actually produced.
A model that is twice the price per token but needs a third of the retries can be the cheaper choice; a cheaper model that fails a fifth of the time is the expensive one. None of this is new economics. It is ordinary price theory, the tradition Chicago is known for: you buy the output, the inputs are substitutes, and you choose by what the next unit of output costs. A token bill prices an input and leaves that arithmetic to you.
There is a harder version of this argument, and it comes from the seller’s own mathematics. Worked out formally, the optimal way to price a token is marginal cost times a markup: p₃(θ) = m(θ) · c₃. The markup term is computed from a private index of what you will pay, aggregated across your tasks. So a published per‑token price is the provider’s cost welded to what they estimate you will bear, and no per‑token telemetry can separate them. Reconciling a metered outcome against the invoice can.
a16z reached the same conclusion from the seller’s side in August 2026: token pricing anchors a product to a cost curve that keeps falling
. Asked how they would rather be billed, 27 of the 50 technical AI buyers they surveyed chose a unit tied to recognisable work; 14 chose tokens. A stated preference from someone else’s sample, not a measurement of what anyone pays.
Nobody buys tokens. They buy resolved tickets, merged PRs, accepted answers. Price the outcome, not the input.
Cost‑per‑outcome is not the last unit either. Its denominator is a count. A 99%‑accurate extraction and an 85%‑accurate one each score as one outcome. So does an outcome worth five hundred of another. And a wrong support answer that loses the account scores as a zero, when what it was worth was negative. The number tells you which route is cheaper per finished job. It can never tell you that a costlier route earned its bill.
That is the difference between minimising and optimising, and it is the difference Margin exists for. Minimising cost per outcome has one answer; price an outcome and there is a frontier instead, and where you sit on it depends on what one bad answer reaching a customer costs you. Nobody outside your company knows that number, and today neither do we. The quality bar our gate holds a route to is a constant we chose, because no customer has yet told us what one of their outcomes is worth. Until one does, the bar stays ours.
Cost per outcome is the honest default of a value-per-outcome model. No customer has priced an outcome yet, so every figure Margin computes is cost per outcome, and it is labelled as one. Price yours and the meter records it beside the cost, every outcome counted once and unweighted. A value you did not supply is never invented.
Measurement is the price of entry, not the product. Everyone will eventually show you a cost‑per‑outcome number. The question that decides whether a tool is a report or a control layer is what happens next.
A dashboard observes. Margin acts.
The gap between a FinOps dashboard and an economic control layer is the difference between a thermometer and a thermostat. Both read the temperature. Only one changes it. Margin runs the full loop:
- Measure: decompose every agentic operation and price it to its real cost‑per‑outcome.
- Recommend: rank the levers by the dollars actually recoverable, biggest bill first.
- Act: apply the change on live traffic, gated on a human ratification.
- Revert: re‑measure the outcome and auto‑back‑out the instant quality dips below the floor.
The last two legs are the entire company. Measuring is table stakes; acting, and proving you can undo the action the moment it stops paying, is what makes an auto‑tune trustworthy enough to leave on. Chicago did it at city scale in 1900, reversing a river that carried its sewage into its own drinking water. The revert leg is the same move made small.
Does quality hold?
The moat is proof, not routing.
Routing a request to a cheaper model is commodity. A dozen gateways do it. The hard, defensible part is proving, on live traffic, per task class, cross‑provider, that the cheaper configuration still holds the real outcome quality. That proof engine stands on three pillars.
Detect the outcome and task class from the trace (a passed test, an accepted answer, a resolved ticket) and ground quality in that real outcome, not a synthetic score. Where the customer already has evals (a coding agent's test suite, or their existing eval framework), use them; where they don't, fall back to reference-free methods.
Never a blind flip. Route a small, statistically-chosen slice to the candidate, measure the real outcome on it, converge to the cost-minimal config that holds parity, and auto-revert on the first dip. A guardrailed live counterfactual, not a shadow guess.
Persist parity results in a task-class-keyed corpus so a brand-new workload starts warm from priors instead of cold. A task class proven ten thousand times transfers a confident prior to customer number one on day one.
Point that same gate the other way and it becomes a migration-safety tool. Providers retire models on their own schedule, so the move to a new one is forced, not chosen. A straight swap is a blind flip: the replacement can clear a benchmark and still miss the outcomes you ship. Run it through the parity engine first and Margin proves the new model holds the real outcome before you commit, and reverts on the first dip. Route down to spend less, or route across to migrate without guessing: one gate, both directions.
And Margin says what its proof depends on. A parity result is measured on your harness and your tasks, because a benchmark score is a claim about the harness it ran in, not the model alone: the same model can post a different score when the orchestration around it changes. A router that picks models from a public‑benchmark prior is trusting exactly such a number. We hold the cheaper config to the real outcome on the work you run, and we publish that condition rather than bury it.
Anthropic reaches the same control from the other side. Its managed‑agents cookbook, which splits a workload between a planner and cheap workers, refuses to claim a saving until it has run the realistic alternative: one frontier agent, the same tools, held to the same verification standard, the same question, real bills against real bills. Two teams arriving at that method independently is the tell that the control, not the routing, is where the work is.
Down a chain, quality compounds.
Cost adds; quality multiplies. Five steps, each holding our 0.97 bar, deliver 0.859 to your user; for the chain to deliver 0.97, each step would have to hold 0.9939. That is the identity 0.97n, not a measurement of your pipeline: it assumes each step’s deviation is independent, so a validating step does better and steps that fail together do worse. Margin advises per step, at the bar your chain’s length demands.
What can it change, and what does each change buy?
Model choice is one lever of many.
“Use a cheaper model” is the obvious move and rarely the biggest one. The act layer is a library of econometric levers, each a safe, gated optimization with a measured cost‑quality tradeoff, each proven at parity before it is applied and each able to auto‑revert:
Each lever is one framework: propose it, prove it holds parity on a real slice, apply it under the gate, keep the revert armed. The science says what to try; the parity engine safely proves it, and only a lever with a shipped benchmark and a proven ledger run is ever applied on its own.
One chain is a wage slip. The estate is a payroll ledger nobody can read by eye.
The catalog above reads like a checklist, and on a single workflow it is one: a few steps, a lever each, an afternoon’s audit. An estate is not a single workflow. Grow one to working size and the graph outruns anyone’s ability to trace it by hand, and the spend follows the structure, not the price list.
You budget for the boxes. You pay for the edges. One supervisor sits between every hand-off, so each worker you add is two more edges — and the tokens live on the edges, not in the boxes.
A support pipeline, drawn to scale. Every worker is one box and two edges through the supervisor; the waste hides on the edges — re-sent context, rework that changes nothing, the round-trip through the planner. That is what Margin meters, and where it acts.
The diagram carries no numbers, by rule: there is no customer behind it, so any figure on it would be an invented one. The callouts name levers from the library above, placed where estates leak. A router prices one stack and stops there; the unit of proof is the whole graph.
Quality has a diminishing-returns curve.
For a given task class, plot outcome quality against tokens spent, or against reasoning effort, and you get a curve that bends. Below some point, spending less costs you outcomes. Past it, spending more buys nothing but latency and bill. The money is in finding that knee for each task class and living exactly on it.
The optimum is not the cheapest config, nor the smartest. It is the cheapest config that still clears the quality floor: the knee of the curve.
Three different things get called token elasticity, and running them together is why most writing on this produces no decision. They are worth keeping apart:
Market elasticity. Unit price falls, volume rises further, and the received story says the bill rises anyway. That story needs demand elasticity near 2. The only estimate from real marketplace transactions puts it just above 1, which means a 20% price cut leaves total spend about 3% lower. When a bill rises after a cut, the cause is elsewhere: more scope, a shift into reasoning work, thinking tokens you cannot see, or a commitment already signed. Three of the four are yours to change.
Instruction elasticity. Tighten a reasoning budget too far and realized consumption goes up. On one question a 50‑token budget produced 86 output tokens; a 10‑token budget produced 157. Five times tighter, 83% more tokens. The curve is not monotone, so nobody can reason their way to the setting. The model cannot either: asked to estimate its own efficient budget it lands in the right range 61% of the time.
Production elasticity. Outcomes with respect to spend. Nobody publishes it for your workload, but the production function is already written down: a standard economics formula for how three inputs — input, output and fine‑tuning tokens — jointly produce value (Cobb‑Douglas), where the exponent on each term is its output elasticity: how much outcome quality moves when that one input moves. The theory assumes those three values; estimating them on your own traffic is the work.
These per‑task‑class elasticity models are the second moat: the science of which lever to pull and how far. They compound with data. The elasticity engine proposes the move; the parity engine proves it is safe; together they turn a guess into a measured, revertible decision.
And the per‑token number gets less useful as you grow. The larger your account, the closer your marginal token price sits to what it costs the provider to serve you, and their margin moves into the part with no tokens attached to it. Anyone shifted onto a seat fee plus a consumption commitment this spring lived it. You cannot divide a commitment by tokens. You can divide it by outcomes.
It is also why “we moved to a cheaper model and saved money” cannot be checked against an invoice. Unit prices fell all year while bills rose, so both are true at once, and only a decomposition separates a price change from a volume change. Our fee is flat by design, so what we build toward is a slope rather than a saving, and a slope tells you to keep spending about as often as to stop.
Status: the parity gate above runs today. The full per‑task‑class curve does not. We will not draw a curve through measurements we did not take.
How is a saving banked, and what is promised?
Every safe act is a datapoint no one else has.
Each change Margin makes and its measured result is a labeled row that nobody else can produce, because nobody else runs the safe act‑then‑re‑measure loop that produces the label: (task‑class, lever, config) → (Δcost, Δquality, held?). Persisted into one growing, task‑class‑keyed ledger, this is where the elasticity curves are learned from, not guessed.
There are two network effects, and they compound:
- Depth: per customer, Margin learns each of their flows better every week. Leaving resets that intelligence to zero. That is the switching cost.
- Breadth: a task class seen ten thousand times across customers transfers a warm prior to a brand‑new one, killing cold‑start. Only abstracted task‑class priors travel, never raw prompts, outputs, or customer data, so breadth compounds while privacy holds.
Running open‑source agents to seed this ledger is therefore not a demo. It is moat‑building: customer number one should arrive to a non‑empty corpus of proven priors.
Meet the customer at their eval maturity.
Whether an outcome is easy to ground depends on whether the customer has evals. So Margin detects eval maturity and adapts:
Mode B is the top‑of‑funnel wedge; Mode A is the higher‑value expansion. And the two connect: a Mode B customer whose evals grow, partly because Margin helped grow them, matures into a Mode A customer, a deeper integration and a bigger deal. Margin compounds with the customer’s own maturation, and they never outgrow it. This is the eval‑coverage flywheel: land on “trust your quality,” expand into “now safely optimize it.”
Both modes assume one thing: the pass or fail has to attach to the run it judges. An eval framework does that by construction. A recorded human acceptance can fail it: a reviewer clears work in a queue that never links back to the run, so there is a decision and no denominator, and reference‑free grounding cannot supply one. Building that link is engineering that comes first, not a setting you flip.
A fabricated saving is the worst failure.
The entire product depends on one discipline: every number is a real metered run, a real outcome, a real parity measurement. A saving that looks better than it is destroys the only thing worth selling: the trust to leave the loop on. So: real spend against real APIs, parity checked against a threshold before any change is kept, an auto‑revert that fires the moment quality dips, and simulated data tagged as simulated, everywhere. And the gate has already fired against our own interest: a route 69.4% cheaper over 116 paired real tasks was refused — quality fell to 79% from 93%. The task‑by‑task read, dated and artifact‑bound, is on the parity gate page.
That is the framework. The console runs it live on real open‑source agents, not a toy estate — watch a route face off, read the parity proof, hit revert yourself.
Cost per outcome is one axis. We name the rest.
Optimising one measured number while the others go unwatched is a known hazard. Holmström and Milgrom (1991): when work has several dimensions and you can only measure some, a strong incentive on the measured one pulls effort off the rest. Cost per outcome is a piece rate on the one axis we can see, so it applies to us with no translation; we name the axes a cut could hurt, and which of them we actually watch.
An axis we do not watch reads “not watched”, never “fine”. Reporting an unwatched axis as passing would be the same fail‑closed breach as reading a missing number as real.
We report the set; we do not silently widen the gate. Only latency refuses a promotion today. Gating every axis would make the system unpromotable, so the order is measurement and disclosure first, one guardrail second — the same discipline the floor tests hold for the number itself.
Where does this end up?
Six things a labor market needs that a spend dashboard cannot give it.
A roadmap presented as a product is the one failure a company selling honest numbers does not survive, so the figure draws the line rather than burying it in a caveat. Solid runs today; hollow is ratified and not built. The last rung is the furthest out, and the one we are most sure about.
Which agent the rest depend on. A cheap retrieval service four specialists share never appears in a cost ranking, and it is the one whose failure stops everything. Read it off the graph: remove this node, and what share of your outcomes becomes unreachable.
The full account of the bill. Every dollar bought a measured outcome, was waste with a named cause, or we cannot see what it bought. The three sum to your bill. Everyone publishes the first two; the third is what makes them believable.
Which agents to retire. Every recommendation in this category tunes an agent to run cheaper. The larger saving is switching one off. A vendor paid on your token spend will never say that; our fee is flat, on purpose, so we can.
The design-time wind tunnel. Before you build an agent, what it will cost per outcome and what it does to everything up and downstream of it. We give you the economics. You design the agent. We never pick your architecture.
What you can stop watching. How much you can safely delegate is set by the task rather than the model: how cheap it is to check the work, and how cheap it is to undo a mistake. We already measure both.
And then it rebuilds the one it condemned. If we tell you an operation is badly built, the fair question is why you are the one fixing it. So we write the replacement and run it against yours on your own traffic. We never grade our own work: your frozen outcome check does, and our version has to beat the incumbent rather than tie it. You keep the architecture.
See the framework run live.
The console is the thesis in motion: measure → recommend → act → revert, on real agent spend, with the parity proof on every change.
or email subh@trymargin.io directly. Email is the door on purpose: it reaches the founder, not a form feeding a queue.