Per‑token price is the wrong unit.
The industry prices intelligence by the token, so that is what every dashboard measures. But a token is an input. Nobody buys tokens; they buy a resolved ticket, a merged pull request, an answer a human accepts. The only unit that matters is the cost‑per‑outcome: dollars spent divided by outcomes actually produced.
This reframes the whole problem. A model that is twice the price per token but needs a third of the retries can be the cheaper choice. A cheaper model that fails a fifth of the time is the expensive one. The failures are invisible in a token bill and ruinous in an outcome bill. You cannot see any of this until you count outcomes, and outcomes are exactly what per‑token telemetry throws away.
There is a harder version of this argument, and it comes from the seller’s own mathematics. Worked out formally, the optimal way to price a token is marginal cost times a markup: p₃(θ) = m(θ) · c₃. The markup term is computed from a private index of what you will pay, aggregated across your tasks. So a published per‑token price is the provider’s cost welded to what they estimate you will bear, in a single number. No amount of per‑token telemetry separates them, because the separation is not in the data you receive. Reconciling a metered outcome against the invoice does separate them.
Nobody buys tokens. They buy resolved tickets, merged PRs, accepted answers. Price the outcome, not the input.
Measurement is the price of entry, not the product. Everyone will eventually show you a cost‑per‑outcome number. The question that decides whether a tool is a report or a control layer is what happens next.
A dashboard observes. Margin acts.
The gap between a FinOps dashboard and an economic control layer is the difference between a thermometer and a thermostat. Both read the temperature. Only one changes it. Margin runs the full loop:
- Measure: decompose every agentic operation and price it to its real cost‑per‑outcome.
- Recommend: rank the levers by the dollars actually recoverable, biggest bill first.
- Act: apply the change on live traffic, gated on a human ratification.
- Revert: re‑measure the outcome and auto‑back‑out the instant quality dips below the floor.
The last two legs are the entire company. Measuring is table stakes. Acting safely, and being able to prove you can undo the action the moment it stops paying, is what a buyer cannot get anywhere else, and what makes an auto‑tune trustworthy enough to leave on.
The moat is proof, not routing.
Routing a request to a cheaper model is commodity. A dozen gateways do it. The hard, defensible part is proving, on live traffic, per task class, cross‑provider, that the cheaper configuration still holds the real outcome quality. That proof engine stands on three pillars.
Detect the outcome and task class from the trace (a passed test, an accepted answer, a resolved ticket) and ground quality in that real outcome, not a synthetic score. Where the customer already has evals (a coding agent's test suite, or their existing eval framework), use them; where they don't, fall back to reference-free methods.
Never a blind flip. Route a small, statistically-chosen slice to the candidate, measure the real outcome on it, converge to the cost-minimal config that holds parity, and auto-revert on the first dip. A guardrailed live counterfactual, not a shadow guess.
Persist parity results in a task-class-keyed corpus so a brand-new workload starts warm from priors instead of cold. A task class proven ten thousand times transfers a confident prior to customer number one on day one.
Point that same gate the other way and it becomes a migration-safety tool. Providers retire models on their own schedule, so the move to a new one is forced, not chosen. A straight swap is a blind flip: the replacement can clear a benchmark and still shift tone or miss the outcomes you actually ship. Run it through the parity engine first and the swap stops being a leap of faith. Margin proves the new model holds the real outcome before you commit, and reverts on the first dip. Route down to spend less, or route across to migrate without guessing: one gate, both directions.
And Margin says what its proof depends on. A parity result is measured on your harness and your tasks, because a benchmark score is a claim about the harness it ran in, not the model alone: the same model can post a different score when the orchestration around it changes. A number measured in another harness does not carry to your stack, and a router that picks models from a public‑benchmark prior is trusting exactly such a number. We hold the cheaper config to the real outcome on the work you run, and we publish that condition rather than bury it. A limit you can see is one you can check.
Anthropic reaches the same control from the other side. Its managed‑agents cookbook, which splits a workload between a planner and cheap workers, refuses to claim a saving until it has run the realistic alternative: one frontier agent, the same tools, held to the same verification standard, the same question, real bills against real bills. Two teams arriving at that method independently is the tell that the control, not the routing, is where the work is.
Model choice is one lever of many.
“Use a cheaper model” is the obvious move and rarely the biggest one. The act layer is a library of econometric levers, each a safe, gated optimization with a measured cost‑quality tradeoff, each proven at parity before it is applied and each able to auto‑revert:
Each lever is one framework: propose it, prove it holds parity on a real slice, apply it under the gate, keep the revert armed. The science says what to try; the parity engine safely proves it, and only a lever with a shipped benchmark and a proven ledger run is ever applied on its own.
Quality has a diminishing-returns curve.
For a given task class, plot outcome quality against tokens spent, or against reasoning effort, and you get a curve that bends. Below some point, spending less costs you outcomes. Past it, spending more buys nothing but latency and bill. The money is in finding that knee for each task class and living exactly on it.
The optimum is not the cheapest config, nor the smartest. It is the cheapest config that still clears the quality floor: the knee of the curve.
Three different things get called token elasticity, and running them together is why most writing on this produces no decision. They are worth keeping apart:
Market elasticity. Unit price falls, volume rises further, and the received story says the bill rises anyway. That story needs demand elasticity near 2. The only estimate from real marketplace transactions puts it just above 1, which means a 20% price cut leaves total spend about 3% lower. A price cut is usually a small saving, not a trap. When a bill rises after one, the cause is elsewhere: more scope, a shift into reasoning work, thinking tokens you cannot see, or a commitment already signed. Each is separable, and three of the four are yours to change.
Instruction elasticity. Tighten a reasoning budget too far and realized consumption goes up. On one question a 50‑token budget produced 86 output tokens; a 10‑token budget produced 157. Five times tighter, 83% more tokens. The curve is not monotone, so nobody can reason their way to the setting. The model cannot either: asked to estimate its own efficient budget it lands in the right range 61% of the time.
Production elasticity. Outcomes with respect to spend. This is the number worth having, and the one nobody publishes for your workload. The production function is already written down: a standard economics formula for how three inputs — input, output and fine‑tuning tokens — jointly produce value (Cobb‑Douglas), where the exponent on each term is its output elasticity: how much outcome quality moves when that one input moves. The theory names those three parameters and assumes their values. Estimating them on your own traffic is the work.
These per‑task‑class elasticity models are the second moat: the science of which lever to pull and how far. They compound with data. The elasticity engine proposes the move; the parity engine proves it is safe; together they turn a guess into a measured, revertible decision.
And the per‑token number gets less useful as you grow. The larger your account, the closer your marginal token price sits to what it costs the provider to serve you, and their margin moves into the part with no tokens attached to it. Anyone shifted onto a seat fee plus a consumption commitment this spring lived it. You cannot divide a commitment by tokens. You can divide it by outcomes.
It is also why “we moved to a cheaper model and saved money” cannot be checked against an invoice. Unit prices fell all year while bills rose, so both numbers are true at once, and only a decomposition separates a price change from a volume change. A vendor paid to shrink your bill finds the flat region and stops looking; ours is flat by design, so what we are building toward is a slope rather than a saving — and a slope tells you to keep spending about as often as it tells you to stop.
Status: the parity gate above runs today. The full per‑task‑class curve does not. We will not draw a curve through measurements we did not take.
Every safe act is a datapoint no one else has.
Each change Margin makes and its measured result is a labeled row that nobody else can produce, because nobody else runs the safe act‑then‑re‑measure loop that produces the label: (task‑class, lever, config) → (Δcost, Δquality, held?). Persisted into one growing, task‑class‑keyed ledger, this is where the elasticity curves are learned from, not guessed.
There are two network effects, and they compound:
- Depth: per customer, Margin learns each of their flows better every week. Leaving resets that intelligence to zero. That is the switching cost.
- Breadth: a task class seen ten thousand times across customers transfers a warm prior to a brand‑new one, killing cold‑start. Only abstracted task‑class priors travel, never raw prompts, outputs, or customer data, so breadth compounds while privacy holds.
Running open‑source agents to seed this ledger is therefore not a demo. It is moat‑building: customer number one should arrive to a non‑empty corpus of proven priors.
Meet the customer at their eval maturity.
Whether an outcome is easy to ground depends on whether the customer has evals. So Margin detects eval maturity and adapts:
Later‑stage teams with mature eval frameworks. Integrate directly, ground quality in their own ground truth, and optimize deep. High signal, high value.
Earlier teams with thin or no evals. Use reference‑free methods, and surface where their eval coverage is thin so they can grow it, a standalone value they will pay for before trusting any auto‑tune.
Mode B is the top‑of‑funnel wedge; Mode A is the higher‑value expansion. And the two connect: a Mode B customer whose evals grow, partly because Margin helped grow them, matures into a Mode A customer, a deeper integration and a bigger deal. Margin compounds with the customer’s own maturation, and they never outgrow it. This is the eval‑coverage flywheel: land on “trust your quality,” expand into “now safely optimize it.”
A fabricated saving is the worst failure.
The entire product depends on one discipline: every number is a real metered run, a real outcome, a real parity measurement. A saving that looks better than it is destroys the only thing worth selling: the trust to leave the loop on. So the guarantees are absolute: real spend against real APIs, outcomes graded on the agent’s own verdict, parity checked against a threshold before any change is kept, and an auto‑revert that fires the moment quality dips. Simulated data is tagged as simulated, everywhere, always.
That is the framework. The console runs it live on real, recognizable open‑source agents, not a toy estate, so you can watch a real route flip, bank a real saving, see the parity proof, and hit revert yourself.
Six things a labor market needs that a spend dashboard cannot give it.
A roadmap presented as a product is the one failure a company selling honest numbers does not survive, so the figure draws the line rather than burying it in a caveat. Solid runs today; hollow is ratified and not built. The last rung is the furthest out, and the one we are most sure about.
Which agent the rest depend on. A cheap retrieval service four specialists share never appears in a cost ranking, and it is the one whose failure stops everything. Read it off the graph: remove this node, and what share of your outcomes becomes unreachable.
The full account of the bill. Every dollar bought a measured outcome, was waste with a named cause, or we cannot see what it bought. The three sum to your bill. Everyone publishes the first two; the third is what makes them believable.
Which agents to retire. Every recommendation in this category tunes an agent to run cheaper. The larger saving is switching one off. A vendor paid on your token spend will never say that; our fee is flat, on purpose, so we can.
The design-time wind tunnel. Before you build an agent, what it will cost per outcome and what it does to everything up and downstream of it. We give you the economics. You design the agent. We never pick your architecture.
What you can stop watching. How much you can safely delegate is set by the task rather than the model: how cheap it is to check the work, and how cheap it is to undo a mistake. We already measure both.
And then it rebuilds the one it condemned. If we tell you an operation is badly built, the fair question is why you are the one fixing it. So we write the replacement and run it against yours on your own traffic. We never grade our own work: your frozen outcome check does, and our version has to beat the incumbent rather than tie it. You keep the architecture.
See the framework run live.
The console is the thesis in motion: measure → recommend → act → revert, on real agent spend, with the parity proof on every change.