Margin
The framework · a working paper

The economics of intelligence, written down.

Every AI‑ops tool can tell you what you spent. None of them can tell you whether you got your money’s worth, or safely spend less without losing it. This is the thesis behind Margin: measure cost‑per‑outcome, close the loop from measurement to action, and prove the quality held before you keep the saving.

01 · The wrong unit

Per‑token price is the wrong unit.

The industry prices intelligence by the token, so that is what every dashboard measures. But a token is an input. Nobody buys tokens; they buy a resolved ticket, a merged pull request, an answer a human accepts. The only unit that matters is the cost‑per‑outcome: dollars spent divided by outcomes actually produced.

This reframes the whole problem. A model that is twice the price per token but needs a third of the retries can be the cheaper choice. A cheaper model that fails a fifth of the time is the expensive one. The failures are invisible in a token bill and ruinous in an outcome bill. You cannot see any of this until you count outcomes, and outcomes are exactly what per‑token telemetry throws away.

There is a harder version of this argument, and it comes from the seller’s own mathematics. Worked out formally, the optimal way to price a token is marginal cost times a markup: p₃(θ) = m(θ) · c₃. The markup term is computed from a private index of what you will pay, aggregated across your tasks. So a published per‑token price is the provider’s cost welded to what they estimate you will bear, in a single number. No amount of per‑token telemetry separates them, because the separation is not in the data you receive. Reconciling a metered outcome against the invoice does separate them.

Bergemann, Bonatti & Smolin, Economics of Large Language Models, ACM EC ’25 (arXiv:2502.07736). Their result is a seller’s pricing mechanism. We read it from the buyer’s side.

Nobody buys tokens. They buy resolved tickets, merged PRs, accepted answers. Price the outcome, not the input.

Measurement is the price of entry, not the product. Everyone will eventually show you a cost‑per‑outcome number. The question that decides whether a tool is a report or a control layer is what happens next.

02 · The closed loop

A dashboard observes. Margin acts.

The gap between a FinOps dashboard and an economic control layer is the difference between a thermometer and a thermostat. Both read the temperature. Only one changes it. Margin runs the full loop:

  1. Measure: decompose every agentic operation and price it to its real cost‑per‑outcome.
  2. Recommend: rank the levers by the dollars actually recoverable, biggest bill first.
  3. Act: apply the change on live traffic, gated on a human ratification.
  4. Revert: re‑measure the outcome and auto‑back‑out the instant quality dips below the floor.

The last two legs are the entire company. Measuring is table stakes. Acting safely, and being able to prove you can undo the action the moment it stops paying, is what a buyer cannot get anywhere else, and what makes an auto‑tune trustworthy enough to leave on.

03 · The parity engine

The moat is proof, not routing.

Routing a request to a cheaper model is commodity. A dozen gateways do it. The hard, defensible part is proving, on live traffic, per task class, cross‑provider, that the cheaper configuration still holds the real outcome quality. That proof engine stands on three pillars.

P1
Reference-free, outcome-grounded quality

Detect the outcome and task class from the trace (a passed test, an accepted answer, a resolved ticket) and ground quality in that real outcome, not a synthetic score. Where the customer already has evals (a coding agent's test suite, or their existing eval framework), use them; where they don't, fall back to reference-free methods.

P2
Safe online optimization

Never a blind flip. Route a small, statistically-chosen slice to the candidate, measure the real outcome on it, converge to the cost-minimal config that holds parity, and auto-revert on the first dip. A guardrailed live counterfactual, not a shadow guess.

P3
Cross-customer parity transfer

Persist parity results in a task-class-keyed corpus so a brand-new workload starts warm from priors instead of cold. A task class proven ten thousand times transfers a confident prior to customer number one on day one.

Point that same gate the other way and it becomes a migration-safety tool. Providers retire models on their own schedule, so the move to a new one is forced, not chosen. A straight swap is a blind flip: the replacement can clear a benchmark and still shift tone or miss the outcomes you actually ship. Run it through the parity engine first and the swap stops being a leap of faith. Margin proves the new model holds the real outcome before you commit, and reverts on the first dip. Route down to spend less, or route across to migrate without guessing: one gate, both directions.

And Margin says what its proof depends on. A parity result is measured on your harness and your tasks, because a benchmark score is a claim about the harness it ran in, not the model alone: the same model can post a different score when the orchestration around it changes. A number measured in another harness does not carry to your stack, and a router that picks models from a public‑benchmark prior is trusting exactly such a number. We hold the cheaper config to the real outcome on the work you run, and we publish that condition rather than bury it. A limit you can see is one you can check.

one gate, both directions
Route downa cheaper modelRoute acrossa forced migrationholds parity?our bar · live sliceholdsAdopt itproven at paritydipsReverton the first dip
route downa cheaper model
route acrossa forced migration
one gate
holds parity at our bar, on a live sliceadopt — proven at parity
parity dips, now or laterrevert — on the first dip
One gate, both directions. Route down to spend less, or across to survive a model you did not choose to move to. The same proof runs either way, and only a candidate that holds the real outcome is kept.

Anthropic reaches the same control from the other side. Its managed‑agents cookbook, which splits a workload between a planner and cheap workers, refuses to claim a saving until it has run the realistic alternative: one frontier agent, the same tools, held to the same verification standard, the same question, real bills against real bills. Two teams arriving at that method independently is the tell that the control, not the routing, is where the work is.

A worked example: live from the estateis_simulated = false
measuring the live estate…
04 · The lever library

Model choice is one lever of many.

“Use a cheaper model” is the obvious move and rarely the biggest one. The act layer is a library of econometric levers, each a safe, gated optimization with a measured cost‑quality tradeoff, each proven at parity before it is applied and each able to auto‑revert:

loading the live act catalog…

Each lever is one framework: propose it, prove it holds parity on a real slice, apply it under the gate, keep the revert armed. The science says what to try; the parity engine safely proves it, and only a lever with a shipped benchmark and a proven ledger run is ever applied on its own.

05 · The elasticity curve

Quality has a diminishing-returns curve.

For a given task class, plot outcome quality against tokens spent, or against reasoning effort, and you get a curve that bends. Below some point, spending less costs you outcomes. Past it, spending more buys nothing but latency and bill. The money is in finding that knee for each task class and living exactly on it.

finished work vs. spend
starvedproductivereboundAI spendfinished work
starvedevery extra dollar buys a lot
productiveit buys less and less
reboundit buys LESS finished work
Three regions, and only one gets modelled. Cut in the third and you bank a real saving; cut in the second and the bill drops this month while the quality regression arrives next quarter, attributed to something else. A per-token dashboard cannot tell you which you are in. Shape only — no units, because we do not fit this curve yet.
The optimum is not the cheapest config, nor the smartest. It is the cheapest config that still clears the quality floor: the knee of the curve.

Three different things get called token elasticity, and running them together is why most writing on this produces no decision. They are worth keeping apart:

MarketPrice −20% → total spend ≈ −3%. Usually a small saving, not a trap.
InstructionBudget 5× tighter → tokens +83%. Non‑monotone — the model can’t self‑tune it either.
ProductionOutcomes vs. spend, per task class. The one Margin measures on your own traffic.

Market elasticity. Unit price falls, volume rises further, and the received story says the bill rises anyway. That story needs demand elasticity near 2. The only estimate from real marketplace transactions puts it just above 1, which means a 20% price cut leaves total spend about 3% lower. A price cut is usually a small saving, not a trap. When a bill rises after one, the cause is elsewhere: more scope, a shift into reasoning work, thinking tokens you cannot see, or a commitment already signed. Each is separable, and three of the four are yours to change.

Demirer, Fradkin, Tadelis & Peng, NBER w34608 (Dec 2025), on OpenRouter and Azure API data. Their words: preliminary, short‑run. We cite it as the best available estimate, not a settled fact.

Instruction elasticity. Tighten a reasoning budget too far and realized consumption goes up. On one question a 50‑token budget produced 86 output tokens; a 10‑token budget produced 157. Five times tighter, 83% more tokens. The curve is not monotone, so nobody can reason their way to the setting. The model cannot either: asked to estimate its own efficient budget it lands in the right range 61% of the time.

Han et al., arXiv:2412.18547, Fig. 1c–d and §A.2. Their evidence is math benchmarks on two models. Generalisation to agent workloads is not established.

Production elasticity. Outcomes with respect to spend. This is the number worth having, and the one nobody publishes for your workload. The production function is already written down: a standard economics formula for how three inputs — input, output and fine‑tuning tokens — jointly produce value (Cobb‑Douglas), where the exponent on each term is its output elasticity: how much outcome quality moves when that one input moves. The theory names those three parameters and assumes their values. Estimating them on your own traffic is the work.

These per‑task‑class elasticity models are the second moat: the science of which lever to pull and how far. They compound with data. The elasticity engine proposes the move; the parity engine proves it is safe; together they turn a guess into a measured, revertible decision.

And the per‑token number gets less useful as you grow. The larger your account, the closer your marginal token price sits to what it costs the provider to serve you, and their margin moves into the part with no tokens attached to it. Anyone shifted onto a seat fee plus a consumption commitment this spring lived it. You cannot divide a commitment by tokens. You can divide it by outcomes.

It is also why “we moved to a cheaper model and saved money” cannot be checked against an invoice. Unit prices fell all year while bills rose, so both numbers are true at once, and only a decomposition separates a price change from a volume change. A vendor paid to shrink your bill finds the flat region and stops looking; ours is flat by design, so what we are building toward is a slope rather than a saving — and a slope tells you to keep spending about as often as it tells you to stop.

Status: the parity gate above runs today. The full per‑task‑class curve does not. We will not draw a curve through measurements we did not take.
06 · The parity ledger

Every safe act is a datapoint no one else has.

Each change Margin makes and its measured result is a labeled row that nobody else can produce, because nobody else runs the safe act‑then‑re‑measure loop that produces the label: (task‑class, lever, config) → (Δcost, Δquality, held?). Persisted into one growing, task‑class‑keyed ledger, this is where the elasticity curves are learned from, not guessed.

proven priors · cold start vs. warm start
flows seenproven priorsfrom scratchwith the corpuscustomer #1 starts stocked
from scratchan empty ledger climbs slowly
with the corpuscustomer #1 arrives already stocked
what travelsabstracted task-class priors only — never prompts, outputs, or your data
The head start is the moat. Every proven safe act is a labeled row nobody else runs the loop to produce, so the ledger only grows. Breadth is why customer number one arrives to a stocked corpus instead of an empty one: a task class proven ten thousand times elsewhere hands over a warm prior, and only the abstracted prior travels — never prompts, outputs, or customer data. Shape only — no units, because the corpus grows on real traffic and we do not fit this curve.

There are two network effects, and they compound:

  • Depth: per customer, Margin learns each of their flows better every week. Leaving resets that intelligence to zero. That is the switching cost.
  • Breadth: a task class seen ten thousand times across customers transfers a warm prior to a brand‑new one, killing cold‑start. Only abstracted task‑class priors travel, never raw prompts, outputs, or customer data, so breadth compounds while privacy holds.

Running open‑source agents to seed this ledger is therefore not a demo. It is moat‑building: customer number one should arrive to a non‑empty corpus of proven priors.

07 · Two modes

Meet the customer at their eval maturity.

Whether an outcome is easy to ground depends on whether the customer has evals. So Margin detects eval maturity and adapts:

Mode A: bring your evals

Later‑stage teams with mature eval frameworks. Integrate directly, ground quality in their own ground truth, and optimize deep. High signal, high value.

Mode B: grow the coverage

Earlier teams with thin or no evals. Use reference‑free methods, and surface where their eval coverage is thin so they can grow it, a standalone value they will pay for before trusting any auto‑tune.

Mode B is the top‑of‑funnel wedge; Mode A is the higher‑value expansion. And the two connect: a Mode B customer whose evals grow, partly because Margin helped grow them, matures into a Mode A customer, a deeper integration and a bigger deal. Margin compounds with the customer’s own maturation, and they never outgrow it. This is the eval‑coverage flywheel: land on “trust your quality,” expand into “now safely optimize it.”

08 · The honest guarantee

A fabricated saving is the worst failure.

The entire product depends on one discipline: every number is a real metered run, a real outcome, a real parity measurement. A saving that looks better than it is destroys the only thing worth selling: the trust to leave the loop on. So the guarantees are absolute: real spend against real APIs, outcomes graded on the agent’s own verdict, parity checked against a threshold before any change is kept, and an auto‑revert that fires the moment quality dips. Simulated data is tagged as simulated, everywhere, always.

That is the framework. The console runs it live on real, recognizable open‑source agents, not a toy estate, so you can watch a real route flip, bank a real saving, see the parity proof, and hit revert yourself.

09 · Where this goes

Six things a labor market needs that a spend dashboard cannot give it.

A roadmap presented as a product is the one failure a company selling honest numbers does not survive, so the figure draws the line rather than burying it in a caveat. Solid runs today; hollow is ratified and not built. The last rung is the furthest out, and the one we are most sure about.

the roadmap ladder · shipped vs. ratified
ratified — not builtrunning todaythe supply-chain mapthe parity gatewhich agent the rest depend onthe full account of the billwhich agents to retirethe design-time wind tunnelwhat you can stop watchingrebuilding the one it condemned
rebuilding the one it condemnednot built
what you can stop watchingnot built
the design-time wind tunnelnot built
which agents to retirenot built
the full account of the billnot built
which agent the rest depend onnot built
the parity gaterunning today
the supply-chain maprunning today
The six rungs above the line are ratified on our roadmap and not built, and they are drawn hollow so this cannot be read as a feature list. Each one takes over a little more of the decision than the rung below it.

Which agent the rest depend on. A cheap retrieval service four specialists share never appears in a cost ranking, and it is the one whose failure stops everything. Read it off the graph: remove this node, and what share of your outcomes becomes unreachable.

The full account of the bill. Every dollar bought a measured outcome, was waste with a named cause, or we cannot see what it bought. The three sum to your bill. Everyone publishes the first two; the third is what makes them believable.

Which agents to retire. Every recommendation in this category tunes an agent to run cheaper. The larger saving is switching one off. A vendor paid on your token spend will never say that; our fee is flat, on purpose, so we can.

The design-time wind tunnel. Before you build an agent, what it will cost per outcome and what it does to everything up and downstream of it. We give you the economics. You design the agent. We never pick your architecture.

What you can stop watching. How much you can safely delegate is set by the task rather than the model: how cheap it is to check the work, and how cheap it is to undo a mistake. We already measure both.

And then it rebuilds the one it condemned. If we tell you an operation is badly built, the fair question is why you are the one fixing it. So we write the replacement and run it against yours on your own traffic. We never grade our own work: your frozen outcome check does, and our version has to beat the incumbent rather than tie it. You keep the architecture.

See the framework run live.

The console is the thesis in motion: measure → recommend → act → revert, on real agent spend, with the parity proof on every change.