Skip to content
Margin
ProductThe Meterevery call, meteredThe Estatethe org chart of your agentsThe Money Mapwhere the spend goesThe Parity Gateprove it before it shipsThe Governorthe control loopDocsPricingCompany
Sign inGet set up →
ProductThe MeterThe EstateThe Money MapThe Parity GateThe GovernorDocsPricingCompanyLive demoSign in

Case study zero: our own agent

Before asking anyone else to trust a recommendation, we took one of our own agents through the whole loop: measure it, test cheaper models on its own evals, open the change as a pull request, merge it, and measure again.

Captured 2026-10-09 from production

The short version

  • 3 face-offs ran; 2 were refused. Neither refused candidate was switched to.
  • gemini-2.5-flash-lite passed 120 of 120 evals at 19.5% less per task. Margin opened that change as a draft pull request, and it was merged.
  • 2 runs after the merge measured the new code and passed every eval, at $0.0000072 per passed outcome against $0.0000072 predicted. 2 more failed before measuring anything.
  • That model retires on 2026-10-20. The next step is a second switch.

The agent and its evals

The agent is oss_ladder/04_fanout/agent.py in our own repository, graded by 120 eval cases. Its baseline run on gpt-4o-mini passed 120 of 120, at $0.0000090 per passed outcome.

Every model it was tested against

Each candidate ran the same 120 cases. A cheaper model is switched to only if it keeps at least 97% of the current pass rate and the result is strong enough at this number of evals to rule out luck.

CandidateWhat movedPassedPer taskResult
mistral-small-3.2-24b-instructthe whole agent72 of 120$0.0000029Refused
gemini-2.5-flash-litethe whole agent120 of 120$0.0000072Held, not proven
mistral-small-3.2-24b-instructthis agent only119 of 120$0.0000048Refused

mistral-small-3.2-24b-instruct, for the whole agent: ruled out on this suite: it missed 48 of the 120 evals; the current model missed none. With 48 misses, a switch needs at least 2021 evals with every other one passed.

mistral-small-3.2-24b-instruct, for this agent only: ruled out on this suite: it missed 1 of the 120 evals; the current model missed none. With 1 miss, a switch needs at least 147 evals with every other one passed.

The gemini-2.5-flash-lite result is early evidence, not proof. Both models passed every case, so these evals cannot tell them apart, and Margin does not promote on a test that recorded no difference. The change was opened for review rather than switched automatically.

The change

Margin opened pull request #21 in our private repository (1 line added, 1 removed, in oss_ladder/04_fanout/agent.py). It was merged on 2026-10-09.

Every run after the merge

RunOutcomePassedPer passed outcome
#30failed: RuntimeError: install step failednonenone
#31measured gemini-2.5-flash-lite120 of 120$0.0000072
#34failed: the worker has no microVM configured, so it will not run customer code on its own hostnonenone
#35measured gemini-2.5-flash-lite120 of 120$0.0000072

The 2 failed runs billed nothing: each stopped on our side before any model was called, for the reason shown.

The prediction was $0.0000072. The measured runs matched it exactly because each one billed the same amount as the face-off did, on the same model and the same 120 cases. 2 runs on one small agent show the change held here. They do not show it would hold on a larger workload.

What happens next

OpenRouter lists google/gemini-2.5-flash-lite as retiring on 2026-10-20. The face-offs above chose candidates from a fixed list of older models, and that is how a retiring model was recommended. Candidates now come from the live model catalogue, newest first, and a model with a retirement date is never tried. The next face-off on this agent tests current models, and this page will be recaptured with the result.

Want the same measurement on your own agents? Start here.

Margin

The economic control layer for your AI workforce. Price an outcome and Margin reports what it returns; a cheaper route ships only when quality holds. Neutral across providers, and every number on this site traces back to a run we can show you.

Product

  • The five parts
  • Console
  • Docs
  • Pricing
  • Compare

Company

  • About
  • Build log
  • Security
  • Contact

Resources

  • How it works
  • Day 0 to 60
  • Essays
  • margin-cost on GitHub
  • margin-meter on PyPI
  • margin-meter on npm

Legal

  • Terms
  • Privacy
  • Cookies
  • DPA
© 2026 MarginBuilt in Chicago41.88° N, 87.63° W