The agent and its evals
The agent is oss_ladder/04_fanout/agent.py in our own repository, graded by 120 eval cases. Its baseline run on gpt-4o-mini passed 120 of 120, at $0.0000090 per passed outcome.
Every model it was tested against
Each candidate ran the same 120 cases. A cheaper model is switched to only if it keeps at least 97% of the current pass rate and the result is strong enough at this number of evals to rule out luck.
| Candidate | What moved | Passed | Per task | Result |
|---|---|---|---|---|
mistral-small-3.2-24b-instruct | the whole agent | 72 of 120 | $0.0000029 | Refused |
gemini-2.5-flash-lite | the whole agent | 120 of 120 | $0.0000072 | Held, not proven |
mistral-small-3.2-24b-instruct | this agent only | 119 of 120 | $0.0000048 | Refused |
mistral-small-3.2-24b-instruct, for the whole agent: ruled out on this suite: it missed 48 of the 120 evals; the current model missed none. With 48 misses, a switch needs at least 2021 evals with every other one passed.
mistral-small-3.2-24b-instruct, for this agent only: ruled out on this suite: it missed 1 of the 120 evals; the current model missed none. With 1 miss, a switch needs at least 147 evals with every other one passed.
The gemini-2.5-flash-lite result is early evidence, not proof. Both models passed every case, so these evals cannot tell them apart, and Margin does not promote on a test that recorded no difference. The change was opened for review rather than switched automatically.
The change
Margin opened pull request #21 in our private repository (1 line added, 1 removed, in oss_ladder/04_fanout/agent.py). It was merged on 2026-10-09.
Every run after the merge
| Run | Outcome | Passed | Per passed outcome |
|---|---|---|---|
| #30 | failed: RuntimeError: install step failed | none | none |
| #31 | measured gemini-2.5-flash-lite | 120 of 120 | $0.0000072 |
| #34 | failed: the worker has no microVM configured, so it will not run customer code on its own host | none | none |
| #35 | measured gemini-2.5-flash-lite | 120 of 120 | $0.0000072 |
The 2 failed runs billed nothing: each stopped on our side before any model was called, for the reason shown.
The prediction was $0.0000072. The measured runs matched it exactly because each one billed the same amount as the face-off did, on the same model and the same 120 cases. 2 runs on one small agent show the change held here. They do not show it would hold on a larger workload.
What happens next
OpenRouter lists google/gemini-2.5-flash-lite as retiring on 2026-10-20. The face-offs above chose candidates from a fixed list of older models, and that is how a retiring model was recommended. Candidates now come from the live model catalogue, newest first, and a model with a retirement date is never tried. The next face-off on this agent tests current models, and this page will be recaptured with the result.
Want the same measurement on your own agents? Start here.