.png)
.png)
Which model is truly most cost-effective when every failed run still lands on your bill? Arize and Fireworks benchmarked 10 open and closed models on 40 real Terminal-Bench tasks. Each model attempted every task six times, producing 2,400 fully traced agent runs. We measured pass rate, reliable task coverage, cost per attempt, cost per successful task, and the retry tax that per-token pricing leaves out. Each run was graded against the task’s own deterministic test suite. The findings challenge several common model-selection shortcuts. In this benchmark, the open-versus-closed label did not predict cost or capability. Lower-cost models handled much of the easy work efficiently, while the hardest tasks still required frontier-class performance. GPT-5.5 and Kimi K3 finished near each other overall, yet they were strongest on different kinds of work. A carefully designed escalation policy also completed more tasks at a lower cost per success than GPT-5.5 alone. The details matter, however. A poorly designed ladder can burn money on ineffective retries and unnecessary model hops. In this research deep dive, the Arize and Fireworks teams will cover: (30 min presentation + 15 min Q&A)
How the benchmark was built, graded, and instrumented
How GPT-5.5, Kimi K3, gpt-oss-120b, Gemini, Claude, DeepSeek, and GLM compared on cost, pass rate, task coverage, and retry tax
How model performance changed across easy, medium, and hard tasks
Why the cheapest model can also reduce what your product can reliably do
When to route by task type, escalate after failure, or escalate when traces show signs of distress
How traces expose token churn, repeated tool failures, budget exhaustion, and confidently wrong answers
How to reproduce the analysis using your own workloads and definition of success
Tracing is particularly important here. The benchmark distinguished runs that exhausted their token budgets, runs that encountered tool-call friction, and silent failures where every intermediate step appeared healthy but the final result was wrong. You will leave with a repeatable framework for evaluating models on your own tasks and deciding when to route, retry, escalate, or stop. You will also understand how to balance inference cost with the reliability and task coverage your product needs. Who this is for AI engineers and developers building agents; technical product managers choosing a model strategy; platform teams and engineering leaders managing inference cost, reliability, or routing. What you'll leave with A repeatable method for benchmarking models on your own workloads, calculating cost per successful task, and designing a routing or escalation policy that balances cost, coverage, and reliability. Format 30-minute research deep dive, followed by 15 minutes of live Q&A. Level Intermediate. The methodology will be explained from first principles, and no prior Arize or Fireworks experience is required.
