Skip to main content
Agentic AI

Tool calling benchmark: the provider is in the score

Tool calling benchmark: the provider is in the score

[!NOTE] TL;DR: On OpenRouter's live τ²-Bench Airline board (141 models, last run Oct 1, 2026, 9:55 AM UTC), Gemini 3.7 Flash leads at 80.6% while the cheapest Pareto point costs 0.006 USD per graded task against 0.077 USD for the leader. Tool-call errors on the models enrolled in OpenRouter's automatic provider optimization moved from 3.9% to 3.2%. A tool calling benchmark score is a photograph of a model plus a provider, so treat anything inside the published error band as a tie and decide on cost.

Gemini 3.7 Flash tops the tau2 tool calling benchmark at 80.6%. It costs 0.077 USD per graded task (last run Oct 1, 2026, 9:55 AM UTC). The model that trails it by 5.6 points costs 0.006 USD, thirteen times less, and both sit on the same Pareto frontier. My first reaction was to distrust the cheap one. My second was to distrust the board.

So I fetched the raw pages, wrote the numbers into a ledger with their as-of dates, and read the methodology sections twice. What follows is only what those pages actually displayed.

The snapshot of October 1

Model Accuracy Std dev Cost / task Time / task Output tok / task
Google Gemini 3.7 Flash 80.6% not published 0.077 USD 2.0 min 14.7k
Anthropic Claude Fable 5 79.6% ±2.8 pp 0.93 USD 2.2 min 5.75k
Anthropic Claude Fable 5.1 79.2% ±1.3 pp 0.72 USD 1.9 min 6.9k
Anthropic Claude Opus 5 78.8% ±2.5 pp 0.49 USD 2.2 min 7.13k
Amazon Nova Micro 1.0 78.7% ±2.0 pp 1.28 USD 1.7 min 5.69k
Z.ai GLM 5 78.0% ±3.4 pp 0.046 USD 3.4 min 9.39k
StepFun Step 3.7 Flash 77.3% not published 0.020 USD 3.8 min 10.9k
Anthropic Claude Sonnet 5.5 76.7% not published 0.18 USD 80 s 6.16k
DeepSeek V4.1 Flash 75.7% ±1.0 pp 0.018 USD 4.0 min 15.1k
Z.ai GLM 5.3 Flash 75.0% ±2.0 pp 0.006 USD 2.4 min 4.22k

Source: the τ²-Bench Airline leaderboard as I fetched it, with default routing where available. OpenRouter's benchmark index reports 13 benchmarks and 1,657,478 task evaluations, all last run Oct 1, 2026.

Read that table twice, because the two columns that matter most are the ones your eyes skip: the error band and the price. Almost nobody ships the top row.

The usage mirage

The reflex is to open the rankings page and take the most-used model as the best one. The numbers say the opposite, and OpenRouter says it out loud: these rankings count tokens, not quality, and varying verbosity between models makes token totals incomparable.

Look at what tops the volume chart. Space Bunny Alpha, authored by "stealth", processed 5.33T tokens on the most recent complete day and 28.4T over the week, with a week-over-week change above 999%. It has no publishing vendor, no weights you can inspect, no documentation I could retrieve. You cannot audit it. A model you cannot name is a model you cannot put in a client-facing refund flow, whatever its rank.

The scale of the numbers is worth a second look too. DeepSeek took 24.4% of all text requests on OpenRouter in the week beginning Sep 21, 2026, ahead of Google at 20.5% and OpenAI at 17.9%. Meanwhile the same page's intelligence index block puts Claude Opus 5.5 first at 57.6. Volume leader and quality leader are two different animals, living on two different scales. A fuel gauge tells you how much you burned, never where you went.

Three mechanical facts complete the picture: variants are ranked separately, so a free tier and a paid tier of the same model occupy two rows; "trending" only counts models above one million tokens in the current window; and brand-new models are listed first, up to five. Every one of those rules pushes new names upward. That is a ranking of attention, not of capability.

Where the agent tokens actually go

The applications consuming the most tokens are not chatbots. On the same rankings page, the top apps are Hermes Agent at 2.07T tokens, Kilo Code at 1.18T, Claude Code at 1.07T, Cline at 1.05T and Codex at 699B. Further down, OpenClaw sits at 204B and Command Code at 379B.

That is the clearest adoption signal in the whole dataset, and it comes with a trap. Coding agents run tool loops. A long turn with a retry storm burns output tokens the way a leaky faucet empties a tank, so an app's token count measures how much machinery it runs as much as how many humans use it. Compare two agents only if you know their step budgets.

The concrete version of that trap was documented this week. A third-party write-up reports that Xiaomi retrained MiMo-V2.6-Pro and MiMo-V2.6-Flash with an extra distillation stage aimed at within-turn repetition, pushed the new weights into the API on Sep 25, 2026 under the same model names, and published the open checkpoints on Sep 27. The reported symptom was a harness locking into the same tool call until the output budget was gone, with one issue citing 148 identical calls in a single turn. Xiaomi's own metric puts the average repetition rate low, 1.02% for Flash in one harness and 0.54% for Pro, and that framing is exactly why averages hide this class of bug: mostly zeros, plus a tail that bills you your entire budget. I did not retrieve Xiaomi's technical note directly, so treat the details as reported rather than vendor-confirmed.

The provider inside the score

Here is the methodological sentence most readers skip. OpenRouter states that it runs τ²-Bench continuously against the same provider endpoints that serve its traffic, so a score reflects both the model and the provider running it, and that the benchmark is used to assess provider variance rather than model capability.

That changes what the table above means. You are not reading a property of Gemini 3.7 Flash. You are reading a property of Gemini 3.7 Flash served by whichever endpoint answered, on that day, at that hour. Call it the provider tax: the share of your measured reliability that no model card ever mentions and no pricing page ever shows.

The error-rate line makes the tax visible. Across the models enrolled in Auto Exacto, OpenRouter's automatic provider optimization for tool-calling requests, the tool-call error rate moved from 3.9% to 3.2%. That metric counts requests where the model called a tool that does not exist, passed arguments that do not match the schema, or emitted arguments that are not valid JSON. A 0.7 point move, obtained by changing routing rather than by changing model.

Then look at the error bands. Rows two through six span 79.6 ±2.8, 79.2 ±1.3, 78.8 ±2.5, 78.7 ±2.0 and 78.0 ±3.4 points. Those intervals overlap each other entirely. On this run, they are a statistical tie, and a leaderboard sorted by accuracy alone will still print them in a strict order, because a page has to render something.

Two anomalies deserve a flag rather than a story. Amazon Nova Micro 1.0 lands at 78.7% and 1.28 USD per task, the most expensive row I extracted, which reads oddly for a model branded Micro. And Xiaomi MiMo-V2.6-Flash sits at 74.7% while taking 8.9 minutes and 21.2k output tokens per task, roughly four times the wall-clock of the leader. Both are consistent with provider or configuration effects on the graded endpoint, and I have no evidence for a specific cause, so I am not going to invent one.

The price of a task

{
  "type": "bar",
  "data": {
    "labels": ["GLM 5.3 Flash", "DeepSeek V4.1 Flash", "Step 3.7 Flash", "Z.ai GLM 5", "Gemini 3.7 Flash", "Claude Sonnet 5.5", "Claude Opus 5", "Claude Fable 5.1", "Claude Fable 5", "Nova Micro 1.0"],
    "datasets": [{ "label": "Cost per graded task (USD)", "data": [0.006, 0.018, 0.020, 0.046, 0.077, 0.18, 0.49, 0.72, 0.93, 1.28], "backgroundColor": ["#3b82f6"] }]
  },
  "options": {
    "responsive": true,
    "plugins": { "title": { "display": true, "text": "tau2-Bench Airline, cost per graded task, OpenRouter, last run Oct 1, 2026" } },
    "scales": { "y": { "title": { "display": true, "text": "USD per task" } } }
  }
}

The spread is two orders of magnitude for 5.6 points of accuracy. That is the single most actionable fact in this radar.

Now apply the error bands to a routing decision. DeepSeek V4.1 Flash: 75.7% with a ±1.0 point band at 0.018 USD. Z.ai GLM 5.3 Flash: 75.0% with a ±2.0 point band at 0.006 USD. The intervals overlap, so the two are indistinguishable on this run, and one costs three times the other. Picking the cheaper one is not a compromise on quality, it is the only reading the statistics support.

Note also that OpenRouter's own detail page calls out a "Best Value" at 0.020 USD per task while its leaderboard carries a Pareto point at 0.006 USD. Two pages, two definitions of value. Vendors are not lying to you; "best" is doing a lot of unlabelled work. If you want the longer version of this argument, I wrote about it in the real cost behind the podium.

The search benchmarks are not ready

The same index lists five search-agent benchmarks, and their sample sizes are a warning label. BrowseComp ran on 4 models with a last run of Aug 18, 2026. DeepSearchQA, WideSearch: 4 models, Aug 18. HLE as a search task: 4 models, Sep 15. The top entries cluster around Perplexity as the search surface, with Claude Opus 5 or GPT-5.6 variants underneath, and BrowseComp shows 89.0% at 0.99 USD per task.

Four models is not a leaderboard, it is a first draft of one. You cannot conclude anything about the field from a comparison of four lanes, and the scores also mix a search engine, a harness configuration and a model into a single number. Wait for the field to fill in.

What it means for your stack

flowchart TD
  A[Agent workload] --> B{Own eval suite passes on a pinned version?}
  B -- no --> X[Reject the model or the provider]
  B -- yes --> C{Accuracy band overlaps the current leader?}
  C -- yes --> D[Treat the candidates as a tie]
  C -- no --> E[Pay the premium only if the task needs those points]
  D --> F{Cost per task under your ceiling?}
  E --> F
  F -- no --> X
  F -- yes --> G[Primary route]
  G --> H[Fallback route: second provider, then deterministic path]

Four decisions follow from the data, and I would defend each of them in a review.

First, gate on your own eval before the public board. A public score is a release signal, not a release gate, and the model you grade today may be a different artefact next week, as the MiMo episode shows. Pin the version string, and record it in telemetry next to every run. When a model changes its weights, the only thing that tells you is your own suite, replayed on real tasks against mocked dependencies. That is how I run evaluations in CI, and it has caught silent quality drift that no error budget would have surfaced.

Second, treat the published band as the decision unit. Winners inside overlapping intervals are ties, full stop, and ties go to the cheap seat.

Third, keep an explicit fallback chain rather than a single primary. In my own agent services the fallback list is configuration, not code: a second provider, then a cheaper model, then a deterministic path that still completes the request. The framework layer matters less than the policy. Pydantic AI ships a FallbackModel, and whichever stack you run, the rule is the same: a provider failure must become a service-level condition you named, never a raw 500 that leaks somebody else's error body.

Fourth, instrument per task, not per token. What you need is cost per completed task, retries, tool-call failures, steps to success and wall-clock, joined to a request identifier and exported as OpenTelemetry spans. Logfire and Langfuse both do this for agent runs, and both will show you the same uncomfortable number: the model on your invoice is not always the model that produced your latency.

If the routing question is still open on the "buy vs compose" axis, my decision grid lives in agents versus deterministic pipelines. The two questions are the same question asked at two levels.

Limits

This snapshot has a short shelf life and I want that on the record. Usage data is dated Sep 30, 2026 and the benchmark ran Oct 1, 2026, so anything below 76% or above 82% will move within weeks.

τ²-Bench Airline tests one thing: executing tool calls under a policy manual in a simulated airline support conversation. It measures no general intelligence, no coding, no long-horizon planning. The authors of the underlying work are explicit that dual control is harder than it looks: in the telecom extension, accuracy drops when a user shares control of the environment, and per-task success on new tasks sat at 34% for one frontier model, 42% for another and 49% for a third. The original paper's pass^k metric makes the same point for consistency: a model above 60% on a single attempt fell below 25% when the same task had to succeed eight times.

Three more holes. Many rows publish no standard deviation at all, including the current leader, so no band comparison is possible for them. The benchmark runs on production provider endpoints, which means the number is co-owned by infrastructure you may never touch. And two sections of the rankings page, top models by task and cost per session, did not render in my fetch, so I cite nothing from them rather than guess.

The surviving principle

More generally, the moment a benchmark is served from production endpoints, the score stops being a property of the model and becomes a property of an infrastructure choice. You are not picking a brain, you are picking a route. Which is why the eval you own, on the tasks you actually run, remains the only number that transfers into your production.

If you want a second pair of eyes on your agent evaluation loop, or a routing layer that survives a silent weight swap, start with our agent engineering offer or book a call. Roughly a week of work, and it usually starts with deleting a leaderboard bookmark.

Rankings data by OpenRouter is licensed under CC BY 4.0; reused here with attribution. Methodology and τ²-Bench details come from the τ-bench and τ²-bench papers.


Processing...
Processing...

Please wait

Secure operation