SupportBench
General benchmarks measure whether a model is smart. This one measures whether it refunds the right amount, keeps private data private, ignores instructions hidden in your help centre, and hands off to a human only when it should.
v4 · 7 models · 31 scenarios · last run 27 August 2026
Leaderboard
Every model we have tested
Score out of 100. Consistency is 100 minus the average swing between repeated runs of the same scenario. Mistake cost weights failed checks by what they would cost a business (money 25, privacy 20, trust 10, inconvenience 3) per 100 conversations - lower is better.
The three takeaways
If you only read one thing on this page:
Gemini 3.7 Flash
88.6 / 100 · #1 overall
Tops the main score because it is the only leader that made no critical mistake in 155 conversations - no data leak, no relayed injection, no bad refund - and it is the cheapest and fastest of the three.
Read the full analysis →
Best individual repliesGrok 4.6
86.5 / 100 · #2 overall
Wins the tiebreaker: graders preferred its transcript in about six of ten decided matchups. But it leaked billing details and relayed a planted instruction in 3 of 5 runs of those scenarios - pick it when you control your content and tools.
Read the full analysis →
Best with untrusted contentClaude Sonnet 5
86.0 / 100 · #3 overall
The most grounded model tested and the only leader never fooled by injection or social engineering - the pick when your knowledge base includes content you don't control. The trade-off is cost: ~7x Gemini per resolved conversation.
Read the full analysis →
| # | Model | SupportBench score 0-100. The mean of two LLM graders from different vendors, each grading eight dimensions against a written answer key - after deterministic checks, which zero any conversation with a wrong refund, a data leak or a claimed action the tool never did. | Tiebreaker The tiebreaker. The top models finish within each other's error bars on the main score, so to split them a grader is shown two models' transcripts of the same conversation side by side and asked which it would rather have sent to the customer. Every pair is judged in both orders (an answer that flips with the order counts as a tie). Win % counts decided matchups only; the rating is a Bradley-Terry fit on an Elo-like scale where 1500 = the average of the compared models. Models outside the top band are not compared - their order is already settled by the main score. | Consistency 100 minus the average swing between repeated runs of the same scenario. 100 = identical handling every time; a model at 80 can score 100 on one run and 60 on the next. | Mistake cost Failed checks per 100 conversations, weighted by what they cost a business: money 25, privacy 20, trust 10, inconvenience 3. Lower is better. | Hard fails Share of conversations zeroed by a deterministic check: an unauthorised refund or credit, private data disclosed, or a claim of an action the tool never performed. | First token Median time from sending the customer's message to the first token of the reply. What the customer perceives as 'is it thinking?'. | Time to resolution Median model-side time for a whole resolved conversation - all turns, all tool calls, excluding the scripted customer's typing. | Tokens / reply Mean output tokens per reply. Around 100 is a short paragraph; 350+ is a wall of text in a chat widget. | Cost / conv. Mean provider cost of one whole benchmark conversation, from OpenRouter usage accounting. | Context Maximum tokens the model can take in one request - your system prompt, retrieved content and conversation combined. | $ / M in · out Provider list price per million tokens, input then output. Chat Thing plans bill in usage points rather than dollars. |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.7 Flash Google | 88.6 95% 86.2–90.9 | #344% wins | 89.6 | 7 | 0.6% | 2.2 s | 6.7 s | 396 | $0.0030 | 1.05M | $0.38 · $1.88 |
| 2 | Grok 4.6 xAI | 86.5 95% 80.9–91.2 | #161% wins | 85.7 | 87 | 3.9% | 2.9 s | 12.2 s | 356 | $0.012 | 500K | $2 · $6 |
| 3 | Claude Sonnet 5 Anthropic | 86.0 95% 80.5–90.8 | #245% wins | 85.1 | 47 | 3.2% | 4.2 s | 10.4 s | 259 | $0.020 | 1M | $2 · $10 |
| 4 | GPT-5.6 Luna OpenAI | 82.5 95% 75.7–87.8 | — | 80.7 | 143 | 5.8% | 2.7 s | 7.7 s | 129 | $0.0011 | 1.05M | $0.2 · $1.2 |
| 5 | GLM 5.3 Flash Z.AI | 81.7 95% 74.2–88.3 | — | 84.3 | 108 | 5.8% | 6.5 s | 24.3 s | 395 | $0.0004 | 1.05M | $0.08 · $0.25 |
| 6 | GPT-4.1 OpenAI | 67.4 95% 56.1–77.5 | — | 77.9 | 418 | 17.4% | 1.5 s | 4.4 s | 93 | $0.0074 | 1.05M | $2 · $8 |
| 7 | GPT-4o mini OpenAI | 51.7 95% 40.1–63.9 | — | 76.3 | 547 | 27.1% | 1.0 s | 3.0 s | 73 | $0.0005 | 128K | $0.15 · $0.6 |
The tiebreaker: splitting the top three
The top three finish within each other's error bars, so the score alone cannot order them. To break the tie, graders were shown two models' transcripts of the same conversation side by side and asked which they would rather have sent to the customer.
Only the top three are compared: the next model, GPT-5.6 Luna, is already 3.5 points off the band on the main score, so the order below them is settled without a tiebreak. 450 matchups over 25 scenarios × 3 repeats, each judged in both orders by 2 graders from different vendors; 9% counted as ties because the grader flipped with the order.
Best at your price point
Budget
under $0.003 per resolved conversation
Also in this tier: GLM 5.3 Flash (82), GPT-4o mini (52)
Premium
over $0.01
Also in this tier: Claude Sonnet 5 (86), GPT-4.1 (67)
Rank by what you care about
Pure SupportBench score. Cost ignored.
- 1Gemini 3.7 Flash88.6 score 88.6 · $0.0035
- 2Grok 4.686.5 score 86.5 · $0.0149
- 3Claude Sonnet 586.0 score 86.0 · $0.0247
- 4GPT-5.6 Luna82.5 score 82.5 · $0.0014
- 5GLM 5.3 Flash81.7 score 81.7 · $0.0005
- 6GPT-4.167.4 score 67.4 · $0.0133
- 7GPT-4o mini51.7 score 51.7 · $0.0014
Value = SupportBench score − weight × log₁₀(cost per resolved conversation ÷ cheapest model). Greyed-out models fall below the preset's quality floor. The score column on every page is always the pure quality number; this only changes the order.
Gemini 3.7 Flash · Grok 4.6 · Claude Sonnet 5 · GPT-5.6 Luna · GLM 5.3 Flash · GPT-4.1 · GPT-4o mini
By scenario category
By judged dimension
Method
How a model is tested
- 1
The same job for every model
Each model is called directly through OpenRouter by a harness that simulates the relevant parts of a Chat Thing bot: the same prompt assembly (operator system prompt, a per-turn context message carrying retrieved help-centre content), the same tool definitions (order lookup, refunds, credits, address changes, escalation to a human) and the same temperature and step limits. It does not run Chat Thing's production retrieval, persistence or billing. Retrieval is frozen per turn and tool results are scripted, so every model sees identical knowledge, identical account data and identical failures. Tools deliberately accept whatever the model sends - a tool that rejected a wrong refund would coach the model - so it is the model's judgment, not the tool's validation, that the checks measure. Two synthetic businesses: a SaaS project tool and a lighting store.
- 2
Scenarios built as traps
31 scripted, multi-turn customers across control (2), grounding (6), tool use (7), policy (6), multi-turn (7), safety (3). Conflicting sources, arithmetic spread across documents, pressure for out-of-policy refunds, data a tool returns but policy forbids sharing, instructions injected into retrieved content, tools that time out, polite social engineering, SMS length budgets, rules only tested on turn seven. A few controls so the floor is visible.
- 3
Hard checks first
Deterministic checks run on every conversation: was a refund or credit issued that policy forbids? Was a customer's private data disclosed? Did the bot claim to have done something the tool never did? Any of these scores the conversation zero, however good the prose. Each failed check carries a severity, which feeds the mistake-cost index.
- 4
Two LLM graders, different vendors
Claude Sonnet 5 and GPT-5.6 Sol each grade eight dimensions - grounding, completeness, policy adherence, tool judgment, knows what it doesn't know, tone & concision, multi-turn coherence, anticipation - against a written answer key for the scenario. The key includes what an ideal handling looks like, and for every dimension the grader must first write down what the ideal did that this transcript did not; if it can name anything, that dimension is capped at 8. Nines have to be earned. They are LLMs, not a human panel; they never see which model produced the transcript, and each grader's mean is published because they differ in how generous they are. The score is their mean.
- 5
Head-to-head for the top band
Absolute scores compress at the top: three models within a few points all read as 'nines with minor polish'. So the leading models are also judged against each other. A grader sees two models' transcripts of the same conversation and says which it would rather have sent to the customer and what decided it. Every pair is judged in both orders - a grader that changes its mind when the order changes is reporting position bias, and that match counts as a tie. The results are fitted with a Bradley-Terry model and published as a rating on an Elo-like scale, alongside the absolute score, not instead of it.
- 6
Floor and frontier
8 scenarios are ones every leading model passes - routine tickets a support bot must never fumble. They stay in the suite as a regression floor and in the overall score, and each model's floor pass rate is published. The frontier score is the same grading over the other 23: the scenarios that still separate the top of the table. It is the better number for choosing between leaders; the overall score is the better number for spotting a model that will embarrass you on easy questions.
- 7
Consistency, uncertainty and cost
Every scenario is repeated, and the swing between repeats becomes the consistency score - a model that is right 70% of the time is a worse support agent than its average suggests. Each overall score carries a scenario-clustered bootstrap interval: when two models' intervals overlap, the benchmark does not separate them. Time to first token, turn latency, tokens per reply and cost are harness measurements through OpenRouter from the provider's usage accounting - not Chat Thing product-stack figures. Cost per resolved conversation is total spend across every attempt divided by the conversations that were correctly resolved.
- 8
Escalation, reported not ranked
How readily a model hands off to a human is preference - some teams want early handoff, most want the bot to try. So we report it: how often, how early, whether it escalates self-serve questions, whether it misses a customer asking for a person, and whether it ever promises a human without actually escalating.
Read before quoting
Limitations
- Synthetic businesses and scripted customers. Your content, prompt and customers will differ; treat scores as relative, not absolute.
- Retrieval is held constant. The benchmark measures the model, not your knowledge base or search quality.
- Judges are LLMs, not a human panel. Two labs are averaged and per-judge means are published; a human reads every excerpt we publish.
- Repeats are few (5 per scenario). Consistency and hard-failure rates are directional.
- Latency and cost are harness measurements through OpenRouter from a developer machine, not Chat Thing product-stack figures. They are comparable between models on the same run; they are not a service level. The p90 shown is per turn, not per conversation.
- Rankings among closely scoring models are not significant where the bootstrap intervals overlap. We report observed scores on this synthetic suite; we do not claim a model is generally the best support model.
- Phrase-based checks can misfire. Every hard failure was reviewed before publication; several checks were corrected on this run after a review found them catching correct answers.
- The head-to-head ranking covers only the top band and the scenarios where that band still differs. It is a preference measure between those models, not a score.
FAQ
Questions about the benchmark
Why build another LLM benchmark?
General benchmarks measure reasoning, coding or trivia. None of them measure whether a model will issue a refund it should not, disclose a customer's data to the wrong person, follow instructions hidden in a help-centre page, or promise a human and never escalate. Those are the failures that cost a support team money and trust, so they are what SupportBench scores.
Do you use an LLM to judge the models?
Two of them, from different labs, averaged - and their individual scores are published so you can see how much judge taste moves a number. They grade against a written answer key for each scenario, not their own opinion, and never see which model produced the transcript. Deterministic checks sit above the judges: a data leak or an unauthorised refund scores zero regardless of how nicely it was phrased.
Is the data real?
No. The businesses, help-centre content, customers, orders and invoices are all synthetic, written for the benchmark. No customer data is used. Retrieval is frozen - every model sees exactly the same knowledge chunks for a given turn - so the benchmark measures the model, not the search.
How often is it updated?
When a model we can offer in Chat Thing is released, and when we add scenarios. Every result carries its run id and date. We do not change a published score without re-running the whole suite for every model.
Can I see the scenarios?
We publish each model's best and worst transcript excerpt and the category and dimension breakdowns, but not the scenario scripts themselves - a benchmark that is trained against stops measuring anything.
All models · AI customer support · Run v4, results generated 27 August 2026.