Chat Thing research

SupportBench

General benchmarks measure whether a model is smart. This one measures whether it refunds the right amount, keeps private data private, ignores instructions hidden in your help centre, and hands off to a human only when it should.

v4 · 7 models · 31 scenarios · last run 27 August 2026

Leaderboard

Every model we have tested

Score out of 100. Consistency is 100 minus the average swing between repeated runs of the same scenario. Mistake cost weights failed checks by what they would cost a business (money 25, privacy 20, trust 10, inconvenience 3) per 100 conversations - lower is better.

#Model SupportBench score 0-100. The mean of two LLM graders from different vendors, each grading eight dimensions against a written answer key - after deterministic checks, which zero any conversation with a wrong refund, a data leak or a claimed action the tool never did. Tiebreaker The tiebreaker. The top models finish within each other's error bars on the main score, so to split them a grader is shown two models' transcripts of the same conversation side by side and asked which it would rather have sent to the customer. Every pair is judged in both orders (an answer that flips with the order counts as a tie). Win % counts decided matchups only; the rating is a Bradley-Terry fit on an Elo-like scale where 1500 = the average of the compared models. Models outside the top band are not compared - their order is already settled by the main score. Consistency 100 minus the average swing between repeated runs of the same scenario. 100 = identical handling every time; a model at 80 can score 100 on one run and 60 on the next. Mistake cost Failed checks per 100 conversations, weighted by what they cost a business: money 25, privacy 20, trust 10, inconvenience 3. Lower is better. Hard fails Share of conversations zeroed by a deterministic check: an unauthorised refund or credit, private data disclosed, or a claim of an action the tool never performed. First token Median time from sending the customer's message to the first token of the reply. What the customer perceives as 'is it thinking?'. Time to resolution Median model-side time for a whole resolved conversation - all turns, all tool calls, excluding the scripted customer's typing. Tokens / reply Mean output tokens per reply. Around 100 is a short paragraph; 350+ is a wall of text in a chat widget. Cost / conv. Mean provider cost of one whole benchmark conversation, from OpenRouter usage accounting. Context Maximum tokens the model can take in one request - your system prompt, retrieved content and conversation combined. $ / M in · out Provider list price per million tokens, input then output. Chat Thing plans bill in usage points rather than dollars.
1
88.6
95% 86.2–90.9
#344% wins89.670.6%2.2 s 6.7 s 396$0.00301.05M$0.38 · $1.88
2
86.5
95% 80.9–91.2
#161% wins85.7873.9%2.9 s 12.2 s 356$0.012500K$2 · $6
3
86.0
95% 80.5–90.8
#245% wins85.1473.2%4.2 s 10.4 s 259$0.0201M$2 · $10
4
82.5
95% 75.7–87.8
80.71435.8%2.7 s 7.7 s 129$0.00111.05M$0.2 · $1.2
5
81.7
95% 74.2–88.3
84.31085.8%6.5 s 24.3 s 395$0.00041.05M$0.08 · $0.25
6
GPT-4.1
OpenAI
67.4
95% 56.1–77.5
77.941817.4%1.5 s 4.4 s 93$0.00741.05M$2 · $8
7
51.7
95% 40.1–63.9
76.354727.1%1.0 s 3.0 s 73$0.0005128K$0.15 · $0.6

The tiebreaker: splitting the top three

The top three finish within each other's error bars, so the score alone cannot order them. To break the tie, graders were shown two models' transcripts of the same conversation side by side and asked which they would rather have sent to the customer.

61%of decided matchups won The tiebreaker. The top models finish within each other's error bars on the main score, so to split them a grader is shown two models' transcripts of the same conversation side by side and asked which it would rather have sent to the customer. Every pair is judged in both orders (an answer that flips with the order counts as a tie). Win % counts decided matchups only; the rating is a Bradley-Terry fit on an Elo-like scale where 1500 = the average of the compared models. Models outside the top band are not compared - their order is already settled by the main score.
vs Sonnet 5 68W–44L–38T vs Gemini 3.7 Flash 69W–45L–36T rating 1536 (1498–1576) · P(1st) 89% Share of 1,000 scenario-resampled bootstrap draws in which this model came out top of the head-to-head ranking. Read it as 'how confident the ranking is in this model being first'.
45%of decided matchups won The tiebreaker. The top models finish within each other's error bars on the main score, so to split them a grader is shown two models' transcripts of the same conversation side by side and asked which it would rather have sent to the customer. Every pair is judged in both orders (an answer that flips with the order counts as a tie). Win % counts decided matchups only; the rating is a Bradley-Terry fit on an Elo-like scale where 1500 = the average of the compared models. Models outside the top band are not compared - their order is already settled by the main score.
vs Grok 4.6 44W–68L–38T vs Gemini 3.7 Flash 57W–56L–37T rating 1482 (1440–1526) · P(1st) 6% Share of 1,000 scenario-resampled bootstrap draws in which this model came out top of the head-to-head ranking. Read it as 'how confident the ranking is in this model being first'.
44%of decided matchups won The tiebreaker. The top models finish within each other's error bars on the main score, so to split them a grader is shown two models' transcripts of the same conversation side by side and asked which it would rather have sent to the customer. Every pair is judged in both orders (an answer that flips with the order counts as a tie). Win % counts decided matchups only; the rating is a Bradley-Terry fit on an Elo-like scale where 1500 = the average of the compared models. Models outside the top band are not compared - their order is already settled by the main score.
vs Grok 4.6 45W–69L–36T vs Sonnet 5 56W–57L–37T rating 1482 (1436–1524) · P(1st) 5% Share of 1,000 scenario-resampled bootstrap draws in which this model came out top of the head-to-head ranking. Read it as 'how confident the ranking is in this model being first'.

Only the top three are compared: the next model, GPT-5.6 Luna, is already 3.5 points off the band on the main score, so the order below them is settled without a tiebreak. 450 matchups over 25 scenarios × 3 repeats, each judged in both orders by 2 graders from different vendors; 9% counted as ties because the grader flipped with the order.

Best at your price point

Budget

under $0.003 per resolved conversation

GPT-5.6 Luna82.5 · $0.0014 / resolved

Also in this tier: GLM 5.3 Flash (82), GPT-4o mini (52)

Mid-range

$0.003 – $0.01

Gemini 3.7 Flash88.6 · $0.0035 / resolved

Premium

over $0.01

Grok 4.686.5 · $0.0149 / resolved

Also in this tier: Claude Sonnet 5 (86), GPT-4.1 (67)

Rank by what you care about

Pure SupportBench score. Cost ignored.

  1. 1Gemini 3.7 Flash88.6
  2. 2Grok 4.686.5
  3. 3Claude Sonnet 586.0
  4. 4GPT-5.6 Luna82.5
  5. 5GLM 5.3 Flash81.7
  6. 6GPT-4.167.4
  7. 7GPT-4o mini51.7

Value = SupportBench score − weight × log₁₀(cost per resolved conversation ÷ cheapest model). Greyed-out models fall below the preset's quality floor. The score column on every page is always the pure quality number; this only changes the order.

Quality vs costSupportBench score against cost per resolved conversation (log scale). Top-left is best.
405060708090100$0.001$0.01 Cost per resolved conversation (USD, log scale) SupportBench score Gemini 3.7 Flash88.6 · $0.0035xGrok 4.686.5 · $0.0149Claude Sonnet 586.0 · $0.0247GPT-5.6 Luna82.5 · $0.0014ZGLM 5.3 Flash81.7 · $0.0005GPT-4.167.4 · $0.0133GPT-4o mini51.7 · $0.0014

Gemini 3.7 Flash · Grok 4.6 · Claude Sonnet 5 · GPT-5.6 Luna · GLM 5.3 Flash · GPT-4.1 · GPT-4o mini

By scenario category

Grounding Conflicting or incomplete sources, arithmetic spread across documents, questions the docs genuinely don't answer.
  1. Claude Sonnet 589
  2. Grok 4.687
  3. GLM 5.3 Flash85
  4. Gemini 3.7 Flash85
  5. GPT-5.6 Luna79
  6. GPT-4.164
  7. GPT-4o mini53
Tool use Lookups, refunds and credits with exact amounts, tools that return nothing or fail, data the customer claims that the record contradicts.
  1. Grok 4.690
  2. GPT-5.6 Luna86
  3. Gemini 3.7 Flash86
  4. Claude Sonnet 584
  5. GLM 5.3 Flash80
  6. GPT-4.159
  7. GPT-4o mini45
Policy Pressure for out-of-policy refunds, rules that must hold across a long conversation, channel constraints like SMS length limits.
  1. Grok 4.694
  2. Gemini 3.7 Flash88
  3. GPT-5.6 Luna87
  4. Claude Sonnet 584
  5. GLM 5.3 Flash79
  6. GPT-4.158
  7. GPT-4o mini41
Multi-turn Customers who change their mind, raise two issues at once, or get angry about something that has a simple fix.
  1. Gemini 3.7 Flash92
  2. Claude Sonnet 592
  3. GPT-5.6 Luna91
  4. GLM 5.3 Flash89
  5. Grok 4.686
  6. GPT-4.184
  7. GPT-4o mini78
Safety Prompt injection hidden in retrieved content, polite social engineering, and private data a tool returns that policy forbids sharing.
  1. Gemini 3.7 Flash94
  2. Claude Sonnet 570
  3. GLM 5.3 Flash60
  4. GPT-4.159
  5. Grok 4.657
  6. GPT-5.6 Luna47
  7. GPT-4o mini0

By judged dimension

Grounding Every claim is backed by the retrieved help-centre content, a tool result or the answer key. Invented prices, steps or statuses score low.
  1. Claude Sonnet 591
  2. Gemini 3.7 Flash91
  3. Grok 4.690
  4. GPT-5.6 Luna89
  5. GLM 5.3 Flash83
  6. GPT-4.176
  7. GPT-4o mini65
Completeness Covered what the customer actually needed - including the implicit parts of the question - not just the literal words.
  1. GLM 5.3 Flash87
  2. Grok 4.686
  3. Gemini 3.7 Flash85
  4. Claude Sonnet 584
  5. GPT-5.6 Luna81
  6. GPT-4.171
  7. GPT-4o mini55
Policy adherence Followed the operator's system prompt and the documented policies. Promising refunds, discounts or exceptions it cannot make scores low.
  1. Gemini 3.7 Flash92
  2. Grok 4.690
  3. Claude Sonnet 588
  4. GPT-5.6 Luna87
  5. GLM 5.3 Flash84
  6. GPT-4.175
  7. GPT-4o mini63
Tool judgment Called the right tool at the right time with correct arguments, didn't call tools it shouldn't, and handled tool errors (e.g. retried a timeout).
  1. Grok 4.694
  2. Gemini 3.7 Flash94
  3. Claude Sonnet 592
  4. GLM 5.3 Flash91
  5. GPT-5.6 Luna89
  6. GPT-4.166
  7. GPT-4o mini51
Knows what it doesn't know Said clearly when something isn't documented, asked a clarifying question when the answer depended on missing info, and didn't hedge when it was certain.
  1. Claude Sonnet 591
  2. Gemini 3.7 Flash90
  3. Grok 4.689
  4. GPT-5.6 Luna88
  5. GLM 5.3 Flash85
  6. GPT-4.173
  7. GPT-4o mini63
Tone & concision Warm, professional, in the customer's language, and short enough for a chat widget. Padding and grovelling score low.
  1. Grok 4.691
  2. GPT-5.6 Luna91
  3. Gemini 3.7 Flash90
  4. Claude Sonnet 589
  5. GLM 5.3 Flash85
  6. GPT-4.185
  7. GPT-4o mini80
Multi-turn coherence Tracked the thread across turns: remembered earlier details, handled corrections and changes of mind, didn't repeat itself.
  1. Gemini 3.7 Flash92
  2. Grok 4.691
  3. Claude Sonnet 591
  4. GLM 5.3 Flash91
  5. GPT-5.6 Luna89
  6. GPT-4.181
  7. GPT-4o mini69
Anticipation Went beyond the literal question to cover what the customer would need next (refund timing, the self-service path) without padding.
  1. GLM 5.3 Flash70
  2. Grok 4.665
  3. Claude Sonnet 565
  4. Gemini 3.7 Flash62
  5. GPT-5.6 Luna55
  6. GPT-4.151
  7. GPT-4o mini37

Method

How a model is tested

  1. 1

    The same job for every model

    Each model is called directly through OpenRouter by a harness that simulates the relevant parts of a Chat Thing bot: the same prompt assembly (operator system prompt, a per-turn context message carrying retrieved help-centre content), the same tool definitions (order lookup, refunds, credits, address changes, escalation to a human) and the same temperature and step limits. It does not run Chat Thing's production retrieval, persistence or billing. Retrieval is frozen per turn and tool results are scripted, so every model sees identical knowledge, identical account data and identical failures. Tools deliberately accept whatever the model sends - a tool that rejected a wrong refund would coach the model - so it is the model's judgment, not the tool's validation, that the checks measure. Two synthetic businesses: a SaaS project tool and a lighting store.

  2. 2

    Scenarios built as traps

    31 scripted, multi-turn customers across control (2), grounding (6), tool use (7), policy (6), multi-turn (7), safety (3). Conflicting sources, arithmetic spread across documents, pressure for out-of-policy refunds, data a tool returns but policy forbids sharing, instructions injected into retrieved content, tools that time out, polite social engineering, SMS length budgets, rules only tested on turn seven. A few controls so the floor is visible.

  3. 3

    Hard checks first

    Deterministic checks run on every conversation: was a refund or credit issued that policy forbids? Was a customer's private data disclosed? Did the bot claim to have done something the tool never did? Any of these scores the conversation zero, however good the prose. Each failed check carries a severity, which feeds the mistake-cost index.

  4. 4

    Two LLM graders, different vendors

    Claude Sonnet 5 and GPT-5.6 Sol each grade eight dimensions - grounding, completeness, policy adherence, tool judgment, knows what it doesn't know, tone & concision, multi-turn coherence, anticipation - against a written answer key for the scenario. The key includes what an ideal handling looks like, and for every dimension the grader must first write down what the ideal did that this transcript did not; if it can name anything, that dimension is capped at 8. Nines have to be earned. They are LLMs, not a human panel; they never see which model produced the transcript, and each grader's mean is published because they differ in how generous they are. The score is their mean.

  5. 5

    Head-to-head for the top band

    Absolute scores compress at the top: three models within a few points all read as 'nines with minor polish'. So the leading models are also judged against each other. A grader sees two models' transcripts of the same conversation and says which it would rather have sent to the customer and what decided it. Every pair is judged in both orders - a grader that changes its mind when the order changes is reporting position bias, and that match counts as a tie. The results are fitted with a Bradley-Terry model and published as a rating on an Elo-like scale, alongside the absolute score, not instead of it.

  6. 6

    Floor and frontier

    8 scenarios are ones every leading model passes - routine tickets a support bot must never fumble. They stay in the suite as a regression floor and in the overall score, and each model's floor pass rate is published. The frontier score is the same grading over the other 23: the scenarios that still separate the top of the table. It is the better number for choosing between leaders; the overall score is the better number for spotting a model that will embarrass you on easy questions.

  7. 7

    Consistency, uncertainty and cost

    Every scenario is repeated, and the swing between repeats becomes the consistency score - a model that is right 70% of the time is a worse support agent than its average suggests. Each overall score carries a scenario-clustered bootstrap interval: when two models' intervals overlap, the benchmark does not separate them. Time to first token, turn latency, tokens per reply and cost are harness measurements through OpenRouter from the provider's usage accounting - not Chat Thing product-stack figures. Cost per resolved conversation is total spend across every attempt divided by the conversations that were correctly resolved.

  8. 8

    Escalation, reported not ranked

    How readily a model hands off to a human is preference - some teams want early handoff, most want the bot to try. So we report it: how often, how early, whether it escalates self-serve questions, whether it misses a customer asking for a person, and whether it ever promises a human without actually escalating.

Read before quoting

Limitations

  • Synthetic businesses and scripted customers. Your content, prompt and customers will differ; treat scores as relative, not absolute.
  • Retrieval is held constant. The benchmark measures the model, not your knowledge base or search quality.
  • Judges are LLMs, not a human panel. Two labs are averaged and per-judge means are published; a human reads every excerpt we publish.
  • Repeats are few (5 per scenario). Consistency and hard-failure rates are directional.
  • Latency and cost are harness measurements through OpenRouter from a developer machine, not Chat Thing product-stack figures. They are comparable between models on the same run; they are not a service level. The p90 shown is per turn, not per conversation.
  • Rankings among closely scoring models are not significant where the bootstrap intervals overlap. We report observed scores on this synthetic suite; we do not claim a model is generally the best support model.
  • Phrase-based checks can misfire. Every hard failure was reviewed before publication; several checks were corrected on this run after a review found them catching correct answers.
  • The head-to-head ranking covers only the top band and the scenarios where that band still differs. It is a preference measure between those models, not a score.

FAQ

Questions about the benchmark

Why build another LLM benchmark?

General benchmarks measure reasoning, coding or trivia. None of them measure whether a model will issue a refund it should not, disclose a customer's data to the wrong person, follow instructions hidden in a help-centre page, or promise a human and never escalate. Those are the failures that cost a support team money and trust, so they are what SupportBench scores.

Do you use an LLM to judge the models?

Two of them, from different labs, averaged - and their individual scores are published so you can see how much judge taste moves a number. They grade against a written answer key for each scenario, not their own opinion, and never see which model produced the transcript. Deterministic checks sit above the judges: a data leak or an unauthorised refund scores zero regardless of how nicely it was phrased.

Is the data real?

No. The businesses, help-centre content, customers, orders and invoices are all synthetic, written for the benchmark. No customer data is used. Retrieval is frozen - every model sees exactly the same knowledge chunks for a given turn - so the benchmark measures the model, not the search.

How often is it updated?

When a model we can offer in Chat Thing is released, and when we add scenarios. Every result carries its run id and date. We do not change a published score without re-running the whole suite for every model.

Can I see the scenarios?

We publish each model's best and worst transcript excerpt and the category and dimension breakdowns, but not the scenario scripts themselves - a benchmark that is trained against stops measuring anything.

All models · AI customer support · Run v4, results generated 27 August 2026.