Message tokens and what happens when you run out
Message tokens measure how much work your bots do when they reply. Every plan includes a monthly allowance of message tokens, shared by every bot on your team. See Compare plans for each plan's allowance.
How a message is counted
When someone sends your bot a message, the AI model reads an input (your bot's instructions, relevant content from your knowledge, the conversation so far and the new message) and writes an output (the reply). Both are measured in tokens, roughly three-quarters of a word each.
- Longer conversations use more tokens per message, because the conversation so far is sent to the model with each new message.
- More knowledge per answer uses more tokens. Retrieval settings such as how many sources are added to each answer change the input size. See Retrieval settings.
- The model you choose changes the count. See Models and message tokens.
What uses message tokens
| Uses message tokens | Doesn't use message tokens |
|---|---|
| Every reply your bot writes, including each step when it uses a power-up | Searching your knowledge for relevant content |
| Rephrasing the question when enhanced retrieval is on | Analytics (topics, sentiment, common questions) |
| AI conversation summaries | Syncing data sources (uses storage tokens instead) |
| Turning a power-up result into a display | |
| Bot test runs | |
| Voice features such as speech replies and transcription | |
| Chats in My Chats and when you test your bot |
Models and message tokens
Models cost different amounts to run, so each model has a multiplier for input tokens and one for output tokens. The tokens counted against your allowance are:
(input tokens × input multiplier) + (output tokens × output multiplier)
A model with a 10× output multiplier uses ten times as many message tokens for the same reply as a model with a 1× multiplier. Output multipliers are often higher than input multipliers, because providers charge more for generated text. In your bot's settings, the model picker shows each model's multipliers.
Anthropic
Name | Input modifier | Output modifier |
|---|---|---|
| Claude Haiku 3 | x 0.5 | x 2.5 |
| Claude Haiku 4.5 | x 2 | x 10 |
| Claude Sonnet 5 | x 4 | x 20 |
| Claude Sonnet 4.6 | x 6 | x 30 |
| Claude Sonnet 4.5 | x 6 | x 30 |
| Claude Sonnet 4 | x 6 | x 30 |
| Claude Opus 4.6 | x 10 | x 50 |
| Claude Opus 4.7 | x 10 | x 50 |
| Claude Opus 5 | x 10 | x 50 |
| Claude Opus 4.5 | x 10 | x 50 |
| Claude Opus 4.8 | x 10 | x 50 |
| Claude Opus 4.8 (Fast) | x 20 | x 100 |
| Claude Opus 5 (Fast) | x 20 | x 100 |
| Claude Opus 4 | x 30 | x 150 |
| Claude Opus 4.1 | x 30 | x 150 |
| Claude Opus 4.7 (Fast) | x 60 | x 300 |
Cohere
Name | Input modifier | Output modifier |
|---|---|---|
| Cohere - Command R | x 0.3 | x 1.2 |
| Command A | x 5 | x 20 |
| Cohere - Command R+ | x 5 | x 20 |
DeepSeek
Name | Input modifier | Output modifier |
|---|---|---|
| DeepSeek V4 Flash | x 0.28 | x 0.56 |
| DeepSeek V3 | x 0.51 | x 2.06 |
| DeepSeek V4 Pro | x 0.87 | x 1.74 |
| DeepSeek R1 | x 1.4 | x 5 |
Name | Input modifier | Output modifier |
|---|---|---|
| Google - Gemini 2.5 Flash Lite | x 0.2 | x 0.8 |
| Gemini 3.1 Flash Lite | x 0.5 | x 3 |
| Gemini 3.5 Flash Lite | x 0.6 | x 5 |
| Google - Gemini 2.5 Flash | x 0.6 | x 5 |
| Gemini 3.7 Flash | x 0.75 | x 3.75 |
| Google - Gemini 3 Flash | x 1 | x 6 |
| Google - Gemini 2.5 Pro | x 2.5 | x 20 |
| Gemini 3.6 Flash | x 3 | x 15 |
| Gemini 3.5 Flash | x 3 | x 18 |
| Google - Gemini 3.1 Pro | x 4 | x 24 |
Meta
Name | Input modifier | Output modifier |
|---|---|---|
| Llama 4 Scout | x 0.2 | x 0.6 |
| Llama 4 Maverick | x 0.4 | x 1.6 |
Mistral
Name | Input modifier | Output modifier |
|---|---|---|
| Mistral - Mistral Small | x 0.1 | x 0.16 |
| Mistral Large 3 | x 1 | x 3 |
| Mistral Medium 3.5 | x 3 | x 15 |
| Mistral - Mistral Large | x 4 | x 12 |
| Mistral - Open Mixtral 8x22b | x 4 | x 12 |
MoonshotAI
Name | Input modifier | Output modifier |
|---|---|---|
| Kimi K2 | x 1.14 | x 4.6 |
| Kimi K2.6 | x 1.9 | x 8 |
| Kimi K3 | x 6 | x 30 |
OpenAI
Name | Input modifier | Output modifier |
|---|---|---|
| GPT-5 Nano | x 0.1 | x 0.8 |
| GPT-5.6 Luna ourdefault | x 0.2 | x 1.2 |
| GPT-5.6 Luna Pro | x 0.2 | x 1.2 |
| GPT-4.1 Nano | x 0.2 | x 0.8 |
| GPT-4o Mini | x 0.3 | x 1.2 |
| GPT-5.4 Nano | x 0.4 | x 2.5 |
| GPT-5 Mini | x 0.5 | x 4 |
| GPT-4.1 Mini | x 0.8 | x 3.2 |
| GPT-3.5 Turbo | x 1 | x 3 |
| GPT-5.4 Mini | x 1.5 | x 9 |
| GPT-5.6 Terra | x 2 | x 12 |
| GPT-5.6 Terra Pro | x 2 | x 12 |
| GPT-5.1 | x 2.5 | x 20 |
| GPT-5.1 Chat | x 2.5 | x 20 |
| GPT-5 | x 2.5 | x 20 |
| GPT-5.2 | x 3.5 | x 28 |
| GPT-4.1 | x 4 | x 16 |
| GPT-5.4 | x 5 | x 30 |
| GPT-4o | x 5 | x 20 |
| GPT-5.6 Sol | x 10 | x 60 |
| GPT-5.6 Sol Pro | x 10 | x 60 |
| GPT-5.5 | x 10 | x 60 |
| GPT-4 Turbo 128k | x 20 | x 60 |
| GPT-5 Pro | x 30 | x 240 |
| GPT-5.2 Pro | x 42 | x 336 |
| GPT-5.4 Pro | x 60 | x 360 |
| GPT-5.5 Pro | x 60 | x 360 |
| GPT-4 | x 60 | x 120 |
Perplexity
Name | Input modifier | Output modifier |
|---|---|---|
| Sonar | x 2 | x 2 |
| Sonar Pro | x 6 | x 30 |
Qwen
Name | Input modifier | Output modifier |
|---|---|---|
| Qwen3.7 Flash | x 0.06 | x 0.26 |
xAI
Name | Input modifier | Output modifier |
|---|---|---|
| Grok 4.3 | x 2.5 | x 5 |
| Grok 4.20 | x 2.5 | x 5 |
| Grok 4.6 | x 4 | x 12 |
| Grok 4.5 | x 4 | x 12 |
Z.ai
Name | Input modifier | Output modifier |
|---|---|---|
| GLM 5.3 Flash | x 0.15 | x 0.5 |
| GLM 4.7 | x 0.8 | x 3.5 |
| GLM 4.6 | x 1 | x 4 |
| GLM 4.5 | x 1.2 | x 4.4 |
| GLM 5.1 | x 1.93 | x 6.07 |
| GLM 5.2 | x 1.93 | x 6.07 |
Model prices change often, usually downwards. When they do, we update the multipliers so your tokens go further.
When tokens reset
Your allowance resets on the 1st of each calendar month (UTC), whatever day you subscribed. Unused tokens don't roll over to the next month.
What happens when you run out
When your team has used its whole monthly allowance:
- Your bots stop replying to new messages until the allowance resets on the 1st, or you add more tokens.
- In the website widget, visitors see No more chats available.
- On Slack, Discord, WhatsApp, Telegram and email, the bot replies that it has reached its message token limit.
- Through the API, requests return an error explaining that the team has used all available message tokens.
Chat Thing emails team owners and admins when your team has used 80%, 90% and 100% of its monthly message tokens, so you have time to act.
To keep your bots running:
- Add tokens straight away with a message tokens add-on (5,000,000 tokens per unit, available on paid plans).
- Upgrade your plan from Billing.
- Enterprise only: add your own OpenAI API key in the OpenAI section of your Account page. It's used only once your included tokens run out.
To stop one busy bot using up the whole team's allowance, set a daily token limit for it. See Usage limits.
See your usage
- Sidebar: the Storage tokens and Message tokens bars at the bottom of the sidebar show how much of this month's allowance your team has used.
- My Bots: the Message tokens this month card shows tokens used against your limit.
- Each bot's overview: Message tokens this month shows that bot's usage, with a chart over time.
Next steps
Related
- Storage tokens and how syncing uses them
What storage tokens are, when syncing and re-syncing use them, what happens when you reach your monthly limit, and how to use fewer.
- Add more tokens, bots or seats with add-ons
Raise one limit on your paid plan without upgrading - extra message tokens, storage tokens, bots, data sources, team members or custom domains.
- Limit how many tokens your bot can use
Cap a bot's daily message token use, get an email before it hits the cap, limit message length, and see what people see when a limit is reached.
- Choose your bot's AI model
See which AI models your bot can use, what each supports, how model choice affects message token use, and how to change it.
Last updated