AI tokens explained: the hidden meter behind every AI bill
You pay for AI in tokens, not words, and tokens quietly decide both your bill and how much a model can read at once. Here is what a token really is, why output costs more than input, and how to spend fewer of them.
Published

If you have used an AI model or its API, you have paid for tokens, whether you noticed or not. Tokens are the unit large language models actually read, generate and bill by, and they are not the same as words. Understanding them is the difference between guessing at your AI costs and controlling them.
What a token actually is
A token is a chunk of text, usually a piece of a word rather than a whole one. Models do not see letters and words the way people do. They convert text into a sequence of numbers, do the math on those numbers, and convert the result back into text. OpenAI's own tokenizer library puts it plainly: language models see a sequence of numbers known as tokens, not the text itself.
Those tokens are produced by a method called byte pair encoding, which starts from raw bytes and repeatedly merges the most common adjacent pairs into single tokens. The result is a vocabulary of frequent characters, subwords and whole words. Common words often become a single token, while rarer words split into pieces. OpenAI's help documentation gives a neat example: the phrase "ChatGPT is great!" encodes into six tokens, split roughly as Chat, G, PT, is, great and the exclamation mark. Notice that spaces attach to the tokens after them, which is why formatting and whitespace also cost tokens.
The four-characters rule of thumb
For everyday English, there is a simple estimate. OpenAI's guidance is that one token is about four characters, or roughly three-quarters of a word, so 100 tokens is about 75 words. Google publishes a compatible rule for its Gemini models, putting a token at about four characters and 100 tokens at 60 to 80 English words. Both vendors agree on the four-characters figure.
“Byte pair encoding (BPE) is a way of converting text into tokens. It's reversible and lossless, so you can convert tokens back into the original text.”
The important caveat, which the vendors state themselves, is that these are estimates, not exact counts. The real number changes with the language, the model and the content. The only way to get an exact count is to run the specific model's tokenizer, such as OpenAI's open-source tiktoken.
Why tokens cost you money
Tokens matter first because you pay for them, and you pay for them twice over. API usage is billed per token, and crucially, input and output are priced separately, with output almost always the expensive half. As a dated snapshot from late September 2026, OpenAI listed GPT-4o at 2.50 dollars per million input tokens and 10 dollars per million output tokens, four times as much for output. Anthropic listed Claude Sonnet 5 at 2 dollars input and 10 dollars output, a five-times gap. Prices change often, so always check the vendor's page, but the shape holds: a long generated answer costs far more than an equally long prompt.
Both halves add up in ways that surprise people. A single request is billed on everything you send, including the system prompt, your instructions, any retrieved context and the entire chat history, plus everything the model writes back. In a multi-turn conversation the history is re-sent every turn, so token cost quietly compounds as the chat grows. The full mechanics of that meter, from why output costs more to how caching cuts it and why one model costs many times another, are the subject of our companion guide to how AI pricing works.
Tokens set the context limit too
The second reason tokens matter is the context window, which is the maximum number of tokens a model can consider at once, input and output together. It is always measured in tokens. Recent models illustrate the range: GPT-4o has a 128,000-token window, while GPT-4.1, Google's Gemini 2.5 Pro and current Claude models reach up to a million tokens. This is closely tied to how much a model can handle in one go, which is also the ceiling that techniques like retrieval-augmented generation are designed to work within. Every token in that window competes for the same budget, and filling a big window costs more because you pay per token to send all of it.
Why code and other languages cost more
Because tokenizers are trained mostly on English, the same meaning in another language, or in source code, frequently breaks into more tokens. That means higher cost and less room in the context window for the same content. A peer-reviewed study at the EMNLP 2023 conference documented large per-language disparities in token counts for equivalent text. Newer tokenizers narrow the gap: OpenAI's o200k encoding, used by GPT-4o, roughly doubles the vocabulary of the older one, with the biggest efficiency gains on non-English text and code. The direction is reliable even if any single multiplier you see online is not.
How to spend fewer tokens
Once you think in tokens, the cost levers are obvious. Trim the prompt of boilerplate and unused examples. Cap the output length and ask for concise or structured answers, since output is the pricier side. In chat apps, summarise or truncate old turns instead of re-sending the full transcript. Right-size the model, using a cheaper option like GPT-4o mini or Claude Haiku for easy tasks, a point that matters whenever you compare options as in our look at pricing across the GPT-6 family. Use prompt caching for stable, repeated context, batch non-urgent jobs for roughly half off where offered, and with retrieval, send only the most relevant chunks rather than whole documents.
The bottom line
A token is the hidden meter behind every AI product: the unit that decides both your bill and how much a model can read at once. It is not a word, it is usually a fragment of one, and it behaves differently for code and for languages other than English. You do not need to count them by hand, but knowing roughly four characters to a token, that output costs several times input, and that the context window is a token budget, is enough to estimate costs before you spend and to cut them afterward.
Frequently asked questions
What is a token in AI?
A token is a chunk of text, usually part of a word, that an AI model reads and generates. Models convert text into tokens, process the numbers, and convert them back into text. Common words are often one token; rarer words split into several.
How many words is a token?
In common English, one token is about four characters, or roughly three-quarters of a word, per OpenAI's rule of thumb. That means 100 tokens is about 75 words. It is an estimate that varies by language, model and content.
Why do output tokens cost more than input tokens?
Providers price the two separately, and generating text is more compute-intensive than reading it. In late September 2026, output ran about four to five times the input rate on models like GPT-4o and Claude, so long answers cost far more than long prompts.
What is a context window in tokens?
The context window is the maximum number of tokens a model can consider at once, counting the system prompt, chat history, any retrieved text and the answer together. Recent models range from 128,000 tokens to about a million.
Why does non-English text or code cost more tokens?
Tokenizers are trained mostly on English, so the same meaning in another language or in source code often splits into more tokens. That raises cost and uses more of the context window. Newer, larger-vocabulary tokenizers narrow the gap.
How can I reduce token costs?
Trim the prompt, cap the output length, and prune old chat history instead of re-sending it. Right-size the model to the task, use prompt caching for repeated context, batch non-urgent jobs, and with retrieval send only the most relevant chunks.
Sources
What each one is, and whose it is.
- 1
What are tokens and how to count them?, OpenAI Help Center (September 27, 2026)
Documentation - 2
tiktoken (byte pair encoding tokenizer), OpenAI / GitHub
Documentation - 3
Understand and count tokens (Gemini API), Google AI for Developers
Documentation - 4
Pricing, Anthropic (September 27, 2026)
Documentation - 5
Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models, EMNLP 2023 (arXiv) (May 23, 2023)
PaperIndependent of the vendor