How AI pricing works, and why the same job can cost pennies or dollars.
AI tools bill by the token, not by the task, which is why the same job can cost pennies in one app and real money in another. Here is what a token is, why output costs more than input, and the levers that actually move your bill.
By Yash Malviya
Published

You are not paying for answers, you are paying for tokens
Almost every AI tool and API bills the same way: by the token, not by the task. A token is a chunk of text the model reads and writes, the small unit we unpack in what an AI token is, and it is smaller than most people expect. As Anthropic's pricing documentation puts it, "As a rough estimate, 1 token is approximately 4 characters or 0.75 words in English." OpenAI's developer guide uses the same rough rule, characters divided by four. So a tidy 100-word email is around 130 tokens, and a 20-page report runs well past 12,000. Providers quote rates per million tokens, and they split the meter in two: input tokens, which is everything you send, including the system prompt, any documents and the running chat history, and output tokens, which is everything the model writes back. Get those two ideas straight and most of the mystery about AI bills disappears.
Why output usually costs more than input
The two halves of the meter are not priced the same. Output tokens typically cost several times more than input tokens, often around four to five times as much on frontier models. The reason is mechanical. A model reads your input in a single pass, but it generates output one token at a time, each new token depending on every token before it, and that step-by-step generation is the expensive part of running the model. So a task that reads a long document and returns a one-line answer is cheap, while one that reads a little and writes a lot, such as drafting long reports from a short brief, is where the bill climbs. If a model charges, hypothetically, ten dollars per million input tokens and fifty per million output, the shape of your task, not just its size, decides what you pay.
The context window is a meter that runs every turn
The context window is the total span of tokens a model can consider at once, and it is easy to mistake it for free headroom. It is not. You pay for every token in the window on every call. In a conversation, the entire history is resent each turn, so a chat that feels effortless is quietly reprocessing everything said so far, over and over. Anthropic describes prompt caching as a way to avoid "reprocessing the same large system prompt, document, or conversation history on every request," which is a plain admission of how the default works. Long contexts can carry a second cost: some providers charge a higher per-token rate above a size threshold. Google, for one, prices prompts over 200,000 tokens at a higher input rate than shorter ones. A bigger context is a capability, and a cost.
“As a rough estimate, 1 token is approximately 4 characters or 0.75 words in English.”
Caching, and why the same prompt can cost a fraction as much
If you send the same long prefix again and again, a fixed system prompt, a knowledge base, a whole codebase, prompt caching is the biggest discount on the menu. Once a prefix has been processed, later reads of it are billed at a steep discount, often around 90% below the normal input rate, and sometimes more on newer models. OpenAI turns this on automatically for supported models, so reused prompts are "discounted up to 90%" with no code changes, while Anthropic lets you mark exactly which parts to cache. There is a catch worth knowing: the first write to the cache can cost a little more than a normal read, so caching pays off when the same context is reused across many calls, which is exactly the pattern in agents and long chats, and does nothing for one-off prompts.
Why one model costs many times another
Prices vary across models for reasons that track real differences, not branding. Larger, more capable models cost more per token than small fast ones, which is why vendors tell you to match the model to the job. Anthropic's own guidance is to use its smallest model "for simple tasks," a mid-tier one "for most production workloads," and its largest "for the most complex reasoning." Speed is a separate lever, and some providers sell a faster response tier at premium rates. The gap between the cheapest and priciest options is wide enough that periodic API price cuts can reset unit economics for teams running at scale. The practical point is that the model you choose is usually a larger cost decision than any prompt you write.
How to actually control the bill
Costs are controllable once you can see them. Estimate before you build, because tokens times the per-million rate, counted separately for input and output, gets you close. Then pull the obvious levers. Right-size the model, since a smaller one often handles routine work at a fraction of the cost. Trim the context you send instead of pasting whole documents into every call. Cap output length so the model cannot wander onto the expensive side of the meter. Cache the prefixes you reuse. For work that is not time-sensitive, batch it, since providers commonly offer around 50% off asynchronous batch jobs. Then watch the usage dashboard, because the token counts that surprise you are the ones worth cutting. None of this is exotic, and together it routinely halves a bill.
Our take
AI pricing looks opaque and turns out to be arithmetic: tokens times a rate, with output dearer than input, the full context billed on every turn, and a caching discount most people forget to claim. Treat the headline per-token numbers with suspicion, because they change constantly and they are not where your money actually goes. The real bill is set by how much context you resend, how much output you ask for, and which model you point at the task. Measure those three and the same workload that felt unpredictable becomes something you can budget to the dollar.
Frequently asked questions
Why is AI billed per token instead of per request?
Almost every AI tool and API meters usage by the token, not by the task. A token is a chunk of text a model reads and writes, roughly four characters or about 0.75 words in English. Providers quote rates per million tokens and split the meter in two: input tokens (everything you send, including the system prompt, documents and chat history) and output tokens (everything the model writes back).
Why does output cost more than input?
Output tokens typically cost several times more than input, often around four to five times as much on frontier models. The reason is mechanical: a model reads your input in a single pass, but it generates output one token at a time, each new token depending on every token before it, and that step-by-step generation is the expensive part.
Does a bigger context window cost more?
Yes. You pay for every token in the context window on every call, and in a conversation the entire history is resent each turn. Some providers also charge a higher per-token rate above a size threshold. Google, for example, prices prompts over 200,000 tokens at a higher input rate than shorter ones.
What is prompt caching and how much does it save?
If you send the same long prefix repeatedly (a fixed system prompt, a knowledge base, a whole codebase), prompt caching bills later reads of it at a steep discount, often around 90% below the normal input rate. The first write to the cache can cost slightly more, so caching pays off when the same context is reused across many calls, as in agents and long chats, and does nothing for one-off prompts.
How do I actually lower my AI bill?
Right-size the model, since a smaller one often handles routine work at a fraction of the cost. Trim the context you send instead of pasting whole documents, cap output length, and cache the prefixes you reuse. For work that is not time-sensitive, batch it, since providers commonly offer around 50% off asynchronous batch jobs. Then watch the usage dashboard, because the token counts that surprise you are the ones worth cutting.
Sources
What each one is, and whose it is.
- 1
Pricing, Anthropic
Documentation - 2
Prompt caching, OpenAI
Documentation - 3
Gemini API pricing, Google
Documentation