An AI token is a small chunk of text — often a word piece, part of a word, or a punctuation mark — that a language model reads and writes in, and because most AI services bill by the token, tokens are the unit that decides what you actually pay.
When you send a prompt to a tool like ChatGPT or Claude, the service does not count words. It counts tokens, and the price on the vendor's page is usually quoted per million of them, split into a cheaper rate for input and a more expensive rate for output.
The mechanism is worth understanding because it explains a lot of confusing bills. Before a model can process your text, a tokenizer breaks it into those chunks and maps each one to a number the model can handle. English text roughly maps to words, but not neatly: short common words are often one token, while long or unusual words split into several, and punctuation and spaces get counted too.
That is why the same paragraph in two languages can cost different amounts — a language the tokenizer handles less efficiently produces more tokens for the same meaning. Output tokens cost more than input tokens at most vendors because generating text is heavier work than reading it.
And every model has a context window, a maximum number of tokens it can hold at once across your prompt and its reply, so a long document plus a long answer can hit a ceiling even when the price per token looks tiny.
Here is a concrete way to think about it. Suppose you run a support inbox and want an AI to draft replies. You paste in the customer's message plus some help-article context, and the model writes back a short reply.
The input side is everything you send — the customer message, the instructions, the pasted context — and the output side is only the draft. If you double the pasted context, you roughly double the input side of the bill while the output stays about the same. That single insight is why trimming context is usually the cheapest optimisation available: cutting a bloated system prompt or dropping irrelevant pasted text reduces input tokens on every single call, not just once.
The exact token count for any specific string is not something you can eyeball — you need the vendor's own tokenizer or its pricing calculator, because different vendors split text differently.
According to our AI tool database, which holds a verified pricing and capability snapshot of 360 AI tools (most recent verification 2026-09-18), the market splits into two billing styles: some tools bill per token through an API, while others bundle usage into a flat subscription. That distinction drives a real decision rule.
If your monthly volume is small and steady — say a handful of drafts a day — a subscription with bundled usage is usually simpler and easier to predict, because you are not watching a meter. If your volume is large or spiky, or you are embedding AI into a product where every user action triggers a call, per-token API billing tends to be cheaper at scale but demands monitoring, because a runaway loop or a giant pasted document can burn through budget fast.
The honest limit: per-token pricing changes frequently, and vendors restructure tiers and rates often, so the only reliable number is the one on the vendor's own pricing page on the day you check. This article cannot give you a current rate, and neither can any snapshot — treat our database entry as a starting point for comparison, not a quote.
One tip that goes beyond the obvious: watch the ratio of input to output in your own use case before choosing a billing model. A summarisation tool is input-heavy and output-light, so it benefits from cheap input rates. A creative writing tool is the reverse — small prompts, long outputs — so output pricing dominates its cost. Knowing which side of the meter your workload leans on tells you which plan structure actually fits, and it is a question most people never ask before signing up.