A token is the basic unit of text that an AI model reads and writes — roughly a word-piece, often a fraction of a word — and because most AI vendors bill by the token, your cost scales with how much text you send in plus how much the model sends back.
If you have ever wondered why one long chat with an AI assistant felt expensive while a short question felt cheap, tokens are the reason. Understanding them turns AI pricing from a mystery into arithmetic you can actually plan around.
The mechanism is simpler than it sounds. Before a model can do anything, software called a tokenizer chops your text into tokens. Common words like "the" or "data" usually become one token each.
Longer or rarer words get split into pieces — "tokenization" might become "token" plus "ization." Punctuation, spaces attached to words, and even emoji all count. The model then processes those tokens, and every token you send (the input) and every token it generates (the output) gets counted.
That is why vendors quote two numbers: an input rate and an output rate. Output usually costs more because generating text is more computationally expensive than reading it. A short prompt with a long answer can therefore cost more than a long prompt with a one-line reply, which surprises almost everyone the first time.
Here is a concrete example. Suppose you paste a 2,000-word report into an AI tool and ask for a 200-word summary. Your input is roughly a few thousand tokens once you account for how words split, and your output is a few hundred.
Now suppose you instead ask the same tool to summarize that report and then rewrite it in three different styles, each 500 words. Your input barely changes, but your output roughly triples. If the vendor charges per token, your bill for the second request is meaningfully higher even though you typed almost the same prompt.
The lesson: in most AI pricing, output length is the lever you control most directly. Asking for "a concise summary" instead of "a thorough rewrite" is a real cost decision, not just a style preference.
Why does the same sentence cost different amounts across different tools? Three reasons. First, different models use different tokenizers, so identical text can split into different token counts.
Second, vendors set different per-token rates based on model size and capability — a small, fast model is typically cheaper per token than a large reasoning model. Third, some tools bundle tokens into subscription tiers with monthly allowances, so your effective cost depends on whether you stay inside the allowance.
According to our AI tool database, which tracks 360 AI tools with pricing and capability snapshots recorded at verification time (most recently 2026-09-18), pricing models vary widely across the category — some charge per token, some per seat, some per generation. That variety is exactly why you cannot assume two tools with similar features cost the same.
The limits matter here. Token math tells you how a vendor's meter runs, but it does not tell you the final bill, because rates change frequently and vendors add minimums, rounding, caching discounts, and free tiers that complicate the picture. We deliberately are not quoting per-token prices or token-per-word ratios, because those figures shift and vary by model — the vendor's own pricing page is the only reliable source.
A useful habit: before committing to a workflow that processes large documents, check whether the tool offers prompt caching or a cheaper model tier for bulk work. Many do, and it can cut costs substantially without changing your output quality. Also watch for hidden token consumers — system prompts, conversation history that gets resent every turn, and retrieved context in tools that search your files.
A long chat thread can quietly resend thousands of tokens on every message. If you want to go deeper on how these systems work under the hood, our guide on what an AI model actually means explains the pieces that sit around the tokenizer.
The practical takeaway: treat tokens like minutes on a phone plan. Short, focused prompts and tight output requests keep the meter low; sprawling threads and open-ended "write me everything" requests run it up fast.