Introduction
Every time you type a message into ChatGPT, Claude, or any other AI tool, something happens behind the scenes that most people never think about: your words get chopped up into pieces.
Those pieces are called AI tokens. And understanding them - even at a basic level - changes how you think about what AI can and can't do.
Tokens are how AI reads
Humans read words. AI reads tokens.
A token is a chunk of text that a language model uses as its basic unit of processing. If you've ever wondered what are tokens in AI, this is the simplest answer: they are the pieces of text the model reads, counts, and generates.
Sometimes a token is a whole word. Sometimes it's part of a word. Sometimes it's a single character or a piece of punctuation.
For example, the word "understanding" might be split into two tokens: "understand" and "ing." The word "cat" is usually one token. A long, uncommon word like "cryptocurrency" might become three or four tokens.
This happens through a process called AI tokenization - where your input text gets converted into a sequence of tokens that the model can process. It's the very first thing that happens when you send a prompt to an AI, and the very last thing that happens when the AI generates a response. The model doesn't think in words. It thinks in LLM tokens.
Why it matters: the context window
Here's where tokens become practically important.
Every AI model has a limit on how many tokens it can process at once. This limit is called the AI context window. It includes both your input - the prompt, any documents you've attached, the conversation history - and the model's output.
If a model has a 128,000-token context window, that's the total budget for everything: what you send in and what comes back out. Once you hit the limit, the model can't process any more information. It either cuts off or starts losing track of earlier parts of the conversation.
This is why AI tools sometimes "forget" things you said earlier in a long conversation. It's not that the AI is being careless. It's that the conversation has exceeded the token budget, and older tokens have been pushed out.
For tools like ChatGPT, these token limits determine how much conversation history the model can remember at once.
Understanding what is context window in LLM explains a lot of frustrating AI behavior. When a model seems to ignore instructions you gave at the beginning of a long conversation, or when it loses track of a document you uploaded, the answer is almost always: tokens ran out.
How big is a token?
A rough rule of thumb for English text: one token is approximately three-quarters of a word. So if you're asking how many words is a token, the practical answer is that 100 words is roughly 130-140 tokens.
But it varies. Common English words tend to be single tokens. Rare words, technical jargon, non-English languages, and code often get split into more tokens. This means the same amount of meaningful content can cost very different amounts of tokens depending on what language it's in or how technical it is.
A simple paragraph of conversational English might be 50 tokens. The same amount of information expressed in dense legal language or Python code might be 80-100 tokens.
This isn't something most casual users need to worry about. But if you're working with AI in a professional context - building applications, processing documents, or running AI at scale - token counts directly affect performance, cost, and what's possible.
Tokens and cost
For AI APIs, where developers build applications using AI models, tokens are how you get charged. You pay per token - both for the tokens you send in, known as input tokens, and the tokens the model generates back, known as output tokens.
This is where AI token cost becomes important.
A long, detailed prompt costs more than a short one. A response that generates a full report costs more than one that gives a two-sentence answer. And processing a 50-page document costs more than processing a one-page summary.
Different models have different pricing per token. More capable models typically cost more per token. This creates a practical trade-off: you can use a powerful model for complex tasks (and pay more), or a lighter model for simple tasks (and pay less).
For individual users of tools like ChatGPT or Claude, this happens behind the scenes. Your subscription covers a certain amount of usage. But for businesses building AI into their products, token economics is one of the most important factors in whether an AI application is financially viable.
Why tokenization isn't perfect
Tokenization works well for common English text, but it has limitations worth knowing about.
Non-English languages often require more tokens. Because many major language models were trained heavily on English text, their tokenizers are often more efficient with English. The same sentence in Japanese, Arabic, or Hindi might require significantly more tokens than its English equivalent, which means higher costs and faster context window exhaustion for non-English users.
Code tokenizes differently than prose. Programming languages have their own patterns - brackets, indentation, variable names - that don't map neatly onto natural language tokenization. This is why specialized coding models often use different tokenizers than general-purpose language models.
Numbers can be unpredictable. Large numbers, dates, and mathematical expressions can tokenize in unexpected ways. The number "123456789" might become multiple tokens, and the model may not "understand" it as a number the way humans do. This is one reason AI models sometimes struggle with precise arithmetic.
Typos and unusual text cost extra. Because tokenizers are built around common word patterns, misspelled words or unusual formatting often get split into more tokens than correctly spelled text. The model still processes them, but less efficiently.
What this means for how you use AI
A few practical implications:
Longer isn't always better. A well-structured, concise prompt often produces better results than a sprawling one - partly because it leaves more of the context window available for the model's response, and partly because shorter, clearer inputs give the model less room for misinterpretation.
Conversation history has a cost. In a long conversation, the model is processing the entire history every time it generates a new response. This is why AI tools can slow down or lose coherence in very long threads. Starting a new conversation when you switch topics isn't just tidier - it can produce better results.
Document size matters. When you upload a large document for an AI to analyze, you're using a significant chunk of the context window. If the document is too large, the model may not be able to process it all at once, or it may not have enough room left to generate a thorough response.
Be aware of the budget. When you're using AI for anything important, think of tokens as a budget. You have a finite amount of space for your input and the model's output. The more wisely you use that space - with clear, focused prompts and relevant context - the better the results.
The bottom line
AI tokens are the invisible currency of every interaction. They determine what the model can process, how much it costs, and where its limits are.
You don't need to count them manually. But understanding that they exist - and that every AI conversation is constrained by a token budget - makes you a more effective user of these tools. Most of the moments where AI seems to "break" - forgetting context, producing incomplete answers, struggling with long documents - trace back to tokens.
The model isn't confused. It's just out of budget.
Mar 10, 2026 - 6 min read
