Skip to content
Cuswpc Tech

70 lettersUpdated

New message
Folders

What Are Tokens in AI? Meaning, Context Windows and Limits

Filed by
Tech
Received
Length
4 min
Media, letter tiles and word

When an AI tool says a document is too long, it is not counting pages or words. It is counting tokens. A token is the small unit of text a large language model reads and writes: sometimes a whole word, sometimes part of one, sometimes a single punctuation mark. The model converts every piece of text into tokens, works with them as numbers and writes its reply one token at a time. Token counts decide how much fits into one conversation, how quickly an answer arrives and, for paid access, how much a request costs.

How text becomes tokens

A component called a tokenizer splits text using a fixed vocabulary that was built from the model's training data. Frequent words usually become a single token. Rare words, names, typos and long technical terms are broken into several fragments. Each fragment is then mapped to an ID number, and those numbers are what the model actually processes.

Because each model family uses its own tokenizer, the same sentence can produce different token counts in different tools. Some kinds of text are reliably more "expensive" than others:

  • Long, rare or invented words, including product names and jargon
  • Strings of digits, such as account numbers, order IDs and phone numbers
  • Many languages other than English, depending on the tokenizer
  • Formatting: tables, code, repeated spaces and decorative characters

Spaces are usually attached to the start of the following word rather than counted on their own, while punctuation often becomes a separate token.

Tokens and the context window

The context window is the maximum number of tokens a model can take into account at once. It includes everything: standing instructions, your prompt, any pasted or attached documents, earlier messages in the conversation and the reply being written. When a conversation grows beyond the limit, the oldest material is cut or condensed. That explains a familiar experience: an instruction given at the start of a long chat quietly stops being followed.

The same limit governs how much of a long file a tool can read in one pass. If you work with lengthy reports, the section-by-section approach in how to summarize a PDF with AI keeps each piece within range.

Why tokens matter in everyday work

SituationWhat tokens affectPractical move
Pasting a long reportPart of it may not fit, or may be skimmedWork section by section, then combine
A chat that has run for hoursEarly instructions fall out of the windowRestate key rules, or start a fresh chat with a short recap
Asking for exactly 200 wordsThe model plans in tokens, not wordsAsk for a range and trim by hand
Paying for access through an APIInput and output tokens are often priced separatelyKeep prompts focused and ask for concise output
Waiting for a replyLonger outputs take longer to generateRequest a short version first and expand if needed

Tokens, words and characters

There is no fixed conversion between tokens and words. In English, a token is on average a little shorter than a word, so a passage usually contains more tokens than words. In other languages, or in text full of numbers and code, the gap can be much wider. When precision matters, use the token counter that many providers publish or display rather than a rule of thumb.

Tokens also explain some odd behavior. Models sometimes stumble when asked to count the letters in a word or spell it backwards, because they never see individual letters, only fragments. Clear instructions help, and our guide on how to write AI prompts lists other habits that avoid wasted tokens and wasted time.

Other meanings of "token" in technology

The word turns up in several other places, which causes confusion:

  • Authentication tokens are temporary digital keys that prove you have already signed in, so you are not asked for a password on every page. Hardware security keys and code generators are sometimes called tokens too; our explainer on two-factor authentication covers them.
  • Payment tokens replace a real card number with a stand-in value, so the actual number does not have to be stored by every shop.
  • Crypto tokens are digital assets recorded on a blockchain.

None of these is related to the units a language model reads. Context usually makes the meaning clear.

Frequently asked questions

Is a bigger context window always better?

It lets a model hold more material, but very long inputs can still lead to details in the middle receiving less attention. A focused prompt containing only the relevant sections often produces a sharper answer.

Does the model remember tokens from yesterday's chat?

Not by default. Each conversation starts fresh unless the tool has a separate memory feature that stores notes and feeds them back in, where they count as tokens again.

Why does the same question sometimes cost more?

Later in a long conversation, earlier messages are sent along as context, so each new reply processes more tokens than the one before.

In the same threadTech