Generative AI is now part of everyday work and learning.
But every prompt begins with a process most users never see: tokenization.
What is a token, and why does it matter?
When you type a prompt into a Generative AI application, the model does not receive the sentence in the same way that you read it.
It does not begin with words, grammar, or meaning. Before the model can process your text, a tokenizer first breaks it into smaller units called tokens.
This conversion affects what the model can receive, how much an API call may cost, how fast the system responds, and how we design prompts, RAG systems, and AI agents.
That is why tokenization is the first core concept I recommend learning when studying text-based Generative AI and Large Language Models.
A language model cannot process raw text directly

A language model cannot process raw text directly.

The text must first be divided into tokens. Each token is then mapped to a numerical identifier called a token ID.

The model’s embedding layer converts these token IDs into numerical representations that the neural network can process.

Although a generative AI application appears to work with sentences and paragraphs, the model receives a sequence of token IDs.

Tokenization bridges human language and model computation.

One token does not mean one word

A common misunderstanding is that one token always equals one word.

It does not.

A token can be a single character, part of a word, a complete word, punctuation, or a common character sequence.

In this example, “generative” may be divided into “gener” and “ative.” The period may become another token, and some tokens may even include a leading space.

The exact result depends on the tokenizer the model uses.

That is why word count, character count, and token count are not the same thing.

Why suppressing splits happen

Why might “generative” become “gener” and “ative”?

The tokenizer is not applying an English grammar rule. Instead, its vocabulary was created from patterns found in large amounts of text.

A character sequence such as “ative” appears in words like creative, innovative, informative, and collaborative. It may therefore become a reusable piece in the tokenizer’s vocabulary.

Frequently occurring character sequences are more likely to become larger, reusable tokens. Less common words, unusual names, and unfamiliar strings may be divided into smaller pieces.

The result may look like linguistic analysis, but it mainly comes from vocabulary construction and the tokenizer’s encoding algorithm.

Tokenization is not semantic understanding

It is important not to confuse tokenization with semantic understanding.

The tokenizer does not first understand a sentence and then decide how to divide it. Instead, it applies a predefined vocabulary and an encoding algorithm.

For example, in the sentence “Two-thirds of my money was spent on Coke,” one tokenizer may represent “-third” as one token and the final “s” as another.

This does not mean the tokenizer understands fractions. It only means that “-third” exists as a pattern in its vocabulary.

With the same tokenizer and the same input, tokenization is generally deterministic.

Language models generate tokens, not words!

Tokenization is not only used to process the input. It is also part of how a language model generates its response.

After receiving the input tokens, the model predicts a probability distribution for the next token.

Based on this probability distribution and the generation settings, one token is selected and added to the sequence. The model then predicts the next token. 

Finally, the generated token IDs are decoded back into readable text.

So a language model does not literally generate the next word. It generates the next token.

Tokens define the model's practical limits.

Tokens also define several practical limits of an AI application.

The context window is the total amount of information the model can work with in a single request. This budget may include system instructions, user input, conversation history, retrieved documents, tool results, and the generated answer.

If these elements use too many tokens, some content must be removed, summarized, or reorganized.

Many AI services also calculate usage and cost based on input and output tokens. Larger token volumes can increase both cost and response latency. 

This means token count is not just a technical detail. It is a resource that must be managed.

Major LLM workflows depend on tokenization

Once we understand tokenization, many other Generative AI concepts become easier to explain.

In prompt engineering, we balance useful information against the available token budget.

In Retrieval-Augmented Generation, documents should be divided using both meaningful boundaries and token limits, rather than character count alone.

In agent workflows, conversation memory, planning steps, and tool results all consume context.

Fine-tuning examples are also converted into token sequences during training.

Instructions from Skills and content returned through Model Context Protocol, or MCP, servers may also enter the model’s context. 

These technologies solve different problems, but they all operate within token-based limits. 

Tokenization changes how you design AI systems

The most important idea is simple.

A generative AI application begins with human language, but the model receives token IDs, converts them into numerical representations, and generates a response one token at a time.

Once you understand this, the context window is no longer an abstract limit. API cost becomes measurable. RAG chunking becomes an engineering decision. Agent memory becomes a resource that must be managed.

That is why tokenization is the first core concept to learn when studying text-based generative AI.

Before learning how to write better prompts, it helps to understand what the model actually receives.

This is AI Espresso, where we explore AI concepts and learn how to build practical AI applications.

See you next time!