TempatGunting

Tokenization Deep Dive: BPE, WordPiece and SentencePiece

Every interaction with a large language model starts with tokenization — the process of converting raw text into a sequence of integer IDs that the model can process. This step determines how much your API calls cost, how much context fits in the model's window, and whether your model handles code, math, or non-English languages well. Yet most engineers treat the tokenizer as a black box. Understanding how BPE, WordPiece, and SentencePiece actually work gives you leverage over model performance that no amount of prompt engineering can match.

Why Subword Tokenization Won

Before subword tokenization, NLP systems used one of two approaches: word-level tokenization (splitting on whitespace/punctuation) or character-level tokenization (one token per character). Both have fatal flaws for language models.

Word-level tokenization creates vocabularies of hundreds of thousands to millions of entries to cover a language. Even then, any word not in the vocabulary becomes an [UNK] token — a complete information blackout. New proper nouns, technical terms, misspellings, and morphological variants all become [UNK]. And the embedding matrix — which stores a learned vector for every vocabulary entry — becomes the dominant parameter cost.

Character-level tokenization solves the vocabulary problem (26 lowercase letters + digits + punctuation covers English) but creates a new one: sequences become extremely long. The word "tokenization" becomes 12 tokens instead of 1-2, and since transformer attention scales quadratically with sequence length, this makes training and inference expensive. Worse, individual characters carry minimal semantic information — the model must learn to reconstruct word meaning from character patterns, which wastes model capacity.

Subword tokenization is the pragmatic middle ground. Common words like "the" and "is" remain single tokens. Rare words decompose into meaningful subparts: "tokenization" might split into "token" + "ization", preserving morphological structure. The vocabulary stays manageable (32K-100K entries), sequences stay short, and no word is ever [UNK]. This is why every major LLM uses some variant of subword tokenization. For understanding how these tokens feed into the transformer, see our attention mechanism guide.

Byte Pair Encoding (BPE)

BPE was originally a data compression algorithm (Gage, 1994) adapted for NLP by Sennrich et al. (2015). The algorithm is elegant in its simplicity:

  1. Start with a vocabulary of individual characters (or bytes).
  2. Count all adjacent pairs of tokens in the training corpus.
  3. Merge the most frequent pair into a new token.
  4. Repeat steps 2-3 for a desired number of merge operations.
BPE Merge Process — "lowest" → ["l","o","w","e","s","t"] Step 0 l o w e s t 6 tokens Step 1 l o w es t 5 tokens (merged e+s) Step 2 l o w est 4 tokens (merged es+t) Step 3 lo w est 3 tokens (merged l+o) Final low est 2 tokens — "lowest"

The vocabulary after training consists of the base characters plus all merged tokens, ordered by their merge priority. At inference time, the tokenizer applies merges in the same order they were learned during training.

import tiktoken

enc = tiktoken.encoding_for_model("gpt-4o")

text = "Distributed gradient accumulation across data-parallel ranks"
tokens = enc.encode(text)
print(f"Text: {text}")
print(f"Tokens: {tokens}")
print(f"Token count: {len(tokens)}")
print(f"Decoded: {[enc.decode([t]) for t in tokens]}")

# Output:
# Text: Distributed gradient accumulation across data-parallel ranks
# Tokens: [85597, 37836, 58229, 4858, 918, 36646, 26759]
# Token count: 7
# Decoded: ['Distributed', ' gradient', ' accumulation', ' across', ' data', '-parallel', ' ranks']

Byte-Level BPE

Modern BPE implementations (GPT-2, GPT-4, Llama) operate on bytes rather than Unicode characters. This eliminates the need for a base character vocabulary that covers all Unicode codepoints — the base vocabulary is always exactly 256 byte values. Non-ASCII characters (Chinese, emoji, special symbols) decompose into their UTF-8 byte sequences and get merged through the same BPE process. The result: every possible input has a valid tokenization, with no [UNK] tokens ever.

OpenAI's tiktoken library implements byte-level BPE with highly optimized Rust internals. It's 3-6× faster than the Hugging Face tokenizers library for the same GPT tokenizers, which matters for preprocessing large training datasets.

WordPiece

WordPiece (Schuster and Nakajima, 2012) is used by BERT, DistilBERT, and ELECTRA. It's similar to BPE but uses a different merge criterion: instead of merging the most frequent pair, WordPiece merges the pair that maximizes the log-likelihood of the training data when treated as a language model.

The practical difference from BPE: WordPiece uses the "##" prefix to denote non-initial subwords. "tokenization" might become ["token", "##ization"]. This prefix makes it explicit which subwords are word-initial and which are continuations, which helps the model distinguish morphological roles.

from transformers import AutoTokenizer

bert_tok = AutoTokenizer.from_pretrained("bert-base-uncased")
tokens = bert_tok.tokenize("tokenization improves efficiency")
print(tokens)
# ['token', '##ization', 'improves', 'efficiency']

gpt_tok = AutoTokenizer.from_pretrained("gpt2")
tokens = gpt_tok.tokenize("tokenization improves efficiency")
print(tokens)
# ['token', 'ization', 'Ġimproves', 'Ġefficiency']
# Note: Ġ represents a leading space in GPT-2's tokenizer

In practice, BPE and WordPiece produce nearly identical vocabularies for the same training data and vocabulary size. The choice between them is more about ecosystem compatibility than algorithm superiority. GPT-family models use BPE; BERT-family models use WordPiece.

Unigram Language Model

Unigram (Kudo, 2018) takes the opposite approach from BPE. Instead of building up a vocabulary by merging characters, it starts with a large vocabulary of candidate subwords and iteratively removes the ones that contribute least to the training data likelihood.

  1. Initialize with a large candidate vocabulary (e.g., all substrings up to length 16 that appear in the training data).
  2. For each candidate token, compute how much removing it would decrease the total log-likelihood of the training corpus.
  3. Remove the bottom 20-30% of candidates by impact.
  4. Repeat until the vocabulary reaches the target size.

The key advantage of Unigram: it produces multiple valid tokenizations for each input and assigns probabilities to each one. During training, you can sample from these tokenizations as a form of data augmentation. The Unigram algorithm is used by T5, ALBERT, and mBART. For understanding how these tokenized sequences are then processed by training optimization techniques, the token sequence length directly affects batch sizing and memory requirements.

SentencePiece: Language-Independent Tokenization

SentencePiece (Kudo and Richardson, 2018) is not a tokenization algorithm — it's a framework that implements both BPE and Unigram as backends. Its distinguishing feature is treating text as a raw byte stream without any language-specific pre-processing. No whitespace splitting, no punctuation rules, no Unicode normalization assumptions.

This matters because pre-tokenization rules are inherently language-dependent. English splits on spaces, but Chinese and Japanese don't use spaces. German compounds like "Donaudampfschifffahrt" shouldn't be split on whitespace. SentencePiece handles all of these uniformly by operating directly on the raw text.

import sentencepiece as spm

# Training a SentencePiece model
spm.SentencePieceTrainer.train(
    input="training_data.txt",
    model_prefix="my_tokenizer",
    vocab_size=32000,
    model_type="bpe",  # or "unigram"
    character_coverage=0.9995,
    byte_fallback=True,  # Handle unknown bytes
    split_by_whitespace=True,
    normalization_rule_name="identity",
)

sp = spm.SentencePieceProcessor()
sp.load("my_tokenizer.model")

text = "Distributed training with FSDP"
tokens = sp.encode(text, out_type=str)
print(tokens)
# ['▁Distribut', 'ed', '▁training', '▁with', '▁F', 'SD', 'P']

ids = sp.encode(text)
print(ids)
# [12543, 287, 3847, 411, 383, 7116, 52]

SentencePiece uses the "▁" (Unicode U+2581, lower one eighth block) character to denote spaces, allowing lossless round-trip conversion between text and tokens. LLaMA, Mistral, and T5 all use SentencePiece tokenizers.

Vocabulary Size: The Hidden Tradeoff

Vocabulary size is one of the most consequential model design decisions, yet it gets less attention than layer count or hidden dimension.

ModelTokenizerVocab SizeAvg Tokens/Word (English)
GPT-2BPE50,2571.3
GPT-4 / GPT-4oBPE (tiktoken)~100,2561.1
BERTWordPiece30,5221.4
LLaMA / MistralSentencePiece BPE32,0001.3
LLaMA 3tiktoken BPE128,2561.0
GemmaSentencePiece256,0000.9
T5SentencePiece Unigram32,0001.4

Larger vocabularies produce shorter token sequences. GPT-4's 100K vocabulary tokenizes English text ~15% more efficiently than GPT-2's 50K vocabulary. This means more text fits in the context window and API calls cost less per word. Gemma's 256K vocabulary pushes this further, particularly for non-English languages where smaller vocabularies create disproportionately long sequences.

The cost: each vocabulary entry adds one row to the embedding matrix and one column to the output projection layer. Going from 32K to 256K vocabulary adds ~450M parameters for a model with 2048-dimensional embeddings. For small models, this overhead is significant — Gemma-2B's vocabulary accounts for ~25% of its total parameters. This relationship between vocabulary and model parameters connects directly to how quantization techniques handle embedding layers.

Tokenization and Non-English Languages

Tokenization efficiency varies dramatically across languages. A 32K BPE vocabulary trained primarily on English text produces approximately 1.3 tokens per English word but 3-5 tokens per word in languages like Chinese, Japanese, Thai, and Arabic. This means non-English users pay 2-4× more per API call and get 2-4× less context window capacity.

Tokenization Efficiency: Tokens Per Word by Language English Spanish German Chinese Japanese 1.3 1.7 2.0 3.8 4.7 LLaMA 32K vocabulary — lower is better

LLaMA 3 and Gemma address this with much larger vocabularies (128K and 256K respectively) that allocate more vocabulary entries to non-English scripts. The tradeoff is increased model size, but for multilingual models this is almost always worth it. Building multilingual capabilities on top of these tokenizers connects to the challenges covered in our embedding models comparison.

Practical Implications for Production

Token Counting for Cost Estimation

API pricing is per-token, so accurate token counting is essential for cost modeling. Use the provider's official tokenizer, not word-count heuristics:

import tiktoken

def estimate_cost(text, model="gpt-4o", per_1k_input=0.005):
    enc = tiktoken.encoding_for_model(model)
    token_count = len(enc.encode(text))
    cost = (token_count / 1000) * per_1k_input
    return token_count, cost

# Compare tokenization across models
text = "The quick brown fox jumps over the lazy dog"
for model in ["gpt-4o", "gpt-3.5-turbo"]:
    count, cost = estimate_cost(text, model)
    print(f"{model}: {count} tokens, ${cost:.6f}")

Tokenization for Retrieval

When chunking documents for vector search, chunk by tokens rather than characters or words. A chunk of 512 tokens maps directly to the embedding model's context window, whereas 512 words might tokenize to anywhere from 400 to 800 tokens depending on vocabulary coverage. For RAG systems, see our RAG architecture guide for how tokenization affects retrieval quality.

Tokenization for Code

Code tokenization is a weak point for many models. Common variable names like getUserById may tokenize into 4-5 subword tokens, consuming context window space. Models with larger vocabularies (GPT-4, LLaMA 3) handle code more efficiently because common programming patterns (function names, syntax tokens, indentation) become single tokens during training. This directly affects the cost and quality of AI code review systems.

Training Your Own Tokenizer

For domain-specific applications — medical text, legal documents, code in specific programming languages — training a custom tokenizer on your domain data can reduce token counts by 20-40% compared to a general-purpose tokenizer. The Hugging Face tokenizers library makes this straightforward:

from tokenizers import Tokenizer
from tokenizers.models import BPE
from tokenizers.trainers import BpeTrainer
from tokenizers.pre_tokenizers import ByteLevel

tokenizer = Tokenizer(BPE())
tokenizer.pre_tokenizer = ByteLevel()

trainer = BpeTrainer(
    vocab_size=32000,
    special_tokens=["", "", "", ""],
    min_frequency=2,
    show_progress=True,
)

tokenizer.train(files=["domain_corpus.txt"], trainer=trainer)
tokenizer.save("domain_tokenizer.json")

The catch: a custom tokenizer means you can't use pre-trained models that expect a different vocabulary. You'll need to either train from scratch or extend an existing model's vocabulary — adding new tokens and fine-tuning the embedding layer. Extending is cheaper but produces suboptimal token embeddings for the new entries. Training from scratch is expensive but gives you a tokenizer perfectly matched to your domain. For managing these training experiments, our experiment tracking guide covers the infrastructure needed.

FAQ

What is the difference between BPE and WordPiece tokenization?

BPE merges the most frequent pair of adjacent tokens iteratively. WordPiece uses a likelihood-based criterion — it merges the pair that maximizes language model likelihood. In practice, both produce similar vocabularies. BPE is used by GPT models; WordPiece by BERT.

Why do LLMs use subword tokenization instead of word-level or character-level?

Word-level creates enormous vocabularies and can't handle unseen words. Character-level creates very long sequences and loses morphological information. Subword tokenization balances both: common words stay single tokens, rare words decompose into meaningful subparts, and the vocabulary stays manageable.

How does vocabulary size affect model performance and cost?

Larger vocabularies produce shorter token sequences, reducing attention computation and API costs. However, the embedding matrix grows linearly with vocabulary size. The sweet spot for most LLMs is 32K-100K tokens. GPT-4 uses ~100K, LLaMA uses 32K, and Gemma uses 256K for multilingual coverage.

What is SentencePiece and how does it differ from Hugging Face tokenizers?

SentencePiece treats text as raw Unicode bytes with no pre-tokenization or language-specific rules. Hugging Face tokenizers apply pre-tokenization (whitespace splitting, regex rules) before BPE or WordPiece. SentencePiece is used by T5, LLaMA, and Mistral.