BPE vs WordPiece vs SentencePiece: Which One Does Your Model Use

A side-by-side comparison of BPE, WordPiece, and SentencePiece: how each picks its merges, how to spot them in the tokens, and which models use which.

Aug 10, 20268 min readFollow

Topics You Will Master

The one rule that separates BPE, WordPiece, and SentencePiece
How to spot each method just by looking at its tokens
Which famous models use which tokenizer
How to pick a method when you train your own

We have now met all three subword methods. BPE glues the most frequent pair. WordPiece glues the pair that best predicts the text. SentencePiece skips word splitting and reads the raw stream. Three methods, one job.

So how do we tell them apart in the wild? The good news is that each one leaves a fingerprint in its output. Once we know the fingerprint, we can name the tokenizer at a glance. In this blog, we will learn the one rule that separates them, how to spot each one, and which models use which.

Bestseller

LLM Fine-Tuning with Hugging Face: LoRA, QLoRA, PEFT

Fine-tune BERT, T5, ViT, LLaMA-style models and Qwen3-TTS using Hugging Face Transformers, custom datasets, LoRA, QLoRA.

Enroll on Udemy →30 day refund, lifetime access

"unbelievable" Split by WordPiece, Byte-Level BPE and SentencePiece


The One Difference That Matters

BPE and WordPiece both start small and build bigger pieces. That part is shared. The difference is the question each one asks before it merges.

BPE asks "which pair appears most often?" WordPiece asks "which pair helps me predict text best?" SentencePiece asks something else entirely, because it changes the input rather than the rule: "what if we never split on spaces at all?"

In simple words, BPE and WordPiece argue about which pair to glue. SentencePiece argues about what counts as a starting piece.

SentencePiece is a toolkit, and it can train two kinds of model on that raw stream. Its default is Unigram, which starts big and prunes the weakest pieces. It can also train BPE, which merges pairs as usual.

BPE vs WordPiece Merge Rules, SentencePiece's Starting Pieces


Advertisement

Fingerprint One: The ## Marker

The easiest tell is the ## prefix. If we see tokens like play and ##ing, we are looking at WordPiece. No other common method uses that marker.

PYTHON
# WordPiece output has ## on inside pieces
# "playing" -> ["play", "##ing"]

The ## says "join me to the piece before me". It only appears inside a word, never at the start. The play and ##ing split is an example. Real BERT keeps "playing" as one token, but rarer words split this way. In the first figure, the DistilBERT tokenizer cuts "unbelievable" into un and ##believable.

The WordPiece ## Marker: play and ##ing Glued Back Into playing


Fingerprint Two: The Space Marker

SentencePiece leaves a different mark. Because it keeps spaces as part of the token, its pieces often begin with a symbol that stands for a space. The real symbol is the character ▁ (U+2581). It looks like an underscore, so this article writes it as "_".

So a sentence tokenized by SentencePiece looks like _New, _York, _City. That leading mark is the space itself, travelling with the token.

Where the Spaces in "New York City" Go in WordPiece and SentencePiece


Fingerprint Three: Plain Pieces

If the tokens have no ## and no space marker, and the pieces look like ordinary chunks of words, we are most likely looking at BPE.

Byte-level BPE, the version GPT uses, has its own small tell. It works on raw bytes, so it marks a leading space with the symbol Ġ. In the figure below, the GPT-2 tokenizer turns "New York City" into New, ĠYork and ĠCity. Accented letters can also show up as odd-looking symbols. Because the space travels inside the token as Ġ, byte-level BPE also decodes back to the exact original text.

Let me tabulate the fingerprints for your better understanding.

What you see in the tokens The method Example
## in front of inside pieces WordPiece play, ##ing
A leading underscore mark SentencePiece _New, _York
Plain pieces, no markers BPE low, est

Naming the Tokenizer From Its Tokens: Check for ##, Then a Space Mark


Advertisement

How Each One Handles a New Word

Give all three the same unseen word and they all survive, but they cut it differently. BPE replays its merge list. WordPiece takes the longest matching piece first. SentencePiece, in its default Unigram mode, picks pieces from its pruned vocabulary over the raw stream.

The splits often look similar. In the figure, WordPiece and BPE cut "tokenization" at the same seam, token plus ization. SentencePiece makes three pieces: _to, ken and ization.

Splitting the Word "tokenization" With WordPiece, BPE and SentencePiece


Which Models Use Which

This is the practical part. BPE, in its byte-level form, powers the GPT family and many open models. WordPiece powers the BERT family. SentencePiece powers T5, Llama 2, and many multilingual models. Llama 3 moved to a BPE tokenizer built on tiktoken.

Let me tabulate the three methods side by side.

BPE WordPiece SentencePiece
Merge rule Most frequent pair Pair that best predicts Prune the weakest pieces
Starting units Characters, split on spaces Characters, split on spaces The raw character stream
Needs spaces? Yes Yes No
Marker in output None ## on continuations Underscore for the space
Rebuilds text exactly By rule By rule By construction
Typical models GPT, RoBERTa, BART BERT, DistilBERT, Electra T5, mT5, Llama, ALBERT

When we load a model from Hugging Face, we rarely choose the tokenizer ourselves. It ships with the model, and it must match, because the token IDs are meaningless to a model trained on a different vocabulary.

Tokenizer Files and Model Weights Side by Side in Three Checkpoints


Why the Tokenizer Must Match the Model

Here is a mistake worth avoiding. A token ID is just a row number in the model's embedding table. Let's say tokenizer A turns "able" into id 5021. Row 5021 means "able" only for model A, because model A learned it that way. Another vocabulary puts a different piece in that row. In bert-base-uncased, for example, id 5021 is "comic".

If we swap in a different tokenizer, the same word turns into different numbers. The model will still run, but it will read nonsense. This is why models on Hugging Face ship their tokenizer files alongside their weights.

PYTHON
# the tokenizer and the model must come from the same checkpoint
# token id 5021 means "able" only for THIS vocabulary

"able" Through the Matching Tokenizer (Id 5021) and a Swapped One


Advertisement

Choosing a Method for Your Own Model

If we ever train a tokenizer from scratch, the choice is simpler than it looks. For English-only text, BPE is a safe default. For a BERT-style encoder, WordPiece fits the family. For anything multilingual, or text with code, or scripts without spaces, SentencePiece is the natural pick.

In practice most teams start from an existing tokenizer and only train a new one when their text looks nothing like ordinary web text.

Picking BPE, WordPiece or SentencePiece When Training a Tokenizer


The Shared Family Tree

When we step back, the three look less like rivals and more like cousins. All three build subwords. All three keep the vocabulary at a fixed size. All three survive words they never saw.

That shared idea, pieces smaller than words but bigger than letters, lets a model read any word with a vocabulary of fixed size.

BPE, WordPiece and SentencePiece: Three Shared Traits, Two Differences


Recap

This is how the three subword methods compare. BPE merges the most frequent pair. WordPiece merges the pair that best predicts real language. SentencePiece drops word splitting, so it can read scripts without spaces, and by default it prunes a big vocabulary down.

We can name them by their fingerprints. ## means WordPiece, a leading "_" mark means SentencePiece, and plain pieces or a Ġ space mark usually mean BPE. Whichever one our model uses, the rule is the same: the tokenizer must match the model. Token IDs only mean something inside the vocabulary that created them.

In the next article, we look at the special tokens that sit alongside these pieces, the ones in square brackets like [CLS] and [SEP], and what each of them is for.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments