We have now met all three subword methods. BPE glues the most frequent pair. WordPiece glues the pair that best predicts the text. SentencePiece skips word splitting and reads the raw stream. Three methods, one job.
So how do we tell them apart in the wild? The good news is that each one leaves a fingerprint in its output. Once we know the fingerprint, we can name the tokenizer at a glance. In this blog, we will learn the one rule that separates them, how to spot each one, and which models use which.

The One Difference That Matters
BPE and WordPiece both start small and build bigger pieces. That part is shared. The difference is the question each one asks before it merges.
BPE asks "which pair appears most often?" WordPiece asks "which pair helps me predict text best?" SentencePiece asks something else entirely, because it changes the input rather than the rule: "what if we never split on spaces at all?"
In simple words, BPE and WordPiece argue about which pair to glue. SentencePiece argues about what counts as a starting piece.
SentencePiece is a toolkit, and it can train two kinds of model on that raw stream. Its default is Unigram, which starts big and prunes the weakest pieces. It can also train BPE, which merges pairs as usual.

Fingerprint One: The ## Marker
The easiest tell is the ## prefix. If we see tokens like play and ##ing, we are looking at WordPiece. No other common method uses that marker.
# WordPiece output has ## on inside pieces
# "playing" -> ["play", "##ing"]
The ## says "join me to the piece before me". It only appears inside a word, never at the start. The play and ##ing split is an example. Real BERT keeps "playing" as one token, but rarer words split this way. In the first figure, the DistilBERT tokenizer cuts "unbelievable" into un and ##believable.

Fingerprint Two: The Space Marker
SentencePiece leaves a different mark. Because it keeps spaces as part of the token, its pieces often begin with a symbol that stands for a space. The real symbol is the character ▁ (U+2581). It looks like an underscore, so this article writes it as "_".
So a sentence tokenized by SentencePiece looks like _New, _York, _City. That leading mark is the space itself, travelling with the token.

Fingerprint Three: Plain Pieces
If the tokens have no ## and no space marker, and the pieces look like ordinary chunks of words, we are most likely looking at BPE.
Byte-level BPE, the version GPT uses, has its own small tell. It works on raw bytes, so it marks a leading space with the symbol Ġ. In the figure below, the GPT-2 tokenizer turns "New York City" into New, ĠYork and ĠCity. Accented letters can also show up as odd-looking symbols. Because the space travels inside the token as Ġ, byte-level BPE also decodes back to the exact original text.
Let me tabulate the fingerprints for your better understanding.
| What you see in the tokens | The method | Example |
|---|---|---|
## in front of inside pieces |
WordPiece | play, ##ing |
| A leading underscore mark | SentencePiece | _New, _York |
| Plain pieces, no markers | BPE | low, est |

How Each One Handles a New Word
Give all three the same unseen word and they all survive, but they cut it differently. BPE replays its merge list. WordPiece takes the longest matching piece first. SentencePiece, in its default Unigram mode, picks pieces from its pruned vocabulary over the raw stream.
The splits often look similar. In the figure, WordPiece and BPE cut "tokenization" at the same seam, token plus ization. SentencePiece makes three pieces: _to, ken and ization.

Which Models Use Which
This is the practical part. BPE, in its byte-level form, powers the GPT family and many open models. WordPiece powers the BERT family. SentencePiece powers T5, Llama 2, and many multilingual models. Llama 3 moved to a BPE tokenizer built on tiktoken.
Let me tabulate the three methods side by side.
| BPE | WordPiece | SentencePiece | |
|---|---|---|---|
| Merge rule | Most frequent pair | Pair that best predicts | Prune the weakest pieces |
| Starting units | Characters, split on spaces | Characters, split on spaces | The raw character stream |
| Needs spaces? | Yes | Yes | No |
| Marker in output | None | ## on continuations |
Underscore for the space |
| Rebuilds text exactly | By rule | By rule | By construction |
| Typical models | GPT, RoBERTa, BART | BERT, DistilBERT, Electra | T5, mT5, Llama, ALBERT |
When we load a model from Hugging Face, we rarely choose the tokenizer ourselves. It ships with the model, and it must match, because the token IDs are meaningless to a model trained on a different vocabulary.

Why the Tokenizer Must Match the Model
Here is a mistake worth avoiding. A token ID is just a row number in the model's embedding table. Let's say tokenizer A turns "able" into id 5021. Row 5021 means "able" only for model A, because model A learned it that way. Another vocabulary puts a different piece in that row. In bert-base-uncased, for example, id 5021 is "comic".
If we swap in a different tokenizer, the same word turns into different numbers. The model will still run, but it will read nonsense. This is why models on Hugging Face ship their tokenizer files alongside their weights.
# the tokenizer and the model must come from the same checkpoint
# token id 5021 means "able" only for THIS vocabulary

Choosing a Method for Your Own Model
If we ever train a tokenizer from scratch, the choice is simpler than it looks. For English-only text, BPE is a safe default. For a BERT-style encoder, WordPiece fits the family. For anything multilingual, or text with code, or scripts without spaces, SentencePiece is the natural pick.
In practice most teams start from an existing tokenizer and only train a new one when their text looks nothing like ordinary web text.

The Shared Family Tree
When we step back, the three look less like rivals and more like cousins. All three build subwords. All three keep the vocabulary at a fixed size. All three survive words they never saw.
That shared idea, pieces smaller than words but bigger than letters, lets a model read any word with a vocabulary of fixed size.

Recap
This is how the three subword methods compare. BPE merges the most frequent pair. WordPiece merges the pair that best predicts real language. SentencePiece drops word splitting, so it can read scripts without spaces, and by default it prunes a big vocabulary down.
We can name them by their fingerprints. ## means WordPiece, a leading "_" mark means SentencePiece, and plain pieces or a Ġ space mark usually mean BPE. Whichever one our model uses, the rule is the same: the tokenizer must match the model. Token IDs only mean something inside the vocabulary that created them.
In the next article, we look at the special tokens that sit alongside these pieces, the ones in square brackets like [CLS] and [SEP], and what each of them is for.