WordPiece: How BERT Splits Words

How WordPiece tokenizes words for BERT, the ## continuation marker, greedy longest-match encoding, and how it differs from BPE by merging on likelihood.

Aug 4, 20267 min readFollow

Topics You Will Master

How WordPiece differs from BPE by merging on likelihood, not raw frequency
What the ## continuation marker means and why it matters
Greedy longest-match encoding, the way WordPiece splits a word
How BERT falls back to [UNK] only when even subwords fail

BPE glues the most frequent pair. Its close cousin WordPiece, the method behind BERT, glues the pair that helps the model predict text best. That one small change is the whole story, and it leads to the familiar ## marks we see in BERT's tokens.

In simple words, WordPiece also builds subwords, but it picks its merges with a smarter rule, and it tags the inside pieces of a word so they can be joined back together. Let's see how WordPiece cuts a word with a small vocabulary.

Training a WordPiece Vocabulary, Then Splitting "playing" With It

Bestseller

LLM Fine-Tuning with Hugging Face: LoRA, QLoRA, PEFT

Fine-tune BERT, T5, ViT, LLaMA-style models and Qwen3-TTS using Hugging Face Transformers, custom datasets, LoRA, QLoRA.

95% off$199.99
$9.99Enroll now →30 day refund, lifetime access

What Makes WordPiece Different from BPE?

WordPiece and BPE start the same way, from single characters, and both merge pieces into bigger ones. The difference is which pair they choose to merge.

BPE is greedy about counts: it merges the pair that simply appears most often. WordPiece is greedy about usefulness: it merges the pair that most increases the chance of the training text under the model. In plain terms, WordPiece asks "which merge best helps predict real language?" rather than "which pair is most common?"

Most of the time the two agree, but WordPiece's rule tends to produce pieces that line up a little better with real word parts, like prefixes and suffixes.

Choosing the Next Merge: BPE Counts Pairs, WordPiece Scores Likelihood


The ## Continuation Marker

Here is the piece of WordPiece we will actually see in BERT's output. When a subword is inside a word rather than at its start, WordPiece tags it with ##. Real BERT keeps "playing" as one token, because the word is common. But with a smaller vocabulary, a word like this might split into "play" and "##ing":

PYTHON
# WordPiece tags inside-pieces with ##
# "playing" -> ["play", "##ing"]
# "##ing" means: attach me to the piece before me

So "play" has no marker because it starts the word, and "##ing" carries ## because it continues one. This tiny tag is important. It lets the model tell "ing" the whole word from "ing" the suffix. It also lets us glue the pieces back into the original word without guessing where the spaces were.

The ## Marker in "playing": "play" Starts the Word, "##ing" Continues It


Advertisement

Building the Vocabulary by Likelihood

During training, WordPiece grows its vocabulary one merge at a time, just like BPE. But at each step it scores every candidate pair by how much gluing it would raise the likelihood of the training text, and it picks the best-scoring one.

Scoring Candidate Pairs by Likelihood Gain: ##t + ##ion Becomes ##tion

We can picture the score as a simple question: does the pair appear together far more than its two parts would on their own? A pair like "##t" and "##ion" scores high because "##tion" is a real, meaningful chunk. Both pieces carry the ## marker because they sit inside a word. A random pair that just happens to co-occur scores low. This is why WordPiece pieces often feel like natural word parts.

Growing a WordPiece Vocabulary From Characters by Repeated Merges


How WordPiece Encodes: Greedy Longest-Match

At tokenize time, WordPiece uses a rule called greedy longest-match-first. For each word, it looks for the longest piece in its vocabulary that matches the start of the word, takes it, then repeats on the rest.

Let's say our small vocabulary has "play" and "##ing" but not "playing". For "playing", it first grabs the longest matching start, "play". Then it continues on "ing" and grabs "##ing". Always taking the longest match keeps the token count low and the pieces meaningful.

Greedy Longest-Match on "playing": Take "play", Then "##ing"


Encoding a Word BERT Knows Well

For a common word that already sits in the vocabulary, WordPiece is done in one step. The word "hello" is a single token, no splitting, no ##. Most everyday words are like this, which keeps sentences short.

Greedy Longest-Match on "hello": the Whole Word Matches in One Step


Encoding a Word BERT Has Never Seen

Now the interesting case. When we give BERT a rare word like "tokenization", greedy longest-match kicks in. It grabs "token", then "##ization", or perhaps "##iza" and "##tion" depending on the vocabulary. Either way, the rare word becomes a handful of known pieces, and its meaning largely survives.

This is the same gift as BPE. WordPiece rarely gets stuck, because it can fall back to smaller and smaller subwords. That only works while the vocabulary holds the characters it needs.

Splitting the Rare Word "tokenization" With Two Different Vocabularies


Advertisement

When Even Subwords Fail: [UNK]

There is one last safety net. Sometimes WordPiece reaches a point in a word where nothing in its vocabulary matches, not even a single character. This is rare. When it happens, WordPiece gives up on the whole word. The entire word becomes one special [UNK] token, meaning "unknown", not just the part it could not match. In the figure, the emoji has no piece to match, so it becomes [UNK]. In practice this almost never happens for normal text, because the vocabulary includes all the common characters. But it is there as a final catch.

Tokenizing "hello", "playing" and an Emoji: Only the Emoji Becomes UNK


Why the Likelihood Rule Matters

We might wonder if choosing merges by likelihood instead of frequency really changes much. It does, in a subtle but useful way. Frequency alone can glue pairs that just happen to sit next to each other a lot. Likelihood favours pairs that carry real meaning together, like true prefixes and suffixes.

The effect is that WordPiece pieces tend to align with the parts of words that matter, which can help a model like BERT understand structure. Same family as BPE, slightly sharper cuts.

Frequency Rule vs Likelihood Rule on Two Kinds of Candidate Pairs


Where WordPiece Is Used

WordPiece is the tokenizer behind the BERT family. BERT, DistilBERT, and ELECTRA all use it. So a large slice of the classification and search systems in production run on WordPiece. Whenever we see tokens with ## in them, we are looking at WordPiece at work.

The next article covers the third big subword method, SentencePiece, which throws out one assumption both BPE and WordPiece quietly make: that we can split on spaces in the first place.


Recap

This is how WordPiece works. It is BPE's close cousin. It also builds subwords by merging small pieces into bigger ones. But it chooses each merge by likelihood, the pair that best predicts real language, rather than by raw frequency. It tags the inside pieces of a word with ## so they can be joined back together. At tokenize time, it uses greedy longest-match to grab the biggest known piece first.

Common words stay whole, rare words split into meaningful subwords, and only in the rarest case does the whole word fall back to [UNK]. This is the tokenizer that powers BERT. Next, we meet SentencePiece, the method that drops the need for spaces entirely.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments