SentencePiece: Tokenization for Multilingual Text and Code

How SentencePiece tokenizes raw text with no spaces to split on, the meaning of the underscore marker, the Unigram model, and why T5 and Llama use it.

Aug 7, 20268 min readFollow

Topics You Will Master

Why SentencePiece treats text as a raw stream instead of splitting on spaces
What the space marker means and how it makes tokenization reversible
The Unigram model: start with a big vocabulary, then prune it down
Why SentencePiece is the go-to for multilingual text and code

BPE and WordPiece both quietly assume one thing: that we can split text into words on spaces first. But many languages do not use spaces at all. SentencePiece, the method behind T5 and Llama 2, throws out that assumption and treats text as one raw stream of characters.

In simple words, SentencePiece does not care about words or spaces. It reads the text as one stream, keeps the spaces as visible markers, and learns its subwords straight from that stream. So one tokenizer can handle many languages, and code too. In this blog, we will learn how the space marker works, how the Unigram model picks its pieces, and where SentencePiece is used.

SentencePiece on "New York": Raw Text to Marked Stream to ▁New, ▁York


The Problem: Languages Without Spaces

BPE and WordPiece start by cutting text into words on spaces, then split each word into subwords. That works fine for English. But Japanese, Chinese, and Thai write sentences with no spaces between words at all. There is nothing to split on.

So, here comes SentencePiece to the rescue. It refuses to depend on spaces. It reads the whole sentence as one stream and finds good subword pieces directly, whether or not spaces exist. So the same tokenizer works for scripts with spaces and scripts without them.

Bestseller

LLM Fine-Tuning with Hugging Face: LoRA, QLoRA, PEFT

Fine-tune BERT, T5, ViT, LLaMA-style models and Qwen3-TTS using Hugging Face Transformers, custom datasets, LoRA, QLoRA.

95% off$199.99
$9.99Enroll now →30 day refund, lifetime access

"New York" in English, Japanese and Chinese: Space Split vs One Stream


Text as a Raw Stream

The key mental shift is this. BPE and WordPiece work on words. SentencePiece works on the raw character stream, spaces included. It does not pre-split anything. The entire sentence, spaces and all, goes in as one sequence, and SentencePiece learns pieces from it.

Because it never assumes where words begin or end, it treats every language the same way. A space is just another character in the stream, no more special than a letter.

BPE Sees "New York" as 2 Words, SentencePiece as One Character Stream


Advertisement

The Space Marker

If SentencePiece keeps spaces in the stream, how does it show them? First, it adds one space at the start of the text. This is a default setting called add_dummy_prefix. Then it replaces every space with a visible marker. The real marker is the character ▁ (U+2581). It looks like an underscore, so in this blog we write it as "_".

PYTHON
# SentencePiece turns each space into a visible marker (shown here as _)
# "New York"  ->  ["_New", "_York"]
# because the space is kept, decoding rebuilds the text exactly

So "New York" becomes two pieces that each begin with the space marker. The marker on "_York" comes from the real space between the words. The marker on "_New" comes from the space added at the start. This looks like a small detail, but it is what makes SentencePiece reversible. The spaces are never thrown away, so we can turn the tokens back into the original text. For ordinary single-spaced text like this, we get it back character for character.

Turning the Spaces in "New York" Into ▁ Markers: ▁New and ▁York


The Unigram Model: Start Big, Prune Down

BPE and WordPiece build up: they start small and merge. By default, SentencePiece works the other way. It uses a method called the Unigram model, which builds down.

It starts with a big pile of candidate pieces, far more than it needs. Then it repeatedly asks, "which pieces are pulling their weight?" and throws away the least useful ones, shrinking the vocabulary until it hits the target size. Instead of gluing pairs together, it prunes a huge set down to the best pieces.

Building a Vocabulary: BPE Merges Up, Unigram Prunes Down to Size


How Unigram Decides What to Keep

For each candidate piece, Unigram asks a simple question: if we removed this piece, how much worse would the tokenizer get at explaining the training text? Pieces that carry a lot of weight, like "ing" or "tion", hurt a lot to remove, so they stay. Pieces that barely help, or that overlap with better pieces, get dropped. In the figure, removing "_Yor" changes almost nothing, because "_York" covers it.

So pruning works in rounds. In each round, Unigram removes the pieces whose removal hurts its fit to the training text the least. It keeps doing this until the vocabulary shrinks to the target size.

Because it scores whole pieces this way, Unigram can also see more than one valid way to cut a word and pick the most likely one. For "New York", the figure shows three valid cuts, and Unigram picks "_New" and "_York". This choice of cuts helps it deal with messy, mixed text.

Unigram Removal Test for ing, tion, ▁Yor and Three Cuts of ▁New▁York


Advertisement

Lossless, Reversible Tokenization

This is one of SentencePiece's best features. It keeps every space as a marker and never pre-splits the text. So for ordinary single-spaced text, we can decode the tokens back into the original string exactly, with the spaces in the right places.

That sounds obvious, but it is not true for every tokenizer. Some throw away spacing details and have to guess when rebuilding text. In the figure, a word-splitting tokenizer cuts "New York." into "New", "York" and ".". It then rejoins them with spaces and gets "New York .", which is not the input. SentencePiece does not need to guess, because its markers record where each space was.

There is one catch. By default, SentencePiece also tidies the spaces before it starts. A setting called remove_extra_whitespaces is on by default. It squeezes a run of spaces into one space, and it strips spaces at the start and end of the text. So the round trip is exact for every space only when we turn this setting off.

Decoding ▁New ▁York Back to "New York" vs Rejoining Split Words


One Vocabulary for Many Languages

Because SentencePiece learns straight from the raw stream, a single model can share one vocabulary across dozens of languages. English, Hindi, Japanese, and Arabic can all live in the same set of subword pieces, with common bits shared where scripts overlap.

This is exactly what multilingual models like mT5 need. Instead of a separate tokenizer per language, one SentencePiece vocabulary covers them all. This keeps the model smaller and lets it transfer knowledge between languages.

English, Hindi, Japanese and Arabic Streams Into One Shared Vocabulary


SentencePiece and BPE, Side by Side

Before we go further, let me tabulate the difference for your better understanding. Both build subwords, but they treat the space in opposite ways.

SentencePiece BPE
The space Kept as a marker inside the token Removed before learning starts
Input it sees The full character stream Words, already split
Word splitting Not needed at all Required first
Rebuilding text Exact, by construction Needs an external rule
Cross-word pieces Can learn them Cannot see them
Best suited to Multilingual text, code, any script English-centric, speed-first work

Here, we can see that one small design choice, keeping the space, is what gives SentencePiece its coverage and its exact round trip on ordinary text.

Advertisement

Tokenizing Code

Code is another place where spaces matter. Programming languages are full of spaces, indentation, and symbols that carry real meaning. In Python, a stray space can change what the code does. This is where the remove_extra_whitespaces setting counts. It is on by default, and with it on, SentencePiece strips the leading spaces before marking, so the indentation is lost. With it off, every leading space survives as a marker, and the line decodes back with its indentation. So, for code, we want this setting off.

Indented Python Line With remove_extra_whitespaces On vs Off


Where SentencePiece Is Used

SentencePiece is used in many well-known models. T5 and its multilingual sibling mT5 use it, and so do Llama 2, ALBERT, and XLNet. Llama 3 moved away from it to a BPE tokenizer built on tiktoken. Whenever we need one tokenizer that works across many languages, SentencePiece is a common choice.

That closes our tour of the three subword methods. The next article steps back and compares BPE, WordPiece, and SentencePiece side by side, so we can tell at a glance which one a model is using.

SentencePiece Tokenizers in T5, mT5, Llama 2, ALBERT and XLNet


Recap

This is how SentencePiece works. It drops the assumption that we can split text on spaces, and instead reads the whole sentence as one raw stream. It adds a space at the start and turns each space into a visible marker. So for ordinary single-spaced text, the tokens decode back to the exact original text. By default, it builds its vocabulary with the Unigram model. It starts from a huge pile of candidate pieces and prunes away the pieces whose removal hurts the least.

Because it never depends on spaces, one SentencePiece vocabulary can cover many languages at once. With remove_extra_whitespaces turned off, it also keeps the exact layout of code. That is why T5, mT5, and Llama 2 rely on it. Next, we put all three subword methods, BPE, WordPiece, and SentencePiece, side by side, so we can spot which one any model uses.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments