What if we ask a large language model how many times the letter R appears in "strawberry"? It may confidently answer two. The correct answer is three. The same model can write a working program and explain quantum tunnelling, yet it fumbles a question a six-year-old gets right.
This is not a reasoning failure. It is a tokenization failure. Once we see why, a whole family of strange model behaviours suddenly makes sense.
In this blog, we will learn why the model cannot see letters, why numbers split in odd ways, and the simple prompt that fixes the strawberry count.

The Model Never Sees Letters
Here is the root of it. By the time text reaches the model, it is no longer text. It is a list of token IDs, and each ID is just a row number.
In cl100k, the tokenizer GPT-4 uses, the word "strawberry" arrives as three tokens: "str", "aw", and "berry". The model receives [496, 675, 15717]. Those numbers carry meaning learned from training, but they do not carry spelling. Nothing in the input says "this piece contains two R's".
# what the model actually receives
# "strawberry" -> ["str", "aw", "berry"] -> [496, 675, 15717]
# the letters are gone by this point

An Analogy: Reading Through Frosted Glass
Let's say someone hands us words printed on cards, but the cards are frosted. We can tell one card from another. We have also learned what each card means from seeing them used millions of times.
Now someone asks us to count the letter R on card number three. We know the card means "berry". We might remember that berry has two R's, because we have read about spelling. But we cannot look and count, because we cannot see through the glass.
That is exactly the model's situation. It can recall facts about spelling, but it cannot look at the letters.

Why the Model Still Sometimes Gets It Right
If the letters are invisible, how does a model ever answer spelling questions correctly? Because spelling appears in its training data as text.
Sentences like "strawberry is spelled s-t-r-a-w-b-e-r-r-y" teach the relationship between a token and its letters. So the model has second-hand knowledge of spelling. That knowledge is patchy, which is why it is right about common words and wrong about odd ones.

Why Splitting Is Inconsistent
There is a second twist. The same letters can be split differently depending on what surrounds them, including whether there is a space in front.
"strawberry" at the start of a sentence and " strawberry" mid-sentence can become different token sequences. So the model's grip on a word's spelling can change with position.

The Same Problem in Arithmetic
Numbers suffer from the same issue. A tokenizer cuts a number into chunks, and a chunk does not say what place value it holds. cl100k reads the digits from left to right in groups of up to three.
So 1234 becomes "123" and "4", 12345 becomes "123" and "45", and 12346 becomes "123" and "46". The same "123" token stands for 1,230 inside 1234 but 12,300 inside 12345. The model has to work out each chunk's real value from how long the whole number is. This is one reason long multiplication often goes wrong even when the reasoning is sound.

Why Rhyming and Wordplay Wobble
Rhyme lives in sound, and sound lives in letters. A model working with chunks has only indirect access to both.
Puns, anagrams, acrostics, and counting syllables all need a level of detail. The model's input does not contain that detail. It can often fake it from memory, and it fails in ways that look careless rather than systematic.

The Simple Workaround
There is an easy fix, and it works because it changes what the model receives.
We ask the model to spell the word out with separators first, and then count. Writing "s - t - r - a - w - b - e - r - r - y" forces each letter into its own token, and suddenly the letters are visible. Chain-of-thought prompting, where we ask the model to write out its steps, helps here for the same reason. It turns a hidden problem into a visible one.
# force the letters into separate tokens, then count
# "spell strawberry letter by letter, then count the r's"
# s - t - r - a - w - b - e - r - r - y -> now each letter is its own token

Why Models Are Not Simply Built on Letters
If letters cause so much trouble, why not tokenize every character? Because the cost is too high.
Character tokens make sequences four to five times longer. Attention compares every token with every other token, so its cost grows with the square of the length. A sequence four to five times longer needs 16 to 25 times as many attention scores. For "strawberry" alone, the grid grows from 3 × 3 to 10 × 10. We would pay that extra cost on every task just to fix a small family of puzzles. The trade is not worth it, so the quirk stays.

A Quirk, Not a Ceiling
Let's be clear about what this does and does not prove. The strawberry failure says nothing about whether a model can reason. It says the model was handed a question about a level of detail its input throws away.
When we give it the letters, it counts fine. That is a strong hint that this is a plumbing problem, not a thinking problem.

Recap
This is how tokenization makes a capable model miscount the R's in strawberry. The word arrives as a few chunks, and those chunks are converted to numbers before the model sees anything. Spelling is thrown away in that step, so the model can only recall what it has read about letters rather than look at them.
The same root cause explains shaky arithmetic, wobbly rhymes, and failed anagrams. All of them need a level of detail that subword tokens do not carry. Asking the model to spell a word out first fixes it, because that puts each letter into its own token.
We keep subwords anyway. Character tokens would make every sequence four to five times longer, and attention would need 16 to 25 times as many scores. That closes our nine days on tokenization. In the next article, we follow the token IDs into the model itself and meet embeddings, the step that turns those row numbers into meaning.