Masked Attention: How GPT Avoids Seeing the Future

Learn how a causal mask blocks future tokens during training, why GPT needs this rule, and how masked softmax keeps text generation honest.

Sep 30, 20269 min readFollow

Topics You Will Master

Why next-token prediction fails if a model sees future words
What a causal attention mask allows and blocks
Why blocked scores receive negative infinity before softmax
How GPT trains many positions at once while following a left-to-right rule

Write the beginning of this sentence: the cat sat on .... The next word should be a guess. If the word mat is visible while making that guess, there is no real prediction task left.

GPT learns by predicting the next token, so it needs a rule that prevents this shortcut. Causal masked attention allows a token to attend to itself and earlier tokens, but never to a token on its right.

This diagram gives the rule in its simplest form.

Causal Mask Grid for "The cat sat on mat" With the sat Row Worked Out

The allowed cells are on and below the diagonal; cells above the diagonal point into the future and are blocked.

Bestseller

LLM Fine-Tuning with Hugging Face: LoRA, QLoRA, PEFT

Fine-tune BERT, T5, ViT, LLaMA-style models and Qwen3-TTS using Hugging Face Transformers, custom datasets, LoRA, QLoRA.

Enroll on Udemy →30 day refund, lifetime access

The Shortcut a Decoder Must Not Take

During training, the data loader puts a complete sentence in front of the model. That is useful: one forward pass can calculate a training loss for many positions. It is also dangerous. When the model processes sat, the tokens on and mat are physically present in the same input matrix.

If attention could read them, the model would learn with an answer sheet that will not exist when it writes text for us. At generation time, it only has the prompt and the tokens it has already selected.

Predicting the Word After "on" With and Without the Causal Mask

The full training sentence is available for fast computation, but future words must remain invisible to earlier positions.

The mask makes training and generation follow the same information rule. Both may use the past. Neither may use a future token to decide the present one.

Advertisement

Reading the Lower Triangle

Use five tokens: The, cat, sat, on, and mat. An attention score grid has one row for every query token and one column for every key token. A row answers, "Which token positions may this query read?"

The first row can see only The. The row for cat can see The and cat. The row for sat can see the first three tokens. Each next row gets one extra allowed position.

Query token Keys it may see
The The
cat The, cat
sat The, cat, sat
on The, cat, sat, on
mat The, cat, sat, on, mat

Those allowed cells form the lower triangle of the grid.

Keys Each Query Row May See Under the Causal Mask

Every row can look left and at itself, while every future column is blocked.

The diagonal is allowed because a token may use its own representation. Some diagrams use a one-token shift when describing next-token labels, but the attention mask itself normally keeps the diagonal available.

The Label Is One Position Ahead

Allowing the diagonal is not cheating. The hidden state for cat uses the known input cat to predict sat. The label is one token ahead of the input position, not the same token.

For a training sequence The cat sat on mat, the paired inputs can be The cat sat on and the labels cat sat on mat. At row 1, the input contains cat while the correct label is sat. The causal mask stops row 1 from reading the input sat at row 2.

This distinction explains why the lower triangle includes its diagonal. Removing the diagonal would hide an input the model really has during generation. Allowing the next column would reveal the answer. The mask and the shifted labels must therefore agree.

How the Mask Changes Attention Scores

Day 14 showed that softmax turns attention scores into weights. The causal mask acts before softmax.

Suppose the query for sat produces already-scaled scores [2, 1, 3, 4, 2] for The, cat, sat, on, and mat. The last two scores belong to future tokens. We add a mask row:

Zero leaves an allowed score unchanged. Negative infinity removes a blocked score before softmax can turn it into a weight.

Adding the Causal Mask to the sat Row Before Softmax

The mask changes only forbidden score cells, then ordinary softmax runs across the row.

Softmax uses exponentials. Since , the blocked positions get exactly zero attention weight. Their values cannot enter the output for sat, no matter how high their original raw scores were.

Advertisement

Why We Add the Mask Instead of Multiplying by Zero

It may seem simpler to multiply every blocked score by zero. That does not work. Zero is an ordinary score, and softmax can still give it a positive share.

For example, if blocked scores become zero, a row such as [2, 1, 3, 0, 0] still gives the final two positions some weight. Future information would leak into the output.

Softmax of the sat Row: Blocked Scores Set to 0 vs Set to −∞

Only a value that becomes zero after the exponential, such as negative infinity, removes a position completely.

In real code, libraries may use a very large negative finite number instead of a literal negative infinity. The purpose is the same: its exponential is so close to zero that the blocked softmax weight is zero for practical computation.

Calculate the Allowed Weights

For the sat row, subtract the largest allowed score, 3. The shifted row is [-1, -2, 0, -infinity, -infinity]. Its exponentials are approximately [0.367879, 0.135335, 1, 0, 0]. Dividing by their sum gives [0.244728, 0.090031, 0.665241, 0, 0].

A small NumPy example makes the blocked positions explicit. We import the array tool before defining the teaching scores.

PYTHON
import numpy as np

The Boolean array says which keys the query may read. np.where selects the score when allowed and negative infinity otherwise. Multiplying a zero-one mask by negative infinity is unsafe because zero times infinity is not a defined floating-point number.

PYTHON
scores = np.array([2., 1., 3., 4., 2.])
allowed = np.array([True, True, True, False, False])
masked = np.where(allowed, scores, -np.inf)
exp_scores = np.exp(masked - masked.max())
weights = exp_scores / exp_scores.sum()
assert np.all(weights[~allowed] == 0)

Keep at least one allowed key in this mathematical example. If a row has no allowed key, every score is negative infinity and there is no valid softmax denominator. Padding and empty sequences need explicit handling rather than hoping the same formula will work.

Training Can Still Process Every Row Together

The mask controls information flow, not whether hardware can calculate a matrix in parallel. During training, the model forms all query-key scores for the full sentence in one matrix operation. It also adds the full triangular mask in one operation.

That means the row for The, the row for cat, and the row for mat can all be computed together. Yet the row for sat has only three allowed cells, while the row for mat has five.

All Five Masked Rows Are Computed in One Matrix Operation

Parallel training computes the whole score grid together, while the mask gives each row its own allowed past.

Focus on the sat row. It can combine the values from The, cat, and sat. It cannot receive a value from on or mat. The same rule holds even though those words sit in the same batch tensor.

Advertisement

The Same Rule During Generation

At generation time, future tokens are not merely hidden. They do not exist yet. After GPT has written The cat sat on, it calculates a probability list for the next token using those four available tokens. If it selects mat, that token becomes part of the next input.

The causal mask describes exactly this growing visible past. It is why a decoder writes from left to right and cannot later revise an old token using a word it has not generated yet.

Generation Adds One Chosen Token to the Visible Past at Each Step

Each chosen token becomes available for later steps, but it was unavailable when earlier tokens were predicted.

This decoder-only rule is different from an encoder's usual self-attention, where every input token can normally read every other input token. We will use that distinction again when we reach encoder-decoder models.

The rule must hold in every attention layer. Otherwise a later token could read the future in one layer, then pass that information backward in the next. Applying causal attention throughout the stack, with token-wise normalization and feed-forward operations, prevents this indirect path.

A useful test is to change only the future suffix of an input and compare the earlier output rows with dropout disabled. Those earlier rows should remain unchanged within numerical tolerance. A model that passes a shape test but fails this test may still be leaking future information.

The Decoder's Information Rule

The mask is not an optional visual decoration around attention. It is what makes a GPT-style decoder a next-token model. Without it, the model could solve training examples by looking right. With it, every prediction uses only information that would be available at the time of generation.

Encoder Self-Attention vs Decoder Causal Attention for "The cat sat on mat"

The triangular mask makes every attention layer obey the same left-to-right information boundary.

Recap

Causal masked attention lets a token attend to itself and earlier tokens, while blocking every future token. The allowed cells form a lower triangle in the attention score grid. Before softmax, blocked scores receive negative infinity, which gives them zero final weight.

On Day 17, we solve a different problem. Attention can read allowed words, but it does not automatically know their order. Positional encoding adds that missing information.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments