On Day 16, a causal mask told a decoder which tokens it may read. That rule introduces a direction: different positions can read different prefixes. Explicit position information adds a direct signal about location and distance.
Self-attention compares a collection of token vectors. By itself, that comparison has no built-in arrow from left to right. Positional encoding adds a numeric clue that tells the Transformer where each token sits in the sequence.
The first diagram shows why that extra clue matters.

A token representation needs information about both the token itself and its place in the sequence.
The Same Words Can Say Different Things
Compare these two sentences:
dog bites man
man bites dog
They use the same three words. Their meaning changes because dog and man traded places. A normal word embedding can describe dog, but the embedding alone does not carry a label saying "I was token 1 in this sentence."

Word identity is not enough when the order of the same words changes the meaning.
This matters for more than short examples. In not good, the position of not changes the feeling of the phrase. In code, moving a parenthesis can change the result. In a story, a pronoun often needs position to connect with the right noun.
Attention Alone Can Be Shuffled
Attention uses queries and keys to compare tokens. If we take every token vector and reorder the rows, attention can make the same kind of comparisons between the reordered rows. Nothing in the dot product itself announces, "this row used to be second."
Mathematicians call this property permutation equivariance. You do not need the name to use the idea: for unmasked self-attention with no positional signal, shuffling the input rows shuffles the output rows in the same way. It does not recover the original reading order by itself.

Without a position signal, the attention calculation sees token vectors and their relationships, not their original places.
The shuffle argument needs its assumptions. If we reorder both keys and values together, the score columns and their matching value rows move together. The weighted mixture still collects the same content. Reordering queries also reorders which output belongs to which input.
A fixed causal triangle changes that situation: permuting tokens without changing the triangle changes who can read whom. Therefore "a decoder without positional embeddings has no order information at all" is too strong. The lesson is that dot products alone do not supply a direct position signal; masks and explicit position methods contribute different structure.
Add a Position Vector to Each Token
In the original additive design, before the first attention layer, the Transformer combines each token embedding with a vector for its position:
Here, means a place in the sequence. The token embedding answers "what is this token?" The position vector answers "where is it?"
Suppose cat has a tiny token embedding [0.2, 0.7, 0.1, 0.4]. For an arbitrary addition example, not the sinusoidal formula below, it could receive a position vector [0.9, 0.2, 0.8, 0.1]. The first layer receives their sum, [1.1, 0.9, 0.9, 0.5].

Addition puts word identity and position information in the same model-width vector.
We add the vectors rather than attach them side by side. Attaching them would make the input wider and force every later weight matrix to change shape. Addition keeps the expected width while letting later query, key, and value projections use both clues.
One Position Creates One Pattern
Every sequence position receives a different position vector. The vectors do not need to look like simple labels such as [0, 0, 1] or [0, 1, 0]. They need to make each position distinguishable in a way the model can learn from.
After addition, even the same token has a different input representation at different locations. cat at the beginning and cat near the end start from the same word embedding, but their combined input vectors differ. Therefore their later queries, keys, and values can differ too.
The Original Sine and Cosine Encoding
The first Transformer paper used fixed sine and cosine waves to make position vectors. For position and vector-pair number , the rule was:
Here is the model width. Each pair of columns changes at a different speed. Some values move quickly as position changes. Others move slowly.

Different wave speeds create a detailed pattern that changes from one position to the next.
At position 0, the first sine value is 0 and the first cosine value is 1. At position 1, they become roughly 0.84 and 0.54. These numbers are not meant for a person to read like a code. They are stable patterns from which a trained model can learn position and distance clues.

One position is represented by many sine and cosine values, not by a single position number passed through the network.
The paired sine and cosine values also help the model compare relative movement between positions. The exact reason involves trigonometry, but the practical point is simple: changing a position produces a structured, predictable change in its vector.
Work Through Four Dimensions
Choose an even model width . Pair 0 uses denominator 1; pair 1 uses denominator . At position 2, the vector is therefore approximately .
This is different from the arbitrary position vector used to introduce addition above. Here every number comes from the sinusoidal rule. At position 0 we get ; at position 1 we get approximately .
We can reproduce all three rows with NumPy. Positions start at zero, and each sine column has a matching cosine column at the same rate.
import numpy as np
The column of positions has shape (3, 1). The two denominators let it form two angles for each position. We allocate four columns because the sine and cosine values each occupy half the width.
positions = np.arange(3)[:, None]
width = 4
denominators = 10000 ** (np.arange(0, width, 2) / width)
angles = positions / denominators
encoding = np.empty((3, width))
encoding[:, 0::2] = np.sin(angles)
encoding[:, 1::2] = np.cos(angles)
Adding a learned embedding to a position vector does not let a later layer uniquely separate two arbitrary vectors from their sum. Instead, token embeddings and the following projections are trained to use the combined signal. Concatenation is possible in other designs, but it changes the dimensions and is not what this original formula does.
Fixed Position Vectors and Learned Position Vectors
Sinusoidal vectors are fixed. We calculate them from the formula and do not update them during training. A learned position embedding is different: it is a table with one trainable vector per supported position.

Fixed encodings come from a formula; learned embeddings are position rows adjusted during training.
Fixed encodings can be calculated for a new position beyond the training length. A learned table gives the model freedom to shape its position vectors, but the table normally has a maximum number of rows. Modern models also use methods such as relative positions and rotary position embeddings. The shared goal is still to give attention useful order information.
Position Is Present Before Attention Begins
In this additive design, position information is added before the first attention layer. Every later layer receives representations that already include a trace of where each token began. A query for cat at position 2 can therefore differ from a query for cat at position 5.
The causal mask and positional encoding solve different problems. The mask says which tokens this position may read. Positional information says where each allowed token is located.

Position enters at the input, so later attention calculations can use order as well as token meaning.
A formula being defined for position 50,000 does not prove that a model trained on 2,048-token sequences will use it well. Test long-distance tasks, not just whether the code accepts a larger position index.
On Day 29, RoPE will place position information directly in query and key rotations inside attention layers. It is not an extra additive input table. Keep the original additive pathway and the later rotary pathway separate when drawing a complete architecture.
Recap
Attention compares token vectors well, but comparison alone does not preserve word order. Positional encoding combines a position vector with every token embedding before attention begins. The original Transformer used fixed sine and cosine patterns; other models learn positions or place position information directly in attention.
Keep the two ideas from Days 16 and 17 separate: the mask controls what is visible, and the position signal tells the model where visible tokens are.
On Day 18, we will connect two token sequences. That is cross-attention, where one sequence asks questions and another sequence supplies the information.