Self-attention lets tokens read other tokens from the same sequence. Cross-attention connects two different sequences. One sequence asks a question, and the other sequence supplies the information that may answer it.
In the original Transformer translator, an encoder first reads the source sentence. A decoder then writes the target sentence. While writing each target token, the decoder uses cross-attention to look back at useful parts of the encoded source.
The diagram below shows this source-to-target bridge.

The encoder prepares source representations, while the decoder repeatedly reads them as it creates the target sequence.
Two Sequences Have Two Jobs
Take a source sentence, the cat sleeps. The encoder reads all three source tokens and turns each one into a context-aware vector. Now suppose a French decoder has already written le and needs to choose its next target token.
The decoder does not need to ask its own previous words what the source says. It needs to ask the encoder's source representations. That is the central split in cross-attention: decoder state creates the question, encoder output supplies the possible answers.

The source sequence is read by the encoder; the growing target sequence is written by the decoder.
This split is useful when input and output have different lengths. A source sentence may use three tokens while its translation needs four or five. Cross-attention lets each target position choose the source information it needs without forcing token 1 to match token 1.
Queries Come From the Decoder
The attention formula stays familiar:
The difference is not the arithmetic. It is where the three inputs originate.
- comes from the decoder's current hidden state.
- comes from the encoder output for each source token.
- also comes from the encoder output and carries source information forward.

The decoder asks the question; the encoder provides the searchable source memory.
The encoder's output is often called a memory because it remains available while the decoder writes. It is not a word-for-word dictionary. Each encoder vector has already mixed context from the source sentence through encoder self-attention.
A Rectangular Attention Grid
Let the source contain three positions and the known target prefix contain two. If each query and key has four features, has shape and has shape . Then has shape : two target questions, each with three source scores.
This grid need not be square. Its row numbers describe target positions; its column numbers describe source positions. A target row 1 is not "later than" source column 2 in the causal-mask sense. These are positions in different sequences.
The encoder output is transformed by separate learned key and value matrices. The decoder state is transformed by a learned query matrix. Source and target hidden widths can differ before projection, provided the resulting query and key feature widths agree.
One Decoder Position Scores All Source Positions
When the decoder is ready to choose the token after le, its current vector becomes one query. That query is compared with the key for the, the key for cat, and the key for sleeps.
The result is one score per source position. A larger score means the decoder's current question matches that source representation more closely.

One target-side question can compare itself with every source token representation.
For a small teaching example, suppose the score row is [0.2, 2.0, 0.5] for the, cat, and sleeps. These are already-scaled scores, after division by the square root of the key width. Softmax will turn them into source weights.
Softmax Makes a Source Focus
Softmax on [0.2, 2.0, 0.5] gives approximately [0.11905, 0.72024, 0.16071]. The second source position, cat, receives the largest share.
That does not mean cross-attention copies cat straight into the French output. It means the decoder's next representation receives mostly the encoder information stored at the cat position. The decoder's output layer still decides which target token is likely.

The decoder can focus mostly on one source position while retaining smaller contributions from others.
The weights then blend encoder value vectors, just as in self-attention. Keys were used to decide the match. Values carry the context selected by the match.
Use teaching value rows [1,0], [0,1], and [1,1] for the, cat, and sleeps. Their weighted mixture is approximately [0.27976, 0.88095]. The second coordinate is large because both the cat and sleeps values contribute to it. The largest attention weight does not mean every output coordinate comes only from that one source position.
We can check the scores and mixture in two small pieces. NumPy gives us arrays and exponentials; the scores are supplied teaching inputs rather than outputs from a trained translator.
import numpy as np
Subtracting the maximum keeps exponentials within a safe range without changing the weights. The matrix product then applies one weight to each source value row.
scores = np.array([0.2, 2.0, 0.5])
exp_scores = np.exp(scores - scores.max())
weights = exp_scores / exp_scores.sum()
source_values = np.array([[1., 0.], [0., 1.], [1., 1.]])
source_context = weights @ source_values
This source-context vector is not a French token ID. More decoder computation and the vocabulary output head still stand between the mixture and the selected word.

Cross-attention uses source weights to build a source-aware vector for the current decoder position.
Self-Attention and Cross-Attention Ask Different Questions
An encoder layer usually uses self-attention: every source token can read other source tokens. That helps an encoder representation for cat include information from the and sleeps.
A decoder layer usually begins with masked self-attention. This asks, "What have I written so far?" It can see earlier target tokens, but not later target tokens that have not been generated.
After masked self-attention, decoder cross-attention asks a different question: "Which source information helps me write the next target token?" It can normally read every encoder output because the complete source sentence was available before decoding began.

The decoder remembers its target-side past, then consults the full source-side memory.
No Unknown Target Token Is Used as a Query
When we say "attention while predicting chat," the query comes from the current known target state, such as the state after le. It does not come from an embedding of chat before the model has chosen it. During training, shifted target inputs enforce the same prediction boundary.
The complete source is already known in ordinary translation, so cross-attention can read all real source positions. It still needs a padding mask if a batch contains shorter source sentences. Streaming translation with an incomplete source would require another visibility policy; the full-source example does not cover that case.
The encoded source also stays fixed while this target answer is generated. Each decoder cross-attention layer can project and retain its own source keys and values. Unlike the growing target self-attention cache, those source arrays do not gain a position each time the decoder emits a token.
Three Attention Jobs in an Encoder-Decoder Transformer
An encoder-decoder model has three related attention jobs:
| Location | Attention type | What it can read |
|---|---|---|
| Encoder | self-attention | all source tokens |
| Decoder | masked self-attention | earlier target tokens and itself |
| Decoder | cross-attention | all encoder source outputs |
The three jobs have the same basic score, softmax, and weighted-value calculation. Their allowed inputs differ because they serve different parts of the task.

Source understanding, target-side history, and source lookup are handled by separate attention steps.
This pattern works beyond translation. A question-answering model can encode a passage and let a decoder look up relevant passage information. A summarization model can encode a long document and look back at different document sections while writing a shorter answer.
Recap
Cross-attention is a source lookup. The decoder provides queries from its current target-side state. The encoder provides keys and values from its complete source representation. Attention weights decide which source values should influence the current target position.
On Day 19, we will compare all three Transformer families: encoder-only, decoder-only, and encoder-decoder models, including the blocks they share.