Cross-Attention: When One Sequence Reads Another

Learn how cross-attention lets a decoder use information from an encoder sequence.

Oct 6, 20268 min readFollow

Topics You Will Master

How cross-attention differs from self-attention
Where its queries, keys, and values come from
How a decoder looks up source information while writing
Why encoder-decoder models use both masked self-attention and cross-attention

Self-attention lets tokens read other tokens from the same sequence. Cross-attention connects two different sequences. One sequence asks a question, and the other sequence supplies the information that may answer it.

In the original Transformer translator, an encoder first reads the source sentence. A decoder then writes the target sentence. While writing each target token, the decoder uses cross-attention to look back at useful parts of the encoded source.

The diagram below shows this source-to-target bridge.

Cross-Attention: the Decoder State After "le" Reads the Encoded Source

The encoder prepares source representations, while the decoder repeatedly reads them as it creates the target sequence.

Bestseller

LLM Fine-Tuning with Hugging Face: LoRA, QLoRA, PEFT

Fine-tune BERT, T5, ViT, LLaMA-style models and Qwen3-TTS using Hugging Face Transformers, custom datasets, LoRA, QLoRA.

95% off$199.99
$9.99Enroll now →30 day refund, lifetime access

Two Sequences Have Two Jobs

Take a source sentence, the cat sleeps. The encoder reads all three source tokens and turns each one into a context-aware vector. Now suppose a French decoder has already written le and needs to choose its next target token.

The decoder does not need to ask its own previous words what the source says. It needs to ask the encoder's source representations. That is the central split in cross-attention: decoder state creates the question, encoder output supplies the possible answers.

Self-Attention Reads One Sequence, Cross-Attention Reads Another

The source sequence is read by the encoder; the growing target sequence is written by the decoder.

This split is useful when input and output have different lengths. A source sentence may use three tokens while its translation needs four or five. Cross-attention lets each target position choose the source information it needs without forcing token 1 to match token 1.

Advertisement

Queries Come From the Decoder

The attention formula stays familiar:

The difference is not the arithmetic. It is where the three inputs originate.

  • comes from the decoder's current hidden state.
  • comes from the encoder output for each source token.
  • also comes from the encoder output and carries source information forward.

Decoder Queries and Encoder Keys Give a 2 × 3 Score Grid

The decoder asks the question; the encoder provides the searchable source memory.

The encoder's output is often called a memory because it remains available while the decoder writes. It is not a word-for-word dictionary. Each encoder vector has already mixed context from the source sentence through encoder self-attention.

A Rectangular Attention Grid

Let the source contain three positions and the known target prefix contain two. If each query and key has four features, has shape and has shape . Then has shape : two target questions, each with three source scores.

This grid need not be square. Its row numbers describe target positions; its column numbers describe source positions. A target row 1 is not "later than" source column 2 in the causal-mask sense. These are positions in different sequences.

The encoder output is transformed by separate learned key and value matrices. The decoder state is transformed by a learned query matrix. Source and target hidden widths can differ before projection, provided the resulting query and key feature widths agree.

One Decoder Position Scores All Source Positions

When the decoder is ready to choose the token after le, its current vector becomes one query. That query is compared with the key for the, the key for cat, and the key for sleeps.

The result is one score per source position. A larger score means the decoder's current question matches that source representation more closely.

One Decoder Query Scores Every Source Key: 0.2, 2.0 and 0.5

One target-side question can compare itself with every source token representation.

For a small teaching example, suppose the score row is [0.2, 2.0, 0.5] for the, cat, and sleeps. These are already-scaled scores, after division by the square root of the key width. Softmax will turn them into source weights.

Advertisement

Softmax Makes a Source Focus

Softmax on [0.2, 2.0, 0.5] gives approximately [0.11905, 0.72024, 0.16071]. The second source position, cat, receives the largest share.

That does not mean cross-attention copies cat straight into the French output. It means the decoder's next representation receives mostly the encoder information stored at the cat position. The decoder's output layer still decides which target token is likely.

Softmax Turns the Source Scores 0.2, 2.0 and 0.5 Into Weights

The decoder can focus mostly on one source position while retaining smaller contributions from others.

The weights then blend encoder value vectors, just as in self-attention. Keys were used to decide the match. Values carry the context selected by the match.

Use teaching value rows [1,0], [0,1], and [1,1] for the, cat, and sleeps. Their weighted mixture is approximately [0.27976, 0.88095]. The second coordinate is large because both the cat and sleeps values contribute to it. The largest attention weight does not mean every output coordinate comes only from that one source position.

We can check the scores and mixture in two small pieces. NumPy gives us arrays and exponentials; the scores are supplied teaching inputs rather than outputs from a trained translator.

PYTHON
import numpy as np

Subtracting the maximum keeps exponentials within a safe range without changing the weights. The matrix product then applies one weight to each source value row.

PYTHON
scores = np.array([0.2, 2.0, 0.5])
exp_scores = np.exp(scores - scores.max())
weights = exp_scores / exp_scores.sum()
source_values = np.array([[1., 0.], [0., 1.], [1., 1.]])
source_context = weights @ source_values

This source-context vector is not a French token ID. More decoder computation and the vocabulary output head still stand between the mixture and the selected word.

Source Weights Blend the Encoder Value Rows Into a Context Vector

Cross-attention uses source weights to build a source-aware vector for the current decoder position.

Advertisement

Self-Attention and Cross-Attention Ask Different Questions

An encoder layer usually uses self-attention: every source token can read other source tokens. That helps an encoder representation for cat include information from the and sleeps.

A decoder layer usually begins with masked self-attention. This asks, "What have I written so far?" It can see earlier target tokens, but not later target tokens that have not been generated.

After masked self-attention, decoder cross-attention asks a different question: "Which source information helps me write the next target token?" It can normally read every encoder output because the complete source sentence was available before decoding began.

Each Decoder Step: Masked Self-Attention, Then Cross-Attention

The decoder remembers its target-side past, then consults the full source-side memory.

No Unknown Target Token Is Used as a Query

When we say "attention while predicting chat," the query comes from the current known target state, such as the state after le. It does not come from an embedding of chat before the model has chosen it. During training, shifted target inputs enforce the same prediction boundary.

The complete source is already known in ordinary translation, so cross-attention can read all real source positions. It still needs a padding mask if a batch contains shorter source sentences. Streaming translation with an incomplete source would require another visibility policy; the full-source example does not cover that case.

The encoded source also stays fixed while this target answer is generated. Each decoder cross-attention layer can project and retain its own source keys and values. Unlike the growing target self-attention cache, those source arrays do not gain a position each time the decoder emits a token.

Three Attention Jobs in an Encoder-Decoder Transformer

An encoder-decoder model has three related attention jobs:

Location Attention type What it can read
Encoder self-attention all source tokens
Decoder masked self-attention earlier target tokens and itself
Decoder cross-attention all encoder source outputs

The three jobs have the same basic score, softmax, and weighted-value calculation. Their allowed inputs differ because they serve different parts of the task.

Which Positions Each of the Three Attention Jobs Can Read

Source understanding, target-side history, and source lookup are handled by separate attention steps.

This pattern works beyond translation. A question-answering model can encode a passage and let a decoder look up relevant passage information. A summarization model can encode a long document and look back at different document sections while writing a shorter answer.

Recap

Cross-attention is a source lookup. The decoder provides queries from its current target-side state. The encoder provides keys and values from its complete source representation. Attention weights decide which source values should influence the current target position.

On Day 19, we will compare all three Transformer families: encoder-only, decoder-only, and encoder-decoder models, including the blocks they share.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments