Qwen2.5-0.5B is the smallest model in the Qwen2.5 family, which Alibaba's Qwen team released in September 2024. It has 494M parameters and runs on almost any GPU. But inside that small file sits the same machine as the big chat models: a decoder that reads tokens and scores every possible next token.
In this blog, we will learn how Qwen2.5-0.5B works by running real code on it. We will follow one short sentence from text to tokens, through 24 decoder layers, to the next-token scores. On the way, we will see what each part does, where the weights sit and how much memory generation needs.
What Is Qwen2.5-0.5B?
Qwen2.5 is built on the Transformer, the design from the 2017 paper "Attention Is All You Need" by Vaswani et al. So before we open Qwen, let's look at that original model as below.

Here, we can see two towers. The left tower is the encoder, which reads the whole input at once. The right tower is the decoder. Its Masked Multi-Head Attention lets each token look only at the tokens before it, and its Linear and Softmax at the top score the next token. Qwen2.5 keeps only the decoder tower and drops the middle attention block that reads from the encoder. The "Nx" next to the tower means the layer is stacked N times, and in Qwen2.5-0.5B N is 24. The two zoom-ins on the right show Multi-Head Attention and the Scaled Dot-Product Attention inside each head, and Qwen still runs that same attention in every layer.
The Qwen2.5 family comes in seven sizes, from 0.5B to 72B parameters. All of them are dense, decoder-only models, pretrained on up to 18 trillion tokens. The 0.5B model is released under the Apache 2.0 license, and its model card lists a context length of 32,768 tokens.
Each size ships as a base model and an instruct model. The base model only continues text. The instruct model is the same network, trained further to follow chat instructions. In this blog, we use Qwen2.5-0.5B-Instruct, the model our course notebook uses. The base model has the same architecture, so every shape and weight count below applies to it too.
Let me tabulate the main settings for your better understanding.
| Setting | Value |
|---|---|
| Parameters | 494,032,768 (0.49B) |
| Parameters without embeddings | 357,898,112 (0.36B) |
| Decoder layers | 24 |
| Hidden size (numbers per token) | 896 |
| Query heads | 14 |
| Key and value heads | 2 |
| Head size | 64 |
| MLP size | 4,864 |
| Embedding rows | 151,936 |
| Tokenizer entries | 151,665 |
| Context length | 32,768 tokens |
| Positions | RoPE, base 1,000,000 |
| Norm | RMSNorm, eps 1e-6 |
| MLP activation | SiLU, inside a SwiGLU MLP |
| LM head | Shared with the embedding table |
| Stored weights | bfloat16 |
The picture below shows the whole model at once. On the left is the full stack: a token embedding, 24 decoder layers, a final RMSNorm and the LM head. The middle panel zooms into one decoder layer, and the right panels zoom into its MLP and its attention.

Here, we can see two things that stand out. There is no position table at the bottom of the model, because positions enter as RoPE rotations inside attention. And attention makes only 2 key heads and 2 value heads for its 14 query heads. We will open every one of these boxes with code.
What Changed Since GPT-2?
GPT-2 is the classic decoder that most tutorials draw, and we took it apart in Transformer Architecture and Tokenization Explained. Qwen2.5 keeps its overall plan: a stack of identical layers, each with masked self-attention and an MLP, joined by residual connections. But almost every part inside the layer has been swapped for a newer idea.
Let me tabulate the differences for your better understanding.
| Part | GPT-2 small | Qwen2.5-0.5B |
|---|---|---|
| Size | 124M parameters, 12 blocks, width 768 | 494M parameters, 24 layers, width 896 |
| Positions | A learned table of 1,024 position vectors, added at the bottom | No table: RoPE rotates Q and K inside attention |
| Norm | LayerNorm, with a mean and a bias | RMSNorm, with neither |
| Attention heads | 12 heads, each with its own K and V | 14 query heads sharing 2 K heads and 2 V heads |
| Q, K, V layers | One fused c_attn layer |
Separate q_proj, k_proj and v_proj, all with a bias |
| MLP | Linear, GELU, Linear | SwiGLU: two input matrices, a SiLU gate, one output matrix |
| Dropout | Yes, during training | None |
| Context | 1,024 tokens | 32,768 tokens |
| Vocabulary | 50,257 tokens | 151,936 embedding rows |
| LM head | Shared with the token embedding | Shared with the token embedding |
Each change has a job. RMSNorm does less work than LayerNorm. RoPE puts the distance between two tokens straight into their attention score. Grouped-query attention shrinks the memory that generation needs. SwiGLU gives every MLP unit its own learned gate. We will meet each one below.
Our Test Setup
We run every example on one NVIDIA GeForce RTX 5090 (32 GB) with PyTorch 2.11.0 and Hugging Face transformers 5.2.0. The model loads in its stored format, bfloat16. We never train anything, so gradients are switched off for the whole blog.
Every number in this blog, and in every figure, comes from the code below. The examples follow the course notebook: the same sentences, the same random seed and the same order of steps.
Let's see the code as below. It loads the tokenizer and the model, then prints the model:
import math
import torch
import matplotlib.pyplot as plt
from transformers import AutoTokenizer, AutoModelForCausalLM, logging
torch.set_grad_enabled(False)
logging.set_verbosity_error()
MODEL = "Qwen/Qwen2.5-0.5B-Instruct"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, device_map="cuda")
model
Qwen2ForCausalLM(
(model): Qwen2Model(
(embed_tokens): Embedding(151936, 896)
(layers): ModuleList(
(0-23): 24 x Qwen2DecoderLayer(
(self_attn): Qwen2Attention(
(q_proj): Linear(in_features=896, out_features=896, bias=True)
(k_proj): Linear(in_features=896, out_features=128, bias=True)
(v_proj): Linear(in_features=896, out_features=128, bias=True)
(o_proj): Linear(in_features=896, out_features=896, bias=False)
)
(mlp): Qwen2MLP(
(gate_proj): Linear(in_features=896, out_features=4864, bias=False)
(up_proj): Linear(in_features=896, out_features=4864, bias=False)
(down_proj): Linear(in_features=4864, out_features=896, bias=False)
(act_fn): SiLUActivation()
)
(input_layernorm): Qwen2RMSNorm((896,), eps=1e-06)
(post_attention_layernorm): Qwen2RMSNorm((896,), eps=1e-06)
)
)
(norm): Qwen2RMSNorm((896,), eps=1e-06)
(rotary_emb): Qwen2RotaryEmbedding()
)
(lm_head): Linear(in_features=896, out_features=151936, bias=False)
)
Here, we can see the whole model in a few lines. embed_tokens is the table of 151,936 token vectors. layers holds 24 copies of the same decoder layer, each with self_attn, mlp and two Qwen2RMSNorm layers. At the end, norm is the final RMSNorm and lm_head turns each vector of 896 numbers into 151,936 scores. rotary_emb holds no weights: it only computes rotation angles.
The config gives us the head counts and the MLP size directly.
model.config.num_attention_heads, model.config.num_key_value_heads, model.config.intermediate_size
(14, 2, 4864)
So every layer has 14 query heads, only 2 key-value heads, and an MLP that widens 896 numbers to 4,864.
How Does Text Become Tokens?
The model never sees letters or words. It only sees token ids, which are row numbers in its vocabulary. Qwen2.5 uses a byte-level BPE tokenizer, so any text, in any language, can be split into tokens it knows.
Let's tokenize the notebook's first sentence.
ids = tok("It's time for tea.")["input_ids"]
ids
[2132, 594, 882, 369, 15243, 13]
The sentence becomes 6 ids. To see what each id stands for, we turn them back into tokens.
tok.convert_ids_to_tokens(ids)
['It', "'s", 'Ġtime', 'Ġfor', 'Ġtea', '.']
Here, we can see 6 tokens. Ġ is how the tokenizer prints a space, so Ġtea means " tea" with its leading space. Common words are one token each.
A rarer word splits into pieces that the vocabulary does have.
tok.convert_ids_to_tokens(tok("unbelievably ")["input_ids"])
['un', 'belie', 'vably', 'Ġ']
Here, "unbelievably" becomes 3 pieces, and the trailing space becomes a token of its own.
Now let's count the tokenizer's entries.
len(tok)
151665
The tokenizer has 151,665 entries: 151,643 learned by BPE and 22 special tokens, such as the markers of the chat template. The figure below shows both examples.

How Do Token Ids Become Vectors?
Each token id picks one row of a large learned table, the embedding table. That row is the token's first vector, and it is what the 24 layers will work on.
E = model.model.embed_tokens.weight.float()
E.shape
torch.Size([151936, 896])
The table has 151,936 rows of 896 numbers each, which is 136,134,656 weights. That is more than a quarter of the whole model. It also has 271 more rows than the tokenizer has entries, so no token ever reads those last rows.
To see what the table has learned, we compare the row for Ġtea with every other row. Cosine similarity measures how closely two vectors point the same way:
Where:
- , : two rows of the embedding table
- : their dot product
- : the length of
- the result: 1 means the same direction, 0 means unrelated
Let's find the 10 closest rows to Ġtea.
tea = E[tok.encode(" tea")[0]]
similarity = torch.cosine_similarity(E, tea[None], dim=-1)
[tok.decode(i) for i in similarity.topk(10).indices.tolist()]
[' tea', ' Tea', '茶', 'tea', ' teas', '茶叶', ' coffee', 'coffee', '喝茶', '红茶']
Here, we can see other forms of "tea", the Chinese words for tea, tea leaves, drinking tea and black tea, and coffee, a related drink. Nobody wrote these links. Training placed tokens that are used in similar ways close together. The figure below shows the lookup and the cosine scores.

What Happens Inside One Decoder Layer?
Every decoder layer does two jobs on a stream of vectors, one vector of 896 numbers per token. This stream is called the residual stream. First, attention lets each token collect information from the tokens before it. Then the MLP works on each token alone.
Each job reads a normalized copy of the stream and adds its result back:
Where:
- : the layer's input, one row of 896 numbers per token
- : the stream after the attention job
- : the layer's output, which is the next layer's input
- , : the modules
input_layernormandpost_attention_layernorm
In simple words, each layer edits the vectors and never replaces them. The added keeps what the token already knew. The norm runs only on the copy that goes into attention or the MLP, not on the stream itself. This order, norm first and add after, is called pre-norm, and GPT-2 uses it too.
What Does RMSNorm Do?
Vectors in the residual stream can grow or shrink from layer to layer. Attention and the MLP work best when their inputs have a steady size. So every layer first rescales its input with a norm.
GPT-2's LayerNorm subtracts the mean, divides by the standard deviation, then multiplies by a learned weight and adds a learned bias. Qwen2.5's RMSNorm keeps only the rescaling. It divides each vector by its root mean square and multiplies by a learned weight:
Where:
- : one token's vector
- : its length, 896 in Qwen2.5
- : a tiny number, 1e-6 in Qwen2.5, that keeps the division safe
- : the learned weight, one number per position in the vector
- : multiply position by position
Let's try both norms on 4 numbers. First the RMSNorm division, without the learned weight:
x = torch.tensor([2.0, -1.0, 0.5, 3.0])
rms = x.pow(2).mean().add(1e-6).sqrt()
rms, x / rms
(tensor(1.8875), tensor([ 1.0596, -0.5298, 0.2649, 1.5894]))
Here, we can see that the root mean square is 1.8875, and every number is divided by it. The signs and the ratios between the numbers stay the same. Only the size changes.
Now LayerNorm on the same numbers:
torch.nn.functional.layer_norm(x, (4,))
tensor([ 0.5773, -1.4021, -0.4124, 1.2372])
LayerNorm first moves the mean to 0, so it gives a different result: 2.0 becomes 0.5773 instead of 1.0596. The figure below puts the two side by side.

How Does Attention Work, Step by Step?
Before we look at Qwen's real heads, let's compute one attention head by hand, exactly as the notebook does. Let's say we have four tokens, It's, time, for and tea, each a vector of 8 numbers, and one head of size 4. The numbers are random, drawn with a fixed seed so that we get the same values every time.
Attention follows one formula:
Where:
- : the queries, what each token is looking for
- : the keys, what each token offers to be matched on
- : the values, what each token hands over
- : the head size, 4 here and 64 in Qwen2.5
First, three weight matrices turn the token vectors into Q, K and V.
torch.manual_seed(0)
X = torch.randn(4, 8)
Wq = torch.randn(8, 4) / math.sqrt(8)
Wk = torch.randn(8, 4) / math.sqrt(8)
Wv = torch.randn(8, 4) / math.sqrt(8)
Q = X @ Wq
K = X @ Wk
V = X @ Wv
Q.shape
torch.Size([4, 4])
Each of Q, K and V has one row per token and 4 numbers per row, the head size. The figure below shows the three projections with their real values.

Next, every query is compared with every key, the scores are scaled by , and softmax turns each row into weights.
scores = (Q @ K.T) / math.sqrt(4)
weights = torch.softmax(scores, dim=-1)
weights
tensor([[0.3552, 0.2201, 0.1022, 0.3224],
[0.1017, 0.3019, 0.1527, 0.4436],
[0.2511, 0.2138, 0.3434, 0.1917],
[0.0810, 0.2949, 0.0874, 0.5367]])
Here, we can see a 4 × 4 grid of weights, one row per token. The row for tea puts 0.5367 on itself and 0.2949 on time. Let's check the row sums.
weights.sum(dim=-1)
tensor([1., 1., 1., 1.])
Every row sums to 1, because softmax turns each row of scores into shares. The output blends the value rows by these shares.
out = weights @ V
out
tensor([[ 0.9502, -0.3241, -0.2500, -0.6852],
[ 1.1327, -0.3864, 0.1902, -0.4594],
[ 0.8674, 0.2653, 0.1599, -0.3794],
[ 1.1864, -0.6340, 0.1394, -0.5130]])
Each output row is a new vector for its token, mixed from all four value rows.
The Causal Mask
A decoder writes left to right, so a token must never look at the tokens after it. A causal mask marks which pairs are allowed.
mask = torch.tril(torch.ones(4, 4)).bool()
mask
tensor([[ True, False, False, False],
[ True, True, False, False],
[ True, True, True, False],
[ True, True, True, True]])
We set the blocked scores to minus infinity before softmax, so they become 0.
causal_weights = torch.softmax(scores.masked_fill(~mask, float("-inf")), dim=-1)
causal_weights
tensor([[1.0000, 0.0000, 0.0000, 0.0000],
[0.2520, 0.7480, 0.0000, 0.0000],
[0.3107, 0.2645, 0.4249, 0.0000],
[0.0810, 0.2949, 0.0874, 0.5367]])
Here, we can see that every weight above the diagonal is now 0. The row for It's can only see itself, so it gets 1.0. The row for tea could already see every token, so its weights do not change.
PyTorch has the same thing built in, as one fast function.
torch.nn.functional.scaled_dot_product_attention(Q, K, V, is_causal=True)
tensor([[ 0.4644, -0.1060, -1.3465, -1.2355],
[ 1.3065, 0.1847, 0.0644, -0.6232],
[ 0.7935, 0.7012, 0.2017, -0.3259],
[ 1.1864, -0.6340, 0.1394, -0.5130]])
Let's compare it with our own masked weights times V.
causal_weights @ V
tensor([[ 0.4644, -0.1060, -1.3465, -1.2355],
[ 1.3065, 0.1847, 0.0644, -0.6232],
[ 0.7935, 0.7012, 0.2017, -0.3259],
[ 1.1864, -0.6340, 0.1394, -0.5130]])
The two results match exactly. The figure below shows all four steps on these numbers: the raw scores, the scaling, the softmax and the causal mask.

How Does Qwen Split 896 Numbers Into Heads?
One head of size 4 is a toy. Qwen2.5 runs 14 heads in every layer, each 64 numbers wide, and 14 × 64 = 896. Splitting is only a reshape of the same numbers.
x = torch.randn(5, 896)
heads = x.view(5, 14, 64)
heads.shape
torch.Size([5, 14, 64])
Here, 5 token vectors of 896 numbers become 5 tokens × 14 heads × 64 numbers. Each head runs the attention we just computed on its own slice. Then o_proj mixes the 14 results back into one vector of 896 numbers.
Why Does Qwen Use Grouped-Query Attention?
While a model generates text, it saves the keys and values of every token it has already read, so that it does not compute them again on every step. This saved memory is called the KV cache. With one key head and one value head per query head, the cache grows quickly with every token.
So, here comes grouped-query attention to the rescue. The query heads are split into groups, and each group shares one key head and one value head. Let's look at the projection shapes in the first layer.
attn = model.model.layers[0].self_attn
attn.q_proj.weight.shape, attn.k_proj.weight.shape, attn.v_proj.weight.shape
(torch.Size([896, 896]), torch.Size([128, 896]), torch.Size([128, 896]))
Here, we can see that q_proj makes 896 numbers, 14 heads of 64. But k_proj and v_proj make only 128 numbers each, which is 2 heads of 64. The figure below shows how the 14 query heads share them.

Query heads 1 to 7 read key head 1 and value head 1. Query heads 8 to 14 read key head 2 and value head 2. Each query head still has its own query weights, so it can still look for its own pattern. Only the keys and values are shared.
Now let's measure what this saves. The cache stores one key and one value per layer, per KV head and per token, at 2 bytes per bfloat16 number.
def kv_cache_bytes(kv_heads, tokens, layers=24, head_dim=64, bytes_per_number=2):
return layers * 2 * kv_heads * head_dim * bytes_per_number * tokens
kv_cache_bytes(2, 1), kv_cache_bytes(14, 1)
(12288, 86016)
Here, one token needs 12,288 bytes of cache, which is 12 KB. If every query head kept its own key and value head, one token would need 86,016 bytes, which is 84 KB. That is 7 times more. Let's scale this up to the full context.
kv_cache_bytes(2, 32768) / 1024**2, kv_cache_bytes(14, 32768) / 1024**3
(384.0, 2.625)
At the full 32,768-token context, the cache needs 384 MB with grouped-query attention. Without it, the same context would need 2.625 GB. The figure below shows both sizes.

How Does RoPE Tell Positions Apart?
Attention compares vectors, and on its own it does not know the order of the tokens. "dog bites man" and "man bites dog" have the same three tokens. GPT-2 fixes this by adding a learned position vector to every token at the bottom of the model. Qwen2.5 has no position table at all.
Instead, it uses RoPE, rotary position embedding. In simple words, RoPE takes the 64 numbers of every query head and every key head, treats them as 32 pairs, and rotates each pair by an angle that grows with the token's position. Each pair turns at its own speed:
Where:
- : how many radians pair turns per position
- : the RoPE base from Qwen2.5's config
- : the head size
- a token at position has pair turned by radians
The model stores these 32 speeds as inv_freq.
inv_freq = model.model.rotary_emb.inv_freq
inv_freq.shape, inv_freq[0].item(), inv_freq[-1].item()
(torch.Size([32]), 1.0, 1.5399265294036013e-06)
Here, we can see 32 speeds. The fastest pair turns 1 radian per position. The slowest turns about 0.0000015 radian per position, so it needs about 4 million positions to go around once. Fast pairs tell nearby tokens apart, and slow pairs still change over long distances.
Why rotate? If we turn a query by one angle and a key by another, their dot product depends only on the difference between the two angles. So the score between a query at position and a key at position depends on , the distance between the two tokens. Let's test this with the model's own rotation code.
from transformers.models.qwen2.modeling_qwen2 import apply_rotary_pos_emb
def rotate(v, position):
cos, sin = model.model.rotary_emb(v, torch.tensor([[position]], device="cuda"))
return apply_rotary_pos_emb(v, v, cos, sin)[0]
torch.manual_seed(0)
q = torch.randn(1, 1, 1, 64, device="cuda")
k = torch.randn(1, 1, 1, 64, device="cuda")
for q_pos, k_pos in [(5, 3), (105, 103), (5, 4)]:
score = (rotate(q, q_pos) * rotate(k, k_pos)).sum().item()
print(q_pos, k_pos, round(score, 4))
5 3 -5.0553
105 103 -5.0553
5 4 -3.7713
Here, we can see that positions 5 and 3 give the same score as positions 105 and 103, -5.0553, because both pairs are 2 positions apart. A gap of 1 gives a different score. The figure below shows the rotations, the 32 speeds and this check.

RoPE touches only the queries and keys, never the values. And inv_freq is a fixed buffer, not a learned weight, so RoPE adds no parameters to the model.
What Do Real Attention Heads Look At?
Now let's look at real heads. Qwen2.5-0.5B has 24 layers × 14 heads = 336 heads. The default attention code in transformers is fast but does not return its weights, so we switch to the eager version, which does.
model.set_attn_implementation("eager")
x = tok("The cat sat on the mat because it was tired.", return_tensors="pt").to("cuda")
response = model(**x, output_attentions=True)
len(response.attentions), response.attentions[0].shape
(24, torch.Size([1, 14, 11, 11]))
Here, we get one grid per layer: 14 heads, each with 11 tokens looking at 11 tokens. Let's ask one head where the token "it" looks.
tokens = tok.convert_ids_to_tokens(x["input_ids"][0])
A = response.attentions[7][0, 7].float().cpu()
tokens[7], tokens[A[7].argmax()], round(A[7].max().item(), 2)
('Ġit', 'Ġcat', 0.84)
In layer 7, head 7, the row for "it" puts 0.84 of its attention on "cat", the noun it refers to. Layers and heads count from 0 here. The figure below shows this head next to three others on the same sentence.

Training found these patterns, and nobody programmed them. Some heads look at the first token, some at the token itself, some one step back, and a few follow meaning, like "it" to "cat".
Note
We keep eager attention on for the rest of this blog, as the course notebook does. The default attention code runs different GPU kernels, so its bfloat16 rounding moves the probabilities slightly: Ġthe below gets 0.2315 with it instead of 0.2259. The generated text stays the same.
What Does the SwiGLU MLP Do?
After attention has moved information between tokens, the MLP works on each token alone. GPT-2's MLP widens the vector, applies GELU and shrinks it back. Qwen2.5 uses SwiGLU, which adds a gate:
Where:
- : one token's 896 numbers
- , : the
gate_projandup_projmatrices, 896 → 4,864 - : the
down_projmatrix, 4,864 → 896 - : the sigmoid function
- : multiply unit by unit
In simple words, up_proj widens the token into 4,864 numbers. gate_proj makes a second set of 4,864 numbers and passes them through SiLU, which works like a soft on/off switch: large positive values pass, and negative values are pushed close to 0. Multiplying the two sets lets each of the 4,864 units be turned up or down by its own gate. Then down_proj squeezes the result back to 896 numbers.
Let's count the weights in one layer's MLP and attention.
mlp = sum(p.numel() for p in model.model.layers[0].mlp.parameters())
attention = sum(p.numel() for p in model.model.layers[0].self_attn.parameters())
mlp, attention
(13074432, 1836160)
Here, the MLP holds 13,074,432 weights per layer, about 7 times attention's 1,836,160. That is three matrices of 896 × 4,864, with no biases. The figure below shows the MLP's path and the SiLU curve next to GPT-2's GELU.

Where Do the 494M Parameters Sit?
Now we can add up the whole model.
embeddings = model.model.embed_tokens.weight.numel()
total = sum(p.numel() for p in model.parameters())
embeddings, 24 * mlp, 24 * attention, total
(136134656, 313786368, 44067840, 494032768)
Here, the 24 MLPs hold 313,786,368 weights and the 24 attention blocks hold 44,067,840. The embedding table holds 136,134,656. Together with 43,904 RMSNorm weights, they add up to exactly 494,032,768.
The LM head does not appear in this sum as a separate matrix. Let's check why.
model.lm_head.weight is model.model.embed_tokens.weight
True
The LM head is the embedding table itself. The same 151,936 × 896 matrix reads tokens in at the bottom and scores them at the top. A separate head would add another 136,134,656 weights. The figure below shows where every weight sits.

How Do Logits Become the Next Token?
After the last layer and the final RMSNorm, the LM head turns each token's vector into 151,936 raw scores, one per vocabulary row. These scores are called logits.
x = tok("It's time for", return_tensors="pt").to("cuda")
logits = model(**x).logits
logits.shape
torch.Size([1, 4, 151936])
Here, we get one row of 151,936 logits for each of the 4 input tokens. Only the last row is used to pick the next token. Softmax turns it into probabilities that sum to 1.
probs = logits[0, -1].float().softmax(dim=-1)
top = probs.topk(5)
tok.convert_ids_to_tokens(top.indices), top.values
(['Ġthe', 'Ġa', 'Ġanother', 'Ġour', 'Ġus'], tensor([0.2259, 0.1759, 0.0537, 0.0504, 0.0393], device='cuda:0'))
After "It's time for", the model spreads its bets. Ġthe gets 0.2259 and Ġa gets 0.1759, and Ġtea is not in the top 5.
Context Changes the Scores
Now let's keep the same last words but give the model more context before them.
x = tok("I'm thirsty and the kettle just boiled. It's time for a cup of", return_tensors="pt").to("cuda")
probs = model(**x).logits[0, -1].float().softmax(dim=-1)
top = probs.topk(5)
tok.convert_ids_to_tokens(top.indices), top.values
(['Ġtea', 'Ġcoffee', 'Ġhot', 'Ġwater', 'Ġcold'], tensor([0.4840, 0.2591, 0.1224, 0.0273, 0.0057], device='cuda:0'))
Now Ġtea leads with 0.4840. Attention carried "thirsty" and "kettle" forward to the last position, and that changed the scores. The figure below shows both prompts side by side.

How Does Generation Work?
Generation is a loop. We take the top token, append it to the input and run the model again. Always taking the top token is called greedy decoding.
ids = tok("I'm thirsty and the kettle just boiled. It's time for a cup of", return_tensors="pt").to("cuda")["input_ids"]
for _ in range(8):
next_id = model(ids).logits[0, -1].argmax()
ids = torch.cat([ids, next_id.view(1, 1)], dim=1)
tok.decode(ids[0])
"I'm thirsty and the kettle just boiled. It's time for a cup of tea. I'll go to the kitchen"
Here, eight passes added "tea. I'll go to the kitchen", one token per pass. Hugging Face's generate runs the same loop. Qwen's saved generation settings turn on sampling and a repetition penalty of 1.1, so we switch both off to get greedy decoding.
x = tok("I'm thirsty and the kettle just boiled. It's time for a cup of", return_tensors="pt").to("cuda")
response = model.generate(**x, max_new_tokens=8, do_sample=False, repetition_penalty=1.0)
tok.decode(response[0])
"I'm thirsty and the kettle just boiled. It's time for a cup of tea. I'll go to the kitchen"
generate gives exactly the same words. It also keeps a KV cache: after the first pass, each pass reads only the newest token and reuses the stored keys and values, so it does less work. The figure below shows the 8 passes of our loop.

How Does a Decoder Learn?
A decoder learns from plain text. Every position is a question with a known answer: what is the next token? The loss is the negative log of the probability the model gave to the real next token, averaged over the positions:
Where:
- : the token at position
- : the probability the model gives the real next token
- : the number of predictions, one fewer than the number of tokens
When we pass the input ids as labels, the model computes this loss for us.
x = tok("It's time for tea.", return_tensors="pt").to("cuda")
response = model(**x, labels=x["input_ids"])
response.loss
tensor(3.4888, device='cuda:0')
The mean loss over our sentence is 3.4888. Let's see the loss at each position.
torch.nn.functional.cross_entropy(response.logits[0, :-1].float(), x["input_ids"][0, 1:], reduction="none")
tensor([1.7237, 3.3228, 0.8647, 8.9251, 2.6079], device='cuda:0')
Here, we can see 5 losses for 6 tokens. The fourth position, Ġfor → Ġtea, costs the most, 8.9251, because after "It's time for" the model gave Ġtea a very small chance. Training nudges the weights to raise the probability of the real next token at every position. The figure below turns each loss back into the model's chance.

Conclusion
This is how Qwen2.5-0.5B works. We started with text, split it into 6 tokens and looked up their rows in a 151,936 × 896 table. We followed those vectors through 24 decoder layers, each one an RMSNorm, grouped-query attention with RoPE, a second RMSNorm and a SwiGLU MLP, joined by residual adds. Finally, the shared LM head turned the last vector into 151,936 scores, and a loop of such passes wrote text.
Key takeaways:
- Qwen2.5-0.5B keeps GPT-2's plan but swaps the parts: RMSNorm, RoPE, grouped-query attention and SwiGLU.
- Grouped-query attention lets 14 query heads share 2 key-value heads, which cuts the KV cache from 84 KB to 12 KB per token.
- RoPE adds no weights: it rotates queries and keys so that attention scores depend on the distance between tokens.
- The MLPs hold 63.5% of the weights, and the shared embedding table holds another 27.6%.
Next steps:
- Read The Attention Formula: What Softmax and the Square Root of d Actually Do for the scaling step in depth.
- Read Masked Attention: How GPT Avoids Seeing the Future for more on the causal mask.
- See this model classify text in Laya 421M vs Qwen3.5 0.8B vs Qwen2.5 0.5B: Small Models for Text Classification.