Prompt Anatomy: The Four Parts of a Prompt and What the LLM Actually Reads

Learn the four parts of a prompt by asking Qwen 3.8 about Amazon's 2024 10-K, then see the chat template and the hidden text the model actually reads.

Oct 2, 202621 min readFollow

Topics You Will Master

The four parts of a prompt: instruction, context, question and output format
How each part changes a real answer about Amazon's 2024 10-K
How a chat template turns our messages into one string of tokens
Why the model can read text that we never wrote

We asked Qwen 3.8 27B a simple question about Amazon's 2024 annual report. It answered in a confident voice, with a source and a neat summary. The answer was wrong. Amazon's 10-K says changes in exchange rates raised AWS operating income by $240 million. The model said they lowered it by $1.1 billion. But here's the real mystery: the same model gets it right when we change nothing but the prompt.

A prompt is all the text a model reads before it writes its answer. In simple words, it is the model's whole view of our task. In this lesson, we will learn the four parts of a good prompt, and we will see the exact tokens the model reads. This is the first lesson of our Prompt Engineering series, and every example in the series comes from real finance filings.

Four Parts of a Finance Prompt and the Tokens in Each

The prompt we build in this lesson. The 10-K excerpt is most of it.

Bestseller

Master Langchain v1 and Ollama - Chatbot, RAG and AI Agents

Master Langchain v1, Local LLM Projects, Ollama, DeepSeek, LLAMA 3.2, Complete Integration Guide.

95% off$199.99
$9.99Enroll now →30 day refund, lifetime access

Setup

We run Qwen 3.8 27B on our own machine with Ollama. Ollama is a free tool that downloads open models and runs them locally. We also load the model's tokenizer from Hugging Face. The tokenizer lets us see the raw text the model reads.

BASH
ollama pull qwen3.8:27b
BASH
pip install ollama transformers

The model file is 16.5 GB, so it needs a GPU or a Mac with enough memory. A smaller model works with the same code. Only the answers and the token counts will change.

Now, we write one small helper. It sends a prompt to the model and prints the answer with two counts. Let's see the code as below:

PYTHON
import ollama

MODEL = "qwen3.8:27b"

def ask(prompt):
    response = ollama.chat(
        model=MODEL,
        messages=[{"role": "user", "content": prompt}],
        think=False,
        options={"temperature": 0, "seed": 7},
    )
    print(f"prompt tokens: {response.prompt_eval_count} | answer tokens: {response.eval_count}")
    print(response.message.content)

A few choices in this helper are worth a word:

  • think=False turns off Qwen 3.8's thinking step. The model can write a long private reasoning before its answer. We skip it here to keep every answer short.
  • temperature and seed make the answer the same on every run. We will learn what temperature does in a later lesson.
  • prompt_eval_count is the number of tokens the model read. eval_count is the number of tokens it wrote. A token is a small piece of text, often part of a word.
Advertisement

What Happens When We Send Only the Question?

Let's start with the way most people use a chatbot. We type the question and nothing else.

Our question comes from Amazon's 10-K for 2024. A 10-K is the yearly report that a US public company must file. AWS is Amazon's cloud business. Operating income is the profit from the main business, before interest and taxes. Exchange rates move the value of money earned in other currencies. Let's test it:

PYTHON
question = "By how much did changes in foreign exchange rates impact AWS operating income in 2024?"
ask(question)
OUTPUT
prompt tokens: 32 | answer tokens: 257
Based on Amazon's financial reporting for the full year 2024, the impact of foreign exchange rate changes on AWS operating income was a **decrease of $1.1 billion**.

### Context and Details:
1.  **Source**: This figure is derived from Amazon’s Q4 2024 earnings release and 10-K filing, where they disclose the impact of foreign currency translation on segment operating income.
2.  **Reason**: The strengthening of the U.S. dollar against major foreign currencies (such as the Euro and British Pound) during 2024 reduced the value of AWS’s international revenues and profits when converted back into U.S. dollars.
3.  **Comparison**: In 2023, the foreign exchange impact on AWS operating income was a decrease of $1.4 billion. Thus, the negative impact was slightly smaller in 2024 compared to the previous year.

### Summary:
- **Impact**: Negative (decrease)
- **Amount**: **$1.1 billion**
- **Year**: 2024

This means that without the adverse foreign exchange effects, AWS’s reported operating income would have been $1.1 billion higher.

Here, we can see a polished answer. It has a source, a reason and a summary. But the number is wrong, and so is the direction. The 10-K says exchange rates raised AWS operating income by $240 million.

So what went wrong? The model never saw the filing. It only knows the text it was trained on. When that memory runs out, it fills the gap with a number that sounds right. In simple words, a question alone asks the model to answer from memory.

What Are the Four Parts of a Prompt?

A good prompt hands the model everything it needs, the way we would brief a new analyst. We tell them their job and the rules. We hand them the page from the report. We ask the question. Then we say how we want the reply written.

Let me tabulate the four parts for your better understanding:

Part Its job Our example
Instruction Says who the model is and which rules to follow You are a financial analyst. Use only the excerpt.
Context Gives the facts the model must use The operating income table and the AWS paragraph from the 10-K
Question Says what we want right now By how much did exchange rates impact AWS operating income?
Output format Says how the answer should look One sentence, amounts in millions of dollars

Now, we write each part as a Python string. The context is copied from page 27 of the 10-K. Let's see the code as below:

PYTHON
instruction = (
    "You are a financial analyst. Answer the question using only the 10-K excerpt below. "
    "If the excerpt does not contain the answer, say so."
)

context = """10-K excerpt (Amazon, fiscal year 2024):
Operating income (loss) by segment is as follows (in millions):
Year Ended December 31,
2023 2024
Operating Income (Loss)
North America $ 14,877 $ 24,967
International (2,656) 3,792
AWS 24,631 39,834
Consolidated $ 36,852 $ 68,593

The increase in AWS operating income in 2024, compared to the prior year, is primarily
due to increased sales, decreased payroll and related expenses, and a reduction in
depreciation and amortization expense from our change in the estimated useful lives of
our servers, partially offset by spending on technology infrastructure that was primarily
driven by additional investments to support AWS business growth. Changes in foreign
exchange rates positively impacted operating income by $240 million in 2024."""

output_format = "Reply in one sentence, with any amount in millions of dollars."

def build_prompt(question, instruction=None, context=None, output_format=None):
    parts = [instruction, context, f"Question: {question}", output_format]
    return "\n\n".join(part for part in parts if part)

The build_prompt function puts the parts in a fixed order: instruction, context, question, output format. It skips any part we leave out. This lets us add one part at a time and watch what changes. A blank line separates the parts, so the model can tell where one ends and the next begins.

Advertisement

How Does Each Part Change the Answer?

First, we add only the context. Let's test it:

PYTHON
ask(build_prompt(question, context=context))
OUTPUT
prompt tokens: 263 | answer tokens: 32
Based on the text provided, changes in foreign exchange rates positively impacted AWS operating income by **$240 million** in 2024.

Here, we can see the right number and the right direction. The context alone fixed the fact. The prompt grew from 32 tokens to 263, and the answer shrank from 257 tokens to 32.

Next, we add the instruction:

PYTHON
ask(build_prompt(question, instruction, context))
OUTPUT
prompt tokens: 295 | answer tokens: 36
Based on the 10-K excerpt provided, changes in foreign exchange rates positively impacted AWS operating income by **$240 million** in 2024.

The answer is the same. Only the wording moved: "the text provided" became "the 10-K excerpt provided". The instruction's main job shows up later, when the answer is missing from the excerpt.

Now, we add the output format:

PYTHON
ask(build_prompt(question, instruction, context, output_format))
OUTPUT
prompt tokens: 309 | answer tokens: 24
Changes in foreign exchange rates positively impacted AWS operating income by $240 million in 2024.

Here, we can see a single plain sentence. There is no bold text and no lead-in. The answer is now 24 tokens. A short, fixed shape like this is easy for our code to use.

Finally, we ask something the excerpt does not cover. The excerpt has no interest expense in it:

PYTHON
ask(build_prompt("What was Amazon's interest expense in 2024?", instruction, context, output_format))
OUTPUT
prompt tokens: 302 | answer tokens: 24
The provided 10-K excerpt does not contain information regarding Amazon's interest expense in 2024.

Here, we can see the instruction at work. The model says the answer is not there. Without a rule like this, nothing in the prompt tells the model what to do when the context runs out. It may then fall back on its memory, and that is exactly what went wrong in our first run.

Let me tabulate the four runs on the exchange rate question:

Prompt Prompt tokens Answer tokens Answer
Question only 32 257 decrease of $1.1 billion (wrong)
+ context 263 32 positively impacted by $240 million
+ instruction 295 36 positively impacted by $240 million
+ output format 309 24 positively impacted by $240 million

One FX Question: Answers as Each Prompt Part Is Added

The context fixed the fact. The instruction and the output format shaped the answer.

Advertisement

What Does the Model Actually Read?

So far, we have sent our prompt as a message with a role, like user. Chat apps also use a system message for the rules and an assistant message for the model's replies. But a model does not read a list of messages. It reads one long row of tokens, from left to right.

So, here comes the chat template to the rescue. A chat template is a small program that ships with the model. It turns our list of messages into one string. In simple words, it wraps each message in special marker tokens that the model learned during training. We met tokens like these in our lesson on special tokens.

Let's load the Qwen 3.8 tokenizer and turn two messages into the string the model reads:

PYTHON
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3.8-27B")

messages = [
    {"role": "system", "content": "You are a financial analyst."},
    {"role": "user", "content": "What was AWS operating income in 2024?"},
]
text = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
print(text)
OUTPUT
<|im_start|>system
You are a financial analyst.<|im_end|>
<|im_start|>user
What was AWS operating income in 2024?<|im_end|>
<|im_start|>assistant
<think>

</think>

Here, we can see the whole conversation as one piece of text:

  • <|im_start|> opens a message, and the role name follows it.
  • <|im_end|> closes the message.
  • add_generation_prompt=True adds an open assistant turn at the end. The model writes its answer from this point on.
  • The empty <think> block tells the model to skip its thinking step. This is what enable_thinking=False adds.

So, the roles are not a separate channel. They are just more tokens in the same row. We will see what this means for system prompts in the next lesson.

Qwen 3.8 Chat Template: Two Messages Become One String

Every message gets an opening marker with its role and a closing marker.

How Many Tokens Does Each Part Cost?

Every token costs something. On a paid API, we pay for each token we send. On our own GPU, each token takes time to read. Every model also has a limit on how many tokens it can read at once.

Let's count the tokens in each part of our finance prompt:

PYTHON
def count(text):
    return len(tokenizer(text, add_special_tokens=False)["input_ids"])

parts = {
    "instruction": instruction,
    "context": context,
    "question": f"Question: {question}",
    "output format": output_format,
}
for name, part in parts.items():
    print(f"{name:14} {count(part):4} tokens")

prompt = build_prompt(question, instruction, context, output_format)
chat_text = tokenizer.apply_chat_template(
    [{"role": "user", "content": prompt}],
    tokenize=False, add_generation_prompt=True, enable_thinking=False,
)
print(f"{'our prompt':14} {count(prompt):4} tokens")
print(f"{'model input':14} {count(chat_text):4} tokens")
OUTPUT
instruction      31 tokens
context         228 tokens
question         22 tokens
output format    13 tokens
our prompt      297 tokens
model input     309 tokens

Here, we can see that the context is 228 of our 297 tokens. That is about three quarters of the prompt. The template adds 12 more tokens, for 309 in total. This is the same number Ollama reported when we ran the full prompt.

This pattern holds for most real finance work. The instruction and the question stay small. The context grows, because Amazon's 10-K runs to 89 pages. Choosing which context to send is a big part of prompt engineering, and later lessons in this series deal with long prompts.

Advertisement

Can the Model Read Text We Never Wrote?

Earlier, we turned thinking off with enable_thinking=False. Now, we render the same two messages with the template's default settings:

PYTHON
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
print(text)
OUTPUT
<|im_start|>system
Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.

You are a financial analyst.<|im_end|>
<|im_start|>user
What was AWS operating income in 2024?<|im_end|>
<|im_start|>assistant
<think>

Here, we can see a whole sentence we never typed. "Reasoning effort is set to xhigh..." sits at the top of the system message, before our own text. Qwen 3.8's template turns thinking on by default, and it adds this line to tell the model how hard to think.

Does Ollama add the same line? Ollama builds its own prompt from the same messages, and it reports how many tokens it sent. So we can compare the two tools on four thinking settings:

PYTHON
settings = [
    ("thinking off", {"enable_thinking": False}, False),
    ("low effort", {"reasoning_effort": "low"}, "low"),
    ("medium effort", {"reasoning_effort": "medium"}, "medium"),
    ("default", {}, True),
]
for name, template_args, think in settings:
    text = tokenizer.apply_chat_template(
        messages, tokenize=False, add_generation_prompt=True, **template_args
    )
    response = ollama.chat(model=MODEL, messages=messages, think=think, options={"num_predict": 1})
    print(f"{name:14} transformers: {count(text):3} | ollama: {response.prompt_eval_count:3}")
OUTPUT
thinking off   transformers:  35 | ollama:  35
low effort     transformers:  59 | ollama:  59
medium effort  transformers:  33 | ollama:  33
default        transformers:  71 | ollama:  33

Here, we can see that the first three rows match exactly. The last row does not. With default settings, the Hugging Face template sends 71 tokens and Ollama sends 33. The 38 extra tokens are the "xhigh" sentence. The template's default is xhigh effort, while Ollama's default for this model is medium, and medium adds no sentence at all.

So, the same model and the same messages can give the model two different prompts. The tool in the middle decides part of what the model reads. When we compare results across tools, we should first check the real model input.

Quick Reference

Setting Where What it does
think=False ollama.chat Turns off the thinking step
options={"temperature": 0, "seed": 7} ollama.chat Makes the answer the same on every run
prompt_eval_count Ollama response Tokens the model read
eval_count Ollama response Tokens the model wrote
apply_chat_template(..., tokenize=False) Hugging Face tokenizer Shows the exact string the model reads
add_generation_prompt=True apply_chat_template Opens the assistant turn at the end
enable_thinking=False Qwen 3.8 template Adds an empty think block, no effort sentence
reasoning_effort="low" Qwen 3.8 template Adds the low effort sentence

Recap

This is how a prompt works. We started with a bare question, and Qwen 3.8 answered from memory with the wrong number. Then we built the prompt from four parts. The context fixed the fact, the instruction handled the missing answer, and the output format gave us one clean sentence. Finally, we looked at what the model actually reads: one row of tokens wrapped by a chat template, sometimes with text we never wrote.

In the next lesson, we will look at system prompts and roles, and how the model treats the system message.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments