We asked Qwen 3.8 for Amazon's segment margins and how much each one changed, all in one prompt. Every margin was right. Three of the four changes were wrong. The model had the right margins on the page, and it still got the differences wrong. But here's the real mystery: the fix was not a better instruction. It was splitting the job.
Prompt chaining means breaking one big task into small steps, where each step is its own prompt or its own piece of code. In simple words, the output of one step becomes the input of the next. In this lesson, we will build a chain for a real finance task, and we will add a checker prompt that reviews the result and sends errors back for a fix.

Each step does one job. The math step is plain Python.
Setup
We use Qwen 3.8 27B in Ollama again, with thinking off and the same answer on every run. This time, the helper takes one prompt string, since each step is a single message. Let's see the code as below:
import json
import ollama
MODEL = "qwen3.8:27b"
def chat(prompt):
response = ollama.chat(
model=MODEL,
messages=[{"role": "user", "content": prompt}],
think=False,
options={"temperature": 0, "seed": 7},
)
return response.message.content.strip()
Our data is two tables from Amazon's 2024 10-K: net sales and operating income for each segment. We copy them as plain text:
tables = """Amazon 10-K, fiscal year 2024 (in millions)
Net sales 2023 2024
North America $ 352,828 $ 387,497
International 131,200 142,906
AWS 90,757 107,556
Consolidated $ 574,785 $ 637,959
Operating income (loss)
2023 2024
North America $ 14,877 $ 24,967
International (2,656) 3,792
AWS 24,631 39,834
Consolidated $ 36,852 $ 68,593"""
An operating margin is operating income divided by net sales. It tells us how many cents of profit each dollar of sales brings in. A change in margin is often given in basis points. One basis point is one hundredth of a percentage point, so a move from 4.2% to 6.4% is about 220 basis points.
What Happens When One Prompt Does Everything?
First, the simple way. One prompt asks the model to compute all the margins, compute the changes, and write a note for an investment committee. Let's test it:
one_prompt = chat(
f"{tables}\n\n"
"Using these tables, compute the operating margin (operating income divided by net sales) "
"for each segment and the consolidated total, for 2023 and 2024, rounded to one decimal place. "
"Then write a three-sentence note for an investment committee about which segment drove "
"the change in Amazon's margin."
)
print(one_prompt)
Here are the calculated operating margins and the investment committee note.
### Operating Margin Calculations
**Formula:** Operating Margin = (Operating Income / Net Sales) × 100
#### 2023
* **North America:** $14,877 / $352,828 = 4.216% → **4.2%**
* **International:** $(2,656) / $131,200 = -2.024% → **-2.0%**
* **AWS:** $24,631 / $90,757 = 27.139% → **27.1%**
* **Consolidated:** $36,852 / $574,785 = 6.411% → **6.4%**
#### 2024
* **North America:** $24,967 / $387,497 = 6.443% → **6.4%**
* **International:** $3,792 / $142,906 = 2.653% → **2.7%**
* **AWS:** $39,834 / $107,556 = 37.034% → **37.0%**
* **Consolidated:** $68,593 / $637,959 = 10.752% → **10.8%**
### Summary Table
| Segment | 2023 Margin | 2024 Margin | Change (bps) |
| :--- | :--- | :--- | :--- |
| North America | 4.2% | 6.4% | +220 bps |
| International | -2.0% | 2.7% | +470 bps |
| AWS | 27.1% | 37.0% | +990 bps |
| **Consolidated** | **6.4%** | **10.8%** | **+440 bps** |
***
### Investment Committee Note
The primary driver of Amazon’s consolidated operating margin expansion from 6.4% to 10.8% was the significant improvement in the AWS segment, which saw its margin surge by 990 basis points to 37.0%. While the International segment also contributed positively by swinging from a loss to a 2.7% profit, the absolute dollar increase in operating income from AWS ($15.2 billion) far exceeded the gains from North America ($10.1 billion) and International ($6.4 billion). This indicates that AWS remains the key lever for Amazon’s overall profitability, with its high-margin growth outpacing the more modest improvements in the retail businesses.
Here, we can see a neat answer. All eight margins are right when we round them to one decimal place. But look at the change column. North America shows +220 bps, International +470 and the consolidated total +440.
The true changes are +223, +468 and +434. Every change in that column equals the difference of the rounded margins. So the model rounded first and subtracted second: 6.4 minus 4.2 gave 2.2 points, or 220 bps. Rounding too early is a classic mistake in finance work, and a committee note with the wrong change can mislead a reader.
The prompt asked the model to do three jobs at once: read the tables, do the math, and write the note. A slip in any job ends up in the final answer, and nothing in the prompt checks it.
What Is Prompt Chaining?
The fix is to give each job to the right worker. Language models are good at reading messy text and writing clear text. Plain code is better at arithmetic. So, here comes prompt chaining to the rescue: we split the task into steps and pass the result of each step to the next one.
Let me tabulate our chain for your better understanding:
| Step | Who does it | Input | Output |
|---|---|---|---|
| 1. Extract | The model | The two 10-K tables as text | The 16 numbers as JSON |
| 2. Compute | Python | The JSON numbers | Margins and basis point changes |
| 3. Write | The model | The computed margin table | A three-sentence note |
| 4. Check | The model, in a new prompt | The margin table and a text to check | A list of wrong numbers, or "ALL NUMBERS MATCH" |
Every step is small, so every step is easy to test on its own.
Step 1: How Do We Extract the Numbers?
The first prompt only copies numbers. It does no math and writes no prose. We ask for JSON, so our code can read the answer directly:
extract_prompt = (
f"{tables}\n\n"
"Copy every number from these two tables into JSON with this shape: "
'{"net_sales": {"North America": [2023, 2024], ...}, "operating_income": {...}}. '
"Use plain integers in millions. A number in brackets is negative. Return only the JSON."
)
numbers = json.loads(chat(extract_prompt))
for name, values in numbers.items():
print(name, values)
net_sales {'North America': [352828, 387497], 'International': [131200, 142906], 'AWS': [90757, 107556], 'Consolidated': [574785, 637959]}
operating_income {'North America': [14877, 24967], 'International': [-2656, 3792], 'AWS': [24631, 39834], 'Consolidated': [36852, 68593]}
Here, we can see all 16 numbers copied exactly, including the International loss of 2,656 as a negative number. The brackets in the table meant negative, and the prompt said so.
Step 2: How Do We Compute the Margins?
Now, plain Python does the arithmetic. We compute the margins from the full numbers and round only at the very end, when we print them.
lines = []
for segment in numbers["net_sales"]:
sales = numbers["net_sales"][segment]
income = numbers["operating_income"][segment]
m23, m24 = (100 * income[i] / sales[i] for i in range(2))
change_bps = round((m24 - m23) * 100)
lines.append(f"{segment}: {m23:.1f}% in 2023, {m24:.1f}% in 2024, change {change_bps:+d} bps")
margin_table = "\n".join(lines)
print(margin_table)
North America: 4.2% in 2023, 6.4% in 2024, change +223 bps
International: -2.0% in 2023, 2.7% in 2024, change +468 bps
AWS: 27.1% in 2023, 37.0% in 2024, change +990 bps
Consolidated: 6.4% in 2023, 10.8% in 2024, change +434 bps
Here, we can see the exact changes: +223, +468, +990 and +434 bps. These come from the unrounded margins, so there is no early rounding.
Step 3: How Do We Write the Note?
The third prompt gets only the finished margin table. We tell the model to use only these numbers and to compute nothing new:
note = chat(
f"Operating margins for Amazon, computed from its 2024 10-K:\n{margin_table}\n\n"
"Write a three-sentence note for an investment committee about which segment drove "
"the change in Amazon's margin. Use only the numbers above. Do not compute new numbers."
)
print(note)
AWS was the primary driver of Amazon's margin expansion, contributing a 990 basis point increase to its operating margin. This was followed by the International segment, which improved by 468 basis points, and the North America segment, which added 223 basis points. Collectively, these segment improvements resulted in a 434 basis point increase in the consolidated operating margin.
Here, we can see every number taken straight from our table. The model only had to write. One sentence still needs a human eye, though. "Collectively, these segment improvements resulted in" the total change is loose wording, because segment margins do not add up to the consolidated margin. Our numbers are right, but a chain does not make the reasoning perfect.
Step 4: How Does a Checker Prompt Work?
A checker is a separate prompt that reviews a finished text against a trusted reference. It does not write anything new. It only reports what does not match. Let's run it on both answers:
def check(text):
return chat(
f"Reference table:\n{margin_table}\n\nText to check:\n{text}\n\n"
"Check every margin and every basis point change in the text against the reference table. "
"List each one that does not match, with the correct value. "
"If all of them match, reply exactly: ALL NUMBERS MATCH"
)
print("one big prompt ->", check(one_prompt))
print()
print("chained note ->", check(note))
one big prompt -> Based on the reference table provided, here are the discrepancies found in the text:
1. **North America Change (bps):**
* Text Value: +220 bps
* Reference Value: +223 bps
* **Correction:** The change should be **+223 bps**.
2. **International Change (bps):**
* Text Value: +470 bps
* Reference Value: +468 bps
* **Correction:** The change should be **+468 bps**.
3. **Consolidated Change (bps):**
* Text Value: +440 bps
* Reference Value: +434 bps
* **Correction:** The change should be **+434 bps**.
*(Note: The AWS change of +990 bps matches the reference table. All individual margin percentages for 2023 and 2024 match the reference table.)*
chained note -> ALL NUMBERS MATCH
Here, we can see the checker catch exactly the three wrong changes in the one-prompt answer, with the right values next to them. The chained note passes with "ALL NUMBERS MATCH".
The checker works well here because its job is narrow. It compares numbers against a table, and it has the true numbers in front of it. A checker with no reference would only be guessing.
What Is Self-Refinement?
Self-refinement means sending the checker's notes back to the model and asking it to fix its own answer. In simple words, it is a draft, review and rewrite loop. Let's fix the one-prompt answer this way:
feedback = check(one_prompt)
revised = chat(
f"Here is an analysis:\n{one_prompt}\n\n"
f"A reviewer found these errors:\n{feedback}\n\n"
"Rewrite the analysis with these errors fixed. Change nothing else."
)
for line in revised.splitlines():
if "bps" in line and line.startswith("|"):
print(line)
print()
print("check again ->", check(revised))
| Segment | 2023 Margin | 2024 Margin | Change (bps) |
| North America | 4.2% | 6.4% | +223 bps |
| International | -2.0% | 2.7% | +468 bps |
| AWS | 27.1% | 37.0% | +990 bps |
| **Consolidated** | **6.4%** | **10.8%** | **+434 bps** |
check again -> ALL NUMBERS MATCH
Here, we can see the change column fixed: +223, +468, +990 and +434 bps. The second check passes.
Self-refinement has one weak spot. The checker is the same model as the writer, so it can share the writer's blind spots. That is why our checker compares against a table computed by Python, not against its own math. We explain this idea in depth in Loop Engineering: the maker cannot grade its own work without an outside reference.

The loop ends when the checker reports that every number matches.
When Is a Chain Worth the Extra Calls?
A chain costs more calls. Our one-prompt answer took one call. The chain took two model calls plus a little Python, or three with the checker. The refinement loop took four calls: the first answer, a check, a rewrite and a second check. Let me tabulate the trade-off:
| Approach | Model calls | Basis point changes right |
|---|---|---|
| One big prompt | 1 | 1 of 4 |
| Chain: extract, compute, write | 2 | 4 of 4 |
| One prompt, then check and rewrite | 4 | 4 of 4 |

Rounding before subtracting moved three of the four changes.
For a quick chat answer, one prompt is fine. When the numbers go into a report, a chain is worth it. The rule of thumb is simple: if a step can be done by plain code, do it in code.
Recap
This is how prompt chaining works. We started with one big prompt that read the 10-K tables, did the math and wrote a note, and three of its four basis point changes were wrong. Then we split the job: the model extracted the numbers, Python computed the margins, and the model wrote the note from our table. Finally, a checker prompt compared each answer with the Python table, and self-refinement fixed the one-prompt answer.
In the next lesson, we will look at prompt injection, and how text hidden inside a document can take over a finance assistant.