Few-Shot Prompting: Teaching an LLM With Examples in the Prompt

Zero-shot vs one-shot vs few-shot on 12 real 10-K sentences with Qwen 3.8: how examples fix amounts and metric names, and how one wrong example gets copied.

Oct 8, 202612 min readFollow

Topics You Will Master

What zero-shot, one-shot and few-shot prompting mean
Why an instruction alone left 4 of 12 amounts wrong
How one worked example fixed every answer
Why a single wrong example gets copied into new answers

We gave Qwen 3.8 a clear instruction: turn each 10-K sentence into one line, with the amount in millions of dollars. For "$2.3 billion", it answered 2. The instruction said millions, and the model still dropped the billions. But here's the real mystery: one worked example, added above the question, fixed every answer.

Few-shot prompting means putting a few solved examples in the prompt before the real question. In simple words, we show the model what a good answer looks like instead of only describing it. In this lesson, we will learn zero-shot, one-shot and few-shot prompting, and we will measure each one on 12 real sentences from the Amazon and Alphabet 10-K filings.

Zero-Shot vs One-Shot Prompt on the Same 10-K Sentence

The few-shot prompt is the same instruction plus solved examples.

Bestseller

Master Langchain v1 and Ollama - Chatbot, RAG and AI Agents

Master Langchain v1, Local LLM Projects, Ollama, DeepSeek, LLAMA 3.2, Complete Integration Guide.

95% off$199.99
$9.99Enroll now →30 day refund, lifetime access

Setup

We keep the setup from the earlier lessons: Qwen 3.8 27B in Ollama, with thinking off and the same answer on every run. Let's see the code as below:

PYTHON
import ollama

MODEL = "qwen3.8:27b"

def chat(messages):
    response = ollama.chat(
        model=MODEL,
        messages=messages,
        think=False,
        options={"temperature": 0, "seed": 7},
    )
    return response.message.content.strip()

What Is Few-Shot Prompting?

There are three common names, and they only count the examples:

  • Zero-shot: the prompt has the instruction and the question. No examples.
  • One-shot: the prompt adds one solved example.
  • Few-shot: the prompt adds a few solved examples, often two to five.

The model's weights do not change at all. It does not train on our examples. It reads them as part of the prompt, and its answer follows the pattern it sees. This is called in-context learning. In our lesson on the attention formula, we saw how every new token can look back at all earlier tokens. Here, the earlier tokens include our worked examples.

Advertisement

Our Test: 12 Sentences From Two 10-K Filings

Our task is a common one in finance work. We read a sentence about a change and turn it into a clean record: which metric moved, which way, and by how much. A clean record can go straight into a spreadsheet.

We took 12 sentences from the 2024 10-K filings of Amazon and Alphabet. Some amounts are in millions and some are in billions. For each sentence, we wrote the correct answer by hand. Let's see the code as below:

PYTHON
instruction = (
    "Read the sentence from a 10-K filing. Return one line in the form: "
    "metric | direction | amount. Direction is up or down. "
    "Amount is in millions of dollars, as a plain whole number."
)

test_set = [
    ("Changes in foreign exchange rates reduced net sales by $2.3 billion in 2024.",
     "net sales | down | 2300"),
    ("Changes in foreign exchange rates reduced North America net sales by $462 million in 2024.",
     "North America net sales | down | 462"),
    ("Changes in foreign exchange rates reduced International net sales by $1.8 billion in 2024.",
     "International net sales | down | 1800"),
    ("Changes in foreign exchange rates reduced cost of sales by $1.7 billion in 2024.",
     "cost of sales | down | 1700"),
    ("Changes in foreign exchange rates reduced technology and infrastructure costs by $244 million in 2024.",
     "technology and infrastructure costs | down | 244"),
    ("Changes in foreign exchange rates reduced sales and marketing costs by $263 million in 2024.",
     "sales and marketing costs | down | 263"),
    ("Changes in foreign exchange rates positively impacted operating income by $240 million in 2024.",
     "operating income | up | 240"),
    ("Google subscriptions, platforms, and devices revenues increased $5.7 billion from 2023 to 2024.",
     "Google subscriptions, platforms, and devices revenues | up | 5700"),
    ("Google Cloud revenues increased $10.1 billion from 2023 to 2024 primarily driven by growth "
     "in Google Cloud Platform largely from infrastructure services.",
     "Google Cloud revenues | up | 10100"),
    ("Google Services operating income increased $25.4 billion from 2023 to 2024.",
     "Google Services operating income | up | 25400"),
    ("Google Cloud operating income increased $4.4 billion from 2023 to 2024.",
     "Google Cloud operating income | up | 4400"),
    ("Other Bets operating loss increased $349 million from 2023 to 2024.",
     "Other Bets operating loss | up | 349"),
]

def score(answers):
    fields = {"metric": 0, "direction": 0, "amount": 0}
    for answer, (_, expected) in zip(answers, test_set):
        got, want = answer.split(" | "), expected.split(" | ")
        if len(got) != 3:
            continue
        for name, g, w in zip(fields, got, want):
            fields[name] += g == w
    return "  ".join(f"{name} {n}/{len(test_set)}" for name, n in fields.items())

The score function splits each answer at the | marks. Then it checks the three fields one by one against our answer. An answer with the wrong shape scores zero on every field.

How Well Does the Instruction Alone Work?

First, the zero-shot prompt: only the instruction and the sentence. Let's test it:

PYTHON
def zero_shot(sentence):
    return chat([{"role": "user", "content": f"{instruction}\n\nSentence: {sentence}\nAnswer:"}])

answers = [zero_shot(sentence) for sentence, _ in test_set]
for answer in answers:
    print(answer)
print(score(answers))
OUTPUT
net sales | down | 2
net sales | down | 462
International net sales | down | 1800
cost of sales | down | 1700
technology and infrastructure costs | down | 244
sales and marketing costs | down | 263
operating income | up | 240
Google subscriptions, platforms, and devices revenues | up | 5
Google Cloud revenues | up | 10100
operating income | up | 25
operating income | up | 4
Other Bets operating loss | up | 349
metric 9/12  direction 12/12  amount 8/12

Here, we can see two kinds of mistakes:

  • Amounts: four billions were cut down instead of converted. "$2.3 billion" became 2, and "$25.4 billion" became 25. It looks like the model read "a plain whole number" as "drop the decimals".
  • Metric names: "North America net sales" became just "net sales". "Google Cloud operating income" became just "operating income". The record no longer says which segment moved.

The directions were all right. So the model understood the task. What it did not get was the exact shape we wanted, and words alone left room for doubt.

How Do We Add Examples to the Prompt?

A few-shot prompt is the same instruction, followed by solved examples in the same form as the real question. Then comes the new sentence with an empty answer. The model fills in the blank.

We pick three examples that are not in the test set. Two are in millions and one is in billions. Two are decreases and one is an increase. Let's print the one-shot prompt for a single sentence:

PYTHON
examples = [
    ("Changes in foreign exchange rates reduced fulfillment costs by $223 million in 2024.",
     "fulfillment costs | down | 223"),
    ("YouTube ads revenues increased $4.6 billion from 2023 to 2024.",
     "YouTube ads revenues | up | 4600"),
    ("Google Network revenues decreased $953 million from 2023 to 2024, primarily driven by "
     "a decrease in Google Ad Manager and AdMob revenues.",
     "Google Network revenues | down | 953"),
]

def few_shot_prompt(sentence, shots):
    worked = "".join(f"Sentence: {s}\nAnswer: {a}\n\n" for s, a in shots)
    return f"{instruction}\n\n{worked}Sentence: {sentence}\nAnswer:"

print(few_shot_prompt(test_set[1][0], examples[:1]))
OUTPUT
Read the sentence from a 10-K filing. Return one line in the form: metric | direction | amount. Direction is up or down. Amount is in millions of dollars, as a plain whole number.

Sentence: Changes in foreign exchange rates reduced fulfillment costs by $223 million in 2024.
Answer: fulfillment costs | down | 223

Sentence: Changes in foreign exchange rates reduced North America net sales by $462 million in 2024.
Answer:

Here, we can see the whole idea. The example ends with "Answer:" and a finished line. The real question ends with "Answer:" and nothing after it. So the most natural next tokens are a line in the same form.

Advertisement

How Much Do the Examples Help?

Now, we run the full test with one example and with three examples:

PYTHON
def few_shot(shots):
    return [chat([{"role": "user", "content": few_shot_prompt(s, shots)}]) for s, _ in test_set]

for k in [1, 3]:
    print(f"{k}-shot: {score(few_shot(examples[:k]))}")
OUTPUT
1-shot: metric 12/12  direction 12/12  amount 12/12
3-shot: metric 12/12  direction 12/12  amount 12/12

Here, we can see every field right in all 12 sentences, even with a single example. That one example was in millions, yet the billions were converted correctly too.

Let me tabulate the results for your better understanding:

Prompt Metric right Direction right Amount right
Zero-shot (instruction only) 9 of 12 12 of 12 8 of 12
One-shot 12 of 12 12 of 12 12 of 12
Three-shot 12 of 12 12 of 12 12 of 12

12 Sentences From Two 10-K Filings: Fields Right With 0, 1 and 3 Examples

The instruction got the idea across. The example fixed the details.

What Happens When an Example Is Wrong?

Examples are strong, so a wrong example is dangerous. Let's break one on purpose. In the YouTube example, we write the amount as 4.6 instead of 4600, like someone who forgot to convert billions.

PYTHON
bad_examples = [
    examples[0],
    ("YouTube ads revenues increased $4.6 billion from 2023 to 2024.", "YouTube ads revenues | up | 4.6"),
    examples[2],
]

answers = few_shot(bad_examples)
for answer in answers[7:]:
    print(answer)
print(score(answers))
OUTPUT
Google subscriptions, platforms, and devices revenues | up | 5.7
Google Cloud revenues | up | 10.1
Google Services operating income | up | 25400
Google Cloud operating income | up | 4.4
Other Bets operating loss | up | 349
metric 12/12  direction 12/12  amount 9/12

Here, we can see the mistake copied into new answers. Four Google sentences are worded like the broken example: "increased $X billion from 2023 to 2024". Three of them now say 5.7, 10.1 and 4.4. Only Google Services kept the right amount, 25400. The Amazon sentences say "reduced ... by" instead, and every one of them kept the correct amount.

So, the model does not only learn the rule from an example. It leans hard on the example that looks most like the new input. One bad example can quietly spoil every answer that resembles it.

One Wrong Example Copied: 4.6 in the Prompt Becomes 5.7, 10.1 and 4.4

The model copied the closest-looking example, mistakes included.

Advertisement

Can the Examples Be Chat Turns Instead?

In System Prompts and Roles, we saw that a chat is a list of messages. We can also give examples that way. Each example becomes a user message with the sentence and an assistant message with the answer. The instruction goes into the system message.

PYTHON
def few_shot_turns(sentence, shots):
    messages = [{"role": "system", "content": instruction}]
    for s, a in shots:
        messages.append({"role": "user", "content": s})
        messages.append({"role": "assistant", "content": a})
    messages.append({"role": "user", "content": sentence})
    return chat(messages)

answers = [few_shot_turns(s, examples) for s, _ in test_set]
print(f"3-shot as chat turns: {score(answers)}")
OUTPUT
3-shot as chat turns: metric 12/12  direction 12/12  amount 12/12

Here, we can see the same perfect score. To the model, our examples look like past turns where it already answered correctly. Both styles work. Chat turns are handy when the app already sends a list of messages. A single block of text is easier to read and store.

How Should We Pick Good Examples?

Our tests point to a few simple habits:

  • Make the examples look like the real inputs. The model copies the closest one, so it should be a good one.
  • Cover the tricky cases: millions and billions, up and down.
  • Check every example by hand. A wrong example spreads, as we saw with 4.6.
  • Keep the set small. Every example is resent with every question, so each one costs tokens on every call.

When the inputs vary a lot, we can also pick the examples for each question on the fly, choosing the ones that look most like it. We cover that idea in Dynamic Few-Shot Learning with Scikit-LLM.

Recap

This is how few-shot prompting works. We started with an instruction alone, and Qwen 3.8 got the idea but cut four billions down and dropped segment names. Then we added one solved example, and all 12 answers came out right. We broke one example on purpose and watched the mistake copy into three of the four sentences that looked like it. Finally, we sent the same examples as chat turns and got the same result.

In the next lesson, we will look at prompt chaining, and how to split one big finance task into small prompts that check each other.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments