Evaluate Scikit-LLM Without Fooling Yourself

Compare Scikit-LLM predictions with a simple text classifier using fixed tickets and honest error counts.

Oct 7, 202615 min readFollow

Topics You Will Master

Keep example tickets away from final test tickets.
Build a simple local text classifier for comparison.
Count correct answers, review cases and failed calls separately.
Explain what a small teaching set cannot prove.

If a model gets one duplicate-charge ticket right, should we trust it with every support message? No. In this blog, we will build a small offline baseline and run the zero-shot classifier from Day 3 on the same held-out tickets.

We will find out how to split tickets fairly and how to count every answer in the right bucket. We will also see why a perfect 3 out of 3 can still be luck.

Decide What Counts as a Fair Test

A fair test needs three groups of tickets. This figure shows the path from raw conversations to those three groups.

How a Fair Evaluation Works for the Day 4 support-ticket example

The training group teaches a local classifier or supplies few-shot examples. The development group is where we try label names, prompts and model settings. The final test group stays closed until those choices are fixed. If we read its errors and then change the system, it has become development data.

Why do we split by conversation and not by message? Because one student may write three messages about the same double charge. If one message trains the system and another tests it, the test repeats words the system has already seen. So we give every message a conversation ID, and all messages with one ID stay in one group.

We divide the IDs before building a vectorizer or choosing prompt examples. A random seed helps us repeat the split. But a seed alone does not stop two messages from one conversation from landing on both sides. After the split, we check the label counts. A tiny final group may hold no cancellation ticket at all, and then we cannot judge cancellation.

The final group gets opened once, after every choice is made.

Keep the Final Test Closed for the Day 4 support-ticket example

Advertisement

Make a Tiny Baseline That Runs Offline

A baseline is a simple method we run before adding a language model. In simple words, it is the yardstick. If a cheap method already sorts common tickets well, the model calls must fix a problem we can measure.

Our baseline has two steps. TfidfVectorizer turns each ticket into a row of numbers, one number per known word. The word list is learned from the training text only.

Turn Words Into Numbers for the Day 4 support-ticket example

LogisticRegression then learns which number patterns go with each label. The scikit-learn Pipeline guide explains how the two steps are joined.

Learn From Labeled Rows for the Day 4 support-ticket example

We import the vectorizer, the classifier and the pipeline that joins them:

PYTHON
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline

Each training text has one human label at the same list position. Six messages are far too few for a real estimate, but they let us trace every step:

PYTHON
X_train = [
    "I was charged twice for the course.",
    "My invoice shows the wrong amount.",
    "My password reset link has expired.",
    "I cannot sign in to my account.",
    "Please cancel my course plan.",
    "Stop renewing my subscription.",
]
y_train = [
    "billing", "billing", "login", "login",
    "cancellation", "cancellation",
]

The vectorizer comes first because the classifier needs numbers, not sentences. We fit the whole pipeline on the training tickets only. max_iter=1000 simply gives the optimizer enough steps to finish on this data.

PYTHON
baseline = Pipeline([
    ("words", TfidfVectorizer()),
    ("classifier", LogisticRegression(max_iter=1000)),
])
baseline.fit(X_train, y_train)

Later, predict reuses the saved word list and the saved classifier. Nothing is fitted again on the check text.

Fit and Predict Stay Apart for the Day 4 support-ticket example

Count Answers on Held-Out Tickets

These three check tickets are not in X_train. We write their human answers before any method sees them. Every method we compare gets exactly these tickets.

Same Tickets for Every Method for the Day 4 support-ticket example

Let's see the code as below:

PYTHON
X_check = [
    "The same payment appears twice on my statement.",
    "I am locked out after changing phones.",
    "Please end my plan before it renews.",
]
y_check = ["billing", "login", "cancellation"]
baseline_pred = baseline.predict(X_check)

A single exact-match number is not enough. Each answer belongs in one of four buckets: correct, wrong_label, needs_review or request_failed. A review marker is not a wrong category, and a failed request is neither. We write the counting once as a function, so both methods use the same rules:

PYTHON
from collections import Counter


def outcome(expected, found):
    if found == "request_failed":
        return "request_failed"
    if found == "needs_review":
        return "needs_review"
    return "correct" if found == expected else "wrong_label"


def show_results(name, predictions):
    matches = sum(a == b for a, b in zip(predictions, y_check))
    print(f"{name} exact matches: {matches}/{len(y_check)}")
    for text, expected, found in zip(X_check, y_check, predictions):
        print(text, "| human:", expected, "|", name + ":", found)
    buckets = Counter(outcome(e, f) for e, f in zip(y_check, predictions))
    print(dict(buckets))


show_results("baseline", baseline_pred)
OUTPUT
baseline exact matches: 3/3
The same payment appears twice on my statement. | human: billing | baseline: billing
I am locked out after changing phones. | human: login | baseline: login
Please end my plan before it renews. | human: cancellation | baseline: cancellation
{'correct': 3}

Here, we can see that the baseline matched all three human labels. The bucket counts add up to three, one per ticket. If the total ever differs from the number of tickets, the report has a bug.

Warning

We wrote these tickets ourselves to teach the steps. So 3/3 describes three designed examples only. It is not an accuracy claim for a real support inbox.

Advertisement

Why the Login Answer Was Luck

A correct label can hide a weak guess. predict_proba shows how sure the baseline was about each label. We also count how many words of each ticket the vectorizer knows:

PYTHON
print(baseline.classes_)
print(baseline.predict_proba(X_check).round(3))
words = baseline.named_steps["words"]
print(words.transform(X_check).getnnz(axis=1))
OUTPUT
['billing' 'cancellation' 'login']
[[0.44  0.28  0.28 ]
 [0.33  0.331 0.338]
 [0.257 0.464 0.279]]
[3 0 3]

Here, we can see the problem in the second row. The login ticket got 0.330 for billing, 0.331 for cancellation and 0.338 for login. The last line explains why: the vectorizer knows zero words from "I am locked out after changing phones." Its row of numbers is all zeros.

So the baseline did not read the ticket at all. With no known words, only the classifier's intercepts decide. These are small starting scores it learned for each label. Login won by less than one percentage point, and a different training set could flip that guess. This is exactly why we read the per-ticket rows and not only the total.

This figure puts the whole 3/3 result together, including the weak login row.

What Three of Three Means for the Day 4 support-ticket example

Setup

The model run needs the same setup as every Scikit-LLM lesson. Day 2 explains each line.

PYTHON
import os

from skllm.config import SKLLMConfig

# Default: Qwen 3.8 27B running on our own machine through Ollama
SKLLMConfig.set_gpt_url("http://localhost:11434/v1")
SKLLMConfig.set_gpt_key("ollama")
MODEL = "custom_url::qwen3.8:27b"

# Option: the free Qwen 3.8 27B on OpenRouter (uncomment these lines)
# SKLLMConfig.set_gpt_url("https://openrouter.ai/api/v1")
# SKLLMConfig.set_gpt_key(os.environ["OPENROUTER_API_KEY"])
# MODEL = "custom_url::qwen/qwen3.8-27b:free"

Note

The local route has no API bill and needs no real key. The OpenRouter route needs OPENROUTER_API_KEY in the environment. Its free models allow 20 requests per minute and 50 per day, or 1,000 per day after buying at least $10 of credits.

Run Day 3's Classifier on the Same Tickets

Now we put Scikit-LLM on the same check set. We build the same zero-shot classifier as Day 3. It gets the three business labels, and needs_review catches any reply outside them.

PYTHON
from skllm.models.gpt.classification.zero_shot import ZeroShotGPTClassifier

labels = ["billing", "login", "cancellation"]
clf = ZeroShotGPTClassifier(model=MODEL, default_label="needs_review")
clf.fit(None, labels)

This figure follows one check ticket through both methods to the same comparison.

Trace One Checked Answer for the Day 4 support-ticket example

We send one ticket per call. If one request fails, Scikit-LLM raises a RuntimeError after its retries. We catch it and record request_failed for that ticket only. The other tickets still get their answers.

PYTHON
llm_pred = []
for text in X_check:
    try:
        llm_pred.append(clf.predict([text])[0])
    except RuntimeError:
        llm_pred.append("request_failed")

show_results("scikit-llm", llm_pred)

This prints the same report as the baseline. Each row shows one of billing, login, cancellation, needs_review or request_failed. The bucket counts again add up to three. We keep failed calls in the total, because dropping them would let a method look good by answering only its easy cases.

Warning

If Ollama is not running, each ticket waits through three retries before the RuntimeError. Scikit-LLM also prints a line that starts with Could not complete the operation after 3 retries. On our PC, that took about 50 seconds per ticket. Start Ollama before running this block.

Qwen 3.8 is a thinking model, so each call takes longer than the local baseline. The baseline predicts on our CPU with no network call. If speed matters, we time both on the same machine with the same tickets.

Here are the four outcomes one ticket can land in.

Count Three Failure Types for the Day 4 support-ticket example

For each disagreement, we ask a plain question. Did the model understand a paraphrase the baseline missed? Did the baseline guess from a word that will not hold in new tickets? The classifier comparison article asks when an LLM earns its extra work. Only results on our own tickets can answer that for us.

Advertisement

Count Precision and Recall by Hand

Accuracy counts all correct answers together, and that can hide a weak category. Let's say we have ten new tickets: six billing, two login and two cancellation. A system that says billing every time gets six right. That is 60 percent accuracy, yet it finds no login or cancellation ticket.

So we look at one category at a time. Suppose the classifier predicts billing for five tickets. Four are truly billing, and one is login. Two real billing tickets were sent somewhere else.

  • True positives: the four correct billing predictions.
  • False positive: the one login ticket sent to billing.
  • False negatives: the two billing tickets sent elsewhere.

Precision and Recall by Hand for the Day 4 support-ticket example

Precision asks, "When we said billing, how often was that right?" Here it is 4 / (4 + 1) = 0.80. Recall asks, "Of all billing tickets, how many did we find?" Here it is 4 / (4 + 2) = 0.67. The F1 score joins the two with 2 * precision * recall / (precision + recall), which is about 0.73. These are hand-worked numbers, not a result from our code.

A macro average scores billing, login and cancellation separately, then gives each one equal weight. So a missed cancellation class cannot hide behind many easy billing tickets. We still show each class row, because an average can hide which category needs work.

Advertisement

Decide What Error Hurts Most

Different mistakes cost different things. A wrong billing label sends a ticket to the wrong queue. A missed login problem leaves a student locked out of a lesson. A needs_review result costs a person some reading time, but it avoids a wrong automatic route.

Let's say two methods each label 80 of 100 tickets correctly. One sends 15 to people and routes 5 wrongly. The other sends 5 to people and routes 15 wrongly. The same correct count hides a very different workload. These numbers are a planning example, not results from our code. The support team decides which mistake costs more.

Let me tabulate the per-ticket record we keep for every run, for your better understanding:

Field Why we keep it
Conversation ID Finds the row again and proves the split was clean.
Human label The answer every method is checked against.
Baseline answer The yardstick result for the same ticket.
Scikit-LLM answer One of the three labels, needs_review or request_failed.
Model name MODEL from the setup block, so the run can be repeated.

A good final test also has different wording and longer threads, not only short phrases. We can then read error groups such as negation, mixed requests and rare wording. For practice, we can write a billing paraphrase with none of the words charged, invoice or payment. We write its human label first, then run both methods on it.

Conclusion

This is how a fair Scikit-LLM evaluation works. We split tickets by conversation and kept the final group closed. We built a TF-IDF baseline, then ran Day 3's zero-shot classifier on the same three check tickets with the same counting code. Along the way, the probabilities showed that the baseline's login answer was a near-equal guess.

  • Split by conversation before choosing examples, prompts or settings.
  • Count correct, wrong_label, needs_review and request_failed separately, and keep failed calls in the total.
  • Read per-ticket rows and probabilities, because a correct label can be luck.
  • Precision, recall and a macro average show weak categories that accuracy hides.
  • Three designed tickets teach the method but cannot estimate real inbox quality.

Next steps:

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments