What Is Scikit-LLM? Build Your First Classifier

Learn what Scikit-LLM does, set up a zero-shot text classifier in Python, and check support-ticket labels against a simple rule baseline.

Sep 28, 202626 min readFollow

Topics You Will Master

Explain what Scikit-LLM does between your Python code and a language model.
Understand zero-shot classification and what fit() means in this example.
Set up a classifier with three support categories on a local Qwen 3.8 27B model.
Compare its role with a small keyword baseline and check results honestly.

Scikit-LLM is a Python library that lets us use a language model through the familiar scikit-learn methods fit() and predict(). In simple words, we hand it some texts and a list of labels, and it asks the model to pick a label for each text.

In this blog, we will learn how that works on Day 1 of a 15-day series. We will sort course-support tickets into billing, login and cancellation, using Qwen 3.8 27B running on our own machine through Ollama. We will also build a tiny keyword rule and find out what our numbers really tell us.

The full path has two stages. First we set the allowed labels. Then we send a ticket and check the label that comes back.

How Scikit-LLM Classifies a Ticket

Start with the Job We Want to Do

Our support inbox has three categories: billing, login and cancellation. For this first lesson, every message has one main issue. The program should return one category for each message.

This task is called text classification: choosing a category for a piece of text. The category is also called a label. A classifier is the part of our program that makes that choice.

Think of sorting letters into named trays. The message is the letter, and billing is one tray. Before sorting anything, we must agree on what belongs in each tray.

Let me tabulate our three trays.

Label Meaning in our teaching example
billing A charge, payment or invoice issue.
login Trouble entering an account or resetting access.
cancellation A request to stop a plan or its renewal.

These meanings are our task policy. The library does not supply them, and another business could draw the lines differently.

A message can also carry two issues, such as "Cancel my plan and refund the duplicate charge". We will handle more than one label on Day 6. Today, we use simple messages so we can follow the whole process.

Advertisement

What Scikit-LLM Actually Does

There are three pieces to keep apart. Our Python code holds the texts and the allowed labels. Scikit-LLM turns them into a request and reads the reply. The language model writes that reply.

An LLM, or large language model, is a model trained on a lot of text. It can write text in reply to an input. Here, we use that skill to ask for a category.

In this series, the model runs on our own computer inside Ollama. Ollama is a free app that downloads models and serves them at a local web address. Scikit-LLM sends each request to that address, much like it would send one to a hosted service.

The route the library uses to reach a model is called a backend. We will look at backends closely on Day 2, including a free hosted option on OpenRouter.

Here, the request starts in our Python code, passes through Scikit-LLM, and reaches Ollama on the same computer.

Your Code, the Library and the Model

Installing scikit-llm installs Python code only. It does not put a language model on our machine. We install Ollama and download the model as separate steps.

Why Use Scikit-LLM, and When Does It Fit?

Scikit-LLM is worth trying when we apply the same language task to many texts. Support categories are one example. Summaries and turning text into numbers are others we will reach later.

An estimator is the Python object that offers methods such as fit() and predict(). A method is simply a function attached to an object. This shape keeps a small task easy to inspect: set it up, pass in texts, then check what comes back.

The interface is handy, but it is still a project choice. Let me tabulate when each approach makes sense.

Situation A sensible first approach
A fixed code such as PAYMENT_FAILED determines the category A direct rule or lookup.
We have many labeled messages and stable categories A local text-classification baseline.
We have clear categories but few labeled examples Try a zero-shot LLM classifier and evaluate it.
Our request needs provider-specific tools or response controls Check whether using that provider's API directly is a better fit.

A small classic model may be faster on our workload. An LLM may cope better with new wording, but we must measure that. The package name is not proof of accuracy.

Advertisement

What Zero-Shot Means Here

Zero-shot classification asks the model to choose a label without any labeled examples of our task. A prompt is the instruction and text we send to the model.

Our request needs three ideas: choose a support category, choose from our three labels, and read the new message. That is enough to ask for a prediction. It is not enough to promise a correct one. Zero-shot classification

The model was trained long before our program uses it. "Zero-shot" only describes the examples we send with the request. It says nothing about how the model learned language.

Here is a short teaching version of the request. It is not the library's exact prompt.

Choose one support category: billing, login or cancellation.

Message: I was charged twice for my course subscription.

A person would pick billing, so that is our reference answer. It is our answer key, not the model's reply.

A zero-shot request holds the task, the allowed labels and the new ticket, with no labeled examples.

A Task, Three Labels and One New Ticket

A Short scikit-learn Refresher

In scikit-learn, fit(X, y) and predict(X) are the two main methods. X holds the inputs and y holds the answers for those inputs. For a normal classifier, fit() learns from those pairs.

Our zero-shot classifier uses the same method names, but its fit() only stores the candidate labels. It does not change the model's weights, the numbers the model learned during training. The model is only called during predict().

Other estimators in this series do different work in fit(). We will trace each one when we meet it.

Advertisement

Install an Isolated Copy

We will use Scikit-LLM 1.4.3, the release we checked this lesson's code with. Pinning a version makes the example easy to repeat. The package needs Python 3.9 or later, and we ran our checks with Python 3.14.4. Package release and requirements

A virtual environment keeps this lesson's packages in one folder, away from other projects. In a new lesson folder, run this Windows command to create one:

POWERSHELL
python -m venv .venv

On Linux/macOS: use python3 -m venv .venv.

Next, we install into that environment. -m pip runs the installer that belongs to that Python, and ==1.4.3 picks the version we checked.

POWERSHELL
.\.venv\Scripts\python.exe -m pip install "scikit-llm==1.4.3"

On Linux/macOS: use .venv/bin/python -m pip install "scikit-llm==1.4.3".

Now we open that environment's Python. We run all the Python blocks below in this one session, in order. In a notebook, we pick this environment as the kernel.

POWERSHELL
.\.venv\Scripts\python.exe

On Linux/macOS: use .venv/bin/python.

Our first import reads the installed package version. It confirms that we are in the environment we meant to use.

PYTHON
from importlib.metadata import version

print(version("scikit-llm"))
OUTPUT
1.4.3

Here, we can see version 1.4.3, so the right package is active.

Install Ollama and the Qwen 3.8 27B Model

Ollama gives us the model itself. We download it from ollama.com and install it like any other app. Once installed, it keeps a small server running in the background at http://localhost:11434.

Next, we download the model in a new terminal. qwen3.8:27b is the model name and tag in Ollama's library.

POWERSHELL
ollama pull qwen3.8:27b

We also pull a small embedding model now. We will not need it until Day 7, where it turns text into numbers.

POWERSHELL
ollama pull nomic-embed-text

Note

qwen3.8:27b is a 16.5 GB download, and nomic-embed-text is 262 MB. The model runs on our own hardware, so there is no API bill and no account key. How fast it answers depends on the machine's memory and GPU.

Advertisement

Our Six-Ticket Teaching Set

We will work with six messages written for this lesson. They are made-up examples, not customer records. The first three contain words our small rule will spot. The other three say similar things in different words.

X is a Python list, and each string is one message. In the figures, we call them T1 to T6 in list order. The order matters, because we compare each prediction with the answer at the same position.

PYTHON
X = [
    "I was charged twice for my course subscription.",
    "My account password reset link has expired.",
    "Please cancel my course subscription.",
    "The same course payment appears twice on my bank statement.",
    "I cannot get into my account after changing phones.",
    "Please stop renewing my plan next month.",
]

We write the human answers in a separate list. y_reference is only for checking. We never send these answers to the zero-shot classifier.

PYTHON
y_reference = [
    "billing", "login", "cancellation",
    "billing", "login", "cancellation",
]

The candidate labels are the categories a prediction may choose. This list tells the model its choices, not which answer belongs to which message.

PYTHON
candidate_labels = ["billing", "login", "cancellation"]

The candidate labels go into the classifier, while each ticket's human answer stays aside for the check.

Text Goes In, Reference Labels Stay Aside

So, the three lists have three jobs. X has six things to read, y_reference has six answers to check, and candidate_labels has three allowed choices.

A Tiny Rule Baseline

A baseline is a simple method we compare a bigger one against. Ours looks for three exact words. We lowercase each message first, so Charged and charged are treated the same.

The if checks run in order. The first match returns a category and ends the function. If nothing matches, the rule returns needs_review, which means a person should look.

PYTHON
def rule_label(text):
    text = text.lower()
    if "charged" in text:
        return "billing"
    if "password" in text:
        return "login"
    if "cancel" in text:
        return "cancellation"
    return "needs_review"

This rule is small on purpose. It only checks for pieces of text, so one failure is easy to see.

Now we call the rule once per ticket and collect the results in a list:

PYTHON
rule_predictions = [rule_label(ticket) for ticket in X]
print(rule_predictions)
OUTPUT
['billing', 'login', 'cancellation', 'needs_review', 'needs_review', 'needs_review']

Here, we can see that the first three tickets get a label and the last three go to review. The fourth ticket is about a payment showing up twice. It has none of our three words, so the rule cannot place it.

These two tickets show the rule catching the word charged but missing the same problem in other words.

Same Billing Problem, Different Words

That gives us a clear case to test with a language model. It does not prove that a model will fix it.

Advertisement

Set the Candidate Labels with Scikit-LLM

Before we create the classifier, we tell Scikit-LLM where the model lives. We will use this same setup block in every lesson of the series:

PYTHON
import os

from skllm.config import SKLLMConfig

# Default: Qwen 3.8 27B running on our own machine through Ollama
SKLLMConfig.set_gpt_url("http://localhost:11434/v1")
SKLLMConfig.set_gpt_key("ollama")
MODEL = "custom_url::qwen3.8:27b"

# Option: the free Qwen 3.8 27B on OpenRouter (uncomment these lines)
# SKLLMConfig.set_gpt_url("https://openrouter.ai/api/v1")
# SKLLMConfig.set_gpt_key(os.environ["OPENROUTER_API_KEY"])
# MODEL = "custom_url::qwen/qwen3.8-27b:free"

The active lines point the library at our local Ollama server. "ollama" is only a placeholder key, since Ollama does not check it. The custom_url:: part of MODEL tells Scikit-LLM to send the request to that URL. The commented lines are the hosted option, which Day 2 explains line by line.

Now we import the classifier from its full module path and create it. Creating clf only stores settings. It does not call the model yet.

PYTHON
from skllm.models.gpt.classification.zero_shot import ZeroShotGPTClassifier

clf = ZeroShotGPTClassifier(model=MODEL, default_label="needs_review")

We set default_label on purpose. The library default is "Random", which picks a real label at random when a reply cannot be used. That would hide failures. A visible needs_review marker is much easier to spot. Classifier parameters

Next, we pass the allowed labels. None means we give no training texts to this zero-shot setup. The classes_ attribute shows the stored labels, and the trailing underscore is part of its name.

PYTHON
clf.fit(None, candidate_labels)
print(clf.classes_)
OUTPUT
['billing', 'login', 'cancellation']

Here, we can see our three labels stored in classes_. No model call has happened yet, so this step works even when Ollama is closed.

This fit call stores three candidate labels without changing the language model's weights.

fit Stores the Candidate Labels

What Happens During predict

For each ticket, the classifier builds a prompt from the stored labels and the ticket text. The prompt asks the model to answer in JSON with one key, label. The library reads that key and checks that the label is allowed. Then it collects one result per ticket into an array. Classifier implementation

Let's say the reply holds billing in its label field. Reading the reply gets us billing. The allowed-label check then confirms that billing is on our list.

These two checks have different jobs. Reading proves the program understood the reply's shape. The list check proves the label is allowed. Neither proves that billing is the right answer.

Our human reference gives that final check. If the model returned login for the double-charge ticket, it would be allowed and still wrong.

So, we first read the returned label, then check it is allowed, and finally compare it with the human reference.

Parsing and Checking a Model Response

Advertisement

Classify the Six Tickets with the Local Model

This step calls the model, so Ollama must be running with qwen3.8:27b downloaded. zip() pairs each ticket with its result in order, and the loop prints each pair.

PYTHON
predictions = clf.predict(X)
for ticket, label in zip(X, predictions):
    print(ticket)
    print("Predicted label:", label)

While it runs, Scikit-LLM shows a progress bar for the six tickets. Then each ticket prints with its label. Every label is one of billing, login, cancellation or needs_review. predictions is a NumPy array of shape (6,), one result per ticket.

Note

Qwen 3.8 is a thinking model: it reasons before it answers. Ollama keeps that reasoning out of the text Scikit-LLM reads, so the label stays clean. The thinking does make each call slower, so six tickets can take a while on a modest machine.

Each result sits at the same position as its ticket, so we can compare them pair by pair.

Preserve the Input Order

Read Failures Before Trusting Labels

Suppose the model replies with payments, while our allowed label is billing. That label is not in classes_, so the library swaps it for our default_label, needs_review. The same happens to Billing with a capital B, because labels must match exactly.

A connection failure is different. If Ollama is not running, there is no reply to check at all. The library tries the call 3 times, waiting 1, 2 and 4 seconds after each failed try. Then it prints this line and raises a RuntimeError:

PLAINTEXT
Could not complete the operation after 3 retries: `APIConnectionError :: Connection error.`

So, an out-of-set label becomes needs_review, while a failed connection raises an error instead.

Make Failed Label Validation Visible

Warning

A failed call is slow. The OpenAI client that Scikit-LLM uses also retries inside each try. On our Windows PC, with Ollama closed, predict took 49.5 seconds before the RuntimeError appeared. If a prediction hangs, check that Ollama is running first.

An error like this says nothing about label quality. Also, a label can look fine and still be wrong without any fallback. We still need reference answers and a look at each mistake.

Advertisement

Check What We Actually Measured

Accuracy is the share of predictions that match the reference answers. For six tickets, we can count it by hand in code. zip() pairs predictions with answers, pred == ref gives True or False, and sum() counts the True values.

PYTHON
correct = sum(
    pred == ref
    for pred, ref in zip(rule_predictions, y_reference)
)
print(f"Exact matches: {correct}/{len(X)}")
print(f"Accuracy: {correct / len(X):.2f}")
OUTPUT
Exact matches: 3/6
Accuracy: 0.50

Here, we can see 3 matches out of 6, an accuracy of 0.50. A needs_review marker counts as a miss, because it is not the human answer. The :.2f format shows two digits after the decimal point.

We also count how often the rule asks for review. A system can avoid wrong answers by handing work to a person, but it may then leave much of the inbox unsolved.

PYTHON
reviews = rule_predictions.count("needs_review")
print(f"Review rate: {reviews / len(X):.2f}")
OUTPUT
Review rate: 0.50

Here, we can see a review rate of 0.50. The rule solves three of our six tickets and sends three to a person.

Three Resolved Tickets Out of Six

These numbers come from a set we designed: three clear keywords, then three paraphrases. It is a teaching device, not a real inbox, so 50% is not a real-world estimate.

We can score the model's labels with the same check. We only swap in predictions:

PYTHON
llm_correct = sum(
    pred == ref
    for pred, ref in zip(predictions, y_reference)
)
print(f"LLM exact matches: {llm_correct}/{len(X)}")

This prints the model's match count out of 6. Before trusting that number, we read every mismatch one by one. On Day 4, we compare methods properly on tickets kept aside for testing.

Practice with Two New Tickets

First, we write a new billing message that has none of the words charged, password or cancel. We decide its reference label before running anything. Then we send it through rule_label() and through clf.predict().

Next, we try a message that asks both to cancel and to dispute a charge. Our one-label policy now forces a choice. If two people disagree about the right answer, we fix the policy before blaming the model.

Conclusion

This is how Scikit-LLM works. We pointed it at Qwen 3.8 27B running locally in Ollama. We stored three labels with fit(), and predict() asked the model for one label per ticket. Along the way, a tiny keyword rule solved 3 of our 6 tickets and gave us a baseline to compare against.

  • Zero-shot classification sends the task and the allowed labels, with no labeled examples.
  • This classifier's fit() only stores labels; predict() makes the model calls.
  • default_label="needs_review" makes a bad reply visible, while a failed connection raises RuntimeError.
  • An allowed label can still be wrong, so reference answers and a look at each mistake matter.

Next steps:

We now have a classifier wired to a local model, a checked rule baseline, and a simple way to score both.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments