Scikit-LLM Setup: Models, Keys and Backends

Set up a Scikit-LLM model route, keep its key out of code, and check where a support ticket is processed.

Oct 1, 202619 min readFollow

Topics You Will Master

Tell a Python package, a model and a backend apart.
Keep the OpenRouter key outside the script.
Switch between local Ollama and hosted OpenRouter on purpose.
Check a connection before judging the model's answers.

On Day 1, a short setup block sent our support tickets to Qwen 3.8 27B on our own machine. Those few lines decide where every ticket goes, which key is used, and which model reads it.

In this blog, we will learn what each line of that block does. We will also switch to the free Qwen 3.8 27B on OpenRouter, switch back, and see exactly what a failed call looks like.

Here is the whole route for one ticket, with the local default and the hosted option.

How the Model Route Works for the Day 2 support-ticket example

Package, Model and Backend

Three words get mixed up in every setup. Scikit-LLM is a Python package: it gives us classes such as ZeroShotGPTClassifier. A model is the trained system that reads the ticket and writes a reply. A backend is the route Scikit-LLM uses to reach that model.

Think of posting a letter. Our script writes the letter, the backend is the delivery route, and the model is the reader at the other end. If the route fails, we have learned nothing about how well the reader understands us.

The Scikit-LLM backend guide groups routes by the request format they use. The gpt family covers any service with an OpenAI-style API. So the name ZeroShotGPTClassifier tells us the request format, not who made the model. Ollama and OpenRouter both accept that format, which is why the same class works with both.

Here, the package runs in our Python process, the route carries the request, and the model weights sit wherever the model runs.

Package, Route and Weights for the Day 2 support-ticket example

Start from the Day 1 Environment

We keep using the virtual environment from Day 1. We open its Python by its full path, so we cannot pick up another Python by mistake:

POWERSHELL
.\.venv\Scripts\python.exe

Now we check the package version inside that session. version() reads the installed package details:

PYTHON
from importlib.metadata import version

print(version("scikit-llm"))
OUTPUT
1.4.3

Here, we can see 1.4.3, the version this series is checked against. If we see a different number, or an import error, we are in the wrong Python.

Advertisement

The Default Route: Ollama on Our Machine

This is the setup block every lesson in the series starts with. The Ollama lines are active, and the OpenRouter lines wait behind # signs:

PYTHON
import os

from skllm.config import SKLLMConfig

# Default: Qwen 3.8 27B running on our own machine through Ollama
SKLLMConfig.set_gpt_url("http://localhost:11434/v1")
SKLLMConfig.set_gpt_key("ollama")
MODEL = "custom_url::qwen3.8:27b"

# Option: the free Qwen 3.8 27B on OpenRouter (uncomment these lines)
# SKLLMConfig.set_gpt_url("https://openrouter.ai/api/v1")
# SKLLMConfig.set_gpt_key(os.environ["OPENROUTER_API_KEY"])
# MODEL = "custom_url::qwen/qwen3.8-27b:free"

Let's go through the active lines one by one:

  • set_gpt_url("http://localhost:11434/v1") sets the address for every request. localhost means this computer, 11434 is the port Ollama listens on, and /v1 is its OpenAI-style entrance.
  • set_gpt_key("ollama") gives the client a key. Without any key, Scikit-LLM stops with RuntimeError: OpenAI key was not found, even for a local model. Ollama ignores the value, so "ollama" is only a placeholder, not a secret.
  • MODEL = "custom_url::qwen3.8:27b" has two parts. custom_url:: tells Scikit-LLM to send the request to the URL we set. qwen3.8:27b is the model name Ollama knows, and it must match what we pulled.
  • import os is only used by the OpenRouter lines. It stays so that uncommenting them just works.

We can confirm the address without calling the model. get_gpt_url() reads back the stored value:

PYTHON
print(SKLLMConfig.get_gpt_url())
OUTPUT
http://localhost:11434/v1

Here, we can see the Ollama address, so the next request will go to our own machine.

The local route keeps everything on this computer: the ticket, the Ollama service and the downloaded model weights.

Local Ollama Route for the Day 2 support-ticket example

The model name is the part that most often goes wrong. The text after custom_url:: must match a model that Ollama has downloaded, letter for letter. The command ollama list shows the models on our machine.

Check the Model ID for the Day 2 support-ticket example

What does the request carry? For a custom_url:: model, Scikit-LLM sends only three things: the messages, the model name and temperature set to 0.0. A temperature of 0.0 asks the model for its most likely words instead of varied ones. There is no special JSON switch. Instead, the prompt itself asks the model to reply in JSON.

Note

Qwen 3.8 is a thinking model: it reasons before it answers. Ollama puts that reasoning in a separate reasoning field of its reply. Scikit-LLM only reads the content field, which holds the final answer, so the label stays clean. The thinking still takes time, so each call is slower than a plain answer would be.

The model itself is the 16.5 GB download from Day 1. It runs on our hardware, so there is no bill per call. The cost is time and memory: a slower machine gives slower answers.

Advertisement

Send One Ticket Through the Route

Now we create the classifier with MODEL and store our three labels. default_label names a marker for a reply that fails the label check. It is not a fourth support category.

PYTHON
from skllm.models.gpt.classification.zero_shot import ZeroShotGPTClassifier

clf = ZeroShotGPTClassifier(model=MODEL, default_label="needs_review")
labels = ["billing", "login", "cancellation"]
clf.fit(None, labels)

For this zero-shot class, fit() stores the choices. It does not train the model.

Store Labels Without Training for the Day 2 support-ticket example

Next, we send one made-up ticket. This is the first line that actually talks to Ollama:

PYTHON
ticket = ["I was charged twice for my course subscription."]
prediction = clf.predict(ticket)
print(prediction)

This prints a NumPy array with one label in it. That label is billing, login, cancellation or needs_review. A reply proves the route works. It does not prove the label is right, so we still compare it with the human answer, billing.

One test ticket checks the route first, and only then do we ask whether its label is right.

Test One Ticket in Two Steps for the Day 2 support-ticket example

The Option: Free Qwen 3.8 27B on OpenRouter

Not every machine can hold a 16.5 GB model. So, here comes OpenRouter to the rescue. OpenRouter is a hosted service that serves many models through one OpenAI-style API. Its model qwen/qwen3.8-27b:free is the same Qwen 3.8 27B, listed at a price of 0 with a 262,144-token context.

Free models have limits. OpenRouter allows 20 requests per minute and 50 requests per day on free models. After buying at least $10 of credits, the daily limit rises to 1,000. Each ticket is one request, so our six-ticket set uses 6 of the 50. OpenRouter limits

Unlike Ollama, OpenRouter needs a real key. We create one in the OpenRouter key settings after signing up. Then we save it as an environment variable, a named value that programs can read from the system. In PowerShell:

POWERSHELL
setx OPENROUTER_API_KEY "PASTE_THE_KEY_HERE"

setx saves the value for new terminals only. So we close the current terminal, open a new one, and start .\.venv\Scripts\python.exe again.

We can check that Python sees the key without showing it. This line prints True or False, never the key itself:

PYTHON
print("OPENROUTER_API_KEY" in os.environ)

If it prints False, the terminal was opened before setx, or the name is misspelled. In that case, the setup line os.environ["OPENROUTER_API_KEY"] stops with KeyError: 'OPENROUTER_API_KEY' before any request leaves our machine. That is the clear failure we want.

Stop Before a Missing Key for the Day 2 support-ticket example

Caution

Never print the OpenRouter key, paste it into a notebook cell, or commit it to Git. Anyone with the key can spend its limits and any credits on the account. If a key leaks, delete it in the OpenRouter settings and create a new one.

To switch, we add # in front of the three Ollama lines in the setup block and remove it from the three OpenRouter lines. Everything else stays the same, because every estimator uses model=MODEL. Calling set_gpt_url() and set_gpt_key() again simply replaces the old values.

There is one real difference to keep in mind. On OpenRouter, the ticket text leaves our computer and goes to a hosted service. With Ollama, it stays on our machine. Our lesson tickets are made up, but real support tickets need that decision first.

The same ticket can take either route, and the reply still needs a correctness check.

Hosted and Local Routes for the Day 2 support-ticket example

Advertisement

Switch Back with reset_gpt_url

SKLLMConfig.reset_gpt_url() removes the custom URL. Without it, Scikit-LLM falls back to its built-in default route, OpenAI's hosted API. We do not use that route in this series, but it helps to see what the reset does:

PYTHON
SKLLMConfig.reset_gpt_url()
print(SKLLMConfig.get_gpt_url())
OUTPUT
None

Here, we can see None, so no custom URL is stored. Now any custom_url:: model has nowhere to go. The library stops at once, before any retry, with a ValueError:

PYTHON
try:
    clf.predict(ticket)
except ValueError as error:
    print(error)
OUTPUT
You are using the `custom_url` backend but no custom URL was provided. Please set it using `SKLLMConfig.set_gpt_url(<url>)`.

Here, we can see the library telling us exactly which call is missing. So, the reset only removes the URL, while the key stays whatever we set last.

Return to the Hosted Route for the Day 2 support-ticket example

To get back to Ollama, we set both values again:

PYTHON
SKLLMConfig.set_gpt_url("http://localhost:11434/v1")
SKLLMConfig.set_gpt_key("ollama")

Our clf still holds MODEL, so its next predict() goes to Ollama again. We never need a reset to move between Ollama and OpenRouter, because setting the URL replaces the old one.

Read Setup Errors as Clues

When a call fails, we first ask how far the request got. Let me tabulate what we might see and where to look first.

What we see What it means First fix
ModuleNotFoundError: No module named 'skllm' This Python does not have the package. Start .\.venv\Scripts\python.exe.
KeyError: 'OPENROUTER_API_KEY' The key is not in this terminal's environment. Run setx, then open a new terminal.
ValueError about no custom URL The URL was reset or never set. Run the setup block again.
RuntimeError: OpenAI key was not found No key was set at all. Run the set_gpt_key() line.
RuntimeError: Could not complete the operation after 3 retries The request never got a usable reply. Check that Ollama runs, the model name, or the key.
needs_review in the results A reply came back, but its label was not allowed. Read the ticket and the label list.
An allowed but wrong label The setup works. Study the task itself, from Day 3 on.

The RuntimeError row needs a closer look. Scikit-LLM wraps every model call in a retry. It tries 3 times, waiting 1, 2 and 4 seconds after each failed try. After the third failure, it prints the reason and raises RuntimeError. With Ollama closed, our run printed:

PLAINTEXT
Could not complete the operation after 3 retries: `APIConnectionError :: Connection error.`

The error type is always RuntimeError. The original cause, here APIConnectionError, only appears inside the message text. So we read the message, not just the type.

A failed stage points to its own fix, and none of these errors says anything about label accuracy.

Find the Failure Stage for the Day 2 support-ticket example

Warning

A failed call is slow, not instant. The OpenAI client inside Scikit-LLM also retries within each of the 3 tries. On our Windows PC, with Ollama closed, one predict call took 49.5 seconds before the RuntimeError.

Time One Prediction

A timer shows how long one call really takes. Python's perf_counter() measures the time between two points. We start it just before the call and read it in finally, which runs whether the call works or fails:

PYTHON
from time import perf_counter

started = perf_counter()
try:
    one_result = clf.predict(ticket)[0]
except RuntimeError as error:
    print("request failed:", error)
else:
    print("returned label:", one_result)
finally:
    print("elapsed seconds:", round(perf_counter() - started, 1))

On success, this prints the returned label and the time taken. That time includes the model's thinking, so it depends on our machine. The very first call can be slower, because Ollama has to load the model into memory.

On failure, the library first prints its own retry message. Then our except line prints the same text after request failed:. The elapsed time includes all the retries and waits, like the 49.5 seconds we measured with Ollama closed. We catch RuntimeError because that is the only type a failed model call raises.

Tip

Before a bigger run, we write down four things: the Python version, the scikit-llm version, the MODEL string and the URL. We never write down the key. This short record explains most differences when a classmate gets another result.

Conclusion

This is how Scikit-LLM backends work. The package builds the request, set_gpt_url() and set_gpt_key() choose the route, and the custom_url:: part of MODEL sends it there. We set up the default route to Ollama on our own machine. We saw how to switch to the free OpenRouter model, with its key kept in the environment. Finally, we used reset_gpt_url() to remove the custom route.

  • The package, the backend and the model are three separate parts.
  • "ollama" is only a placeholder key; the OpenRouter key lives in OPENROUTER_API_KEY and is never printed.
  • A custom_url:: model needs a URL, or it stops at once with a ValueError.
  • A failed call ends in RuntimeError after 3 tries, and the real cause is in the message text.
  • A working route is a setup check, not proof that a label is right.

Next steps:

  • Go back to Day 1 to see this setup inside the full six-ticket example.
  • Continue with Day 3, where we design clearer zero-shot labels.
  • Learn more Ollama commands in the Ollama setup guide.

We now have a model route we can read, switch and test before asking the model to sort more tickets.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments