Jev vs Laya: Hosted Jev 1.13 vs Local Laya on an RTX 5090

Hosted Jev 1.13 scored 0.970 and Laya on a local RTX 5090 scored 0.875 on the same 10-intent question over 200 bank card messages, while Laya answered one request in 20.5 ms against 654 ms for Jev.

Oct 11, 202620 min readFollow

Topics You Will Master

How one TypeSafe SDK script talks to hosted Jev and to Laya on our own GPU
How Jev 1.13 and Laya compare on 200 bank card messages with 10 intents
Why a local model wins on one request but not on 200 at once
How far we can trust Laya's confidence on a store inbox

What if the same Python code could call a hosted model or a model on our own GPU, and only the address changed? We tried it. On 200 bank card messages, hosted Jev 1.13 picked the right intent 97% of the time. Laya, running on our RTX 5090, got 87.5%. But here's the twist: Laya answered a single request in 20.5 ms, and Jev took 654 ms.

Jev and Laya are decision models. They do not write text. We give them some text, called the state, and a question with a fixed list of answers. In simple words, they fill in a multiple choice sheet and tell us how sure they are about each option. Jev is TypeSafe's hosted model. Laya is an open-weights model from Convai Innovations that runs on our own machine.

In this blog, we will learn how to point one script at both models, how they compare on the same 200 messages, how fast each one answers, and how much we can trust Laya's confidence on a store inbox.

Bestseller

Jev AI and Laya Fundamentals: Decision AI for Beginners

Build fast, low-cost AI decisions with Jev, Laya and System One models: Choice, Score, confidence and calibration.

Enroll on Udemy →30 day refund, lifetime access

Our Test Setup

Let me tabulate the setup for your better understanding:

Item Jev Laya
Model that answered typesafe/jev-1.13-20260917 laya:en (ModernBERT-large)
Where it runs Hosted, reached through OpenRouter Our RTX 5090 desktop, on cuda:0
Server TypeSafe's API Ollaya 0.9.0 on http://localhost:11435
Client TypeSafe SDK The same TypeSafe SDK

The notebook ran on 4 October 2026, the date stamped in Ollaya's responses. Ollaya is a small local server for Laya, a bit like Ollama is for chat models. It speaks the same request format as hosted Jev, and that is what makes this test fair: the question code is the same for both models.

Advertisement

One Script for Both Models

The TypeSafe SDK has three question types. A Choice picks one option from a list. A Score picks a level on a scale. A Noul is a yes or no question, and its answer is the probability of yes. Let's see the imports as below:

PYTHON
import asyncio, time
import httpx
import pandas as pd
from dotenv import load_dotenv
from typesafe_sdk import TypeSafeClient, AsyncTypeSafeClient, Choice, Score, Noul

load_dotenv()
OLLAYA = "http://localhost:11435"

load_dotenv() reads our .env file. That file holds TYPESAFE_BASE_URL and TYPESAFE_API_KEY, and we point them at OpenRouter. So TypeSafeClient() with no arguments talks to hosted Jev.

For Laya, only three things change: base_url, api_key and model. Ollaya ignores the key, but the SDK will not start without one, so we pass the word "local". Let's ask both models the same yes or no question:

PYTHON
local = TypeSafeClient(base_url=OLLAYA, api_key="local", model="laya")
jev = TypeSafeClient()

REFUND = Noul(instructions="The customer asks for money back.")
state = "I was charged twice for order 20966. Please refund the second payment today."

local.system_one(state, {"refund": REFUND})
OUTPUT
SystemOneResponse(model='laya:en', usage=Usage(input_tokens=49, output_tokens=0), answers={'refund': NoulAnswer(type='noul', noul=0.8915)})

Here, we can see Laya gave a probability of 0.8915 that the customer wants money back. Now the same call on hosted Jev:

PYTHON
jev.system_one(state, {"refund": REFUND})
PYTHON
SystemOneResponse(model='typesafe/jev-1.13-20260917', usage=Usage(input_tokens=293, output_tokens=20), answers={'refund': NoulAnswer(type='noul', noul=0.98)})

Jev said yes with 0.98. Both models agree, and the answer comes back in the same shape. Our code does not care which model is on the other end.

The Bank Card Question

Now we need a harder job. We use 200 real bank card messages from the Banking77 dataset, with 10 intents and typos included. Each message has a gold label, the intent a person gave it. Let's load them as below:

PYTHON
cards = pd.read_csv("data/card_messages.csv")
cards.head()
OUTPUT
idtextintent
01I'm starting to think my card is lost because ...card_arrival
12What do I need to do to change my PIN?change_pin
23What do I do if I think my card was improperly...compromised_card
34I cannot find my credit card.lost_or_stolen_card
45How can I change my Tholepin ?change_pin

Here, we can see real customer wording, typos and all, like "Tholepin" for PIN. We write one Choice question with a short description for every option. The description tells the model what kind of message belongs there. Let's see the code as below:

PYTHON
CARD_INTENT = Choice(
    instructions="What does the bank customer need help with?",
    criteria={
        "activate_my_card": "activating a card that has arrived",
        "card_about_to_expire": "a card that expires soon, getting a replacement before expiry",
        "card_arrival": "a new card that has not arrived yet",
        "card_not_working": "a card that does not work at all",
        "card_payment_fee_charged": "an unexpected fee on a card payment",
        "card_payment_wrong_exchange_rate": "a wrong exchange rate on a card payment abroad",
        "change_pin": "changing or resetting the PIN",
        "compromised_card": "someone else may have used the card or its details",
        "declined_card_payment": "a card payment that was declined",
        "lost_or_stolen_card": "a card that is lost or stolen",
    },
)

Both models get this exact question. Nothing else changes between them.

Advertisement

Which Model Is More Accurate?

We send all 200 messages to each model at the same time with asyncio.gather. For that we need the async clients. alocal is AsyncTypeSafeClient(base_url=OLLAYA, api_key="local", model="laya"), the async twin of our local client, and ajev is the hosted one. We also time each batch, so we get speed and accuracy from the same run:

PYTHON
ajev = AsyncTypeSafeClient()

start = time.perf_counter()
laya_results = await asyncio.gather(*[alocal.system_one(t, {"intent": CARD_INTENT}) for t in cards.text])
laya_seconds = time.perf_counter() - start

start = time.perf_counter()
jev_results = await asyncio.gather(*[ajev.system_one(t, {"intent": CARD_INTENT}) for t in cards.text])
jev_seconds = time.perf_counter() - start

laya_results[0].model, jev_results[0].model
OUTPUT
('laya:en', 'typesafe/jev-1.13-20260917')

Here, we can see the exact models that answered: Laya's English checkpoint and the dated Jev 1.13 build 20260917. Now we score both against the gold labels:

PYTHON
cards["laya"] = [r.answers["intent"].choice for r in laya_results]
cards["jev"] = [r.answers["intent"].choice for r in jev_results]

pd.DataFrame({
    "accuracy": [(cards.laya == cards.intent).mean(), (cards.jev == cards.intent).mean()],
    "seconds_for_200": [laya_seconds, jev_seconds],
}, index=["laya", "jev"])
OUTPUT
accuracyseconds_for_200
laya0.8752.275112
jev0.9702.078333

Here, we can see Jev got 0.970 and Laya got 0.875 on the same 200 messages. That is 97% against 87.5%. Let me tabulate every number from this run for your better understanding:

Measure Jev 1.13, hosted Laya, local RTX 5090
Accuracy, 200 bank card messages, 10 intents 0.970 (97%) 0.875 (87.5%)
One request, wall time 654 ms, through OpenRouter, network included 20.5 ms
All 200 requests at once 2.08 s 2.28 s

Where Laya Loses Points

So where does the gap come from? We break the accuracy down by intent:

PYTHON
cards.assign(laya_ok=cards.laya == cards.intent, jev_ok=cards.jev == cards.intent).groupby("intent")[["laya_ok", "jev_ok"]].mean()
OUTPUT
intentlaya_okjev_ok
activate_my_card0.950.95
card_about_to_expire0.800.85
card_arrival0.551.00
card_not_working0.800.95
card_payment_fee_charged0.951.00
card_payment_wrong_exchange_rate0.900.95
change_pin1.001.00
compromised_card0.851.00
declined_card_payment0.951.00
lost_or_stolen_card1.001.00

Here, we can see one intent does most of the damage. On card_arrival, Laya got 0.55 and Jev got 1.00. Laya sent many "where is my new card" messages to lost_or_stolen_card. For example, "How do I track my card?" and "still waiting on that card" both went to lost_or_stolen_card with Laya, and to card_arrival with Jev.

That makes sense once we read the messages. A card that never arrived does sound a bit like a lost card. Jev separates the two. Laya does not, at least not with these option descriptions.

Advertisement

Which Model Is Faster?

Speed has two sides. One user waits for one answer. A batch job waits for all of them. Let's time one request first, the way a single user would feel it:

PYTHON
%time local.system_one(cards.text[0], {"intent": CARD_INTENT}).answers["intent"].choice
OUTPUT
CPU times: total: 0 ns
Wall time: 20.5 ms
'lost_or_stolen_card'

Laya answered in 20.5 ms on our GPU. Now the same call on hosted Jev:

PYTHON
%time jev.system_one(cards.text[0], {"intent": CARD_INTENT}).answers["intent"].choice
OUTPUT
CPU times: total: 0 ns
Wall time: 654 ms
'card_arrival'

Jev took 654 ms. This number includes the trip over the internet to OpenRouter and back, not just the model's own work. Notice the answers too. The gold label of this first message is card_arrival, so Jev was right and Laya was wrong, the same pattern we saw in the table above.

Now look back at the batch timings. All 200 at once took Laya 2.28 s and Jev 2.08 s. So when we send many requests at the same time, the gap closes, and Jev was even a little faster. A hosted service can work on many requests in parallel, so the network wait overlaps instead of adding up.

Note

Each timing here is a single run, not an average of many. The 654 ms also depends on where we sit and on OpenRouter's route that day. Treat these as one honest snapshot, not a fixed speed.

Advertisement

Laya on a Store Inbox

Laya is not only for bank messages. Its real strength is that it runs on our own machine, so private text never leaves it. Let's try it on an inbox: 40 labelled emails to a mid-size online home goods store, 5 each of 8 intents. We load them from data/emails.csv into a DataFrame called emails, with a subject, a body and the gold intent for each email.

We write one Choice with a description per option. The state is JSON, so the model reads the subject and the body together:

PYTHON
INTENT = Choice(
    instructions="What does the sender of this email want from the store?",
    criteria={
        "order_status": "a customer asks where an order is or when it will arrive",
        "return_request": "a customer wants to send an item back, exchange it or get a return label",
        "invoice_request": "someone needs an invoice, a receipt or a corrected invoice",
        "partnership": "a business, creator or event proposes working together or buying wholesale",
        "job_application": "someone applies for a job or an internship",
        "spam": "unsolicited marketing, scams and phishing",
        "complaint": "a customer is unhappy about service, quality, a charge or a person",
        "supplier_update": "a supplier announces prices, delays, new products or changes",
    },
)

Now we send all 40 emails at once with the async client and check the accuracy:

PYTHON
alocal = AsyncTypeSafeClient(base_url=OLLAYA, api_key="local", model="laya")

results = await asyncio.gather(*[
    alocal.system_one({"subject": s, "body": b}, {"intent": INTENT})
    for s, b in zip(emails.subject, emails.body)
])
emails["pred"] = [r.answers["intent"].choice for r in results]
emails["confidence"] = [r.answers["intent"].confidence for r in results]
(emails.pred == emails.intent).mean()
PYTHON
np.float64(0.875)

Here, we can see Laya got 0.875, so 35 of the 40 emails went to the right place. Let's look at the five misses with their confidence:

PYTHON
emails[emails.pred != emails.intent][["subject", "intent", "pred", "confidence"]]
OUTPUT
subjectintentpredconfidence
15Delay on PO 4471, linen napkinssupplier_updateorder_status0.5382
23New AW cataloguesupplier_updateorder_status0.4508
27Listing your products on Greenshelfpartnershipspam0.8856
29Invoice #INV-88213 overdue, action requiredspaminvoice_request0.9947
38Promo code didn't work and nobody helpscomplaintspam0.5355

Three of the five misses sit near 0.5 confidence. So, can confidence tell right answers from wrong ones? Let's compare the average confidence of each group:

PYTHON
emails.groupby(emails.pred == emails.intent).confidence.mean()
OUTPUT
False    0.680960
True     0.912631
Name: confidence, dtype: float64

Here, we can see wrong answers averaged about 0.68 confidence and right answers about 0.91. That is a useful signal. If we send low-confidence emails to a person, we catch most of the misses.

But it is not a perfect signal. The fake "overdue invoice" email is a phishing attempt, and Laya filed it as invoice_request with 0.9947 confidence. A confidence cut-off would never stop that one. This is why we test on our own labelled data before trusting any cut-off.

Which One Should We Use?

Let me tabulate the answer from this run:

If we need Pick Why
The most right answers on this card question Jev 1.13 0.970 against 0.875
The fastest single answer Laya on a local GPU 20.5 ms against 654 ms
A big batch sent all at once Either 2.08 s and 2.28 s for 200 messages
Text that must stay on our machine Laya it never leaves the local server

Since the code is the same for both, we do not have to choose once and for all. We can start with one and switch by changing base_url, api_key and model.

Advertisement

Limits of This Test

This is one desktop, one network and one run:

  • The single-request times come from one %time call each, not an average.
  • Jev's time includes the network to OpenRouter. Calling from another place, or through TypeSafe directly, could be faster or slower.
  • We wrote the option descriptions once and did not tune them for either model. Different wording can change both scores.
  • 200 messages and 40 emails are small samples. A gap of a few messages could move on another run.

If you run Jev and Laya on your own data, please share your accuracy and your time per request in the comments.

Conclusion

This is how hosted Jev 1.13 and Laya on a local RTX 5090 compare on the same question. We used one TypeSafe SDK script for both, and only the address, key and model name changed. On 200 bank card messages, Jev was more accurate, 0.970 against 0.875, mostly because Laya mixed up cards that had not arrived with lost cards. Laya answered one request in 20.5 ms against 654 ms for Jev, but with all 200 sent at once the two finished in about the same time.

Key takeaways:

  • Ollaya serves Laya with the same request format as hosted Jev, so one script runs both.
  • Jev 1.13 scored 0.970 and Laya 0.875 on the same 10-intent card question.
  • Most of Laya's misses came from one intent: card_arrival at 0.55 against 1.00 for Jev.
  • Local Laya wins on one request; under load the network wait overlaps and the gap closes.
  • On the store inbox, wrong answers averaged about 0.68 confidence against 0.91 for right ones, but one phishing email was wrong at 0.9947.

Next steps:

In short, we sent the same 10-intent question to hosted Jev and to local Laya, and measured how often each was right and how long each took.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments