Jev 1.13 Intent Routing on 200 Bank Messages: Picking a Confidence Threshold

We tested Jev 1.13 as a router for 200 Banking77 card messages with 10 intents, reached 0.985 accuracy with option descriptions, and used its confidence to pick an auto-answer cut-off that still covers 90% of messages.

Oct 11, 202620 min readFollow

Topics You Will Master

How to test a Jev 1.13 Choice question on 200 labelled bank messages
Why one-line option descriptions lifted accuracy from 0.965 to 0.985
How to read a reliability table and spot overconfidence
How to pick a confidence cut-off from coverage and accuracy

Before a message router answers customers on its own, we need two answers. How often is it right? And when it says it is sure, can we believe it? We tested Jev 1.13 on 200 real bank card messages. With a one-line description for each option, it routed 98.5% of them correctly. A cut-off at 0.95 confidence still let it answer 90% of the messages by itself, and 98.3% of those were right.

Intent routing means reading a customer message and deciding which team or flow should handle it. In simple words, it is a sorting desk for messages. Jev is TypeSafe's hosted decision model. It does not write text. It picks one option from our list and returns a probability for every option.

In this blog, we will learn how to measure Jev's accuracy on labelled data, how much option descriptions help, how to check whether its confidence is honest, and how to pick a confidence cut-off for automatic answers.

Bestseller

Jev AI and Laya Fundamentals: Decision AI for Beginners

Build fast, low-cost AI decisions with Jev, Laya and System One models: Choice, Score, confidence and calibration.

Enroll on Udemy →30 day refund, lifetime access

Our Test Setup

Let me tabulate the setup for your better understanding:

Item Value
Model that answered typesafe/jev-1.13-20260917
How we reached it TypeSafe SDK, through OpenRouter
Data 200 Banking77 card messages, 20 for each intent
Question One Choice with 10 intents
Run October 2026

The TypeSafe SDK reads TYPESAFE_BASE_URL and TYPESAFE_API_KEY from our .env file, and we point them at OpenRouter. client sends one request at a time, and aclient sends many at once. Let's see the setup code as below:

PYTHON
import asyncio
from datetime import date

import pandas as pd
import matplotlib.pyplot as plt
from dotenv import load_dotenv
from typesafe_sdk import TypeSafeClient, AsyncTypeSafeClient, Choice, Noul

load_dotenv()
client = TypeSafeClient()
aclient = AsyncTypeSafeClient()

df = pd.read_csv("data/card_messages.csv")
df.head()
OUTPUT
idtextintent
01I'm starting to think my card is lost because ...card_arrival
12What do I need to do to change my PIN?change_pin
23What do I do if I think my card was improperly...compromised_card
34I cannot find my credit card.lost_or_stolen_card
45How can I change my Tholepin ?change_pin

Here, we can see each row has the customer's text and a gold intent, the label a person gave it. Some messages have typos, like "Tholepin" for PIN. That is good, because real customers type like this.

Advertisement

Labels Only vs Labels With Descriptions

First, we ask the question with option names only. None means the option has no description:

PYTHON
INSTRUCTIONS = "What does the customer need help with?"

BARE = Choice(
    instructions=INSTRUCTIONS,
    criteria={
        "activate_my_card": None,
        "card_about_to_expire": None,
        "card_arrival": None,
        "card_not_working": None,
        "card_payment_fee_charged": None,
        "card_payment_wrong_exchange_rate": None,
        "change_pin": None,
        "compromised_card": None,
        "declined_card_payment": None,
        "lost_or_stolen_card": None,
    },
)

df.text[0], df.intent[0]
OUTPUT
("I'm starting to think my card is lost because it still hasn't arrived, can you help?", 'card_arrival')

This first message is a good test. It says "lost", but the real problem is a card that has not arrived. Now we write the same question with one line per option. Each line says what a message in that class is about:

PYTHON
CRITERIA = {
    "activate_my_card": "activate a card the customer already has",
    "card_about_to_expire": "the card is expiring, or getting the replacement card",
    "card_arrival": "a new card has not arrived yet, delivery time or tracking",
    "card_not_working": "the card does not work at all, not one declined payment",
    "card_payment_fee_charged": "an extra fee was charged on a card payment",
    "card_payment_wrong_exchange_rate": "a payment in another currency used the wrong exchange rate",
    "change_pin": "change or reset the card PIN",
    "compromised_card": "someone else may have used the card or its details",
    "declined_card_payment": "a card payment was declined or did not go through",
    "lost_or_stolen_card": "the card is lost, missing or stolen",
}

DESCRIBED = Choice(instructions=INSTRUCTIONS, criteria=CRITERIA)

r = client.system_one(df.text[0], {"intent": DESCRIBED})
r.answers["intent"]
PYTHON
ChoiceAnswer(type='choice', choice='card_arrival', confidence=0.92, probabilities={'activate_my_card': 0.0, 'card_arrival': 0.93, 'lost_or_stolen_card': 0.07, 'declined_card_payment': 0.0, 'card_about_to_expire': 0.0, 'card_not_working': 0.0, 'card_payment_wrong_exchange_rate': 0.0, 'change_pin': 0.0, 'compromised_card': 0.0, 'card_payment_fee_charged': 0.0})

Here, we can see Jev picked card_arrival with 0.93 probability and kept 0.07 on lost_or_stolen_card. The answer carries the full list of probabilities, not just the winner.

Running All 200 Messages

Now we send all 200 messages through both versions of the question. asyncio.gather sends every request at once and keeps the input order:

PYTHON
async def ask_all(question):
    return await asyncio.gather(*[aclient.system_one(text, {"intent": question}) for text in df.text])

bare = await ask_all(BARE)
described = await ask_all(DESCRIBED)
len(bare), len(described)
OUTPUT
(200, 200)

We get 200 answers for each version. Next, we save the answers as columns. p_top is the probability of the chosen option. confidence rescales it so that an even split across all options gives 0:

PYTHON
df["pred_bare"] = [r.answers["intent"].choice for r in bare]
df["pred"] = [r.answers["intent"].choice for r in described]
df["confidence"] = [r.answers["intent"].confidence for r in described]
df["p_top"] = [max(r.answers["intent"].probabilities.values()) for r in described]
df["correct"] = df.pred == df.intent
df.head()
OUTPUT
idtextintentpred_barepredconfidencep_topcorrect
01I'm starting to think my card is lost because ...card_arrivalcard_arrivalcard_arrival0.940.95True
12What do I need to do to change my PIN?change_pinchange_pinchange_pin1.001.00True
23What do I do if I think my card was improperly...compromised_cardcompromised_cardcompromised_card1.001.00True
34I cannot find my credit card.lost_or_stolen_cardlost_or_stolen_cardlost_or_stolen_card1.001.00True
45How can I change my Tholepin ?change_pinchange_pinchange_pin0.991.00True
Advertisement

How Accurate Is Jev?

Accuracy is the share of rows where the prediction matches the gold label. Let's compare the two versions:

PLAINTEXT
(df.pred_bare == df.intent).mean(), df.correct.mean()
PLAINTEXT
(np.float64(0.965), np.float64(0.985))

Here, we can see labels only gave 0.965, and labels with descriptions gave 0.985. One short line per option was enough to lift the score. Let's see which rows the descriptions changed:

PYTHON
df[df.pred_bare != df.pred][["text", "intent", "pred_bare", "pred"]]
OUTPUT
textintentpred_barepred
50My card didn't work in one of the shops.declined_card_paymentcard_not_workingdeclined_card_payment
92I need to know the cost and when I will receiv...card_about_to_expirecard_arrivalcard_about_to_expire
117Can I receive a new card while I am in China?card_about_to_expirecard_arrivalcard_about_to_expire
129in china, need new cardcard_about_to_expirelost_or_stolen_cardcard_about_to_expire

All four changed answers moved to the gold label. The description "not one declined payment" pulled the shop message to declined_card_payment. The line "getting the replacement card" pulled the "new card" messages to card_about_to_expire.

The Three Wrong Answers

With descriptions on, Jev got 3 of the 200 messages wrong. Let's read them with their confidence:

PYTHON
df[~df.correct][["text", "intent", "pred", "confidence"]]
OUTPUT
textintentpredconfidence
123Nothing goes through on my card.card_not_workingdeclined_card_payment0.95
175Why am I being charged more ?card_payment_wrong_exchange_ratecard_payment_fee_charged0.97
185Is there a way my new card can be renewed?activate_my_cardcard_about_to_expire0.98

Here, we can see each miss is a close neighbour of the right intent. "Why am I being charged more?" could be a fee or an exchange rate, and even a person might hesitate. Notice the confidence too: all three misses were at 0.95 or above. Keep that in mind for the cut-off section.

Advertisement

Can We Trust Jev's Confidence?

A model is well calibrated when its confidence matches how often it is right. In simple words, if it says 90% sure, it should be right about 9 times in 10. We check this with a reliability table. We group the answers into bins by p_top. Then, for each bin, said is the average top probability and right is the share that was correct:

PYTHON
df["bin"] = pd.cut(df.p_top, [0.1, 0.5, 0.8, 0.9, 0.95, 0.99, 1.0], include_lowest=True)

reliability = df.groupby("bin", observed=True).agg(
    said=("p_top", "mean"),
    right=("correct", "mean"),
    rows=("correct", "size"),
)
reliability
OUTPUT
binsaidrightrows
(0.099, 0.5]0.490001.0001
(0.5, 0.8]0.600001.0004
(0.8, 0.9]0.873751.0008
(0.9, 0.95]0.935000.8758
(0.95, 0.99]0.985000.87516
(0.99, 1.0]1.000001.000163

Here, we can see 163 of the 200 answers landed in the 0.99 to 1.0 bin, and every one of them was right. The low bins were all right too, so there Jev was more careful than it needed to be.

The weak spot is the middle-high range. In the 0.9 to 0.95 and 0.95 to 0.99 bins, Jev said 0.935 and 0.985 on average but was right only 0.875 of the time, on 24 rows in total. So Jev was slightly overconfident there. With only 8 and 16 rows in those bins, one miss moves the number a lot, so we read this as a warning sign, not a final answer.

Note

A bin with few rows says little. Before trusting a reliability table, always look at the rows column next to it.

Advertisement

Coverage vs Accuracy per Threshold

Now the practical question. If Jev answers only when its confidence is at or above a threshold, how many messages does it handle, and how accurate are those answers? Coverage is the share of messages at or above the threshold, the ones answered automatically. Accuracy here is measured on those rows only:

PYTHON
thresholds = [0, 0.5, 0.7, 0.8, 0.9, 0.95]

pd.DataFrame({
    "threshold": thresholds,
    "coverage": [(df.confidence >= t).mean() for t in thresholds],
    "accuracy": [df[df.confidence >= t].correct.mean() for t in thresholds],
})
OUTPUT
thresholdcoverageaccuracy
00.001.0000.985000
10.500.9900.984848
20.700.9750.984615
30.800.9750.984615
40.900.9350.983957
50.950.9000.983333

Here, we can see a 0.95 threshold still answers 90% of the messages automatically, and those answers are 98.3% right. Let me tabulate the key numbers of this run for your better understanding:

Measure Jev 1.13, 200 bank card messages
Accuracy, labels only 0.965
Accuracy, with option descriptions 0.985
Coverage at a 0.95 confidence threshold 0.900
Accuracy on answers at or above 0.95 0.983
Answers in the 0.99 to 1.0 bin 163 of 200, all right
Share right in the 0.9 to 0.99 bins 0.875, on 24 rows

But there is a surprise in this table. Raising the threshold did not make the answers more accurate. Accuracy stays near 0.98 at every threshold. Why? Because all three wrong answers had a confidence of 0.95 or more, and every answer below that was right. On this sample, the cut-off sent correct answers to people and kept the mistakes.

So is the cut-off useless? No. On new messages, low confidence is still where unclear text shows up, and a person should look at those. But this run shows that a cut-off cannot catch confident mistakes. To find those, we need better descriptions, or a person who spot-checks a sample of the automatic answers.

Advertisement

Choosing a Confidence Threshold

We set THRESHOLD = 0.9 and read every message below it before deciding anything:

PYTHON
THRESHOLD = 0.9

df[df.confidence < THRESHOLD][["text", "intent", "pred", "confidence"]]
OUTPUT
textintentpredconfidence
14I think I lost my card . I dont know how long ...lost_or_stolen_cardlost_or_stolen_card0.82
29How can I get my physical card to work?card_not_workingcard_not_working0.45
50My card didn't work in one of the shops.declined_card_paymentdeclined_card_payment0.83
58I broke my cardcard_not_workingcard_not_working0.88
62going to need a new card what are the fees and...card_about_to_expirecard_about_to_expire0.85
71My card was taken from melost_or_stolen_cardlost_or_stolen_card0.87
75Some idiot stole my card.lost_or_stolen_cardlost_or_stolen_card0.88
87How do I unblock my card using the app?card_not_workingcard_not_working0.81
92I need to know the cost and when I will receiv...card_about_to_expirecard_about_to_expire0.88
117Can I receive a new card while I am in China?card_about_to_expirecard_about_to_expire0.43
129in china, need new cardcard_about_to_expirecard_about_to_expire0.55
141How do I freeze my account?compromised_cardcompromised_card0.60
194Can I freeze my card right now?compromised_cardcompromised_card0.60

Here, we can see the low-confidence messages are the unclear ones: a card in China, a frozen account, a card that will not work. Every one of them was answered correctly, but they are exactly the messages where a human check is cheap insurance.

Now we turn the confidence into three bands. Below 0.5 goes to a human, 0.5 up to 0.9 asks for a quick confirm, and above 0.9 lets the system act:

PYTHON
df["band"] = pd.cut(df.confidence, [0, 0.5, THRESHOLD, 1.0], labels=["human", "confirm", "act"], include_lowest=True)
df.groupby("band", observed=True).correct.agg(["mean", "size"])
OUTPUT
bandmeansize
human1.0000002
confirm1.00000012
act0.983871186

Here, we can see 186 messages fall in the act band at 0.984 accuracy. Only 2 go to a person and 12 ask for a confirmation. The right band edges depend on the cost of a mistake. A wrong answer about a PIN change costs little. A wrong answer about a stolen card costs a lot, so that intent may deserve a higher bar.

Advertisement

Does Option Order Change the Answer?

A good router should not care which option we list first. So we build REVERSED with Choice(instructions=INSTRUCTIONS, criteria=dict(reversed(CRITERIA.items()))). It has the same options and descriptions in reverse order. Then we count how many answers change:

PYTHON
reordered = await ask_all(REVERSED)
df["pred_reversed"] = [r.answers["intent"].choice for r in reordered]
(df.pred != df.pred_reversed).sum()
PYTHON
np.int64(1)

Here, we can see only 1 of the 200 answers changed. The one message that flipped was "Can I receive a new card while I am in China?", which also had one of the lowest confidences, 0.43.

Pinning the Model Version

Our thresholds only hold for the model we tuned them on. If the model behind the alias changes, the confidence numbers can shift. So we pin the exact version:

PYTHON
pinned = TypeSafeClient(model="typesafe/jev-1.13")
pinned.system_one(df.text[0], {"intent": DESCRIBED}).model
OUTPUT
'typesafe/jev-1.13-20260917'

Here, we can see the pinned client answered with typesafe/jev-1.13-20260917, the same version as the whole run. We save this version with our results, and we run the test again whenever it changes.

Limits of This Test

This is one run on one dataset:

  • 200 messages with only 3 misses is a small sample for judging calibration. The 0.9 to 0.99 bins hold just 24 rows.
  • Jev was reached through OpenRouter, and every number belongs to the version typesafe/jev-1.13-20260917.
  • The option wording matters. Our Jev vs Laya comparison uses the same 200 messages with differently worded descriptions, so its Jev accuracy is not comparable with the numbers here.

If you test Jev on your own messages, please share your accuracy and your threshold in the comments.

Conclusion

This is how we test a Jev 1.13 intent router before launch. We started with option names only and got 0.965 accuracy. One line of description per option lifted it to 0.985. The reliability table showed 163 of 200 answers at 0.99 or above, all right, and a slightly overconfident middle range on a small sample. Finally, a 0.95 cut-off still answered 90% of messages automatically at 0.983 accuracy, but it could not catch the three confident mistakes.

Key takeaways:

  • Always measure a router on labelled data before it answers customers.
  • Short option descriptions lifted Jev 1.13 from 0.965 to 0.985 on these 200 messages.
  • Read the rows column of a reliability table before trusting any bin.
  • A confidence cut-off sends unclear messages to people, but it cannot catch confident mistakes.
  • Pin the model version, because thresholds belong to one version.

Next steps:

In short, we routed 200 bank card messages with Jev 1.13, checked its confidence, and picked an auto-answer cut-off from real coverage and accuracy numbers.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments