What Are Jev and Laya? Fine-Tuning Judge-Only Models in 4 Languages – A Verification Report

Jev is a judgment-specialized model that does not generate text but only returns probabilities for "which option" or "yes or no," and Laya is an OSS (Apache 2.0) that can be used the same way. For our AI employee service AGTO's post-judgment use case, we fine-tuned the multilingual version of Laya using synthetic data in Japanese, Thai, English, and Lao. Before fine-tuning, the accuracy for distinguishing whether a post was a work request was around 50%, but after fine-tuning it rose to about 0.90, with response times shorter than an LLM's. However, the accuracy did not reach that of top-tier LLMs, and the combination with the best balance was routing only ambiguous posts to the LLM.
Are you making the LLM wait about one second every time just to answer "yes" or "no"? For engineers and product managers who want to reconsider the speed and cost of judgment processing, we introduce, in order, the mechanisms of Jev and Laya, our measured results, methods we tried that didn't work, and points to verify before adoption.
What Are Jev and Laya? Classification-Only Models That Don't Generate Text
**Jev and Laya are models that receive input text and a "structured question" and return the probability of the answer in a single computation.**Unlike LLMs, they do not generate text word by word, so answers come back quickly, and the returned probability can be used directly for program branching. The difference is that Jev is a hosted API from TypeSafe AI, while Laya is an OSS that you run on your own server.
Jev: A Hosted Classification API by TypeSafe AI
Jev is TypeSafe AI's flagship model, and the first model in what the company calls the "System One model" category. When you pass in the main text to be judged (officially called the state; hereafter "state") together with a question, it returns an answer for each question. There are three types of questions it can handle: Choice, which selects one from candidates; Score, which evaluates on a scale; and Noul, which returns a yes/no probability. According to the official documentation, Choice and Score return a confidence value in addition to the answer and the probability for each candidate. Even if multiple questions are mixed into a single call, the questions are evaluated independently and in parallel, so response time is said to barely increase.
Noul is the one that's hard to understand from the name alone. For a question that can be answered with yes or no, it returns a single value between 0 and 1 representing the probability that the answer is "yes." In the official example, for the question "Does the customer want a human representative?", "Thank you, it's fixed!" scored 0.02, "Are you a bot?" scored 0.40, and "I've asked three times already. Let me talk to a person" scored 0.99. Since there are only two possible answers, this value itself represents the strength of confidence, and no separate confidence value is returned as with Choice or Score. Also, the value is not a scale of degree. For example, if "Are you good at Python?" returns 0.6, that is not a degree of skill but rather "the probability of being skilled." When you want to measure degree, you should use Score.
It is provided as a hosted API, called via POST https://api.typesafe.ai/v1/systemone using an API key issued in the console. The pricing and main constraints are described on the model page as follows (as of the time of writing; please check the official page for the latest information).
| Item | Official description |
|---|---|
| Pricing | $42 per 1 billion input tokens. Output is free |
| Limit per request | 64k tokens (state plus the longest question combined, up to 32k) |
| Input | Text only (images, audio, and video are not supported) |
| Number of Choice candidates | Up to 255 per question |
| Language | Primarily trained on English. Other languages are supported but accuracy is not equivalent |
An easily overlooked point is that Jev cannot be fine-tuned per customer. All accounts use the same weights, and adjustments for your own use are made through the information included in the state and the way question instructions and candidate descriptions are written. Regarding speed, the official site presents a demonstration stating "for System One-oriented tasks, 0.114 seconds compared to 8.566 seconds for an LLM," but this is a comparison made by TypeSafe itself.
Laya: A Jev-Compatible OSS Running on Your Own Server
Laya is a judgment model released by Convai Innovations under Apache 2.0. It comes in an English version (ModernBERT-large, 421 million parameters) and a multilingual version supporting over 100 languages (mmBERT-base, 322 million parameters, 1,024 input tokens), among others, and handles Choice, Score, and Noul using the same concepts as Jev. Since the included HTTP server accepts requests in the same POST /v1/systemone format as Jev, you can even redirect a client written for Jev to your own server.
The README states a p50 of 32.8ms per question on a GPU (T4), comparing it to Jev's third-party-reported figure of 236–276ms. However, it explicitly notes that the Jev figures were not measured by the author themselves. Regarding accuracy, it states that the fine-tuned version achieves 0.766 on the public benchmark typed-decisions (compared to Jev's published figure of 0.727), while also noting that Jev significantly outperforms on intent classification with more than 70 options (Banking77).
An easily overlooked point is that the multilingual version is distributed without temperature calibration (a correction for how sharp the probabilities are). The README itself indicates the assumption that you should calibrate it with your own data before use.
Differences from LLMs: Which Classifications Fit and Which Don't
You can also ask an LLM "Is this post a request? Answer yes or no," but judgment-specialized models and LLMs are good at different jobs. In conclusion, judgment-specialized models excel in speed and ease of handling probabilities, while they fall short of LLMs in question flexibility and explanatory power.
| Aspect | Judgment-specialized models (Jev / Laya) | LLM |
|---|---|---|
| Output | Probability for each candidate | Text (the answer is read from the string) |
| Speed | Completes in a single computation | Slows down depending on generation length |
| Question flexibility | Limited to the three question types | Can ask anything |
| Number of candidates | For Laya, descriptions get truncated when there are many (Jev supports up to 255 choices) | Can handle many, since it can read them |
| Explaining reasoning | Not returned | Can be made to write one |
These models are suited to scenarios where the same type of question needs to be judged repeatedly, in large volume, in a short time. On the other hand, judgments that require an answer with reasoning, classification with dozens of candidates, or judgments requiring long context are better suited to LLMs. TypeSafe itself, on its known weaknesses page, advises that counting, numerical calculation, and date comparison are weak points and should be handled on the code side instead.
The division of roles between judgment-specialized models and LLMs can be organized more clearly using the same way of thinking as in routing between cloud LLMs and on-device SLMs.
Why We Tried Classification-Only Models: An AI Employee's "Is This a Work Request?" Classification
To cut to the conclusion, our goal was to verify whether the entry-point judgment—by which an AI employee picks up a post—could be made faster and cheaper than with an LLM. In AGTO, an AI employee reads chat posts and begins work based on them. We split this judgment into two questions and started measuring with Laya, without any training.
Two Questions We Wanted to Classify
The first question is a yes/no question (Noul): "Does this post request someone to do a task?" We refer to this hereafter as the "request" judgment. We draw the line as follows: "It would help if you could " counts as a request, " is finished" does not, and "I'd like to ask Mr./Ms. A to ~" counts as a request. The second is a six-way multiple choice (Choice) that selects, from five tasks plus "none of the above," which task the requested work corresponds to. Since the candidate tasks vary by tenant and channel, rather than having the model memorize specific task names, what's required is the ability to read the candidates' descriptions and choose among them.
The languages involved are four: Japanese, Thai, English, and Laotian. Asking an LLM can produce the judgment, but each post requires waiting about 0.6 to 2 seconds, and as the volume increases, costs pile up as well. We reasoned that if a model dedicated to this judgment could do the same thing, we could improve on both of these points.
The Starting Point Measured with Untrained Laya
First, we tried the multilingual version of Laya without any training. In a small test using 114 parallel sentences in Japanese, Thai, and English, the ability to distinguish "requests" (AUROC) was high, at 0.89–0.95, but the probabilities were heavily skewed toward the "no" side. Cutting at 0.5, only 5–25% of answers came out as "yes," and accuracy stayed in the range of 0.55–0.70. For the six-way choice of identifying the responsible department (a separate test task from the six-way task selection), accuracy was 0.53–0.72, and for the three-level urgency task, performance was no better than random.
Even in an evaluation set we created later (100 items per language), the pre-training accuracy of the "request" judgment was 0.47 for Japanese, 0.54 for English, 0.53 for Thai, and 0.52 for Laotian—barely different from flipping a coin.
The reason AUROC is high while accuracy is low is that the ranking order is assigned correctly, but the position of the boundary is off. Re-determining the threshold using our own data raised accuracy to 0.75–0.93, but adjusting temperature alone does not move the 0.5 boundary. We realized this wasn't a model we could just plug in and use as-is—training and calibration tailored to our own questions were a prerequisite—so we moved on to fine-tuning.
How We Fine-Tuned with Synthetic Data
We did not use customer posts; instead, we trained on synthetic data of workplace chat that we had an LLM write. What surprised us was that what determined accuracy was not the amount of data, but rather "how we designed the evaluation" and "how we assigned the correct answers."
How to Create Training Data
The training data was created by having the LLM write workplace chat posts one at a time, each under specified conditions. The conditions specified were the type of post (7 types of requests, 9 types of non-requests), the task category (one of 60 types), length, tone, and industry. For casual chat, unless we specified the topic as well, the posts ended up being all about the weather and lunch.
For the correct answers, we had labeling done twice without showing the original generation instructions, and kept only the items where both labelings agreed. In the second pass, we shuffled the order of the candidates. We also removed sentences in the wrong language and near-duplicates whose character 3-gram Jaccard coefficient exceeded 0.8.
For the six-choice candidates, each time we randomly selected 2 to 8 items from the 60 task types, always placing "none of the above" at the end. For some requests, we deliberately excluded the correct task from the candidates, making "none of the above" the correct answer. Of the 60 task types, 12 were withheld from training and used only for evaluation, to measure whether the model could correctly classify tasks it had not seen during training.
The final training data comprised about 10,000 rows (roughly 32,000 judgment questions), also mixing in 600 rows from the public dataset typed-decisions. We also included auxiliary questions such as "does it mention a date?" Initially, we didn't have these, and accuracy on judgments not covered in training dropped by 10 points.
Separating the Evaluation Set from Training Data
If we create the evaluation set using the same LLM and the same instructions as the training data, we end up measuring "whether the LLM's own quirks were memorized." So we had the evaluation set written as "one week's worth of a certain team's channel" using a different model and different instructions than those used for the training data. The correct answers were reassigned with the original answers hidden, and only the cases where they diverged were adjudicated and finalized.
We reconsidered the scale partway through. Initially we had 100 items per language, but when we retrained on the same data changing only the random seed, 87 out of 1,200 questions (about 7%) had their answers flip. With this, a 3–5 point difference couldn't be distinguished from random fluctuation. So we increased it to 300 items per language (1,200 items across 4 languages) and switched to comparing using McNemar's test, which counts discrepancies in answers to the same question. After this change, we found that much of what had looked like "slight improvement" was actually within the range of fluctuation.
Since synthetic data alone diverges from actual writing style, we also use 121 actual posts from internal chat (anonymized by replacing names, company names, and URLs) as a separate evaluation.
Training Environment and Costs
Training was done on a single in-house GPU machine. Laya's official notebook assumes 2 Kaggle GPUs, so we rewrote it to run on a single GPU. Peak memory usage is about 6.5GB, and training on roughly 10,000 rows for 4 epochs takes about 1 hour per run.
We also tried "model soup," training 3 runs with different random seeds and averaging their weights. Since inference cost stays the same as a single model while the 6-choice accuracy improved by 1–2 points, we build the final weights using this method.
API costs were about $17 for generating and labeling the initial ~5,600 training data items, and about $26–29 per attempt to relabel the entire training data's correct answers using a higher-tier model (reference values based on pricing at the time of writing). Since there's no GPU usage fee, most of the cost goes into data creation.
Our Measurements: How Much Faster and How Accurate Is Laya Compared to LLMs?
Laya alone achieved accuracy comparable to Claude Haiku on "request" classification, with shorter latency. However, a gap remains between the 6-choice classification and the top-tier LLMs, and the best-balanced approach was a combination that routes only ambiguous posts to an LLM.
Comparing Laya Alone vs. LLM Alone
In conclusion, Laya slightly outperformed Claude Haiku on "request" classification and had the shortest latency, but fell short of Claude Sonnet in accuracy. Below, we compare the trained Laya (final version) against results from having LLMs solve the same questions. The times include the round trip from the Thailand office to the Tokyo region, measured by sending requests one at a time in sequence.
| Configuration | Eval Set Request / 6-choice | Actual Posts Request / 6-choice | p50 | p95 |
|---|---|---|---|---|
| Laya (trained) | 0.903 / 0.892 | 0.876 / 0.851 | 431ms | 620ms |
| Claude Haiku | 0.887 / 0.928 | 0.835 / 0.909 | 618ms | 1,335ms |
| gpt-oss-120b | 0.914 / 0.953 | 0.826 / 0.884 | 483ms | 928ms |
| Claude Sonnet | 0.948 / 0.965 | 0.950 / 0.917 | 1,885ms | 3,032ms |
Laya's p95 is less than half that of Haiku. For the 6-choice task, Haiku and gpt-oss-120b score higher, and while Claude Sonnet is the most accurate on every metric, it takes about 1.9 seconds at p50. Looking only at Laya's server-side processing time, p50 is about 310ms, with the remaining ~120ms being the round trip for communication and encryption.
Broken down by language, Laya's accuracy on "request" was lowest for English at 0.860 (compared to 0.927 for Thai and 0.903 for Lao), tending to miss indirect requests or questions that merely ask for information, like "Does anyone know...?" The overall 4-language figure of 0.903 is an average smoothing out this difference.
Laya's values in the table include a small rule to fix discrepancies between the two questions. Since a contradiction can occur where the 6-choice task selects a task but the "request" probability is low, we apply a rule that raises the "request" probability up to the probability of the selected task. With the final weights, this alone raised the "request" accuracy from 0.897 to 0.903. Jev's official documentation also warns that answers to separate questions are not guaranteed to be mutually consistent.
Note that the sentences and correct answers in the evaluation set were also created by an LLM. Since this could favor LLMs from the same lineage as the model used for labeling, the rankings among LLMs should be read with this caveat in mind.
A Combined Approach: Routing Only Uncertain Posts to the LLM
Since Laya returns probabilities, we can build a configuration where "only posts with confidence below a threshold are routed to an LLM." The threshold was determined via cross-validation so that, for each question, the accuracy is maximized while keeping the proportion routed to the LLM under 10%. In the evaluation set, 16–17% of posts were routed (measured with Haiku and Sonnet). Since the cap is applied per question, and a post is routed if it's ambiguous on either of the two questions, the per-post rate exceeds 10%. In conclusion, p50 remains nearly unchanged while accuracy on the evaluation set improves over Laya alone regardless of which LLM it's paired with. However, "request" accuracy on actual posts dropped in all cases except when paired with Sonnet.
| Configuration | Eval Set Request / 6-choice | Actual Posts Request / 6-choice | p50 | p95 | LLM cost per 1,000 posts |
|---|---|---|---|---|---|
| Laya only | 0.903 / 0.892 | 0.876 / 0.851 | 431ms | 620ms | $0 |
| Laya + Claude Haiku | 0.922 / 0.919 | 0.851 / 0.901 | 450ms | 1,155ms | $0.087 |
| Laya + gpt-oss-120b | 0.926 / 0.926 | 0.851 / 0.884 | 449ms | 1,119ms | $0.019 |
| Laya + Claude Sonnet | 0.933 / 0.925 | 0.901 / 0.901 | 450ms | 2,503ms | $0.156 |
Even in combination, p50 stays nearly the same as Laya alone, and "request" accuracy on the evaluation set exceeds that of Haiku alone in every combination. This is because the questions Laya gets wrong and the ones the LLM gets wrong don't overlap much. Since p95 is determined by the LLM's time for the routed posts, pairing with Sonnet brings it to 2.5 seconds. We prioritized accuracy and chose the configuration paired with Sonnet.
If you want to shrink p95, one option is to reduce the number of posts routed in the first place. This involves using logistic regression to estimate "posts Laya is likely to get wrong" based on confidence, language, discrepancy between the two questions, and sentence length, and routing only the top-ranked ones. Limiting routed posts to 5% or less brought p95 back down to 0.65–0.77 seconds on the evaluation set, at the cost of a 0.6–2.4 point drop in accuracy.
Response Speed Depends on the Server's CPU
Laya runs on CPU. The verification server was hosted on a single AWS EC2 instance with Docker, and we measured under the same conditions for each instance type. In conclusion, for the same monthly cost, choosing an instance type with more physical cores was the most effective approach.
| Instance | Physical Cores | Monthly Cost (Tokyo, On-Demand) | In-server p50 | p95 |
|---|---|---|---|---|
| m7i.large | 1 | ~$95 | 583ms | 935ms |
| c7a.large | 2 | ~$94 | 486ms | 776ms |
| c7a.xlarge | 4 | ~$189 | 313ms | 507ms |
The monthly costs are reference values calculated by multiplying the unit price at the time of writing by 730 hours. Even at nearly the same monthly cost, switching to c7a.large—where 1 vCPU corresponds to 1 physical core—reduced the time by 17%, and increasing to 4 cores cut it nearly in half.
On the other hand, quantization (converting weights to INT8) made things about 2x faster, but the accuracy on the 6-choice task dropped by 1.4 to 3.1 points, so we did not adopt it. Exporting to ONNX Runtime (fp32) kept the answers unchanged, but the speed only changed by 5–8%. Under our conditions, what actually improved speed was increasing the number of CPU cores.
What Didn't Work? The Limits of Synthetic Data
The numbers so far may look promising. However, the improvements from synthetic data eventually hit a plateau, and a gap remained between the data and actual posts that adding more data could not close.
Approaches We Tried and Rejected
The approaches that seemed promising but were ultimately not adopted taught us more than the ones we did adopt.
We had assumed that adding the channel name and description as input would provide more grounding and improve the 6-choice accuracy. The result was the opposite: accuracy dropped from 0.787 to 0.706. Even for posts whose correct answer was "not applicable," the model increasingly misclassified them by selecting the channel's typical task—the model had become overly reliant on context.
We also found that the correct labels for borderline posts differed between the training data and the evaluation set. For English queries that simply ask for information, such as "does anyone know …," 59% were labeled as requests in the training data, whereas only 33% were labeled as requests in the evaluation set. We tried relabeling the entire training dataset using the same model used for the evaluation set (at a cost of about $26), but overall accuracy did not improve. This turned out to be less a training issue than a product decision that needed to be made first: whether we want the AI employee to treat "does anyone know…?" as something to pick up on.
We also tried other approaches, such as routing English-only posts to the English version of Laya, softly assigning correct answers as probabilities, and excluding ambiguous examples—but none of these outperformed the final version. Among the six training runs using different synthetic data generation methods, none exceeded the previous best by more than 1 point.
The Gap Between Synthetic Data and Real Posts
Even though the accuracy for the Japanese "request" category was 0.92 on the synthetic data evaluation set, it dropped to 0.80 on actual posts from internal chat (comparison made at an intermediate stage of training). The gap was especially pronounced for posts like "I fixed it, so please check" in development channels.
To isolate the cause, we fixed the candidate options and compared using our own custom sentences. For a sentence like "I fixed the bug on the login screen, so please check," the model, even after training, selected "bug fix" with a probability of 0.98 or higher. However, for a sentence like "I made fixes in 4 places, so please check," which lacks a term indicating the specific task, the model selected "not applicable" after training. Before training, the model had inferred "bug fix" from the word "fix" alone; training shifted it toward a more cautious judgment that avoids selecting an answer unless there is textual evidence. In actual development channels, however, context makes it unnecessary to explicitly state the task name in the text. Humans and LLMs can infer this and assign the correct label, but Laya—which only looks at the text itself—has no such clue to rely on.
Adding 189 lines of examples reflecting how development channels are actually written raised the 6-choice accuracy on real posts from 0.818 to 0.851. Improvements from synthetic data alone hit a ceiling, and what's needed to improve further is real posts. However, since our company has a policy of not using customer data for training, how much we can gather from internal posts alone remains our next challenge.
What to Verify Before Adopting a Classification-Only Model
A dedicated classification model is effective in situations where the volume of judgments is high, labeled data can be prepared, and LLM latency is a problem. Conversely, if even one of these conditions is missing, it may be more rational to continue querying the LLM directly, as is currently done. Before deciding to adopt this approach, check the following six points.
- Does the judgment volume exceed the break-even point? With our query format, Claude Haiku cost about $0.45 per 1,000 requests. This is on par with a 4-core verification server (about $189/month) at around 420,000 requests per month (both are reference values based on unit prices at the time of writing). When volume is low, querying the LLM directly is cheaper and requires less operational effort.
- Can you prepare labeled data? Before training, Laya could not be used as-is for our own queries. Both training and evaluation require data labeled in the format of your own queries.
- Can you build an evaluation set separate from the training data, at a scale of about 300 items per language? With only 100 items, we could not distinguish between random fluctuation and actual improvement.
- Have borderline decisions been settled as a product matter? Whether "does anyone know…?" counts as a request or not changes what the correct answer actually is.
- Are you trying to reduce p50 or p95? Which one you target changes how you combine the model with the LLM and what proportion of traffic you route through each.
- Have you checked the terms of service? If you're using LLM output to train another model, make sure the LLM's terms of service don't restrict this kind of training use.
The criteria for deciding whether to handle fine-tuning in-house are covered in Introduction to Fine-Tuning.
Frequently Asked Questions
We answer the questions that often come up when considering the introduction of Jev and Laya, to the extent clarified through this evaluation.
Q1: Can Laya (Multilingual Version) Be Used As-Is for Japanese Classification?
It could not be used as-is. The accuracy of Japanese "request" before training was 0.47, and it rose to around 0.9 after training. Re-determining thresholds with in-house data or fine-tuning is a prerequisite.
Q2: How Should You Choose Between Jev and Laya? Differences in Pricing and Fine-Tuning
The choice depends on whether you want to train on your own data. Since Jev is hosted, it requires no effort for training or server operation. In exchange, per-customer fine-tuning is not possible, and adjustments must be made through the question instructions and candidate descriptions. The official documentation also recommends testing with your own text, since accuracy is not equivalent across non-English languages.
Laya, on the other hand, runs on your own server and can be trained without sending data outside, but preparing training data, evaluation, and operation become your own responsibility. A GPU is not required; in our evaluation, the in-server p50 was about 310ms on a 4 vCPU server. If Japanese or Thai is central and the classification boundaries are company-specific, Laya is realistic; if English is central and you want to try something quickly, starting with Jev is realistic.
Summary
Jev and Laya are models that give up the text generation LLMs excel at, focusing instead on classification speed and ease of handling probabilities. In our evaluation, Laya before training could not be used as-is for our own questions, but after fine-tuning with synthetic data, the accuracy of "request" across all 4 languages rose to about 0.90, and response time became shorter than that of LLMs. While accuracy does not reach that of top-tier LLMs, routing only the roughly 20% of ambiguous posts to an LLM keeps p50 nearly the same as Laya alone while achieving accuracy close to that of an LLM.
What became clear through this evaluation is that accuracy is determined less by model selection than by how evaluation and ground truth are constructed. Separate evaluation from training data, compare at a scale of around 300 cases, and decide the boundary judgment as a product decision beforehand. Once this is in place, whether to introduce a dedicated classification model can also be decided based on numbers. For the overall approach to building small models, please also see our SLM Distillation Guide.
We also support the design and PoC of AI-based classification processing. Please reach out via AI/DX Support for consultations.
References (as referenced at the time of writing): Laya's GitHub repository and Hugging Face model card, and TypeSafe AI's official documentation. All are linked within the body text.
Author & Supervisor
Yusuke Ishihara
Started programming at age 13 with MSX. After graduating from Musashi University, worked on large-scale system development including airline core systems and Japan's first Windows server hosting/VPS infrastructure. Co-founded Site Engine Inc. in 2008. Founded Unimon Inc. in 2010 and Enison Inc. in 2025, leading development of business systems, NLP, and platform solutions. Currently focuses on product development and AI/DX initiatives leveraging generative AI and large language models (LLMs).

