Jev vs Laya: Testing Decision Models on Configuring Search
Should this field match by sound?
At Relevan’s core is a tuning machine we call ASRE (Automated Search Relevance Engine). This is our secret sauce for tuning the relevance of each customer’s indexes. Not every customer needs every relevance feature, so ASRE’s core job is deciding which features to enable, disable, or tune to make search more relevant.
When TypeSafe released Jev, we immediately realized that it could seriously simplify how ASRE picks which features to enable. Right now we use complicated heuristics, or even LLM calls that can fail or have output parsing issues. With Jev (and similar models), you get a discrete output that matches the exact data shape you need.
To test it, we decided to use a straightforward case. We have a phonetic analyzer that we can enable for customers. The analyzer helps when people search for things they can only spell phonetically. Take, for example, the name Caitlin, which has over a hundred recognizable spellings . As long as you can spell it phonetically, Relevan can find it.
We tested three models on this task, and we measured two things: how often the model was correct, and how long each answer took. The eval is open source at relevan-dev/decision-model-eval so you can reproduce or modify it for your own experiments.
The test
We created 10 sample indexes that are similar to some we’ve seen at Relevan. Each index has a mapping, a short description, and some sample values for each field. They include recruiter candidates, medical providers, English parish registers, and application logs. There’s a nice mix of fields that can benefit from phonetic matching and ones that don’t.
For each index, a search engineer wrote a reasonable configuration, which we use as the label. Then we asked two questions about each of the 35 text fields:
- Should this field get phonetic matching? The model answers with a probability.
- Which encoder should it use:
double_metaphone,metaphone, orsoundex?
Here are the three models we tested:
- Laya, an open-weights model that answers typed questions. We ran it on a CPU on our own machine.
- Jev, the hosted model from TypeSafe. Laya and Jev use the same request format, so both got the exact same requests.
- Claude Opus 5, through the Claude API. Claude gets the same state and questions as a prompt, and it replies with structured output that matches a schema built from those questions. We set the effort to
low, since this is a short classification task.
We sent the requests one at a time, and each model got one warm-up request before we started the clock.
The results
| Laya | Jev | Claude Opus 5 | |
|---|---|---|---|
| Enable decision correct | 29 of 35 | 35 of 35 | 35 of 35 |
| Precision | 71% | 100% | 100% |
| Recall | 100% | 100% | 100% |
| Encoder correct | 9 of 15 | 15 of 15 | 15 of 15 |
| Whole index config correct | 4 of 10 | 10 of 10 | 10 of 10 |
| Brier score (0 is perfect) | 0.173 | 0.020 | 0.010 |
| Median time per request | 1.10 s | 0.16 s | 2.27 s |
| 95th percentile time | 1.14 s | 0.24 s | 3.42 s |
| Time for the full test | 38 s | 6 s | 84 s |
Jev and Claude both nailed it. Every field, every encoder, all ten indexes. Claude’s probabilities were a little sharper than Jev’s. Fields that needed phonetic matching came back close to 1, and fields that didn’t came back close to 0. That’s what the Brier score captures. Lower is better, and 0 means every probability was spot on.
Funny enough, the first mistake we found was ours. In an earlier run, Jev and Claude both picked double_metaphone for the English parish registers. Our label said metaphone, since every name in that index is English. We assumed the narrower encoder would mean fewer false matches, but we’d never actually tested that. Double Metaphone handles English names just fine, so we changed our rule: use double_metaphone unless there’s a compatibility requirement or a measured improvement. All the results in this post use the updated labels.
The real difference is speed. Jev answered in about 160 ms, which is fast enough to run while a customer is still creating their index. Claude took about 2.3 seconds per answer. That’s still fine for a background job, and we didn’t have to train or host anything. But it’s roughly 14 times slower than Jev for the same answers.
Laya struggled. It caught every field that needed phonetic matching, but it also turned it on for six fields that didn’t, like summary, email, and cuisine. Most of its probabilities landed between 0.5 and 0.75, which isn’t much better than a coin flip. It also picked soundex for four indexes that had no reason to use it.
Architecting with decision models
One request per field is the obvious approach, but it isn’t the only one. We wanted to know if grouping the questions differently would get us better answers or fewer requests, so we tried four setups.
- Per field. Each field gets one request that asks both questions.
- Encoder first. The encoder usually applies to the whole index, so one request per index picks it. Then each field request only asks whether to enable it.
- Index gate. One request per index asks whether any field holds names. If not, we skip every field in that index.
- Gate and encoder first. One index request asks the gate question and the encoder question together.
| Architecture | Requests | Laya: enable, encoder, configs | Jev and Claude | Jev total | Claude total |
|---|---|---|---|---|---|
| Per field | 35 | 29 of 35, 9 of 15, 4 of 10 | all correct | 5.9 s | 84 s |
| Encoder first | 45 | 29 of 35, 8 of 15, 3 of 10 | all correct | 7.4 s | 111 s |
| Index gate | 38 | 27 of 35, 7 of 12, 3 of 10 | all correct | 6.4 s | 102 s |
| Gate and encoder first | 38 | 27 of 35, 7 of 12, 3 of 10 | all correct | 6.7 s | 121 s |
A quick note on the table: when the gate closes an index, its fields never get an encoder question. That’s why Laya’s encoder score is out of 12 instead of 15 for the two gated setups.
For Jev and Claude, the setup didn’t matter. They gave the same answers every time, and only the number of requests changed. Claude’s total time also jumped around between runs, with its median time per request ranging from 2.3 s to 3.1 s.
With Laya, every alternative made things worse. Encoder first spread soundex across entire indexes instead of just a few fields. The gate correctly shut off the application logs, which fixed one wrong field. But it also shut off three indexes that do hold names, and Laya’s recall dropped from 100% to 80%.
The gate didn’t save us any time either. It only pays off when a lot of indexes don’t need phonetic matching. In our set, eight of ten indexes hold names, so the gate added more index requests than it saved in field requests. This is probably something we should have modeled by keeping the distribution of phonetic enabled indexes the same as in production but that’s work for a later date.
Shaping the state
Laya’s biggest improvements didn’t come from the architecture. They came from how we worded the request. We made two changes during development (against an earlier version of the labels), and both made a big difference.
Describe when to use it, not what it does. Our first encoder options described what each encoder does, like “short English surnames only; it keeps the first letter and is coarse.” That’s a great description for a search engineer, but the model has nothing in the index description to match it against. Laya picked soundex for 27 of 28 fields. So we rewrote each option as the situation that calls for it: “another system already runs a soundex index over the same records, and this index must return the same results.” Laya’s encoder accuracy jumped from 21% to 71%.
Show the model sample values. Our index request originally included only the field names. With just that, Laya picked the right encoder 57% of the time. Adding a few sample values per field bumped that to 79%, and whole configs correct went from 5 to 6 of 10.
Here’s what a single field request looks like:
{
"index": "restaurant-listings",
"index_contents": "Restaurant listings for a city guide. Diners search for a place a friend recommended out loud, so they rarely have the spelling.",
"field": "restaurantName",
"declared_type": "text",
"search_capabilities": ["searchable", "autocompletable"],
"analyzer_language": "none",
"sample_values": ["Pizzeria Bianco", "Le Bernardin", "Sqirl"]
}And here’s the index request, including the sample values that made the difference:
{
"index": "restaurant-listings",
"index_contents": "Restaurant listings for a city guide. Diners search for a place a friend recommended out loud, so they rarely have the spelling.",
"fields": {
"restaurantName": "Pizzeria Bianco, Le Bernardin, Sqirl",
"neighbourhood": "Kreuzberg, Shoreditch, Ravenswood",
"cuisine": "italian, vietnamese, seafood",
"phone": "+1 212 554 1515"
}
}What this means for Relevan
Every feature in ASRE comes down to two questions: should this be on for this customer, and how should it be set? Today, answering that takes heuristics, LLM calls we have to babysit, and a search engineer double-checking the result. Phonetic matching was our first test, and honestly, it went better than we expected.
The part I’m most excited about is the probabilities. We already tune with a human in the loop. A well-calibrated probability tells that person which decisions are easy and which ones aren’t. With a model like Jev making the first pass, they can skip the fields the model is sure about and spend their time on the handful it isn’t.
Next up, we want to put our labels to the test with real spoken-name queries and count how many false matches each encoder actually produces. This is how we can ultimately be certain that we’re producing a net-positive effect.
Small note: the day before we were set to publish this, Contrastive-LM (a collaboration between Nvidia and Stanford) released CLM-8B . As of this writing, its performance wasn’t good enough to include in this post, but we’ll be keeping our eyes open for new open models.
Want to try this with your own use cases? Clone decision-model-eval , write one function that answers the typed questions, and see how it stacks up against Laya, Jev, and Claude (and other models we add).
And if you’d rather have search that tunes itself, try our free tier, or book a call to see what ASRE can do for your data.