A week in, JEV gets competition
JEV now has challengers: KEV runs locally, and Span-01 does well on yes/no questions. I tested all four models on insurance submission evidence.

Summary
Barely a week ago, I was getting to know JEV and OpenRouter's new decisioning category. Now we have actual competition.
The competitors are different, but have their advantages. KEV has the same output style as JEV, the same pricing per token, and was slower and less accurate on my benchmark. It ended up slightly cheaper, for an interesting reason I'll get to below. BUT: it's open weights. You can run it locally on consumer hardware.
Span-01 isn't as broad through OpenRouter: it answers yes/no questions. That's it. It's “Bit” from Tron. But it does them well, does them cheaply, and while not quite as wicked fast as JEV, is less spendy. Worth looking at if your questions are yes/no, or can sensibly be rephrased that way.
Span-01 Lite is free, but struggled on this test. It may only accommodate easier use cases. At roughly half a second per request, it remains fast compared with the conventional LLMs I've tested, and the price makes it worth experimenting with.
JEV remains my strongest overall performer here. But a week in, there are already some interesting choices.
What I tested
I've been looking at how these models can surface evidence for insurance underwriting. What does the application actually say? What does the website advertise? Who did the work? Do the supporting records agree?
For this round, I used 180 fictional contractor submission cases grounded in 30 archived business websites. Each website profile supplied six scenarios, including straightforward submissions, understated work, subcontracted work, revised scope, and missing or conflicting evidence.
These are document-evidence questions, rather than another round of picking workers' compensation class codes. An application might describe related work while a completed job record describes the target activity. Or a website might advertise a service without establishing that the applicant's employees performed it on the job being reviewed.
I ran two comparisons:
- JEV vs. KEV: ten questions per case, covering yes/no answers, choices, and scores. That's 1,800 answers per model.
- All four models: the three existing yes/no questions from each case, giving 540 answers per model. I reused JEV and KEV's saved answers and ran both Span models on those questions.
That distinction matters when reading the hero: JEV's rounded 95% and KEV's 88% are their full-benchmark results. Span's 92% and Lite's 65% cover the yes/no subset. The charts below keep those comparisons separate.
KEV: the alternative you can run locally
First there's KEV. JEV to KEV, we see what you did there.
Jared Palmer's KEV-4B is built on Qwen3.5-4B-Base, with a trained adapter and a pointer head. It takes evidence and typed questions, and returns probability distributions without generating a prose answer. So while its foundation is a conventional language model, this is more than asking a chatbot to behave itself and return JSON.
That doesn't settle whether it's good at my use case. Testing does.
JEV got 1,715 of 1,800 answers correct. KEV got 1,575. JEV was about 2.8 times faster by median successful response time.
KEV did have a win within the test: it was better on the evidence-strength question, getting 156 of 180 right versus JEV's 141. The overall winner doesn't necessarily win your particular question.
I also passed the answers through the same procedural rules. JEV's answers led to the expected action on 180 of 180 cases; KEV's did on 144. KEV's 36 misses were unnecessary holds. Neither model's answers caused the rules to incorrectly clear a case.
That is a limited submission-checking rule, not a policy acceptance decision. JEV still made 85 individual answer errors; correct routing doesn't make those answers right.
Same token price, different bill
Both providers charged $0.042 per million input tokens. I sent identical evidence and questions to each model.
But JEV reported 471,926 input tokens, while KEV reported 361,194. KEV's reported bill was therefore about 23.5% lower: roughly 8.4 cents per 1,000 cases, versus 11 cents for JEV.
Different tokenization or provider formatting could explain that. The logs don't establish which. What they do establish is that the same submitted text did not produce the same billed token count. Comparing the price per million tokens alone would have missed it.
And yes, it really runs locally
I downloaded KEV and tested it on both my CPU and an AMD Radeon RX 6800 XT, a consumer GPU with 16 GB of video memory. A newer iPhone might be next, but mine's a 13.
The model files took about 9.5 GB of disk space. The GPU run used about 10.2 GiB of GPU memory at peak. On eight cases with ten questions each, the CPU configuration took a median 66 seconds per case; the initial GPU configuration took 10 seconds. After closing background apps, I reran the GPU tests and got substantially better times; the updated table is at the end of this post. The GPU run matched all 80 answer labels from the hosted KEV run on those cases.
This was a working local setup, with reference-kernel fallbacks, not a fully optimized serving deployment. The hosted API was much faster. But the point is that you can host KEV yourself, and keep the submission evidence on your own hardware.
If you're going via API and want speed, JEV is your friend. If you want local, KEV is the option among these models. It may be accurate enough for your use case, and it's small enough to test without a data center.
Span-01: good for yes/no use cases
Then there's Respan's Span-01, from a company focused on agent evaluation and observability. It evaluates specified behaviors against supplied evidence. Respan's native interface has probabilities for present, absent, and not observable; the OpenRouter interface I tested exposes a single probability of yes.
That OpenRouter interface rejected the choice and score questions, so I used only the three compatible questions. I kept their wording and evidence unchanged, apart from serializing the evidence object into the text format Span requires.
Span-01 landed between JEV and KEV on this subset. Lite was quite a bit weaker.
The errors are more useful than the overall score:
| Model | Correct answers | False positives | False negatives |
|---|---|---|---|
| JEV 1.13 | 513 / 540 | 27 | 0 |
| KEV 4B | 453 / 540 | 87 | 0 |
| Span-01 | 497 / 540 | 43 | 0 |
| Span-01 Lite | 350 / 540 | 141 | 49 |
JEV, KEV, and Span-01 each got 180 of 180 subcontractor-disclosure questions right. Span-01 Lite got 153.
The harder question was whether supporting completion records made unresolved, incompatible claims about the same job. JEV got 154 of 180 right, Span-01 147, KEV 129, and Lite 127.
For example, an application can understate the work while two completion records agree about what actually happened. That is a discrepancy between the application and the records. It does not mean the completion records contradict each other. Span-01 sometimes flagged the latter anyway.
This is where I want more experiments: clearer source boundaries, narrower questions, and testing whether wording improvements carry across models. I haven't optimized these questions specifically for Span.
All fairly fast. All fairly cheap.
Across the full 180-case yes/no run, Span-01's median completion time was 0.494 seconds, and Lite's was 0.508 seconds. Span-01's reported charge for the entire run was $0.00734. Lite reported no charge. Both returned valid answers on all 180 requests without retries.
For a fairer comparison of cost and speed across all four, I had eight saved JEV/KEV cases where the same three questions had been asked on their own. Here is that small matched subset:
JEV is still the speed winner. Span-01 is cheaper than either JEV or KEV on these calls. Lite is free, though it was no faster than paid Span in this test.
These are all small model bills. I would choose based on what the model gets right, what happens when it gets something wrong, and whether I need to run locally. The price difference is worth measuring, but it wouldn't be my first concern here.
What I take from this
I started with JEV because it felt like a different way to put models into insurance workflows. Now there are alternatives, just a few days into exploring it.
JEV gives me the best combination of accuracy and speed in this test. KEV gives me local deployment. Span-01 gives me another capable, inexpensive option for yes/no evidence questions. Lite may fit simpler questions, and costs very little to try.
My recommendation keeps coming back to the same thing: develop and test on the individual use case, and measure against your own trade-offs. KEV winning the evidence-strength question is a useful reminder of why.
And for insurance, I still prefer to think of these as evidence-surfacing models. The model tells me what the evidence supports under the question I asked. The underwriting platform decides whether that means ask another question, request records, refer, or decline.
I'm interested in how much more of that evidence we can afford to check, how quickly, and where we can run the checks. This is already getting more interesting.
Methodology and limits
The cases combine archived public website evidence with synthetic applications and supporting records. They are not actual customer submissions, and no ACORD PDFs were parsed. The 180 cases reuse six scenario variants per source business; 1,800 answers are not 1,800 independent risks.
Answer keys were frozen before these runs and remain provisional, without independent reviewer signoff. I made no prompt or label changes during the scored runs. Yes/no answers use a 0.5 threshold; choice answers use the selected category; score questions are graded against the most likely rubric category. These are answer-accuracy measurements, not probability-calibration claims.
JEV and KEV's yes/no answers were reused from ten-question requests. Span received three questions per request. The eight-case cost/speed chart uses the same three-question requests across models, measured on different dates. The local KEV test is a separate eight-case smoke test, not a repeat of all 180 cases.
Timings are successful-request completion times, excluding evidence gathering, model loading, queue waits, and retry backoff. The JEV/KEV run encountered throttling and other failures; its charted prices are reported successful-request model charges and exclude unpriced failures. Two initial Span compatibility probes returned HTTP 400 without usage information; the 360 scored Span requests all succeeded. Local hardware and electricity costs were not measured.
The hero rounds JEV/KEV's full-benchmark results and Span's yes/no results. The detailed charts show more precise figures. Span's native three-way uncertainty output was not tested through OpenRouter.
Models and earlier experiments
- JEV 1.13 on OpenRouter
- KEV-4B model card and open weights
- KEV source code
- Span-01 documentation
- Span-01 on OpenRouter and Span-01 Lite
- Earlier: JEV benchmarking learnings, ask away
- Earlier: getting JEV to implementable
Update: KEV on consumer hardware
I also tried KEV locally at 16-bit, 8-bit, and 4-bit precision. It works well on consumer hardware: these runs used my AMD Radeon RX 6800 XT with 16 GB of VRAM, in a desktop with an Intel Core i5-12600K and 64 GB of RAM.
I closed background apps and reran the same eight packets, with ten questions per packet. These are the updated results:
| Precision | Median per packet | Correct answers | Peak GPU memory |
|---|---|---|---|
| BF16 (16-bit) | 5.87s | 90.00% | 10.20 GiB |
| HQQ 8-bit | 4.97s | 91.25% | 6.99 GiB |
| HQQ 4-bit | 4.68s | 87.50% | 5.40 GiB |
The 8-bit version looks like a useful balance here. The 4-bit version saves more GPU memory, but was only a little faster and lost three correct answers compared with 8-bit. It also caused one unnecessary referral. All three repeated their earlier answers exactly.
Memory includes more than the compressed weights. GPU figures are peak PyTorch allocations during inference and exclude host RAM, operating-system and driver overhead. This setup loads the original model before quantizing, so startup still needs roughly 8.5 GiB of GPU memory for the quantized versions. The table is not a minimum-device-memory specification.
CPU only works too. My earlier FP32 CPU run took a measured median 66.2 seconds per packet, with 88.75% correct answers and 21.1 GiB peak host RAM. That is a different precision and runtime configuration, not an estimate for 4-bit or 8-bit CPU inference. Slow, but usable for experiments or work that can wait.
And the iPhone idea? An iPhone 17 Pro is a candidate to test at 8-bit, and an iPhone 16 at 4-bit, with a mobile implementation. The 17 Pro has 12 GB of RAM, while the iPhone 16 family has 8 GB. I haven't run KEV on either. Desktop GPU memory doesn't translate directly into an iOS app budget, and Apple says additional app memory is not guaranteed. A compatible mobile runtime, direct loading of compressed weights, and an on-device test are still needed before I can say it fits.
These are small local checks: 80 answers per configuration, with one warmed timing pass. Loading and warm-up are excluded. The HQQ backend and reference kernels are not a fully optimized deployment. But the practical point holds: KEV gives me a way to check insurance evidence on hardware I already own, while keeping that evidence local.
Addendum: I also tried Laya
I also tested Laya, another open-weight model in this space. It needs more work for this use case. JEV and KEV worked well out of the box on these insurance questions; Laya did not.
I checked the implementation, tried three checkpoints, rephrased questions, and gave it shorter, more focused evidence. The best result across those exploratory checks was 46 correct answers out of 80, using eight submission packets. JEV had answered 78 correctly and KEV 72 on those same packets with the original inputs. The Laya adaptations changed the inputs, so this is a diagnostic comparison rather than another entry in the main benchmark.
Why keep an eye on it? Ryan Porter's Anthus study explains how fine-tuning can help Laya. On a constructed sentiment dataset, training on 140 labeled examples brought Laya to 89.6% accuracy. That is a useful reason to consider a small local model when you have your own labels and are prepared to train it. The study also found changes to answers outside the training task, so those need testing too.
Could Laya ever be a better choice than KEV? Its strongest argument is size: Laya's checkpoints have 322 to 421 million parameters, compared with KEV's 4 billion. If a trained Laya could match KEV on a narrow, recurring task, that much smaller footprint could make it an attractive choice at high volume. I haven't demonstrated that accuracy or measured a speed advantage on matched hardware.
Fine-tuning and keeping data local are options with KEV too. To establish an accuracy advantage, I'd need to give both models the same training examples and evaluate them on an untouched test set. The evidence so far supports considering Laya as a small specialist you can train, with no demonstrated advantage over KEV on my submission-review questions.
For my insurance workflow, Laya would need more development and possibly fine-tuning. I haven't tested that, and I don't know whether it would close the gap. For now, I'm parking it and continuing with the models that worked well without task-specific training.
With all the excitement about JEV, it feels like every lab is looking at this. These results are a snapshot in time. By the time you read this, they're almost guaranteed to be out of date. But don't worry, I'll keep testing.
Credit to Astra for helping build, run, and analyze the experiments, and put this write-up together.


