
JEV: easy, fast, cheap, accurate
A 50-case workers’ compensation classification test: JEV had the highest accuracy, second-lowest cost, and fastest reviewed workflow.
Essays on underwriting, insurance products, AI systems, and the economics of putting them into production—plus occasional notes on the technology and policy shifts changing the industry around them.


A 50-case workers’ compensation classification test: JEV had the highest accuracy, second-lowest cost, and fastest reviewed workflow.

The result was interesting. How it happened is more important.
I gave low-cost models actuarial methods, Python tools, operating rules, native tool delivery, and answer cleanup, then measured accuracy, completion, cost, and latency on repeated insurance tasks.

Chinese labs and hybrid systems are catching the flagship US models on the coding leaderboards. Washington is slowing its own labs at the same time.

A cheap, open, self-hostable model reads real insurance forms at ~96% for about $1.50 per 1,000. What makes it deployable isn't the model — it's the model-and-harness pairing, guided by a human expert who knows which errors are safe, which to refer, and which are dangerous.

InsureBench V1 is live — a public benchmark scoring model-plus-harness combinations on insurance work. The headline finding: the best model depends on the job, and a low-cost model beat the frontier on actuarial tasks at roughly 1/400th the cost.

Early InsureBench results show why insurance AI evaluation needs to measure model plus harness performance, cost per task, and task-specific reliability. The charts are preliminary, but the deployment question is already clear.

Labs is for things that work: prototypes, demos, and tools that make the argument clickable. Two are live today, including an agentic Workers' Comp underwriting workstation.

Dario and Elon don't exchange birthday cards. They just signed a 220,000-GPU deal anyway. The math, the digs, and what could still go wrong.

A model is the talent. The harness is the operation. A tiny model went from 60% to 94.8% on MATH-500 with nothing but harness changes. The harness is where the leverage is.

A 15-cent AI matches a $40 one on simple tasks and dies on hard ones. The harness, not the model, is usually where the gains live.

Claude Mythos Preview prices intelligence at a level where reading every word of every document for every account starts to pencil out. Insurance should pay attention.

Manus AI positions itself as the first truly autonomous AI agent. If DeepSeek and ChatGPT are like McKinsey, Manus demonstrates the ability to actually get things done.

As AI gets dramatically cheaper, Jevons' Paradox predicts usage will grow even faster. What happens when underwriting intelligence stops being scarce?

Unlike traditional AI chatbots that focus on fluency and creativity, Deepseek R1 prioritizes logical consistency and structured reasoning. That has big implications for insurance.

Generative AI is reshaping underwriting. Here are practical approaches for using it without losing control.

AI agents go beyond chatbots. They interact with external systems, perform multi-step tasks, and represent a natural evolution of automation in insurance.