Research
Jev, Laya and the rest of the field-coverage roster answer millions of bounded questions a day, so we run the experiments that check what they're actually doing: where their verdicts move on a name or a pronoun, how well their confidence tracks reality, and what survives when you fine-tune, distill or gate them. Every piece here comes with the method and the numbers, not just the headline.
![Decision Models Are Not Calculators]()
Decision Models Are Not Calculators
September 25, 2026Jev chose one in every ordering of a fair-die question; Kev leaned toward one, and Laya overwhelmingly chose six. Benford's Law remains a hypothesis.
![Can You Fix It? Gating, Averaging and Fine-Tuning Against a Gendered Verdict]()
Can You Fix It? Gating, Averaging and Fine-Tuning Against a Gendered Verdict
September 22, 2026The verdict moves on "she". We put a gate on the learning loop that rejects any question whose answer moves with the pronoun. It kept out the obvious ones and still made a hiring shortlist worse, because a question can ignore the pronoun and read gender from the words. Here's what worked instead.
![We Told the AI She Was a Woman. It Demoted Her]()
We Told the AI She Was a Woman. It Demoted Her
September 22, 2026A new kind of AI makes instant decisions about people: who gets an interview, whose claim gets flagged. It's cheap and fast, and companies are adopting it without checking what it does. We checked. It quietly ranks women lawyers below men, and every company using it makes the same mistake.
![The One-Word Test: How Jev and Laya Read Gender, Race and Age]()
The One-Word Test: How Jev and Laya Read Gender, Race and Age
September 22, 2026Two thousand real bios per job, first names blanked, one pronoun swapped, both models asked twice. Laya's verdict moved on 8 to 18 bios in 100, Jev's on 1 to 4, almost always toward the stereotype. Here is the method, the jobs, the shortlist, and what names and ages do.
Earlier
![Distilling an Aligned Jev System into a Classifier You Own]()
Distilling an Aligned Jev System into a Classifier You Own
A Jev scorer aligned with 140 human labels taught a 66M-parameter DistilBERT that scored 0.912 against the human label, two points above its teacher, at 5.6 to 15 ms an item. The accuracy is the least general part; the per-slice ship gate is the part to copy.
![Fine-Tuning Jev: You Can't. Here's What Gets You the Same Effect]()
Fine-Tuning Jev: You Can't. Here's What Gets You the Same Effect
You can't fine-tune Jev: TypeSafe serves the same weights to everyone. So we kept it frozen and adapted the questions and a small fitted head around it. On 140 labels, one new plain-English question bought 10 points of accuracy, and it finds the hidden pattern about a quarter of the time.
![Jev vs Laya: Same Labels, Same Questions, One Variable]()
Jev vs Laya: Same Labels, Same Questions, One Variable
We ran Jev and the open-weights Laya through one harness where only the engine changes. Jev led by 4.7 points alone and 6.8 with our feedback layer, and then fine-tuning Laya on the same 140 labels beat both.
![Can You Trust Jev's Confidence?]()
Can You Trust Jev's Confidence?
Jev's confidence values are useful: higher really does mean more likely right. But on 8,801 labeled examples they ran overconfident, and how you phrase the request changes the pattern. Calibration fixes most of it.
![Making Decisions Instead of Generating Text]()
Making Decisions Instead of Generating Text
Jev turns a scorecard’s bounded questions into the model’s native output—making high-volume QA faster, cheaper, and easier to operate.
![The Turn Detection Trap: When 100% Accuracy is Wrong]()
The Turn Detection Trap: When 100% Accuracy is Wrong
We fine-tuned MobileBERT for turn detection and achieved 100% accuracy. Then we realized we hadn't solved turn detection at all—we had just built a very expensive punctuation detector.
![The Dominance of Ones: A Handy Quirk of Numbers]()
The Dominance of Ones: A Handy Quirk of Numbers
From money launderers to breeding rabbits, this fascinating quirk of numbers is lurking on the sidelines, ready to surprise us.










