Articles

  • It's Not a System 1 Model If It Takes Ten Seconds

    October 2, 2026

    Fastino's GLiDE beats every open model on our reasoning benchmark, but not Jev: less accurate, five times slower, ten times the cost, and a confidence that barely tells right from wrong.

  • Can Averaging Reduce Noise in an AI Decision Model?

    October 1, 2026

    We asked Jev the same question ten times in one request and expected ten identical answers. They split, four votes to six. Kahneman, Sibony and Sunstein call that kind of variation noise. Averaging the copies didn't make Jev more accurate, but it cut changed answers on a rerun by about a third.

  • The OpenAI Decisions API Needs a Confidence You Can Trust

    October 1, 2026

    We've routed decisions on LLM confidence since 2023, and calibration is what we most hope OpenAI's new Decisions API gets right. The model underneath, GPT-6 Luna, isn't there yet: on 3,600 reasoning problems its 99% meant 68%, and by five inference steps it was at a coin flip.

  • Plexus Is a Classifier Lab: Operationalizing Classifiers at Scale

    September 27, 2026

    A hosted decision model sat at 77 percent accuracy through 87 rounds of reviewer feedback, then one approved question took it to 87. Here is the loop that did it, every figure published, and the offer to run it on your judgment task.

Earlier