Benchmarks
We publish two public benchmarks for decision models, the fast AI models that answer a bounded question with a verdict. Each one shows its method, every model's results, and the cases behind them, so you can check a model against the decision you're about to hand it.
Biased-Decisions
Does an AI model's answer change when one personal detail about a person changes?
Would AI identify a different job for the same person if “he” became “she”? Would it remove the same online comment after its author said they were gay? Biased-Decisions changes one personal detail in a text, asks the same question again, and ranks models by how much more often the answer changes than it does after a neutral control edit of the same size.
Hard-Decisions
How far can a decision model follow a chain of reasoning before it starts guessing?
On true-or-false questions, every model tested except Jev falls to coin-flip accuracy by five chained inferences: the open decision models Kev, GLiNER2.5-Decide and Laya, GPT-6 Luna used as a one-shot classifier, and GLiDE. Jev still answers 89.3% correctly at that depth.
The articles that explain these findings are on the research page, and the guide to aligning a decision model to your own data is here.
We do this for clients, at scale
Anthus has run this kind of loop in production for years: reviewers correct the model and say why, the explanation becomes policy, and the system gets more trustworthy month over month across hundreds of scorecards and millions of interactions. Bring us the judgment task and we'll run it on Plexus with your reviewers in the loop, and hand you a scorecard you can inspect after the first month.
See how an engagement works