It's Not a System 1 Model If It Takes Ten Seconds
An agent is about to act, and it asks a decision model one question: approve, escalate or reject. A decision model is supposed to answer in a fraction of a second, with a probability you can set a threshold on, so the software keeps moving. Jev, the decision model we use most, answered our benchmark's questions in a median of about 180 milliseconds. On October 1, Fastino Labs released GLiDE, which it calls the first thinking decision model, and reported that it beats Jev. We gave GLiDE the same 3,600 reasoning problems. Its median answer took a second, 243 answers took ten seconds or more, and the slowest took two minutes and fifteen seconds.
In Daniel Kahneman's Thinking, Fast and Slow, System 1 is fast, automatic judgment and System 2 is slow, deliberate reasoning. TypeSafe calls Jev a System One model, one "built to make fast, structured decisions that software can use directly." Jev "generates all outputs in a single query," in parallel, with no text generated token by token, and TypeSafe puts its response time at 70 to 500 milliseconds. Its /v1/systemone endpoint takes a text and a list of typed questions and returns answers with probabilities.
Fastino is upfront that GLiDE works differently. It calls GLiDE a thinking decision model, and says its API "conforms to the System One model schema": GLiDE answers on the same endpoint, with the same request and response, and Fastino publishes a guide to migrating from TypeSafe. That made our test simple, because GLiDE got exactly the request Jev gets. It also makes a critical difference easy to overlook if you don't evaluate your options carefully. The interface is the same, but what happens behind it isn't. A model that sometimes stops to reason for a minute is a System 2 model behind a System 1 interface.

Time per decision on the same 3,600 problems. Jev was timed one request at a time, GLiDE with eight requests in flight.
GLiDE's times form two humps. About three quarters of its answers come back in half a second to two seconds. The other quarter take three seconds to two minutes, and those are GLiDE thinking. Fastino describes the design in its launch post: "Traditional decision models score a set of options in a single fixed pass." GLiDE instead "first produces a fast probability distribution, then allocates additional reasoning when the leading result is uncertain," so it can "spend more computation on difficult decisions while keeping straightforward ones fast." That's the pattern usually called test-time compute: spending more time and money at answer time to get a better answer. Fastino's API documentation recommends a read timeout of at least 300 seconds.
The extra work shows up on the bill. Fastino charges $0.30 per million input tokens, and its documentation says the reported input tokens sum the model's internal passes. The slow requests reported about 2,000 input tokens each against about 290 for the fast ones. Across the 3,600 problems, at list prices, GLiDE cost $213 per million decisions and Jev $23: 9.3 times as much, and 10.4 times as much per correct answer, because GLiDE was also less accurate. The cost breakdown has every model with an API price.
One caveat on the timing. We sent GLiDE eight requests at a time, because sending them one at a time would have taken more than three hours, so some of its fast-path time may be queueing on Fastino's side. That doesn't explain the second hump: those requests carried seven times as many tokens.
Deeper problems cost GLiDE more time and money
If the extra time is reasoning that gets billed, time and cost should rise together on the problems GLiDE finds hard. They do: across all 3,600 requests, GLiDE's billed tokens and its response time correlate at 0.94.

Average time and list-price cost per decision by proof depth, about 300 problems per bar. Jev's times are from a one-at-a-time rerun.
Jev does the same work on every problem: about a fifth of a second and $23 per million decisions, whether the answer is written in the text or needs five chained rules. On the open-world task GLiDE spends more as problems get deeper, from $97 per million decisions and 1.1 seconds on average at depth zero to $391 and 6.6 seconds at depth four. On the closed-world task it spends about the same, $160 and 2.4 seconds, at every depth past zero, and that's the task where its accuracy fell furthest.
A tie on one task, a 17-point loss on the other
Hard-Decisions is our benchmark of decision models on multi-step logic. Each problem is a short list of facts and if-then rules from ProofWriter, a statement, and a question: does the statement follow? Proof depth counts how many rules you have to chain to answer it, from zero (the answer is written down) to five. In the open-world task the answer is true, false or unknown. In the closed-world task anything you can't prove is false, so the answer is true or false. Every model answers every problem once, with the same wording and no examples. The GPT-6 Luna write-up walks through a problem at each depth.
| Model | Open world | Closed world |
|---|---|---|
| Jev | 83.8% | 89.3% |
| GLiDE | 82.9% | 71.8% |
| GPT-6 Luna, reasoning off | 64.1% | 65.0% |
| Kev-9B, the best open model | 58.6% | 63.8% |
| Laya | 42.4% | 55.6% |
We expected GLiDE to beat the open models, and it did, by more than 20 points on the open-world task. There it also matched Jev: 0.9 points behind on the same 1,800 problems, inside the margin of error, and level at every depth up to four.
The closed-world task went the other way. Jev was 17.5 points ahead (95% interval 15.4 to 19.6), and the gap grew with depth: about 20 points at depths two to four and 40 points at depth five, where GLiDE got 49% of the true-or-false questions right. A coin flip gets 50%.

Accuracy by proof depth, about 300 problems per point. The dashed line is chance: one in three for open world, one in two for closed world.
Fastino's claim comes from its own Decision Index, a suite of 38 benchmarks on which it reports GLiDE at 64.81 and Jev at 57.91. Hard-Decisions isn't in that suite, and on Hard-Decisions GLiDE didn't beat Jev.
GLiDE's open-world answers also lean toward "unknown." It got 96% of the unknown problems right but only 76 to 77% of the true and false ones, and 282 of its misses were true or false statements it called unknown.
When GLiDE says 80 to 90%, it's right 73% of the time
A decision model's probability is what lets you route: act on the confident answers, send the rest to a person. So we asked of GLiDE what we asked of GPT-6 Luna: when it states a probability, how often is it right?

Each point is a band of stated probability; the dashed diagonal is where a perfectly calibrated model would sit. Both tasks pooled.
GLiDE's least confident answers were its most accurate. The 339 answers it stated at 50 to 70% were right 93% of the time. The 1,278 answers it stated at 80 to 90%, its most common band, were right 73% of the time. Its stated probability only matches its accuracy at the very top: answers at 99% or more were right 98% of the time, but only 85 of its 3,600 answers were stated that high.
That shape matters more than the average. The standard test is AUROC: how well the stated probability separates right answers from wrong ones, where 0.5 means it carries no information and 1.0 means it separates them perfectly. GLiDE scored 0.56 on the open-world task and 0.52 on the closed-world one. Jev scored 0.85 and 0.86. Even GPT-6 Luna's log-probabilities scored 0.66 and 0.68. We'd predicted in writing that GLiDE would beat Luna on this, and it didn't.
In practice: if you only let answers stated at 90% or more through, Jev passes 70% of its decisions and gets 95.9% of those right. GLiDE passes 26% and gets 84.5% right, close to its 77% overall.
Its least sure answers are the ones it thought about
The fast and slow humps explain the strange shape. We split GLiDE's answers at two seconds:
| Answers | Right | Average stated probability | |
|---|---|---|---|
| Open world, fast | 1,157 | 79% | 85% |
| Open world, slow | 643 | 91% | 72% |
| Closed world, fast | 1,559 | 68% | 86% |
| Closed world, slow | 241 | 95% | 76% |
The thinking works. When GLiDE stopped to reason, it was right 91 to 95% of the time, and it reported lower probabilities for those answers. The problem is how it decides when to think: by its own first probability, the same probability that can't tell its right answers from its wrong ones. On the closed-world task its fast answers averaged 86% stated confidence and were right 68% of the time, so it thought on only 13% of the problems and answered the rest quickly and often wrong.
You can see it in the timing by depth. On the open-world task GLiDE took more than two seconds on 3% of depth-zero problems and on 56% of depth-four ones, and its median time climbs from 0.9 to 5.0 seconds, so it notices that deep problems are hard. On the closed-world task it took more than two seconds on 15 to 16% of problems at every depth from one to five, and its median stays at 0.9 seconds throughout. It treats a five-step true-or-false question as if it were a one-step one. That's the difference between tying Jev and losing to it by 40 points.
We think the fix GLiDE needs is the one Luna needs: a probability that's higher when the answer is right. GLiDE's design depends on that more than most models do, because the same probability decides when to spend the extra time.
What we'd tell a team choosing a decision model
GLiDE is better at multi-step reasoning than any open model we've tested, by more than 20 points on the open-world task, and its thinking answers were right more than nine times in ten. But it isn't real competition for Jev. It's less accurate, by 17.5 points on the closed-world task. It's slower, five times at the median, with a tail that reaches minutes. Its confidence barely separates its right answers from its wrong ones. And it costs about ten times as much per decision.
The reason is the design Fastino chose. A decision model competes with Jev where it can sit inside every step of an agent and answer in a fraction of a second, which is what a System One model is for. A model that sometimes takes a minute to respond isn't a System One model, however its API is shaped. Fastino says plainly that GLiDE thinks; the risk is for teams that see the same endpoint and the same request format and assume they're getting the same kind of model.
- Plan for the slowest answers. A one-second median hides a 6.8% chance that a decision takes ten seconds or more, and an agent makes many decisions per task.
- Check the probability against labeled answers before you set a threshold. On our problems, GLiDE's high numbers weren't more likely to be right than its middling ones.
- Test the question formats you'll actually use. The same logic asked as true, false or unknown and as true or false gave GLiDE a tie with Jev on one and a coin flip on the other at depth five.
- Price the thinking. Its slow answers bill about seven times the tokens of its fast ones. Here GLiDE cost 9.3 times what Jev did per decision, and more on the harder problems.
How we measured
GLiDE got exactly Jev's request through Fastino's /v1/systemone endpoint, model fastino/GLiDE, one request per problem, no examples and no retries of an answer. We wrote down four predictions before the first GLiDE request: it beats the best open model overall and at depths three to five, its accuracy falls with depth, and its probability separates right from wrong better than Luna's. The first three held; the last failed. We skipped the rerun we'd planned, so there's no measure of how often GLiDE changes its answer when asked twice, and its times come from the scored run with eight requests in flight. Two requests returned server errors and succeeded on the retry. GLiDE launched the day before we tested it, and Fastino may change it. ProofWriter is synthetic, templated logic, so these numbers describe multi-step reasoning of that kind. The comparison, the preregistration and every record are on hard-decisions.anth.us.
If you're choosing a decision model, run this kind of test on your own decisions: labeled cases from your own work, hard ones included, every model asked the same way, accuracy by difficulty, the slowest answers as well as the median, and a check that each model's probability separates its right answers from its wrong ones. That's work we do with teams.