

180 billion tokens of production LLM workload. AI in production, not just in demos.
What We Do
![Knowledge bases with learned ontologies and taxonomies]()
Knowledge Bases
Your agents are only as good as what they can look up. We build knowledge bases with ontologies and taxonomies that learn from your data, refining their own structure instead of going stale the week after someone hand-built them.
![Unattended automation with a human in the loop]()
Self-Aligning Automation
Unattended business process automation with a human in the loop. Reviewers correct it and say why; the system turns the explanation into a stated policy and applies it from then on.
![Agent systems, tools, and orchestration]()
Agent Systems
A model is half a system. The other half is the harness: the tools it can reach, the procedures it follows, the limits it runs inside, and the ability to work for hours without losing the plot.
![Custom and fine-tuned machine learning models]()
Machine Learning
Custom classifiers, fine-tuned models, decision models like Jev, and calibrated confidence that tells you which decisions to trust and which to escalate. We find the cheapest model that clears your bar, prove that it clears it, and run it in production on AWS—training, serving, and evaluation included.
Clients usually arrive asking about one of these. The work rarely stays in one box.
Case studies
Call Criteria
- Call-center QA, scored by an outside service's fine-tuned classifiers.
- Too expensive and too slow to scale to every new scorecard.
- We built their classifier lab on Plexus in their own AWS account and ran it for years.
- Hundreds of scorecards, millions of calls scored.
Venue Driver
- Ticketing and reservations backbone for Las Vegas nightlife.
- An AWS data center failed catastrophically.
- We relocated the entire system within hours — ticket scanning never stopped.
- In continuous operation since 2007.
How we work
You might be wondering whether AI-built means nobody checked. For us it means the opposite, and we learned it the hard way: in February 2026 the nightly routine was pasting "Continue." into four Codex sessions before bed, because nothing else kept the agents going or told us what they'd done. That night is on the record. Everything we ship now runs under specs, tests, staged rollout, and a person who can say no. We call that cybernetic development: we use AI to write code the same way we use it to classify calls, inside a governor of constraints, feedback loops, and human judgment that keeps systems reliable in production.
Modern failures increasingly look less like isolated “bugs” and more like operational, multi-system breakdowns. Great unit tests help—but they don’t cover every emergent scenario. So we build layered defenses and close the loop with real-world feedback.
- Specs first: define behavior before implementation.
- Defense in depth: sandboxed tools, CI gates, staged rollouts, and fast rollback.
- Operational feedback: telemetry and incident-driven regressions that tighten the loop over time.
- Simplify and delete: reduce degrees of freedom to eliminate entire classes of failure.
The Anthus Platform
Solve complex business problems with AI and ML using a proven, reusable technology stack that grew out of real delivery work — runtime, agent execution, knowledge, observability, and media, with the enterprise controls that matter in production.
Recent Articles
![It's Not a System 1 Model If It Takes Ten Seconds]()
It's Not a System 1 Model If It Takes Ten Seconds
2026-10-02Fastino's GLiDE beats every open model on our reasoning benchmark, but not Jev: less accurate, five times slower, ten times the cost, and a confidence that barely tells right from wrong.![Can Averaging Reduce Noise in an AI Decision Model?]()
Can Averaging Reduce Noise in an AI Decision Model?
2026-10-01We asked Jev the same question ten times in one request and expected ten identical answers. They split, four votes to six. Kahneman, Sibony and Sunstein call that kind of variation noise. Averaging the copies didn't make Jev more accurate, but it cut changed answers on a rerun by about a third.![The OpenAI Decisions API Needs a Confidence You Can Trust]()
The OpenAI Decisions API Needs a Confidence You Can Trust
2026-10-01We've routed decisions on LLM confidence since 2023, and calibration is what we most hope OpenAI's new Decisions API gets right. The model underneath, GPT-6 Luna, isn't there yet: on 3,600 reasoning problems its 99% meant 68%, and by five inference steps it was at a coin flip.![Plexus Is a Classifier Lab: Operationalizing Classifiers at Scale]()
Plexus Is a Classifier Lab: Operationalizing Classifiers at Scale
2026-09-27A hosted decision model sat at 77 percent accuracy through 87 rounds of reviewer feedback, then one approved question took it to 87. Here is the loop that did it, every figure published, and the offer to run it on your judgment task.











