

180 billion tokens of production LLM workload. AI in production, not just in demos.
What We Do
![Knowledge bases with learned ontologies and taxonomies]()
Knowledge Bases
Your agents are only as good as what they can look up. We build knowledge bases with ontologies and taxonomies that learn from your data, refining their own structure instead of going stale the week after someone hand-built them.
![Unattended automation with a human in the loop]()
Self-Aligning Automation
Unattended business process automation with a human in the loop. Reviewers correct it and say why; the system turns the explanation into a stated policy and applies it from then on.
![Agent systems, tools, and orchestration]()
Agent Systems
A model is half a system. The other half is the harness: the tools it can reach, the procedures it follows, the limits it runs inside, and the ability to work for hours without losing the plot.
![Custom and fine-tuned machine learning models]()
Machine Learning
Custom classifiers, fine-tuned models, and calibrated confidence that tells you which decisions to trust and which to escalate. We find the cheapest model that clears your bar, prove that it clears it, and run it in production on AWS—training, serving, and evaluation included.
Clients usually arrive asking about one of these. The work rarely stays in one box.
Case studies
Call Criteria
- Call center QA, scored by human reviewers.
- Their QA couldn't scale without scaling headcount.
- We built a self-evolving RLHF system: reviewers correct the AI and say why, and it turns the explanation into policy.
- 100% of calls reviewed, up from a sample.
Venue Driver
- Ticketing and reservations backbone for Las Vegas nightlife.
- An AWS data center failed catastrophically.
- We relocated the entire system within hours — ticket scanning never stopped.
- In continuous operation since 2007.
How we work
Everything we ship runs under specs, tests, staged rollout, and a person who can say no. We call that cybernetic development: we use AI to write code the same way we use it to classify calls, inside a governor of constraints, feedback loops, and human judgment that keeps systems reliable in production.
Modern failures increasingly look less like isolated “bugs” and more like operational, multi-system breakdowns. Great unit tests help—but they don’t cover every emergent scenario. So we build layered defenses and close the loop with real-world feedback.
- Specs first: define behavior before implementation.
- Defense in depth: sandboxed tools, CI gates, staged rollouts, and fast rollback.
- Operational feedback: telemetry and incident-driven regressions that tighten the loop over time.
- Simplify and delete: reduce degrees of freedom to eliminate entire classes of failure.
The Anthus Platform
Solve complex business problems with AI and ML using a proven, reusable technology stack that grew out of real delivery work — runtime, agent execution, knowledge, observability, and media, with the enterprise controls that matter in production.
Recent Articles
![Decision Models Are Not Calculators]()
Decision Models Are Not Calculators
2026-09-25Jev chose one in every ordering of a fair-die question; Kev leaned toward one, and Laya overwhelmingly chose six. Benford's Law remains a hypothesis.![Can You Fix It? Gating, Averaging and Fine-Tuning Against a Gendered Verdict]()
Can You Fix It? Gating, Averaging and Fine-Tuning Against a Gendered Verdict
2026-09-22The verdict moves on "she". We put a gate on the learning loop that rejects any question whose answer moves with the pronoun. It kept out the obvious ones and still made a hiring shortlist worse, because a question can ignore the pronoun and read gender from the words. Here's what worked instead.![We Told the AI She Was a Woman. It Demoted Her]()
We Told the AI She Was a Woman. It Demoted Her
2026-09-22A new kind of AI makes instant decisions about people: who gets an interview, whose claim gets flagged. It's cheap and fast, and companies are adopting it without checking what it does. We checked. It quietly ranks women lawyers below men, and every company using it makes the same mistake.![The One-Word Test: How Jev and Laya Read Gender, Race and Age]()
The One-Word Test: How Jev and Laya Read Gender, Race and Age
2026-09-22Two thousand real bios per job, first names blanked, one pronoun swapped, both models asked twice. Laya's verdict moved on 8 to 18 bios in 100, Jev's on 1 to 4, almost always toward the stereotype. Here is the method, the jobs, the shortlist, and what names and ages do.











