
A deployed decision includes the choices people made around the model: what it reads, where its threshold sits, and what it's allowed to do.
An engineer ships a model that sorts job applications. A few weeks later, someone runs a paired test: same résumé, same experience, one protected cue changed. The recommendation changes more often than ordinary noise checks say it should.
The model didn't choose its job. It didn't choose the applicant pool, the question, the score that moved someone forward, or the production threshold. It can't answer a customer, a reporter, an applicant, or a regulator. The engineer and the company that put it between a résumé and a decision can.
FTC Chair Andrew Ferguson made the same point at a Reuters event in September. He rejected the idea that an AI agent which causes harm has become an independent actor with its own will. If someone tells a tool to do something and it does it, he said, the question is about the person who gave the instruction. Reuters reported that he pointed to the people who instruct the agent, and to audit trails that showed supposedly uncontrolled systems carrying out the instructions they'd been given.
It was a chair's remark, not a new FTC rule. But it names a problem engineers already have: responsibility for a decision stays with the people who deployed the system that makes it.
The decision system is bigger than the model
A decision model can classify a support ticket, rank a candidate, route a claim, flag an account, or choose who gets a closer look. The model may be one API call. The decision system around it includes everything that turns that call into an outcome for a person.
Someone chose the model version. Someone wrote the rubric or prompt. Someone selected the data and features. Someone set the confidence threshold, shortlist size, escalation path, and permission to act. Someone described the product to buyers. And someone kept the system running after the first strange result.
Those choices are why "the model made the decision" is a thin answer. It describes the mechanism and leaves out the person or organization that made the mechanism consequential.
Federal agencies don't all regulate the same workflow or use the same legal test. But they've repeatedly refused the idea that software complexity creates a zone where nobody is responsible.
The CFPB says a creditor using a complex model still has to give accurate, specific reasons for an adverse credit decision, and can't cite the model's complexity as the reason it can't explain. The EEOC argued that a vendor whose software screens and refers candidates can be doing the work of an employment agency. And DOJ and HUD have said that housing providers and tenant-screening companies are still accountable when algorithmic screening disproportionately denies people housing.
The statutes differ, and the lesson for engineers is the same under each of them: if your system makes a regulated decision, "it was an algorithm" doesn't end the inquiry.
How an outside finding becomes your operating problem
Here's how it usually goes. A civil-rights organization publishes a report. A customer forwards it to compliance. Someone posts two nearly identical applications that landed on different sides of a screen, and the post reaches more people than the model's documentation ever will.
None of that requires anyone to prove a legal case on social media. It needs a credible piece of evidence: a repeatable input change, an observed output change, a named version, and a plain description of what the system was allowed to decide. Affected people can bring it to an agency. Organizations can help file or support charges. Buyers can ask a vendor to account for a claim that sounded settled in a demo.
Two cases show how that plays out.
Meta had to rebuild how it delivered ads
Civil-rights groups including the ACLU, the National Fair Housing Alliance, and the Communications Workers of America challenged Facebook's systems for targeting housing, employment, and credit ads. Facebook announced settlements and product changes in 2019. The issue then moved from the press to the regulators. HUD filed a Fair Housing Act charge, and in 2022 DOJ's settlement required Meta to stop using its Special Ad Audience tool for housing ads and develop a Variance Reduction System to address disparities caused by personalization algorithms.
That's an engineering consequence. The company had to retire a tool and build a new system that changed how it delivered ads. The DOJ case record shows that the consequential decision lived in the ad-delivery optimization, where no buyer ever clicked a box to make it.
HireVue dropped a feature before any enforcement order
In 2019, the Electronic Privacy Information Center filed an FTC complaint over HireVue's use and description of facial analysis in automated hiring. HireVue later said it had removed visual analysis from new assessment models in 2020, announcing the change in 2021 and attributing it to research on predictive value and improving natural-language processing.
The available record doesn't say the FTC compelled that change. It shows a common kind of pressure: a product feature gets hard to defend once an advocacy organization, applicants, researchers, and customers are all asking what it measures, whether it's fair, and why it's necessary. For a hiring-tech founder, that pressure can change the roadmap.
"Fair" is a factual claim
The FTC has a direct route to the company marketing a decision product. Claims such as "bias-free," "validated," "fair," "compliant," or "98% accurate" are factual claims about what a system does, and the company making them needs evidence that matches each one.
The FTC's Workado action wasn't a hiring case. It concerned an AI-content detector advertised as 98% accurate. The agency alleged the claim was unsupported and required competent, reliable evidence for future effectiveness claims. The same principle applies to every product page that turns a limited benchmark, a friendly pilot, or an internal audit into a broad promise.
Don't call a system fair because nobody has found the failure mode yet. Say what you measured: which version, which decision, which population, which counterfactual, which control, and which result.
Build the record before someone else does
Our Biased-Decisions benchmark changes one protected signal on the same text, keeps everything else fixed, and compares the effect with an equally trivial edit that changes no protected signal. It reports the result in terms of the decision: a verdict flip, a probability shift, or a shortlist ratio.
A benchmark can't decide whether a hiring, housing, lending, or benefits system has broken a law. What it can give you is a versioned, replayable record of what the system did when a cue changed, which is worth more than any general promise of responsible AI.
If the effect clears the control floor, the next questions are operational. Can you name the model version? Can you reproduce the input and output? Can you show the threshold that turned a score into a decision? Can you show who reviewed the result, what the system was allowed to do while it was under review, and what changed afterward?
If you can answer those questions, you find the failure mode in your own testing. If you can't, you find out from a viral screenshot, a customer escalation, or an agency demand.
We build that record for clients
The paired test and the audit trail are the small, reproducible version of a loop we've run in production for years. We've run it for a call-center QA operation across hundreds of scorecards and millions of interactions: model versions tracked, reviewer disagreements logged with the reason, and each new version scored against held-out human judgments before it was promoted. The case study shows what that looks like at scale.
Plexus is where we run it. It keeps the per-slice checks, the approval gate, and the record of who approved what, so the questions above have answers on file before anyone asks them.
If a model is about to make decisions about people on your behalf, we'll run the paired test on your data and build the record around it. Here's how an engagement works.
Sources
- Reuters: FTC chair suggests AI developers should be liable for conduct of agents
- Consumer Financial Protection Bureau: Circular 2022-03
- EEOC: amicus brief in Mobley v. Workday
- Department of Justice: United States v. Meta Platforms
- EPIC: In re HireVue
- HireVue: decision on visual analysis
- Federal Trade Commission: Workado order
- AnthusAI: Biased-Decisions benchmark
We do this for clients, at scale
Anthus has run this kind of loop in production for years: reviewers correct the model and say why, the explanation becomes policy, and the system gets more trustworthy month over month across hundreds of scorecards and millions of interactions. Bring us the judgment task and we'll run it on Plexus with your reviewers in the loop, and hand you a scorecard you can inspect after the first month.
See how an engagement works