Most conversations about automated decision making in hiring start at the wrong end. They start at the hiring decision, which is the one moment a human is almost always present. By then, software has already made a series of smaller decisions about who was visible, who was ranked, who was parsed correctly, and who quietly stopped moving.
That gap became the subject of a useful debate among talent practitioners recently, prompted by a LinkedIn post from Glen Cathey mapping automated decision making across the full recruiting and hiring lifecycle. The responses to it were as instructive as the post, and they converged on something worth writing down: the question is not whether you use AI in hiring. It is where you have handed judgment to a system, and whether anyone can still describe the rule it follows.
What Counts as Automated Decision Making
The definition matters more than the label on the software. As Cathey puts it:
"Automated decision making is a computational process that makes a consequential decision about a person, or that substantially replaces or materially influences the human who does."
Nothing in that requires artificial intelligence. A default sort order is an algorithm making a decision about who is number ten and who is number forty-eight. A knockout question is an algorithm deciding who moves forward. Resume parsing that silently drops a certification is making a decision about a person's qualifications. None of these would be described as AI by the teams running them, and all of them determine who advances.
The boundary is not the technology. It is the moment a system starts influencing who gets seen, who gets filtered out, and who gets an opportunity. That has a practical consequence worth stating plainly. "We don't use AI" is the most common reason a talent team gives for never examining any of this, and it is not an answer to the governance question. It describes what a vendor calls its product. It says nothing about how decisions actually get made in your funnel.
Definitions do differ across regulations in ways that matter legally, and the terms are not interchangeable: an Automated Employment Decision Tool under NYC Local Law 144 is not the same as an automated-decision system under California's civil rights regulations, or Automated Decision-Making Technology under its privacy rules. Our guide to automated employment decision tools covers how each framework draws its line.
The Map Is the Easy Part
Listing where automated decisions happen takes an afternoon. Deciding which ones to govern first is the harder problem, and it is the one the conversation tends to stop short of.
Industry analyst Chris Havrilla of AI Alchemists frames the underlying issue as one of consequence. Her Systems of Consequence™ work centers on who owns the consequence of a decision, where human judgment actually sits, and whether the impact can be traced after the fact. Applied to a hiring funnel, that resolves into three questions we use to rank what needs attention now against what can wait.
- Consequence. Does the system remove someone from consideration, or only reorder them? Removal is final. Reordering is survivable, until it isn't, because nobody scrolls to the bottom of a ranked list.
- Visibility. Does the decision leave a trace that anyone reviews? An auto-rejection is logged. An ad that was never served to a segment leaves nothing behind at all.
- Ownership. Can you name the person who set the rule and the date they set it? If not, nobody is watching it drift.
Run the funnel through those three and the priorities invert. Knockout questions feel like the biggest risk because they are the most obviously decisive, but they are also visible, logged, and easy to test, which makes them a quick win rather than an emergency. Ad delivery is the opposite: equally consequential, almost entirely invisible, and owned by a media budget rather than by talent. And the sleeper is the default sort order in your applicant tracking system, which shapes what every recruiter sees, leaves no record, and in most organizations has no owner at all.
Triage also means deciding where automation should stay. Parsing and knockout questions have largely earned their place, with two conditions: the knockout questions have to match the role's real minimum qualifications, and someone has to keep checking what the parser captures. Earned is not the same as unexamined.
Search, matching, and ranking are a different case. Deciding what matters for a role, and therefore who is returned and in what order, is a judgment the people accountable for the hire should own. Leaving it to automation is more efficient, but as Cathey put it, "I don't think efficiency in hiring is really what any company's business strategy is founded on." Business results come from having the right people in the right roles. Speed matters. Quality of hire matters more.
Tina Shah Paikeday, Responsible AI Senior Advisor at Findem, a Warden AI customer, draws the same line from the vendor side of the problem:
"When AI is part of hiring, it's important to understand how it's behaving in real use, not just how it was designed."
That gap between design and behavior is what testing exists to close. An inventory of intended rules is not the same thing as evidence of what the system did.
Warden AI's audit data supports treating this as a testing problem rather than a technology problem. Across more than 150 AI system audits covering over a million test samples, analyzed in our State of AI Bias in Talent Acquisition research, automated systems measured up to 45% fairer for racial minorities by impact ratio than the human-led decisions they replaced. Tina Shah Paikeday has argued the same point from her own analysis, that structured processes can outperform intuitive human judgment on bias. Automation is not the problem. But 85% of systems tested met fairness thresholds, which means roughly 15% failed at least one demographic threshold, and a system failing a threshold nobody tests is indistinguishable from one that passes.
Where Automated Decision Making Operates, Stage by Stage
Here is the same funnel most talent teams work in every day, read as a sequence of automated decisions rather than a sequence of tasks.
Advertising: who sees the job at all
Before a single application exists, automation has decided the size and shape of the audience. Which boards receive budget, how programmatic job advertising platforms deliver impressions, which audience segments the posting reaches and which it never reaches. The decision here is not about a candidate's qualifications. It is about whether a qualified person ever learns the role exists, and it leaves no trace in the applicant tracking system.
Sourcing: who surfaces in the search
Search relevance ranking decides which profiles appear on the first screen. "Recommended candidates" features decide which names get pushed toward a recruiter unprompted. Talent pool segmentation decides who belongs to which group, and silver medalist resurfacing decides who gets a second look from a prior req. Each of these is recruiting automation shaping the pool before a human forms any opinion.
Cathey singles out ranking as the decision that moved from human to system most quietly. With basic keyword search, the results tracked what the searcher explicitly asked for. Modern matching interprets that input and expands it with inferences that can go well beyond what the recruiter intended. Many tools go further and retrieve and rank candidates from the job description alone, a document the recruiter often didn't write and that is rarely a fully accurate picture of the role. Unless the recruiter adds their own view of what matters most, and actually uses that ability, the system is deciding who is returned and in what order.
That is the sourcing version of the default sort problem: a rule nobody chose, applied to every search. Vendors building in this space are starting to test it directly. Juicebox built its AI sourcing platform around that principle, and Findem put its applicant-matching through an independent audit examining whether candidates were evaluated against uniform standards and how outcomes distributed across protected demographic groups.
Applying: knockout questions and what parsing drops
This is the stage most people recognize as automated, and still underestimate. Knockout questions and auto-rejection on minimum qualifications make explicit, immediate, and usually final decisions. Resume parsing makes a quieter one: what the parser captures becomes the candidate's record, and what it fails to capture effectively did not happen. A misread certification, a non-standard degree format, or an unusual job title can erase a qualification the person actually holds. Duplicate detection decides whether two records are one person, which sounds administrative until it merges or discards the wrong history.
Screening: scores, rankings, and the default sort
Match scores and fit scores assign a number to a person. Ranked applicant lists turn that number into an order. Auto-scored assessments generate their own numbers on top. And then there is the most overlooked piece of candidate screening automation in the entire stack: your applicant tracking system's default sort order. Nobody chose it deliberately, nobody documented the rule behind it, and it determines what a recruiter sees in the first screen of a list they will realistically never scroll to the end of. Where a vendor can explain its scoring, that changes the conversation entirely, which is what SquarePeg set out to prove with explainable, continuously assured models.
Interviewing: scheduling, transcription, and scorecard rollups
Who receives a scheduling link first is a decision. So is how transcription and summarization tools represent what a candidate said, since the summary, not the conversation, is often what a hiring manager actually reads. Structured scorecard rollups apply weighting to individual interviewer inputs and produce a composite, and that weighting is a rule someone configured, whether or not anyone remembers configuring it. Interview platforms carrying this load can and do submit to independent testing, as Alex.com has with its AI interviewing product.
Selecting: comp engines, adjudication matrices, and routing
Compensation recommendation engines produce an offer range from comparators and rules. Background check adjudication matrices decide which records disqualify and which do not, applying a policy consistently but not necessarily correctly. Approval routing decides which offers require which sign-offs, which shapes how much scrutiny a decision receives on the way out.
Onboarding and beyond: routing, classification, and mobility
Automated decision making does not stop at the offer. Task routing decides what a new hire encounters and when. Worker classification decides employment status, with consequences reaching well beyond the hiring team. Internal mobility and promotion recommendations decide who is surfaced for the next opportunity, which is the same ranking problem as sourcing applied to people who already work for you. This is where hiring automation becomes employment automation, and the same governance questions follow it. Agentic systems that act across several of these steps at once raise the stakes again, which is why Jack & Jill stress-tested its agentic AI for bias risk before scaling it.
The Exercise: Count Them on a Real Req
Cathey's post closes with a test more revealing than any framework: take one requisition you have worked recently and count how many of the items above touched it before a human formed an opinion. Put in front of a leadership team, it lands hardest as a single number, which is how many automated steps stood between a candidate and their first human reader.
- How was the audience for the posting determined, and by what?
- What ranked the sourcing results you reviewed, and in what order did they arrive?
- Which questions could auto-reject, and what were the thresholds?
- What did the parser capture, and did anyone check what it missed?
- What was the default sort on the applicant list, and who set it?
- What weighting produced the composite scorecard?
- Which of those rules can you describe from memory, and which would you have to go and look up?
Most teams find the count is higher than expected and the documentation thinner. That gap is the finding. It is not evidence that anything went wrong. It is evidence that nobody would know if it had.
A Human in the Loop Is Not the Same as a Human in Charge
The standard reassurance is that a human makes the final decision. In most funnels that is true, and it answers less than it appears to. Martyn Redstone, Head of Responsible AI at Warden AI, puts the question the count exposes:
"Yes, a human makes the final hiring decision, but can that human stand behind that decision? Do they genuinely know how those last few candidates got there, how they were filtered, screened, ranked and assessed by systems in place? If the human making the final decision cannot defend the decision and the entire decision making process behind it, then is the human really making the decision, or just rubber stamping the AI decisions?"
The test is not whether a person is present at the end. It is whether that person could explain every step that produced the shortlist in front of them. Regulation is moving the same way: Colorado's SB 26-189 gives candidates a right to meaningful human review of adverse decisions, and a review only means something when the reviewer can see what the system did.
Where Regulation Actually Reaches
A reasonable question at this point is how much of this the law already covers. The honest answer is a narrow slice, unevenly.
- NYC Local Law 144 covers automated employment decision tools and is the only US framework that mandates a bias audit by name.
- Illinois prohibits discriminatory use of AI in employment decisions and requires notice when it is used.
- California's privacy rules attach notice, access, and opt-out duties to automated decision-making technology.
- Colorado's SB 26-189, which takes effect January 1, 2027, centers pre-use notice and human review of adverse decisions.
Our multi-state AI hiring compliance guide tracks what applies where, and it moves often enough that a static summary in an article like this one would be wrong within months.The differences between those frameworks are real, technology recruiter Brian Fink made the point that the differences are not the part most teams should be planning around:
"Here's what the compliance crowd keeps whiffing on: the regulations don't agree with each other. Colorado, the EU, New York, they draw the lines in different spots, and those differences will absolutely matter the day the lawyers show up. But they rhyme. Every one of them lands in the same zip code: if the software determines who advances, it counts. Nobody has to utter the magic letters A-I for the statute to bite."
That is the operating principle. Wait for a definition to settle and you will be waiting through several legislative sessions. Build to the shared premise, that software determining advancement is in scope, and the jurisdictional detail becomes a documentation question rather than a strategy question.
It also exposes the gap. Most of the automated decisions listed above sit outside any named mandate. Ad delivery, sourcing rank, default sort order, parsing fidelity, and scorecard weighting are rarely the explicit subject of a statute, and they are precisely the decisions with no human in the room. If your program only covers what a law names, it covers the smallest part of your exposure.
From Awareness to Evidence
Awareness is the first step and not the last one. An inventory tells you what is running. It does not tell you what rule each system follows, whether that rule does what you assume, or whether you could demonstrate any of it to someone who asked.
The progression is straightforward. Inventory every stage where software makes or materially influences a decision about a person. Document the rule behind each one, including who set it and when it last changed. Test the ones that affect who advances, across the demographic groups that matter, starting with the high-consequence and low-visibility decisions the triage surfaces. Then keep the record, because a rule correct at configuration can drift as data, roles, and vendors change.
Independent testing is what turns an inventory into evidence. Warden AI's independent bias audits examine these systems as hiring tools rather than as products, and a versioned audit trail preserves what was tested, what was found, and what changed, so the answer to "how does your process work" is a record rather than a recollection. It is also what lets the person making the final call defend it. It is the principle Warden AI was founded on: systems that shape people's careers must be subject to meaningful accountability.
None of this requires treating automation as a threat. Every system in the funnel embeds an assumption about which signals predict success, and most of those assumptions were set carefully by people trying to hire better. The audit data says these systems often produce fairer outcomes than the processes they replaced. The point is narrower and harder to argue with: you cannot govern what you have not noticed.
Ready to Find Out What Is Already Deciding for You?
Building an inventory takes an afternoon, but discovering an ungoverned rule after a candidate complaint is a severe risk. A knockout threshold nobody set deliberately, a parser dropping a credential, or a default sort nobody documented can shape who advances for years without producing a single visible error. Mapping these decisions now, then testing the ones that decide who moves forward, tells your team what is actually running and gives you something to show when a regulator, a client, or a candidate asks. Warden AI's independent AI bias audits and continuous assurance cover automated decisions across the hiring lifecycle, not only the tools that carry an AI label. Schedule a consultation to map the automated decisions running in your hiring process.
Related Articles
Automated Decision Systems in Hiring: Frequently Asked Questions
What is automated decision making in hiring?
It is a computational process that makes a consequential decision about a person, or that substantially replaces or materially influences the human who does. In hiring that covers ad delivery, sourcing rank, knockout questions, resume parsing, match scores, default sort order, scheduling priority, scorecard weighting, compensation recommendations, and internal mobility suggestions.
Does automated decision making require AI?
No. A rules-based knockout question, a filter, and a default sort order are all automated decision making, and none involve machine learning. The test is whether software determines who advances, not what technology sits behind it. This matters because teams routinely exempt themselves from scrutiny on the grounds that they "don't use AI."
If a human makes the final hiring decision, is the process still automated?
Usually, yes. The final decision is made from a shortlist that automated systems filtered, screened, and ranked. The useful test is whether the person making the call can explain and defend how each candidate reached them. If they cannot, the human is approving the system's decision rather than making their own.
Where should humans stay in the loop?
Search, matching, and ranking are the strongest case. Parsing and knockout questions tied to real minimum qualifications have largely earned their place. Deciding what matters for a role, and therefore who is returned and in what order, is a judgment the people accountable for the hire should own rather than delegate to a job description.
Which stages of hiring are actually regulated?
Only a narrow slice. NYC Local Law 144 is the only US framework mandating a bias audit by name, and it applies to covered automated employment decision tools. Illinois prohibits discriminatory use and requires notice, California attaches privacy duties to automated decision-making technology, and Colorado's SB 26-189 adds notice and human-review duties when it takes effect in January 2027. Ad delivery, sourcing rank, and default sort order are rarely named in any of them.
Which automated decisions should we govern first?
Rank them by consequence, visibility, and ownership. Decisions that remove candidates rather than reorder them, that leave no reviewable trace, and that no named person owns should come first. By that test, ad delivery and default sort order usually outrank the knockout questions most teams start with.
Should we govern automated decision making if our jurisdiction has no ADM law?
Yes. Anti-discrimination law applies regardless, private claims do not depend on an ADM statute, and enterprise buyers increasingly ask how automated hiring decisions are made. Regulation determines your documentation obligations. It does not determine whether an automated decision can disadvantage someone.



