The demo goes fine until you ask what the number means. The rep says the model is 94% accurate at identifying strong candidates. Accurate against what? Agreement with recruiter decisions on a held-out set. So the product reproduces the choices your recruiters already make. That is not the same as finding better people, and if your recruiters ever advanced a weaker candidate for a bad reason, the model learned that too, at volume.
Most searches for the best AI recruiting software want a ranked list of vendors, and a ranked list would not help you. “AI recruiting” covers at least six unrelated products whose risk levels differ enormously, and the legal exposure here sits with the employer rather than the vendor: a tool can work exactly as advertised and still be unlawful in how you use it. That inverts normal software buying. Compliance and validation are not procurement hygiene for after you choose. They are the product questions.
The definition to hold onto: an automated employment decision tool is software that uses statistical modelling, machine learning or AI to substantially assist or replace human judgement about who gets hired or advanced, and under US employment law the employer using it answers for its effects, not the company that built it. General information, not legal advice.
Six products wearing one label
Buyers compare a scheduling assistant against a video scoring engine as though they carry the same risk. They do not.
| Category | What it does | Who really decides | Risk |
|---|---|---|---|
| Sourcing and outreach | Finds profiles, writes and sequences messages | A human, on who to contact | Low to moderate. Risk is in targeting, not scoring |
| Matching and ranking | Scores and orders applicants against a req | The tool, because order controls who gets read | High |
| Assessment and video scoring | Grades a test, recorded answer, voice, affect | The tool. It is a selection procedure | Highest |
| Scheduling | Books interviews, reschedules, reminds | Nobody, beyond calendar logic | Low |
| Chat screening | Knockout questions, availability, FAQs | Depends on whether it auto-rejects | Low if it collects, high if it filters |
| Interview note-taking | Transcribes, fills a structured scorecard | The interviewer, if designed right | Low to moderate. Recording consent is the issue |
Buy these separately, with separate approval paths. Scheduling and plain CV parsing can be bought on convenience. A ranking or scoring tool should get the scrutiny you would apply to introducing a new pre-employment test, because legally that is close to what it is.
Why the vendor’s accuracy claim is not your defence
US federal law here predates the technology and does not care that a decision was automated. Under Title VII, a neutral-looking procedure that disproportionately excludes people by race, colour, religion, sex or national origin is unlawful unless the employer shows it is job related for the position in question and consistent with business necessity. The EEOC’s fact sheet on employment tests and selection procedures says exactly that, and points to the Uniform Guidelines on Employee Selection Procedures, adopted 1978.
The Uniform Guidelines sit at 29 CFR Part 1607 and contain the rule everyone half-remembers. Section 1607.4(D): “A selection rate for any race, sex, or ethnic group which is less than four-fifths (4/5) (or eighty percent) of the rate for the group with the highest rate will generally be regarded by the Federal enforcement agencies as evidence of adverse impact.” Two things get misread. It is an enforcement guideline rather than a statutory bright line, and the same section notes that smaller differences can still count while larger ones may not if the numbers are too small to be reliable. And it bites at the point of selection, which for a ranking tool means the threshold you actually cut at. Where adverse impact exists, Part 1607 requires validation and recognises criterion-related, content and construct validity. Validity is a property of a procedure used for a particular job in a particular setting, not a PDF about a vendor’s model in general.
The ADA regulations make ownership plainer still. 29 CFR 1630.10(a) makes it unlawful to use “qualification standards, employment tests or other selection criteria that screen out or tend to screen out an individual with a disability” unless the criterion, “as used by the covered entity,” is job related and consistent with business necessity. As used by the covered entity. The rule is about your use, not the builder’s design.
One piece of live context. The EEOC published two AI-specific technical assistance documents, one on the ADA and one on adverse impact under Title VII, and both URLs now return 404 on eeoc.gov, checked September 2026. Withdrawn guidance repeals nothing. Title VII, the ADA, 29 CFR 1607 and 29 CFR 1630 all stand, and a private plaintiff or a state agency can still bring a disparate impact claim. What you lost is free interpretive help.
The rules that name this technology
New York City Local Law 144. You cannot use an automated employment decision tool for an NYC job unless it has had a bias audit within one year of use. DCWP defines an AEDT as a computer-based tool that uses machine learning, statistical modelling, data analytics or AI, helps employers make employment decisions, and substantially assists or replaces discretionary decision-making. The audit must be by an independent auditor: someone exercising objective and impartial judgement who does not work for the employer or the vendor, has no prior involvement with the tool and no financial interest in the parties. A vendor’s own audit therefore cannot satisfy it. You publish a summary of results, which per DCWP’s FAQ includes the audit date, the data source and an explanation of it, the number of applicants assessed, the number in unknown demographic categories, selection or scoring rates, and impact ratios for all categories. Candidates get notice at least 10 business days before use, describing the job qualifications the tool will assess and how to request a reasonable accommodation. DCWP began enforcing on 5 July 2023.
And note where the impact ratios end up: published on your careers site, under your name.
Colorado. SB24-205 regulates high-risk AI used in consequential decisions and puts real duties on deployers: impact assessments, an annual review that the deployment is not causing algorithmic discrimination, notice to the person when the system substantially factors into a consequential decision about them, a chance to correct inaccurate data and appeal for human review where technically feasible, a public statement about systems in use, and disclosure to the Attorney General within 90 days of discovering algorithmic discrimination. Its requirements moved to 30 June 2026 under SB25B-004, signed 28 August 2025. States run their own bills every session, so treat any state list as a prompt to check rather than a citation.
The EU AI Act. Regulation (EU) 2024/1689 entered into force 1 August 2024 and became generally applicable 2 August 2026. Employment is explicitly high risk: Annex III point 4 covers recruitment and selection, promotion and termination, task allocation based on behaviour or traits, and monitoring and evaluation of workers, and the Commission names CV-sorting software as its example. Obligations for Annex III high-risk uses now apply from 2 December 2027, extended by the AI Omnibus package, adopted 19 November 2025, agreed on 7 May 2026 and in force from 27 July 2026.
Two parts already bind. Prohibited practices applied from 2 February 2025 and include AI used to infer emotions of a person in the workplace, with a narrow carve-out for medical or safety reasons. If you hire in Europe and a vendor offers sentiment or enthusiasm scoring on interview video, that is not a feature to weigh against price. Design for the rest early: a deployer must inform workers’ representatives, or workers directly where there are none, before putting a high-risk system into service at the workplace, and affected people get a right to an explanation of the decision. The provider carries one set of duties, you carry the deployer set, and you cannot contract out of yours.
The accessibility question nobody asks in the demo
A timed gamified assessment measures reaction speed alongside reasoning. A one-way video interview measures fluency and facial expressiveness alongside content. A chat screener with a response window measures typing. Each can screen out a candidate whose disability affects the medium rather than the job. 29 CFR 1630.11 requires that tests be selected and administered so results reflect what the test purports to measure “rather than reflecting the impaired sensory, manual, or speaking skills” of the applicant, except where those are the skills being measured. A stutter is not a competency unless the job is voice acting.
So you owe accommodation in how the assessment is administered, which means a real alternative route with someone staffing it, and you must tell candidates it exists beforehand in plain language. Settle it before you buy: “we can extend the timer” and “we offer a live interview instead” are very different answers, and only one works for someone who cannot use the interface at all. Ask for the accessibility conformance report, then test the flow yourself with a screen reader and keyboard only, during the trial.
What to ask a vendor, in the order that sorts them fastest
Send these in writing and keep the answers. The written record is part of what good faith looks like later on.
- What does it output? A score, a rank, a pass/fail, a flag? Ranking is a decision, because order controls who gets read.
- Can it run advisory-only? Auto-rejection fully off, scores hidden until a reviewer forms a view, and the configuration locked so an admin cannot switch filtering back on.
- What was it trained on? Our hiring history, your cross-client pool, public web data, or a general-purpose language model with a prompt on top? Each fails differently, and the last is often what “proprietary AI model” means.
- Which fields does it see? Names, photos, postcodes, school names, graduation years, employment gaps. Several carry demographic signal. Ask how they verified what removing them does.
- Bias audit by whom, on what data, at which threshold? It needs to be independent of both of us, on our pool, at the cut-off we will really use.
- What validation evidence exists, of which type? Criterion-related, content or construct, for which job families. Testimonials and a drop in time-to-screen are efficiency evidence, not validity.
- Can we export raw scores for every applicant? With demographics joined on our side. No scores, no audit, and no way to meet a disclosure rule later.
- What happens to candidate data? Retention, deletion on request, subprocessors, storage location, and whether our candidates train your models for other clients. Contract the last one as a no.
- Will you indemnify us for discrimination claims arising from the output? Almost nobody will, and the refusal is informative. The winnable ask is carving breaches of their own bias and data representations out of the liability cap.
Test it in shadow mode, then keep a real human in the loop
The most useful evaluation costs nothing in candidate risk, and few vendors volunteer it. Take a requisition you closed six to twelve months ago with a decent pool, feed the original applications unedited, and ask for scores on every applicant rather than a shortlist.
Then check four things. Where does it rank the hires who are still performing? If your two best land in the bottom half, it is not measuring what you think. Read twenty candidates it downranked that humans advanced, because the disagreements tell you what it keys on. Compute selection rates by group at the threshold you would genuinely use, then the impact ratio against the highest-rate group. And rescore the same pool a week later to see whether the order moves.
The ratio, worked, because people compute it wrong. You cut at the top 25% of scores. Group A: 400 applicants, 100 above the cut, a 25% selection rate. Group B: 200 applicants, 30 above the cut, 15%. The impact ratio is 15 divided by 25, or 0.60, well under four-fifths. That is not proof of unlawful discrimination. It is exactly the evidence that triggers the job-relatedness question, and you would rather meet it on a historical pool than in a charge.
One trap. High agreement with past recruiter decisions is imitation, not validity: the model learns your previous pattern, including the parts you would not defend. And a model trained only on people you hired never saw how the rejected would have performed, which is why “trained on your top performers” is weaker than it sounds. Watch accuracy figures on rare outcomes too. If 3% of applicants get hired, rejecting everybody is 97% accurate.
Then the operating design. Most setups that claim a human in the loop do not have one, because the human sees a pre-sorted list of twelve and approves it. Four rules make it real. No tool removes a candidate from the pool a human reviews, so it annotates and the human sorts. Reviewers write a reason for each rejection at the stages that matter, which forces a human decision and creates the audit trail. Candidates get a stated route to a human alternative, and someone owns that queue. And you track reversal rate, the share of cases where the reviewer disagrees. Near zero reversal is not oversight. It is a rubber stamp with a job title.
Measure accordingly. Time-to-fill and cost-per-screen will improve, which is what vendors sell on and says nothing about hiring better or lawfully. Track per stage and per group: selection rate, pass-through, offer-accept, 90-day and 12-month retention, hiring manager satisfaction, drop-off inside the automated steps. Hold selection rates by group as a time series, because a model that passed in March can drift by September. And count recruiter hours honestly: if screening falls 20 hours a week while review, appeals and accommodations add 12, the saving is 8. Worth having. Just do not report 20.
Where AI recruiting software genuinely helps, and where claims outrun evidence
My read, and I will defend it. The reliable wins are unglamorous: scheduling and rescheduling, reminders that cut no-shows, transcription and structured scorecard capture that makes interviews comparable, deduplication and search across your own ATS so you stop re-sourcing people you already know, and outreach a human edits and sends. Those share one property. Errors are visible, reversible, and not selection decisions. They buy back coordinator time, the real bottleneck in most mid-market teams.
The weak end is candidate quality prediction. Fit scores, personality or culture inference from writing or video, and affect scoring all assume the signal in an application or a recorded answer beats a structured human process at predicting performance. That is the claim needing validation evidence for your job on your data, and the one almost never supplied. Emotion inference at work is already prohibited in the EU. Treat it as a red flag everywhere.
So buy aggressively in the coordination layer, buy carefully and with validation evidence in the assessment layer, and never let a tool own the cut. A vendor comfortable with shadow-mode testing, independent audits, raw score export and advisory-only configuration has already told you more about product quality than any feature grid.
Questions that come up next
Does Local Law 144 apply if we are not in New York? The location of the job and the candidate matters, not your headquarters. Screening for an NYC role, or a candidate resident in the city, should be treated as in scope and confirmed with counsel. Remote roles open to NYC residents catch people out.
Is it safer to use AI only after the human shortlist? It reduces exposure without removing it. Anything that reorders or deprioritises candidates at any stage influences a selection decision, and a small late-stage pool makes your statistics unreliable, which cuts both ways.
The vendor says it is not an AEDT because a human makes the final call. Not sufficient on its own. The test is whether the tool substantially assists or replaces discretionary decision-making, and a human approving a machine-ordered list may well meet it. Document what the human actually does, reversal rate included, or you have no evidence for the claim.
If you want help with this from the delivery side rather than the slide-deck side, it is a fair bit of what we do at AB7 Solutions: recruitment and RPO desks where humans keep the screening decision, AI and automation built with human-in-the-loop review designed in rather than bolted on, and the data plumbing that lets you answer an impact-ratio question in an afternoon instead of a quarter. For a shadow-mode test against a hiring round you have already closed, call +1 321 341 7733, email ab@ab7solutions.com or director@ab7solutions.com, or start at www.ab7solutions.com.
Sources: U.S. EEOC, Employment Tests and Selection Procedures; Uniform Guidelines on Employee Selection Procedures, 29 CFR Part 1607 (§1607.4(D), §1607.5); ADA regulations, 29 CFR 1630.10 and 29 CFR 1630.11; NYC Department of Consumer and Worker Protection, Automated Employment Decision Tools and the DCWP AEDT FAQ; Colorado General Assembly, SB24-205 and SB25B-004; European Commission, AI Act regulatory framework, and Regulation (EU) 2024/1689, Article 5(1)(f), Annex III point 4, Articles 26 and 86. General information, not legal advice.