AI engineering for HR & Workforce Technology

AI that supports the decision, never the verdict

We build AI that prepares hiring and workforce decisions, and the controls that make each one defensible: the criteria applied, the evidence behind them, a named reviewer, and the notice the candidate is owed.

See the decision record

A screening decision is prepared by a system and decided by a person. The basis, the reviewer, and the notice stay on the record.

Screening decision req #2291 · senior engineer
Decision Advance to a structured interview
Assessed against
  • Distributed systems experience 7 years, evidenced
  • Production on-call ownership current role
  • Kubernetes depth not in the record
  • Location and work authorisation meets the requirement
Reviewed by

R. Okafor, hiring manager · 14 minutes with the criteria and the evidence in front of them

On record

The basis, the outcome notice, and the contest route, kept for the retention period.

This is the one industry where the AI itself is regulated

Hiring and workforce systems sit in the high-risk tier by law, which changes what you have to build, document, and be able to prove.

The AI itself is the regulated object

Employment and workers' management is a high-risk category under the EU AI Act, covering recruitment and selection as well as promotion, termination, task allocation, and performance monitoring. That pulls in the full high-risk stack, including impact assessment, record-keeping, and oversight designed into the system.

Employers carry duties as deployers

An organisation using one of these systems is a deployer, which means using it per instruction, ensuring real oversight, monitoring it, reporting incidents, and informing workers' representatives and affected workers before it goes into use. The vendor's paperwork does not discharge those duties.

The audit requirement already exists in one market

New York City has required an independent bias audit before use, a published summary, and candidate notice ten business days ahead since 2023. The calculation is prescribed: selection rates by category, compared to the most-selected group to produce an impact ratio.

The litigation is about the maths and the liability

The largest AI hiring case has around fourteen thousand people in an age-discrimination collective, and its central questions are whether screening tools produced a disparate impact and whether the vendor or the employer answers for it. A court has already allowed the argument that the vendor can act as the employer's agent.

Volume rose, outcomes did not

Recruiters now handle roughly twice the applications they did in 2021 while hires per recruiter have fallen sharply, and most candidate rejections happen at the screening stage before any human contact. HR teams use AI daily while only a minority fully trust it, and candidates are more negative about it than positive.

Agentic AI fits work that supports a decision about a person

An agent can screen against stated criteria, assemble the evidence, schedule the next step, and answer employee questions. The decision stays with a named person, and the reason stays on the record.

Screening and selection

Applications assessed against written criteria with the evidence attached, so the shortlist is a documented judgement rather than a score nobody can explain.

  • Criteria-based screening
  • Evidence summaries
  • Structured interviews
  • Shortlist rationale

Employee support and service

Policy questions, benefits, leave, and payroll queries answered from the current handbook and the employee's own record, with the source attached.

  • Policy answers
  • Leave and benefits
  • Payroll queries
  • Case routing

Workforce operations

Onboarding, scheduling, and internal mobility work that reads and writes to the HRIS under the same bounds as everything else.

  • Onboarding tasks
  • Scheduling
  • Internal mobility
  • Data hygiene

HR products and platforms

AI features shipped inside an HR or workforce product, where the high-risk duties travel with the feature into every customer's environment.

  • In-product screening
  • Assessment scoring
  • Workforce analytics
  • Coaching tools

No client results published

What the public record already shows

We are not publishing outcomes from HR or workforce engagements. The case for governing this work is public: employment AI is a high-risk category in the EU, one market has required an independent bias audit since 2023, and the largest AI hiring case has around fourteen thousand people in a collective with a class-certification hearing set for March 2027. What follows is the standard every deployment has to meet.

The figures above are published regulatory requirements, court filings, and market research. They are not client results, and no engagement outcome is claimed here.

  • Classification documented before build
  • Every judgement carries its evidence
  • Named reviewer who can disagree
  • Selection rates and impact ratios measured live

Our path to a defensible decision

We start from one decision that affects a person, then build the classification, the evidence, the reviewer, and the record around it.

  1. 1

    Classify the system honestly

    Map what each tool actually does against the high-risk categories. Adding a reviewer to a screening tool does not move it out of scope.

  2. 2

    Name the criteria before the candidates

    State what the decision is made on, in writing, so the model is matching against a standard rather than an impression.

  3. 3

    Attach the evidence to every judgement

    Each criterion is met, unclear, or unmet, and the record shows what supported it rather than a single opaque score.

  4. 4

    Make the human step real

    A named person reviews with the criteria and the evidence in front of them, and the time spent is part of the record.

  5. 5

    Measure who advances, by category

    Selection rates and impact ratios by category are computed continuously, so a disparity shows up as a number rather than a claim.

  6. 6

    Operate and revalidate

    Notices, retention, reviewer behaviour, and drift are reviewed on a cadence, with revalidation when models, criteria, or roles change.

The controls sit outside the model

Classification, criteria, reviewer assignment, notices, and the record are enforced by a service the model cannot reason around.

Review our governance approach
01

Classification and documentation

Each tool is mapped to its intended purpose and the obligations that follow, with the technical documentation produced as the system is built rather than assembled for a review.

02

Criteria as configuration

Role requirements live in versioned configuration, so a change to what the system screens on is a reviewed change rather than an untracked model update.

03

Evidence per criterion

Every judgement carries what supported it, and an unclear criterion is routed to a person instead of being averaged into a score.

04

Named reviewer with authority

The reviewer is assigned, has the criteria and evidence in front of them, and can disagree. Time on task is recorded, because oversight that takes seconds is oversight in name.

05

Notices and retention

Candidate notices, adverse-outcome explanations, and the retention period are properties of the workflow, so the duty is discharged by the run rather than by a manual process.

06

Continuous impact measurement

Selection rates and impact ratios by category are computed from the live process, which makes the audit a report on the system rather than a project that reconstructs it.

07

Employee-facing surfaces

Policy answers cite the handbook version they came from, and an employee question that touches their own record is scoped to their own record.

Standards and regulation

The rules that shape what an HR or workforce deployment has to produce, and where each one lands in the architecture.

EU AI Act Annex III
Employment and workers' management is a high-risk category, covering recruitment and selection as well as promotion, termination, and performance monitoring.
EU AI Act Article 26
Deployer duties: use per instructions, ensure oversight, monitor, report incidents, and inform workers' representatives before use.
NYC Local Law 144
An independent bias audit before use, a published summary, and candidate notice ten business days ahead, enforced since 2023.
Automated employment decision tools
The statutory category that triggers the audit and notice duties, and the state laws that followed it.
EEOC Uniform Guidelines
The selection-rate calculation behind the impact ratio, including the four-fifths reference point.
GDPR and impact assessments
A data protection impact assessment is expected before a screening tool goes live, and solely automated decisions need meaningful review.

Questions HR, legal, and platform leaders ask

Does adding a human reviewer take us out of the high-risk category?

No, and this is the most common misconception in the sector. Classification turns on the intended purpose, and the Commission's own guidance says human involvement cannot change the purpose or the area of use. Oversight is a requirement you have to meet as a high-risk system, and it does not move you out of the category.

What does the bias audit actually require?

A bias audit is an independent evaluation of the tool: it calculates selection or scoring rates by sex and race and ethnicity, including intersectional categories, and compares each to the most-selected group to produce an impact ratio. New York City requires one within a year of use, with the summary published and candidates notified ten business days ahead. What we provide is the measurement underneath: those rates and ratios computed from live decisions, so the auditor verifies a running system rather than rebuilding one.

We buy the tool rather than build it. What are our obligations?

You are a deployer, which is its own set of duties: use the system as instructed, ensure oversight that is genuinely capable of disagreeing, monitor the operation, report incidents, and inform workers' representatives and affected workers before putting it into use. The vendor's documentation supports those duties but does not discharge them.

How do we keep screening from producing a disparate impact?

By writing the criteria first, attaching evidence per criterion, and measuring outcomes by category as the process runs. A model trained on your historical hires will reproduce your historical pattern, so the measurement has to be continuous rather than annual. Where a ratio drops below the reference point, that is a design question you can act on before it becomes a filing.

What does a candidate see?

The basis for the outcome, and a route to contest it. That is a legal requirement in some jurisdictions and a trust requirement everywhere: candidates increasingly cannot tell whether a job is real or whether a person ever read their application, and a documented basis is the cheapest answer to that.

How do we handle employee-facing assistants?

Policy answers cite the handbook version they came from, and anything touching an individual's own record is scoped to that record. An assistant that improvises a leave policy is a compliance problem, and one that quotes a superseded handbook is a different one.

How do we measure whether this works?

At the decision level: criteria coverage, reviewer time on task, override rate, notice and retention compliance, selection rates and impact ratios by category, and candidate experience scores. Quality is reviewed with HR and legal together, not inferred from throughput.

Can you review an HR AI vendor we are already evaluating?

Yes. An oversight engagement reviews the architecture, classification, criteria design, evidence basis, oversight step, bias measurement, and vendor claims while the platform is being implemented, with findings delivered at agreed milestones.

Find the first decision worth documenting

We will examine the decision, its classification, the evidence basis, the reviewer step, and the failure modes before recommending a build.

Reviewed by Jayaveer Bhupalam, Founder & CTO Last updated September 29, 2026