AI engineering for HR & Workforce Technology
AI that supports the decision, never the verdict
We build AI that prepares hiring and workforce decisions, and the controls that make each one defensible: the criteria applied, the evidence behind them, a named reviewer, and the notice the candidate is owed.
A screening decision is prepared by a system and decided by a person. The basis, the reviewer, and the notice stay on the record.
This is the one industry where the AI itself is regulated
Hiring and workforce systems sit in the high-risk tier by law, which changes what you have to build, document, and be able to prove.
The AI itself is the regulated object
Employment and workers' management is a high-risk category under the EU AI Act, covering recruitment and selection as well as promotion, termination, task allocation, and performance monitoring. That pulls in the full high-risk stack, including impact assessment, record-keeping, and oversight designed into the system.
Employers carry duties as deployers
An organisation using one of these systems is a deployer, which means using it per instruction, ensuring real oversight, monitoring it, reporting incidents, and informing workers' representatives and affected workers before it goes into use. The vendor's paperwork does not discharge those duties.
The audit requirement already exists in one market
New York City has required an independent bias audit before use, a published summary, and candidate notice ten business days ahead since 2023. The calculation is prescribed: selection rates by category, compared to the most-selected group to produce an impact ratio.
The litigation is about the maths and the liability
The largest AI hiring case has around fourteen thousand people in an age-discrimination collective, and its central questions are whether screening tools produced a disparate impact and whether the vendor or the employer answers for it. A court has already allowed the argument that the vendor can act as the employer's agent.
Volume rose, outcomes did not
Recruiters now handle roughly twice the applications they did in 2021 while hires per recruiter have fallen sharply, and most candidate rejections happen at the screening stage before any human contact. HR teams use AI daily while only a minority fully trust it, and candidates are more negative about it than positive.
Agentic AI fits work that supports a decision about a person
An agent can screen against stated criteria, assemble the evidence, schedule the next step, and answer employee questions. The decision stays with a named person, and the reason stays on the record.
Screening and selection
Applications assessed against written criteria with the evidence attached, so the shortlist is a documented judgement rather than a score nobody can explain.
- Criteria-based screening
- Evidence summaries
- Structured interviews
- Shortlist rationale
Employee support and service
Policy questions, benefits, leave, and payroll queries answered from the current handbook and the employee's own record, with the source attached.
- Policy answers
- Leave and benefits
- Payroll queries
- Case routing
Workforce operations
Onboarding, scheduling, and internal mobility work that reads and writes to the HRIS under the same bounds as everything else.
- Onboarding tasks
- Scheduling
- Internal mobility
- Data hygiene
HR products and platforms
AI features shipped inside an HR or workforce product, where the high-risk duties travel with the feature into every customer's environment.
- In-product screening
- Assessment scoring
- Workforce analytics
- Coaching tools
Featured decision workflow
A governed workforce decision
The system assembles the evidence, a named person decides, and the outcome is explained to the candidate. Classification, criteria, oversight, and notices are part of the run rather than a document written afterwards.
-
Classify before you build
Map the intended purpose against the high-risk categories. That classification decides which obligations attach, and a human reviewer in the loop does not change it.
-
State the criteria in writing
The decision is made against published requirements for the role, so the system matches a standard rather than an impression of a good candidate.
-
Assemble the evidence per criterion
Each requirement is met, unclear, or unmet, with what supported it. An unclear criterion goes to the reviewer rather than being resolved by the model.
-
Route to a named reviewer
A person with the authority to decide reviews the criteria and the evidence. Human oversight here means the ability to disagree, and the time spent is recorded.
-
Record the decision and the basis
The decision record holds the criteria version, the evidence, the reviewer, and the outcome, retained for the period the rules require.
-
Measure the outcomes by category
Selection rates and impact ratios are computed as the process runs, so a disparity is visible while it is still a design question.
-
Revalidate on a cadence
Criteria, reviewer behaviour, notices, and drift are reviewed on a schedule, with revalidation when models, criteria, or roles change.
No client results published
What the public record already shows
We are not publishing outcomes from HR or workforce engagements. The case for governing this work is public: employment AI is a high-risk category in the EU, one market has required an independent bias audit since 2023, and the largest AI hiring case has around fourteen thousand people in a collective with a class-certification hearing set for March 2027. What follows is the standard every deployment has to meet.
The figures above are published regulatory requirements, court filings, and market research. They are not client results, and no engagement outcome is claimed here.
- Classification documented before build
- Every judgement carries its evidence
- Named reviewer who can disagree
- Selection rates and impact ratios measured live
Our path to a defensible decision
We start from one decision that affects a person, then build the classification, the evidence, the reviewer, and the record around it.
- 1
Classify the system honestly
Map what each tool actually does against the high-risk categories. Adding a reviewer to a screening tool does not move it out of scope.
- 2
Name the criteria before the candidates
State what the decision is made on, in writing, so the model is matching against a standard rather than an impression.
- 3
Attach the evidence to every judgement
Each criterion is met, unclear, or unmet, and the record shows what supported it rather than a single opaque score.
- 4
Make the human step real
A named person reviews with the criteria and the evidence in front of them, and the time spent is part of the record.
- 5
Measure who advances, by category
Selection rates and impact ratios by category are computed continuously, so a disparity shows up as a number rather than a claim.
- 6
Operate and revalidate
Notices, retention, reviewer behaviour, and drift are reviewed on a cadence, with revalidation when models, criteria, or roles change.
The controls sit outside the model
Classification, criteria, reviewer assignment, notices, and the record are enforced by a service the model cannot reason around.
Review our governance approachClassification and documentation
Each tool is mapped to its intended purpose and the obligations that follow, with the technical documentation produced as the system is built rather than assembled for a review.
Criteria as configuration
Role requirements live in versioned configuration, so a change to what the system screens on is a reviewed change rather than an untracked model update.
Evidence per criterion
Every judgement carries what supported it, and an unclear criterion is routed to a person instead of being averaged into a score.
Named reviewer with authority
The reviewer is assigned, has the criteria and evidence in front of them, and can disagree. Time on task is recorded, because oversight that takes seconds is oversight in name.
Notices and retention
Candidate notices, adverse-outcome explanations, and the retention period are properties of the workflow, so the duty is discharged by the run rather than by a manual process.
Continuous impact measurement
Selection rates and impact ratios by category are computed from the live process, which makes the audit a report on the system rather than a project that reconstructs it.
Employee-facing surfaces
Policy answers cite the handbook version they came from, and an employee question that touches their own record is scoped to their own record.
Standards and regulation
The rules that shape what an HR or workforce deployment has to produce, and where each one lands in the architecture.
- EU AI Act Annex III
- Employment and workers' management is a high-risk category, covering recruitment and selection as well as promotion, termination, and performance monitoring.
- EU AI Act Article 26
- Deployer duties: use per instructions, ensure oversight, monitor, report incidents, and inform workers' representatives before use.
- NYC Local Law 144
- An independent bias audit before use, a published summary, and candidate notice ten business days ahead, enforced since 2023.
- Automated employment decision tools
- The statutory category that triggers the audit and notice duties, and the state laws that followed it.
- EEOC Uniform Guidelines
- The selection-rate calculation behind the impact ratio, including the four-fifths reference point.
- GDPR and impact assessments
- A data protection impact assessment is expected before a screening tool goes live, and solely automated decisions need meaningful review.
Flytebit products fit the team building your HR platform
Use these products when review, tests, documentation, or release control constrain the engineers behind your HRIS, ATS, and workforce products.
Review AI-generated changes to HRIS, ATS, and workforce products across eight categories before they merge.
Generate tests around changed functions in screening, scheduling, and payroll code, then run them in your existing pipeline.
Regenerate technical and architecture documentation from the repository so the written record matches the deployed system.
Reshape requirements, review, testing, and release around AI-assisted engineering so throughput follows generation speed.
Choose the first decision the engagement must produce
The classification is unclear.
AI feasibility study
The decision set, high-risk classification, evidence basis, and a Go or No-Go verdict. Assess the workflowWe know what to build.
Architecture and delivery
A production system with criteria, reviewer gates, notices, and an operating model. Design the systemA vendor is selling us a tool.
Independent oversight
Architecture, classification, bias measurement, vendor claims, and oversight design reviewed. Review the active buildThe system is already live.
LLMOps and operations
Drift reviews, criteria freshness, reviewer behaviour, retention, and incident response. Review the operating modelTechnical context for workforce AI
Use these guides to examine the governance, evaluation, and oversight work behind the page.
Questions HR, legal, and platform leaders ask
Does adding a human reviewer take us out of the high-risk category?
No, and this is the most common misconception in the sector. Classification turns on the intended purpose, and the Commission's own guidance says human involvement cannot change the purpose or the area of use. Oversight is a requirement you have to meet as a high-risk system, and it does not move you out of the category.
What does the bias audit actually require?
A bias audit is an independent evaluation of the tool: it calculates selection or scoring rates by sex and race and ethnicity, including intersectional categories, and compares each to the most-selected group to produce an impact ratio. New York City requires one within a year of use, with the summary published and candidates notified ten business days ahead. What we provide is the measurement underneath: those rates and ratios computed from live decisions, so the auditor verifies a running system rather than rebuilding one.
We buy the tool rather than build it. What are our obligations?
You are a deployer, which is its own set of duties: use the system as instructed, ensure oversight that is genuinely capable of disagreeing, monitor the operation, report incidents, and inform workers' representatives and affected workers before putting it into use. The vendor's documentation supports those duties but does not discharge them.
How do we keep screening from producing a disparate impact?
By writing the criteria first, attaching evidence per criterion, and measuring outcomes by category as the process runs. A model trained on your historical hires will reproduce your historical pattern, so the measurement has to be continuous rather than annual. Where a ratio drops below the reference point, that is a design question you can act on before it becomes a filing.
What does a candidate see?
The basis for the outcome, and a route to contest it. That is a legal requirement in some jurisdictions and a trust requirement everywhere: candidates increasingly cannot tell whether a job is real or whether a person ever read their application, and a documented basis is the cheapest answer to that.
How do we handle employee-facing assistants?
Policy answers cite the handbook version they came from, and anything touching an individual's own record is scoped to that record. An assistant that improvises a leave policy is a compliance problem, and one that quotes a superseded handbook is a different one.
How do we measure whether this works?
At the decision level: criteria coverage, reviewer time on task, override rate, notice and retention compliance, selection rates and impact ratios by category, and candidate experience scores. Quality is reviewed with HR and legal together, not inferred from throughput.
Can you review an HR AI vendor we are already evaluating?
Yes. An oversight engagement reviews the architecture, classification, criteria design, evidence basis, oversight step, bias measurement, and vendor claims while the platform is being implemented, with findings delivered at agreed milestones.
Find the first decision worth documenting
We will examine the decision, its classification, the evidence basis, the reviewer step, and the failure modes before recommending a build.