Product

What Annex III actually asks of a hiring team

What Annex III actually asks of a hiring team

Employment sits in the high-risk tier of the EU AI Act. Here is what that means in practice, minus the legal throat-clearing.

Published

Most explanations of the EU AI Act start with the risk pyramid, spend four paragraphs on it, and stop before they reach anything a recruiter could act on. This one starts where the work is. It is a plain reading of what the Act says about hiring, not legal advice, and the specifics for your organisation are a conversation with your own counsel.

Where hiring sits in the Act

The Act sorts uses of AI rather than products. A handful of practices are banned outright. A large middle group is classed as high-risk. Most other things carry light transparency duties or none at all.

Annex III is the list of high-risk uses, and employment is one of its headings. The Act describes it in two halves: systems used for recruitment or selection, including targeting job ads, filtering applications and evaluating candidates, and systems used once someone is hired, for decisions about promotion, termination, task allocation and monitoring.

That is broader than most teams expect. It is not limited to a tool that autonomously rejects people. A system that ranks a longlist, or scores an application against a role, sits inside the description even when a person signs off every decision at the end.

High-risk is a label about consequences, not quality

Worth saying early, because it is the most common misreading in a procurement thread: being classed high-risk is not a finding that a tool is dangerous or badly built. It is a statement about where the tool is pointed. Hiring decides who gets work, the applicant usually cannot see the process, and the effects compound quietly across a whole intake. The Act treats that consequence as the reason for the extra obligations.

So the answer to "are we high-risk?" is usually yes, and that is not the interesting question. The interesting question is what follows from it.

What follows, in practice

The Act splits duties between the provider, who develops the system and puts it on the market, and the deployer, who uses it. Most of the heavy engineering obligations sit with the provider. The deployer's side is smaller but real, and it is the side a hiring team lives on.

Read across the employment provisions and the same few themes keep returning:

  • A record of what happened. Systems are expected to log their operation, and deployers are expected to keep those logs for a period. In hiring terms: which candidates were assessed, against what, and what came back.

  • Human oversight that is real. A named person, with the standing and the information to disagree with the output. Oversight that cannot see the reasoning is not oversight.

  • Telling people. Candidates are meant to know when a high-risk system is being used on them, and people affected by a decision can ask for an explanation of the role it played.

  • Using the thing as intended. Feeding a system data it was never designed for, or stretching it to a use the provider never described, moves responsibility toward the deployer.

  • Fundamental-rights impact work. Certain deployers, particularly public bodies, carry an assessment duty before they start using a high-risk system.

None of that is exotic. Most of it is what a careful hiring team would want anyway, written down.

The part that bites is evidence, not intent

Most hiring teams already believe a person makes the call. Far fewer could show it a year later.

That gap is the practical problem. Oversight you cannot evidence looks identical, from the outside, to no oversight at all. If a candidate asks why they did not progress, or a works council asks how a shortlist was produced, or an auditor asks what the system weighed, the answer has to be retrievable rather than remembered. "A human reviewed it" is a claim. A timestamped record of what the human saw, what the system proposed, and where the two differed is an answer.

This is why we built Yardstick the way we did. Every rating carries the passage it came from, the reasoning is written before anyone sees a rank, bias controls are set and documented per role rather than claimed once for the product, and the whole trail exports. Yardstick screens and recommends. A person decides. The evidence layer exists so that the person deciding can check, and so that the checking leaves a trace.

Questions worth putting to any vendor

Not a compliance checklist. Just the questions that separate a tool you can explain from one you cannot.

Can you show the reasoning behind a single ranking, in full, a year after it was produced? Does the system tell you what it weighed, or only what it concluded? Is bias testing documented for the specific opening, or asserted once in a brochure? Can you export the record in a form somebody outside the product can read? What does the candidate see, and when?

A vendor who can answer those in writing is a vendor you can defend. One who answers with a certification logo is not answering.

Where this stops

This is a description of what the Act says, not a ruling on what applies to your organisation. Which duties land on you depends on your role in the chain, the sector you hire in, where your candidates sit, and what your system actually does. Timelines for the high-risk tier are staged, and how the detail is interpreted is still settling.

Take the shape of it from here. Take the specifics to your counsel, early, while it is still cheaper to design for than to retrofit.

Planning a rollout like this?

Elin Ahlberg, Sales lead

Planning a rollout like this?

Elin Ahlberg, Sales lead

Start with one open role

Create a free website with Framer, the website builder loved by startups, designers and agencies.