Engineering

How a screening model gets tested before it ships

How a screening model gets tested before it ships

What a ranking model has to pass before a new version reaches your roles, and which results stop a release.

Published

A screening model makes the same judgement thousands of times a week, and nobody reads most of them. That is why it is useful. It is also why testing matters more here than in most software. A bug on a checkout page shows up as a complaint. A bug in a ranking shows up as a shortlist that looks fine and is quietly wrong.

This post walks through how a model that ranks applicants should be tested before a new version reaches your roles, and what should stop it. It describes how Yardstick does it, and it doubles as a list of questions to put to any vendor.

Test on applications the model has never seen

The first rule is old and still broken often. A model has to be scored on applications it was never tuned on. Otherwise you are measuring its memory.

Yardstick keeps a held-out evaluation set for each role family: terminal and warehouse work, care and nursing, retail floor, contact centre, engineering. Each set is built from real briefs and real applications with identifying details removed. It is never used to adjust the model.

A few things make a held-out set worth trusting:

  • It holds at least 500 applications per role family, so one odd batch cannot swing the result

  • It includes the hard cases: applicants with no CV, answers in a second language, career changers, people returning after a gap

  • It is refreshed on a schedule, because roles change and a stale set flatters the model

  • It is locked, so nobody building the next release can see its labels

The hard cases matter most. A model that does well on tidy CVs and badly on a forklift driver who answered five questions in Polish fails the people high-volume hiring depends on.

Measure agreement with reviewers, one criterion at a time

The labels on a held-out set come from people. Two trained reviewers score each application against the same written rubric the model uses. They work independently, without seeing each other’s marks or the model’s. Where they disagree, a third reviewer settles it and writes down why.

That gives you two numbers worth watching. One is how often the two reviewers agree with each other. The other is how often the model agrees with the settled human score. If the reviewers cannot agree on a criterion, the model is not the problem yet. The rubric line is.

Agreement is checked per criterion as well as overall. An overall figure can hide a model that is excellent at shift availability and poor at judging a written answer about a difficult night on a care ward. The criterion view is where you find that.

Rank agreement matters too. A shortlist is the top of a list, so the test also asks whether the people reviewers would put in the top twenty actually end up there.

Because every rating in Yardstick carries an evidence link, reviewers check the reasoning as well as the number. A correct score reached from the wrong passage counts as a failure. It will not stay correct for long.

Check adverse impact for each role

A model can agree closely with reviewers and still disadvantage a group. Agreement tells you it reproduces human judgement. It says nothing about whether that judgement was fair.

So each release is tested for adverse impact: whether applicants from one group move past the cut line at a noticeably lower rate than another. A common reference point is the four-fifths comparison. It flags any group advancing at less than 80% of the rate of the group that advances most. Treat it as a signal to look closer, not a legal test. This post is not legal advice, and the standard that applies to you is a question for your counsel.

The test runs per role, because impact depends on the brief. A requirement for night availability or a driving licence can be entirely legitimate and still land unevenly. The bias test report per role, on the Team plan and above, records which groups were compared and the pass rates at each stage. It also shows which criteria drove any gap, and what the hiring team decided to do about it.

Group data is used for testing only and never feeds the ranking. Where you do not collect it, the report says so. The test then compares groups the brief can see, such as the language an applicant answered in and whether a CV was attached.

Run the last release’s cases through the new one

Most model changes are meant to improve one thing. The risk is what they change by accident.

Before a release, Yardstick runs a fixed regression suite of 300 cases per role family. Each case has an agreed outcome, and every case that has ever caused a problem is added to it. The new version and the current one score the same cases side by side, and three things get compared:

  • Scores that moved, and by how much

  • Applicants who crossed the cut line in either direction

  • Evidence links that now point to a different passage

That last check catches a quiet class of failure. A score that stays the same while its reasoning changes usually means the model is right for a new and worse reason.

Movement on its own is not a defect. A person reviews every change, and any change nobody can explain is treated as one.

“I do not need the model to be perfect. I need to know which version ranked my shortlist, and what it was tested against.”

Head of engineering hiring, Hexlid

What blocks a release

A release ships only when all of these hold. Any one failing stops it:

  • Agreement with reviewers on the held-out set is no lower than the current version, overall and per criterion

  • No role family shows a new adverse-impact flag, or a wider gap on an existing one

  • Every regression case that crossed the cut line has a written explanation a reviewer has accepted

  • Every sampled rating has an evidence link that resolves to a real passage in the application

  • The same application, scored twice, gets the same result

When a release is blocked, the current version stays in place and customers see no change.

A release that passes still reaches roles gradually. It runs alongside the current version on live briefs first, and the two shortlists are compared before anything switches over. Nothing reaches an applicant from a comparison run. Every send still waits at the approval step.

What testing does not tell you

Testing before release shows that the model behaves on cases like the ones in the set. It cannot promise that for your role, your applicants and your brief. That is why the bias report runs on your opening, and why the evidence link stays on every rating in production.

It also cannot make a decision for you. Yardstick ranks and recommends. A person on your team reads the shortlist, checks the reasoning and makes the call. The testing exists so that what they are checking was worth checking.

Planning a rollout like this?

Elin Ahlberg, Sales lead

Planning a rollout like this?

Elin Ahlberg, Sales lead

Start with one open role

Create a free website with Framer, the website builder loved by startups, designers and agencies.