Research
You never see the people a shortlist left out. Blind second reviews and sampling below the cut line get you closer.
Published

Every shortlist has two halves. You see what happens to the people on it: who interviewed well, who was hired, who was still there after a season. You never see what would have happened to the people who were not.
That is the hard part of measuring screening. Applicants you declined do not come back with a performance review. If the shortlist missed a strong night-shift nurse because her answers were brief, nothing in your data will say so. The miss is silent, and silence reads as success.
This post is about measuring shortlist quality anyway. None of these methods closes the gap entirely. Used together, they narrow it enough to act on.
Why hire outcomes are not enough
The obvious measure is what happens after the hire. Retention at 30 or 90 days, time to competence, a manager’s rating at the end of probation. Those are worth tracking, and they answer a narrower question than it looks.
Outcomes only exist for people you chose. They can tell you whether the shortlist contained good people. They cannot tell you whether it contained the best people available, or whether it left out a group who would have done just as well.
There is a second trap. Outcomes depend on onboarding, the team, the rota and the manager as much as on the screen. A strong season at one terminal and a poor one at another can have nothing to do with how either shortlist was drawn.
So hire outcomes are one signal. The others have to come from the shortlist itself, measured before anyone is hired.
Run a blind second review
The simplest check is a second reader who does not know the first answer. For a sample of applications on each role, a reviewer scores them against the same rubric. They do this without seeing the rank, the score or anyone else’s marks.
Then you compare. Where the blind reviewer and the shortlist agree, you have some confidence. Where they disagree, you have a specific application to open and a specific criterion the disagreement sits on.
Keep it blind in practice as well as on paper:
Hide the rank and the score until the reviewer has finished
Mix shortlisted and declined applications in the same batch, in random order
Give reviewers the brief and the rubric, and nothing else about the role’s history
A reviewer who can tell which pile an application came from will drift toward agreeing with it. People do this without meaning to, which is exactly why the batch has to be mixed.
Sample below the cut line
This is the step most teams skip, and it is the one that speaks to the missing counterfactual. Each week, pull a small random sample of applicants who ranked just below the cut line, and a smaller one from further down. Put them through the blind review alongside the shortlisted ones.
Yardstick suggests 20 applications per role per week from just below the line and 10 from the rest of the list. On a high-volume role that is one reviewer’s afternoon. Within a month it is enough to notice a pattern.
What you are looking for is a declined applicant the blind reviewer would have advanced. One of those is an individual case. Several that share a reason are a finding. Perhaps they all applied without a CV, or answered in a second language. Perhaps they described forklift work without naming the certificate. That shared reason usually points to a line in the brief.
When the sample turns up someone who should have been shortlisted, a person can still advance them. The sample is a measurement and a second chance at once.
Check how much your reviewers agree with each other
A blind review is only as good as the people doing it. If two reviewers reading the same answer with the same rubric reach different scores, their disagreement is noise. You cannot use noise to judge a shortlist.
So measure reviewer agreement first. Have two reviewers score the same set independently, and compare their marks per criterion. Use a measure that corrects for chance, such as Cohen’s kappa for a pair of reviewers, rather than raw percentage agreement. Raw agreement looks high on any criterion where nearly everyone gets the same mark.
Low agreement on a criterion is useful. It usually means the rubric line is vague. “Communicates well” invites five readings. “Explains a shift handover in order, naming what is still open” invites one. Rewrite the line, run the pair again, and only then use that criterion to judge the shortlist.
This protects you from a common mistake: blaming the ranking for disagreement that was already there between people.
What the evidence link makes auditable
All of this depends on seeing why an application landed where it did. In Yardstick every rating carries an evidence link to the passage it rests on. That might be a line in the application, an answer given in screening or a structured reference.
That changes what an audit can do. When a blind reviewer disagrees with a score, you open the link and read the sentence the rating rests on. The cause is usually one of these:
The passage was read correctly, and the rubric line is wrong
The passage was read wrongly
There was no passage, and the rating leaned on a claim with no artifact behind it
Each has a different fix, and you can only tell them apart with the passage in front of you.
The link also keeps the below-the-line sample defensible later. Every decision, including a decision not to advance, is logged with the reviewer, the rubric and the evidence. The audit export carries that record out of the product, for a works council or anyone else who asks how a shortlist was produced.
“We stopped asking whether the shortlist was right in general. We pull twenty names from below the line every Monday and ask whether any of them belonged on it.”
Talent acquisition lead, Bergwall
What this still cannot tell you
These checks measure the shortlist against your rubric. They cannot tell you the rubric is the right one. If the brief asks for the wrong thing, a blind reviewer will happily agree with a shortlist that finds it.
They also cannot fully replace the counterfactual. A declined applicant who might have thrived stays unknown, and no sample will surface every one. What sampling gives you is an estimate of how often it happens and a pattern to look for. A hire-outcome report alone never offers that.
The decision stays with a person throughout. Yardstick ranks and recommends, and every advance or decline is made and recorded by a named person on your team. Measuring the shortlist is how that person knows what they are signing.


