Company
A single fit score is easy to sort by and hard to defend. What we built in its place, and what we chose to leave out.
Published

The first version of Yardstick that anyone outside the team saw had a number next to every name. It was a good number, in the sense that it sorted people in a sensible order. The problem showed up the first time a recruiter asked us what it meant.
We could explain how it was produced. We could not point at anything a recruiter could check. A candidate ranked 81 and a candidate ranked 77 looked different, and nobody could say why in words a hiring manager would accept.
So we took the number apart. This post is about what replaced it, the things we decided never to build, and what those choices cost.
A single score cannot be argued with
Hiring teams disagree with shortlists all the time, and they should. The site lead knows the night rota. The clinical manager knows which certificate is a formality and which one is not. Their disagreement is the most useful signal in the whole process.
A single fit score gives that disagreement nowhere to go. You can accept it or ignore it. If you ignore it often enough, you stop looking at it, and the tool has quietly become a sorting column that nobody trusts.
We wanted the opposite. A shortlist should be something a person can check, argue with and sign. That means showing the case for each name in sentences, with the source of every claim one click away.
What the evidence link is
Every rating in Yardstick carries a link to the thing it came from. That is the whole idea, and most of the product is built around it.
In practice, a row on a shortlist looks like this:
The claim. “Carried an on-call rotation for two years.”
The source. An incident retro on the candidate’s team blog, dated.
The strength. Whether the claim is backed by an artifact, a reference or a structured answer, or only stated in an application.
That last part matters as much as the link. “Led a Postgres migration, claimed in application, no artifact yet” is an honest line. It tells you what to ask about in the interview. A single score would have folded it into the total and hidden the gap, and the interviewer would never have known to ask.
The reasoning is written before anyone sees a rank. The rank is an order, built from ratings you can open. You will still see a number beside each name, and every part of it opens to a passage you can read.
“Our partners stopped asking what the ranking meant once they could open the answer it came from.”
Head of recruitment, Bergwall
What we chose not to build
Some features were requested more than once, and we said no to each of them. We would rather be clear about that than let the product drift.
No autonomous rejections
Yardstick does not decline anyone. It ranks and recommends, and a named person on your team makes and records every decision to advance or decline.
This was the easiest call to make and the one we get asked about most. At 400 applications, an auto-reject threshold is tempting. It would also mean the person who never hears back was turned away by a line of configuration, with nobody who looked. We do not think that is a defensible way to hire, and we do not want the interface to ever look like it did.
The other three were quieter decisions:
No opaque fit score. Any number on a shortlist is a summary of ratings, and each rating points at its evidence.
No sends without approval. Yardstick drafts outreach, and someone on your team approves each send before it goes. The approval is logged.
No product-wide bias claim. Bias testing is documented for the opening it applies to. A single test and a badge would say nothing about your warehouse role in Aarhus.
The trade-offs we accepted
These choices are not free, and we would be misleading you to pretend otherwise.
Reasoning takes longer to read than a number. A recruiter scanning a list of scores can move faster than one reading the case for each name. We think the extra minutes per shortlist are cheaper than a decision you cannot explain later, but they are real minutes.
Evidence also exposes weak records. When a candidate’s case rests on one claimed line, the shortlist says so. Some teams find that uncomfortable at first, because the old process hid the same gaps behind a single impression.
And a human approval on every send is slower than automation. On a high-volume role, that step has to fit into someone’s afternoon. We built approvals to be quick to review, and we left the step itself in on purpose.
What this asks of your team
A shortlist with reasoning works best when someone reads it. That sounds obvious, and it changes how teams work.
The brief matters more, because the reasoning quotes it back to you. When a candidate is ranked highly for a reason you did not intend, you can see the sentence responsible and change it.
Reviewers matter more too. The person approving a shortlist is doing the oversight the process depends on, and their notes become part of the record. If a candidate asks why they did not progress a year later, the answer is retrievable rather than remembered.
We built Yardstick to do the reading and write down why. The deciding stays with you, and the reasoning is there so you can do it well.


