Skip to content
Surhires

Scorecards

Shipping, being extended

Structured feedback instead of an opinion

Two interviewers who both say the candidate was strong have not agreed on anything until you know what each of them was measuring.

A scorecard asks each interviewer to rate a candidate against the requirements the requisition listed, with a comment supporting every rating. Ratings roll up across the panel so agreement and disagreement are visible. Feedback capture and the score dial ship today; structured scorecard templates are being extended.

By Surhires Editorial · Published · Reviewed

In the product: InterviewFeedbackDialog.tsx, InterviewFeedbackList.tsx and InterviewScoreDial.tsx ship today; structured scorecard templates are being extended

Comparable feedback is the whole objective

Structured scoring exists for one reason: to make four interviews about the same candidate comparable with one another, and to make one candidate comparable with the next. Free-text feedback fails both tests. It cannot be aggregated, it cannot be counted, and it cannot be defended.

The failure is not obvious in the moment. A debrief with four thoughtful paragraphs feels rigorous. But when the paragraphs disagree there is no way to tell whether the interviewers saw different things or measured different things, and when a decision is questioned six months later, four paragraphs of impression are not a record of why a person was hired.

Structured scoring is what makes a panel comparable and a decision defensible. That is the entire argument.

What ships today

Interview feedback capture works now. An interviewer records feedback against the interview, the feedback list shows every submission against a candidate, and the score dial gives a visual summary of where the panel landed. A recruiter reviewing a shortlist can see who has submitted, who has not, and how the scores compare.

This is real, usable structure and it is already better than a reply-all thread. What it does not yet have is full template configurability, which is the part being extended.

What is being extended

The work in progress is scorecard templates: defining a criterion set per requisition or per role type, assigning criteria to particular stages, and setting a rating scale with written anchors describing what each point on the scale means.

Anchors are the part that does the work. A one-to-five scale with no anchors produces a distribution centred on four, because interviewers are reluctant to write down a two about somebody they have just met. A scale where three is defined as meets the requirement as written and four as exceeds it with evidence produces ratings that mean something. Until templates land, criteria are captured against the interview rather than configured centrally.

  • Criterion sets defined per role type rather than retyped per requisition
  • Criteria assigned to specific stages so the panel divides coverage between them
  • Rating scales with written anchors describing each point
  • A mandatory supporting comment on every rating
  • Aggregated panel view showing agreement and disagreement per criterion

Every rating needs evidence attached

A rating on its own is a number somebody chose. A rating with a supporting comment is a claim with evidence, and requiring the comment changes what interviewers do: writing down why forces the interviewer to notice whether they actually have a reason.

It also makes the debrief shorter and better. Instead of reconstructing four impressions from memory, the panel reads what each person recorded at the time and spends the meeting on the criteria where they disagreed. Most panel disagreement, once the evidence is visible, turns out to be about coverage rather than judgement.

Timing is part of the structure. Feedback submitted before an interviewer has spoken to the rest of the panel is worth more than feedback written after a corridor conversation, because a panel that talks first converges on whoever spoke first. Outstanding feedback is flagged as an exception with the interview date attached, so a debrief two days later does not quietly become the moment when four people decide together what they each thought separately.

Disagreement is a signal worth keeping

Averaging a panel into one number destroys the most useful information the panel produced. Two fives and two twos average to a three and a half, which describes nobody's view and hides the fact that half the panel would not hire this person.

Panel views show the spread per criterion alongside any aggregate, so a split is visible rather than smoothed away. A candidate with wide variance is a candidate to discuss, not a candidate to rank. The aggregate exists because people want one, and the spread sits next to it so the aggregate is not mistaken for a verdict.

The record a challenged decision needs

Where a hiring decision is questioned, whether by a candidate, a client or a regulator, the useful artefact is a contemporaneous record showing what was assessed, by whom, against what criteria, with what evidence. Scorecards produce exactly that as a by-product of doing the interview properly.

Surhires stores that record, keeps demographic data separately from it so it cannot influence a decision it is meant to measure, and exports the whole trail per requisition. Whether your process is lawful in a given jurisdiction depends on how you run it. What the product does is make sure the evidence exists.

What you get

Feedback capture

Structured interview feedback recorded against the interview rather than sent as a reply.

Feedback list

Every submission against a candidate in one view, including who has not submitted yet.

Score dial

A visual summary of where the panel landed, readable before opening individual submissions.

Criterion templates

Criterion sets per role type, being extended so they are configured once rather than retyped.

Stage-assigned criteria

Criteria allocated to stages so a panel divides the brief instead of overlapping on it.

Anchored rating scales

Each point on the scale described in words, so a three means the same thing to everyone.

Mandatory comments

A rating without a supporting comment cannot be submitted; evidence travels with the number.

Panel spread

Disagreement shown per criterion rather than averaged away into a single misleading figure.

Outstanding feedback flags

An interview whose feedback has not landed is surfaced instead of silently delaying a decision.

Requisition linkage

Criteria derive from the requirements the requisition listed, not from a generic template.

Demographic separation

Equal-opportunity data is held apart from the scoring record and is not visible to interviewers.

Exportable trail

The whole scoring record per requisition, exportable as CSV or JSON when a decision is questioned.

Questions recruiters ask

What can we use today?

Interview feedback capture, the feedback list and the interview score dial all ship now, so a panel can record structured feedback and a recruiter can compare it. Fully configurable scorecard templates with per-role criterion sets and anchored scales are being extended, which is why this page is marked as shipping and being extended rather than as finished.

Should scores decide the hire?

No. A scorecard structures evidence for a human decision. Ranking candidates purely by aggregate score reintroduces the problem structured interviewing was meant to solve, because it treats four ratings from four differently prepared interviewers as though they were measurements on one instrument. Use the scores to focus the conversation, not to replace it.

Why is a comment required on every rating?

Because a number without a reason cannot be checked, and requiring the reason changes the interviewer's behaviour. Writing down why forces them to notice whether they have evidence or an impression. It also makes the debrief shorter, since the panel reads what was recorded at the time instead of reconstructing it from memory.

Can clients score candidates on the same scorecard?

That is the intent once client portal feedback lands in Wave 1: a client contact leaves structured feedback against the same criteria, so agency and client assessments are comparable. Today a client's feedback is recorded by the recruiter on their behalf, which is workable but adds a transcription step and a delay.

Does this help with bias?

It helps by fixing what is assessed and requiring evidence, which reduces the room for an unexamined impression to carry a decision. It does not eliminate bias, and we will not claim it does. Demographic data is stored separately from the hiring record and is not visible to recruiters or interviewers, so it cannot influence what it is meant to measure.

How many criteria should a scorecard have?

Four to six per stage is usually right. More than that and interviewers rate everything the same because they have not gathered evidence on all of it. Assigning criteria across stages, so each interview covers a few properly, produces better data than asking every interviewer to rate twelve things after one conversation.

See it against your own reqs

Bring one live role and three resumes. In twenty minutes you will see the match scores, the shortlist and the placement invoice that comes out the other end.