Skip to content
Surhires

AI matching

Shipping today

Matching that shows its working

A match score nobody trusts is worse than no score at all. Every point in ours traces back to a requirement in the brief.

Surhires scores candidates against a live requisition using the structured requirements in the brief and the parsed content of the candidate record. Each score breaks down by requirement, showing which must-haves were met, which were missing and which were inferred, so a recruiter can accept or overrule the ranking on evidence.

By Surhires Editorial · Published · Reviewed

In the product: match-candidate-to-job edge function, recruitmentMatching.ts, matchBatch.ts, CandidateJobMatches.tsx, CandidateMatchDialog.tsx and MatchScoreBadge.tsx

The problem with a bare score

Most matching tools return a percentage and nothing else. The recruiter has no way to tell whether an eighty-seven means the candidate has every must-have and a weak location fit, or four of six must-haves and an excellent title match. So the score gets ignored, and the tool becomes a slow way to browse.

A score is only useful if it is auditable at the level of a single requirement. Ours decomposes: this candidate met the Kubernetes requirement through two years at a named employer, met the seniority band, missed the security clearance, and is forty minutes outside the stated radius. That is a ranking a recruiter can act on and, when it is wrong, correct.

The breakdown also changes what happens when the tool is wrong, which it will be. A bare percentage that misses gives a recruiter no way to diagnose anything, so the conclusion is that the matching does not work and the feature gets abandoned. A breakdown that misses points at its own mistake: the requirement was scored as met because the resume named the tool once, in a list, six years ago. That is a fixable problem, and fixing it usually means fixing the brief.

The brief has to be structured before matching means anything

Matching quality is set by intake, not by the model. A requisition written as a paragraph of prose can only be matched loosely, because nothing in it distinguishes a hard requirement from a preference.

The Job Intake Agent turns a client brief, an emailed spec or a pasted job description into structured requirements: must-haves, nice-to-haves, seniority band, location and radius or remote policy, salary or rate band, work authorisation and the interview stages. A recruiter reviews and corrects that structure in under a minute, and every subsequent match runs against it.

The review step is where most of the value is created, and it is the step teams try to skip. Client job descriptions routinely list eleven must-haves, of which three are genuine and the rest are a wish list somebody pasted from the last advert. A recruiter who has spoken to the hiring manager knows which is which. Moving two requirements from must-have to nice-to-have takes fifteen seconds and changes the ranking more than any model choice will.

  • Must-haves are scored as pass or fail, never averaged away by a strong nice-to-have
  • Nice-to-haves contribute weighted points and are shown separately
  • Seniority is matched on a band, not on a job title string
  • Location matches on travel time or a stated remote policy, not on a city name
  • Work authorisation and clearance are hard gates by default

Both directions, because desks work both ways

Given a requisition, the system ranks candidates. Given a candidate, it ranks the open requisitions they fit, which is what a recruiter needs when a strong person comes off a finished search and the question is where else they land.

Reverse matching runs across the whole database rather than only against candidates already attached to a job, which is where database rediscovery actually pays for itself.

Reverse matching is also the honest test of whether a database is worth what it costs to maintain. A desk that runs it on a strong candidate and gets three plausible open requisitions back has a working book of business. A desk that gets nothing back is holding a database of candidates for jobs it no longer wins, and that is worth knowing in a week rather than in a year.

Where the model helps and where it does not

Semantic comparison is genuinely good at recognising that a candidate who has run distributed systems on AWS has done the thing the brief calls cloud infrastructure experience, even though neither phrase appears in the other document. That is the work worth automating.

It is much weaker at judgement calls that depend on context it does not have: whether this particular client will accept a career changer, whether a two-year gap matters, whether a contractor will convert. Those stay with the recruiter, and the interface is built so overruling the ranking takes one click and is recorded.

There is a third category worth naming, which is the things the resume simply does not contain. Whether somebody wants this job, whether they are interviewing elsewhere, whether they will move for the money on the table, and what their current employer will do when they resign are the facts that decide most placements, and none of them are in the document being scored. A match score ranks fit. It has nothing at all to say about intent.

Feedback that improves the next brief

When a recruiter dismisses a high-scoring candidate, they pick a reason from the same coded disposition list the pipeline uses. Those reasons accumulate against the requisition.

Eleven dismissals for salary above budget is a signal about the brief rather than the matching, and the requisition health view surfaces it. That loop is more valuable than incremental model tuning, because it changes the conversation with the client.

The distinction the coded reason makes is between a bad match and a good match for a bad brief, and those look identical on a board. Dismissals clustered on one requirement mean the requirement is wrong or the market does not contain it. Dismissals scattered across every reason mean the ranking is genuinely off. The first is a client conversation, the second is a configuration problem, and telling them apart takes the codes.

Cost, credits and control

Matching consumes AI credits, and the cost is visible rather than buried. Professional includes one hundred credits per user per month, an operation-weight table shows what each action costs, and additional credits are an add-on with a published price rather than a quote.

Matching can also be disabled per tenant or per requisition. Where a client or a jurisdiction requires that automated tools not be used to screen candidates, turning it off is a setting, and the requisition records that it was off.

Publishing the weights is a deliberate choice about how the product gets used. Where AI cost is opaque, teams either avoid the feature in case it is expensive or run it on everything and get an unexpected invoice. Neither is good. A visible per-operation weight, shown before a batch runs, lets a desk decide that scoring a whole pool against a requisition is worth it and scoring the entire database against every open job is not.

What the first fortnight of scoring actually looks like

The first scores usually look wrong, and the reaction is almost always to distrust the tool. In practice the brief is what needs fixing. A requisition imported as prose, with eleven must-haves and a location written as a city name, will rank people who are technically qualified and unavailable, and the score is faithfully reporting a bad question.

The sequence that works is small and specific. Pick one live requisition. Score it. Then open the requirement breakdown on five candidates you already have a firm opinion about, including two you would definitely submit and two you definitely would not. Where the breakdown disagrees with you, read which requirement drove the difference. In most cases it is a must-have that should be a nice-to-have, a seniority band set one level too high, or a radius that was never converted from a city name into a travel time.

Then set the weightings for that job type and leave them alone for a fortnight. A contract desk usually weights availability and rate far above title fit; a retained search inverts that. Changing the weightings every time a ranking surprises you produces a configuration nobody understands and no way to tell whether anything improved.

Where matching does not help

It cannot tell you who wants the job. Fit and intent are different questions, and the second one decides the placement. A candidate scoring in the nineties who is three weeks from a counter-offer they will accept is worth less than a candidate in the seventies who has already told you they are leaving. Nothing in a resume carries that, which is why the ranking orders a call list rather than a shortlist.

It is also bounded by what you hold. A score is a comparison against your own database, so on five hundred records matching mostly saves reading time, and the value compounds with size and with how much of that size was captured with real structure. Scoring a thin database produces confident rankings of insufficient data, and the confidence is the dangerous part.

The regulatory boundary is deliberate and worth stating plainly. The tool ranks; it never rejects. A person makes every decision to progress or decline, matching can be turned off per requisition where a client or a jurisdiction requires it, and the requisition records that it was off. The controls that sit around this, including the separated demographic store and the scored-candidate register an independent bias audit would need, are in build for the Wave 1 release rather than shipping today.

What you get

Requirement-level breakdown

Every point in the score traces to a named requirement in the brief, met, missed or inferred.

Hard gates

Must-haves, work authorisation and clearance fail the match rather than being averaged away.

Structured intake

The Job Intake Agent turns a prose brief into scoreable requirements in about a minute.

Semantic skill matching

Recognises equivalent experience described in different words across resume and brief.

Seniority bands

Matched on a seven-level career band rather than on job title text.

Travel-time location

Radius expressed as commute time or an explicit remote policy.

Reverse matching

Rank open requisitions for a given candidate across the whole database.

Batch scoring

Score an imported file or a whole pool against a requisition in one run.

Per-job-type weightings

A contract desk can weight availability and rate above title fit, with the weighting shown on the requisition.

One-click overrule

Dismiss or promote a ranking with a coded reason, recorded on the requisition.

Explanation on the submittal

The requirement breakdown is stored with the submittal, so the reason a candidate was presented survives the session.

Brief feedback loop

Dismissal reasons roll up into requisition health so a bad brief surfaces early.

Published credit weights

An operation-weight table shows what each AI action costs before you run it.

Per-requisition off switch

Automated scoring can be disabled where a client or jurisdiction requires it, and the record says so.

Questions recruiters ask

Does the score decide who gets rejected?

No. Ranking orders a list; a person decides. That is a deliberate design choice and also the safer position where automated employment decision rules apply, because a tool that only ranks and never rejects has a much narrower compliance surface than one that screens people out.

How is this different from keyword search?

Keyword search finds documents containing a string. Matching compares a candidate against a structured requirement set and reports which requirements were met and how. A keyword search cannot tell you that four of six must-haves were satisfied.

What data is sent to the model?

The parsed candidate record and the structured requisition. Demographic fields are stored separately and are never included. The subprocessor list names every model provider involved, and Enterprise accounts can restrict which providers are used.

Can we tune the weightings?

Yes, per job type. A contract desk usually weights availability and rate far higher than a retained search does. Weightings are visible on the requisition rather than buried in a settings page.

What does a match cost?

Scoring one candidate against one requisition is a single credit under the published operation-weight table. Professional includes one hundred credits per user per month and further credits are an add-on at a published price.

Does matching work on a small database?

Yes, but the value compounds with size. On five hundred records matching mostly saves reading time. On fifty thousand it is the only realistic way to find the person you already know.

How do we know this is not a keyword count with extra steps?

Read the breakdown on a candidate whose resume never uses the words in the brief. If the requirement was met through equivalent experience, the breakdown names the employer and the period it credited. Where a requirement was met by a literal term appearing once in a skills list, it says that too, which is usually the point at which a recruiter demotes the requirement.

See it against your own reqs

Bring one live role and three resumes. In twenty minutes you will see the match scores, the shortlist and the placement invoice that comes out the other end.