Skip to content
Surhires

Data quality

Shipping, being extended

Stop paying twice for the same candidate

Two recruiters calling the same person about two different jobs on the same afternoon is not a data problem. It is a reputation problem.

Surhires detects duplicates as records arrive, using exact matches on email and normalised phone number plus fuzzy matching on name, employer history and resume content. Matches are shown side by side before the record is created, and existing duplicates can be merged without losing activity history from either copy.

By Surhires Editorial · Published · Reviewed

In the product: matchBatch.ts carries the fuzzy-match logic today; the dedicated merge interface is in build for Wave 1

Duplicates are created faster than they are cleaned

Every route into the database creates them. A candidate applies to two jobs with two email addresses. A sourcer adds someone from LinkedIn who was already parsed from a job board six months ago. A bulk import of nine hundred resumes overlaps with four hundred records you already held. Nobody sets out to duplicate anything; the database does it while the desk is working.

The cost is rarely a storage cost. It is the candidate who gets three approaches in a week and concludes the agency is disorganised, the search that returns the stale copy without the recent conversation, and the placement fee argued over because two consultants each have a record showing first contact.

The only durable fix is to catch it at the moment of entry, when a human is present and can make the call in two seconds.

How the match is scored

Matching runs in layers. An exact match on a normalised email address or a phone number reduced to its international form is treated as near-certain. Below that, a fuzzy comparison of name variants, current and previous employer, education and the shingled text of the resume produces a confidence score.

Name matching alone is deliberately weak evidence. There are a great many people called David Smith, and a system that merges on name will destroy records. The score only becomes actionable when two independent weak signals agree, such as a matching employer history and a matching resume fingerprint.

The asymmetry in the scoring is intentional and it is worth understanding before you tune it. A missed duplicate costs you an awkward second phone call and a record that has to be merged later. A wrong merge silently combines two people, and the damage shows up months afterwards in a submittal to a client with somebody else's career history attached. The defaults are set to make the first mistake rather than the second, and that is the trade a conservative threshold is buying.

  • Email compared after lowercasing, dot-stripping and plus-address removal
  • Phone numbers normalised to E.164 before comparison, so a leading zero is not a new person
  • Employer and education history compared as sets rather than as strings
  • Resume text compared by shingled fingerprint, which survives reformatting
  • Score thresholds configurable per tenant, with a conservative default

Caught at entry, not in a monthly clean-up

When a recruiter adds a candidate manually, a likely duplicate is shown before the record is created, with the two profiles side by side and the differences highlighted. The recruiter can open the existing record, merge the new information into it, or confirm that these really are two different people.

Bulk imports run the same check. The import preview reports how many rows matched existing records, and lets you decide once for the whole file: skip, merge into the existing record, or create anyway and flag for review. That decision is recorded against the import so the choice can be audited later.

Entry time is the only moment when the decision is cheap. The person adding the record has the resume open, remembers the conversation, and can tell in two seconds whether the David Smith already in the system is the one they just spoke to. A month later that context is gone, and the same decision takes ten minutes of reading two activity streams. That is why a monthly clean-up never actually happens: the work grows and the information needed to do it shrinks.

Merging without losing history

A merge is designed to keep both activity streams. Emails, calls, notes, submittals, interviews and stage changes from both records land on the surviving profile in chronological order, each still attributed to the recruiter who logged it.

Field conflicts are resolved explicitly rather than by last-write-wins. Where two records disagree on current employer or salary expectation, the merge screen shows both values and asks. The discarded value is kept in the merge record, so a bad merge can be understood and, within a retention window, reversed.

Silent field resolution is where most merge tools do their quiet damage. Taking the most recently updated value looks sensible until the most recently updated record is the thin one a sourcer created from a public profile, and the salary expectation the candidate actually stated in a conversation eight months ago is overwritten by a blank or a guess. Asking is slower by a few seconds per conflict and it is the difference between a merge that improves a record and one that degrades it.

Ownership and fee attribution survive the merge

In an agency, the question behind most duplicates is who owns the candidate. The merge preserves the earliest genuine first-contact event and its owner, rather than whichever record happened to survive.

Ownership rules are configurable: first contact wins, most recent meaningful contact wins, or ownership is held by the desk rather than the individual. Whichever rule you use, the merge applies it consistently instead of leaving it to whoever clicked the button.

The word genuine is doing real work in that first sentence. A record created by a bulk import of a purchased list is not first contact, and neither is a scrape of a public profile. The event that counts is one the candidate took part in, which is the same definition the retention clock uses, and it is deliberately narrow because the alternative is an ownership system that rewards whoever imports the largest file.

Finding the duplicates already in there

A background pass scores the existing database and produces a review queue ordered by confidence, so the highest-certainty pairs can be cleared quickly and the ambiguous ones can be left alone.

The queue is worked in bulk. Pairs above a confidence threshold you set can be merged as a batch, with a report of what was merged and a window in which the batch can be undone. Nothing is merged silently.

The pragmatic advice is to stop before the queue is empty. In any database of real size there is a long tail of pairs that sit at middling confidence, where two people genuinely might be one person and nothing in the record settles it. Working that tail costs hours and produces the merges most likely to be wrong. Clearing the top of the queue and leaving the ambiguous middle for entry-time checks to resolve naturally is the better use of an afternoon.

What ships today and what is still being built

The matching logic ships today. Records arriving through parsing, import and manual entry are scored against the existing database using the layered comparison described above, and likely matches are surfaced rather than silently created.

The dedicated merge interface is in build for the Wave 1 release. That is the side-by-side conflict screen, the reversible bulk merge, the merge record that keeps the discarded values, and the background sweep that produces a ranked review queue over the existing database. We describe those in the design tense because they are being written now, not because they are hypothetical.

In the meantime a flagged match is a prompt to a recruiter, who can open the existing record and work it rather than creating a second one, which prevents most of the duplicates that would otherwise be created but does not clean up the ones already there. If deduplication of a large legacy database is the reason you are evaluating us, ask for the current build state before you commit to anything. That is a better conversation than discovering the gap in week three.

Where duplicate detection does not help

It cannot tell you whether two records are two people. It can tell you how much evidence there is either way, and the honest ceiling on that evidence is low for common names in large markets with thin records. Two contractors with the same name, the same city and no email on either record are not resolvable by any amount of scoring, and the correct behaviour is to refuse to guess and put the pair in front of somebody who can ring one of them.

It also does not fix the process that produces duplicates. Two desks working the same market with no shared ownership rule will keep colliding, and merging the results every month treats the symptom. The merge is a clean-up; the fix is a decision about who owns which market and what happens when two recruiters reach the same person in the same week, and that decision is not something software can make for a business.

And nothing here reaches copies that have left the tenant. A shortlist forwarded to a client, a spreadsheet a recruiter keeps for their own pipeline, an export taken before the merge ran: those are all still out there with the old record in them. That is an argument for exporting less rather than an argument against merging, but it is worth saying plainly because deduplication inside one system is frequently sold as though it tidied the whole world.

What you get

Normalised email match

Case, dots and plus-addressing removed before comparison.

E.164 phone match

Numbers compared in international form so formatting differences do not hide a match.

Fuzzy name variants

Nicknames, transliterations and reversed given and family names handled.

Employer history match

Career history compared as a set, which survives job title rewording.

Resume fingerprint

Shingled text comparison that still matches after a resume is reformatted.

Cross-source scoring

Job board, import and manually entered records scored against each other, not only against manual ones.

Entry-time warning

Likely duplicates shown side by side before a new record is created.

Import-time dedupe

Bulk files report matches in the preview and apply one decision to the batch.

History-preserving merge

Both activity streams survive, still attributed to the original recruiter.

Explicit conflict resolution

Disagreeing fields are shown and chosen, never silently overwritten.

Match audit record

Every proposed match keeps the signals that produced its score, so a decision can be explained later.

Ownership rules

First contact, latest contact or desk ownership applied consistently on merge.

Reversible batches

Bulk merges produce a report and can be undone inside a retention window.

Confidence thresholds

Per-tenant tuning, with a conservative default that prefers a review over a merge.

Questions recruiters ask

Will it merge two different people with the same name?

Not on a name alone. A name match on its own is treated as weak evidence and never triggers an automatic merge. Two independent signals have to agree before the score reaches a threshold where anything is proposed, and even then a person confirms.

Can a merge be undone?

Yes, within a retention window. The merge record keeps the discarded field values and the original record boundaries, so the split can be reconstructed. After the window closes the merge is permanent, which is stated on the merge screen.

What happens to attachments from both records?

All of them survive. Resume versions are kept in date order rather than deduplicated, because a two-year-old resume is often the only record of a role the newer one has dropped.

Does this affect candidate ownership and commission?

It respects it. The merge applies the ownership rule your agency has configured and preserves the earliest genuine first-contact event, rather than leaving ownership to whichever record happened to survive.

How does it handle candidates who apply with a personal and a work email?

That case is caught by the secondary signals rather than by the email. Matching phone number, employer history or resume fingerprint will surface the pair, and both addresses end up on the merged record as separate contact points.

Is the existing-database sweep safe to run on a live tenant?

Yes. The sweep only scores and queues; it never merges on its own. You choose what to action, and bulk actions are reported and reversible.

How often does it get it wrong, and which way does it fail?

We do not publish an accuracy figure, because the number would be measured on our data rather than yours and the difference is large. What we will state is the direction of the failure. The defaults are tuned to miss duplicates rather than to merge people who are not the same, so expect some pairs to reach the review queue instead of being caught, and expect very few wrong merges.

See it against your own reqs

Bring one live role and three resumes. In twenty minutes you will see the match scores, the shortlist and the placement invoice that comes out the other end.