Skip to content
Surhires

Parsing

Shipping today

Resumes become records, not attachments

A resume sitting in a storage bucket is a file. A parsed resume is a record you can search, match and report on.

Surhires extracts text from PDF, DOCX and scanned resumes, then structures it into normalised fields: contact details, current and previous employers with dates, education, certifications, skills split from tools, seniority band, location and stated salary or rate expectation. The recruiter reviews the result before it is written.

By Surhires Editorial · Published · Reviewed

In the product: parse-resume and extract-resume-text edge functions, CandidateResumeSection.tsx, ResumeAutofillDialog.tsx and ResumeAIDialog.tsx

Extraction and structuring are two different problems

Getting text out of a document is the easy half, and it still fails often enough to matter. Two-column layouts interleave, tables lose their rows, scanned pages need optical recognition, and a resume exported from a design tool can be a single image with no text layer at all.

Structuring is the harder half. A block of correct text still has to be resolved into which employer goes with which date range, whether a phrase names a skill or a product the person used, and what a job title implies about seniority in that industry. Both halves are separated in Surhires so a failure in one is visible rather than producing a plausible-looking wrong record.

Keeping them separate is what makes a failure diagnosable. When a record comes out wrong, the question is whether the text was read incorrectly or read correctly and interpreted badly, and those have completely different fixes. The first is a document problem you solve by re-uploading a text version or accepting lower confidence. The second is a synonym or a title-mapping problem you solve once, for every resume that arrives afterwards.

Normalisation is what makes the record searchable

Raw extraction gives you the words on the page. Normalisation maps them to values you can filter on. React, ReactJS and React.js become one skill. Senior Software Engineer II resolves to a seniority band. A location becomes a place that can be measured for travel time rather than a string.

Skills, tools and certifications are held in separate fields rather than one blended list, because they answer different questions. AWS Solutions Architect is a certification with an expiry; AWS is a platform; Terraform is a tool. Collapsing them is why so many databases cannot answer a precise brief.

Seniority is the field that deserves the most suspicion, and the interface treats it that way. A title means different things in different companies and different markets, and a five-person startup and a bank do not use the word senior to describe the same job. The band is derived from title, scope and years together, it is always displayed as an inference rather than as a fact, and it is the field recruiters correct most often, which is exactly as it should be.

  • Skill synonym mapping with a per-tenant override list
  • Employer names resolved and deduplicated across records
  • Date ranges parsed into a continuous career history with gap detection
  • Seniority inferred from title, scope and years, and always shown as an inference
  • Currency detected on salary and rate expectations rather than assumed

The recruiter reviews before the record is written

Parsed values are presented next to the source text they came from. Anything the parser inferred rather than read is marked as an inference, and low-confidence fields are highlighted rather than silently accepted.

This is not a formality. A parser that writes confidently wrong values into a database quietly poisons every search that runs afterwards, and nobody notices for months. A ten-second review at capture prevents the kind of error that is very expensive to find later.

The economics of that ten seconds are worth spelling out. A wrong seniority band on one record costs nothing until the day a search for mid-level candidates excludes exactly the right person, and at that point the cost is a placement nobody knows they lost. Errors in a candidate database are not noticed as errors; they are noticed as an absence, which is why they survive for years.

The original file never goes away

The uploaded document is kept alongside the parsed record, and every subsequent version is kept too. A two-year-old resume frequently holds a role the newest one has dropped, and clients periodically ask for the original rather than a formatted version.

A branded, formatted version can be generated for submittal, with contact details masked where the client relationship requires it. The original, the parsed text and the formatted output are three artefacts on one record, not three copies in three places.

Keeping the original also means a parse is never a one-way door. When the synonym list improves, or a title mapping is corrected, or a better extraction path is added for a format that used to fail, the stored document can be parsed again. A system that discards the source after extraction has permanently capped the quality of its own records at whatever the parser could do on the day the file arrived.

Parsing at volume

A folder of resumes, an inbox of applications or an export from a previous system can be processed as a batch. Each row reports its own outcome: parsed cleanly, parsed with low confidence, matched to an existing candidate, or failed with a reason.

Batches are reviewable before they are committed, and duplicate detection runs inside the batch as well as against the existing database, so importing nine hundred resumes does not create four hundred duplicates.

In-batch checking is the part that is easy to leave out and expensive to omit. A folder exported from an older system routinely contains the same person three times, under three filenames, from three different years. Checking each row only against the existing database would let all three through as new records, and the clean-up afterwards costs more than the import saved.

Languages, formats and the limits

Parsing handles the common English-language resume formats well, including the dense multi-page format usual in Indian IT staffing and the one-page format usual in the United States. Non-English resumes are extracted and stored but structured with lower confidence, and the record says so rather than pretending otherwise.

Where a document cannot be parsed at all, it is stored, flagged and queued for manual entry rather than dropped. A silent failure is worse than a visible one.

The formats that reliably cause trouble are worth naming so you can test them rather than discover them. Resumes exported from design tools as a single image, heavily tabular consultancy CVs where a project grid carries the actual experience, federal-style documents that run to fifteen pages, and any layout where the dates sit in a narrow column beside the employer. All of those are readable; all of them are more likely to arrive with fields flagged for review.

What a first bulk import week looks like

The first real test of a parser is never the demo file, it is the backlog: a folder of several thousand documents accumulated over years, in every format anybody ever sent. The mistake is to budget for the runtime and not for the review, because the batch runs overnight and the queue of low-confidence rows is what actually takes the week.

The order that saves the most time is to parse a sample of fifty first, deliberately picked across your worst formats rather than your cleanest, and read what came out. Almost every desk finds two or three synonym problems specific to its market, where an industry term or an internal job title maps to the wrong skill or the wrong band. Promoting those corrections to the tenant override list before the full batch runs applies them to every subsequent document instead of to fifty records by hand.

Then run the batch and work the outcomes by category rather than by row. Everything that failed extraction goes to manual entry. Everything that matched an existing candidate goes to the duplicate queue. Everything flagged low-confidence gets reviewed for the two or three fields that were actually uncertain, not re-read end to end. The credit cost for the batch is shown before it runs, with scanned pages weighted higher, so the size of the job is a decision rather than a surprise.

Where parsing does not help

It cannot recover what the document never contained. Reason for leaving, notice period, salary expectation, right to work, willingness to relocate and whether the person is actually looking are the fields that decide whether a candidate is submittable, and almost none of them appear on a resume. Those come from a conversation, and a parsed record is the thing that makes the conversation shorter, not the thing that replaces it.

It also cannot judge truth. A resume is a claim, and the parser reads claims exactly as written. Overstated seniority, a contract described as a permanent role, dates rounded to hide a gap: all of that is structured faithfully and none of it is verified. Screening, referencing and verification are separate pieces of work, and a confidently structured record can make an unverified claim look more solid than it did as prose, which is worth being alert to.

And it is bounded by the document quality you actually receive. Non-English resumes and unusual layouts are extracted and stored at lower confidence, and the record says so. A document that cannot be parsed is flagged and queued rather than dropped. We do not publish a headline accuracy number, because one measured on clean single-column English resumes would describe our sample rather than your inbox, and it would be the most misleading number on this page.

What you get

Multi-format extraction

PDF, DOCX, DOC, RTF and plain text, plus optical recognition for scanned pages.

Layout handling

Two-column and table-based resumes reconstructed in reading order rather than interleaved.

Skill normalisation

Synonyms collapsed to one value, with a per-tenant override list you control.

Separate skill classes

Languages, tools, platforms and certifications held as distinct fields.

Career history

Employers and date ranges resolved into a continuous history with gaps flagged.

Seniority inference

A band derived from title, scope and years, always labelled as an inference.

Currency-aware pay

Salary and rate expectations captured with the currency they were quoted in.

Confidence flags

Low-confidence fields highlighted for review instead of silently written.

Source-text linking

Every parsed value shown next to the text it came from.

Version history

Every uploaded resume kept in date order alongside the parsed record.

Re-parse from source

The stored original can be parsed again when synonyms or extraction improve.

Formatted output

Branded submittal version with optional contact masking for blind CV workflows.

Batch parsing

Folders and mailbox imports processed with per-row outcomes and in-batch dedupe.

Failure queue

A document that cannot be read is stored, flagged and queued for manual entry rather than dropped.

Questions recruiters ask

How accurate is it?

Accuracy varies by format and we do not publish a single headline number, because a number measured on clean single-column English resumes would not describe what your inbox contains. What we do instead is surface confidence per field so you can see where the parser was unsure.

Does it handle scanned resumes?

Yes, through optical character recognition, at lower confidence than a text-layer document. Scanned pages are marked as such on the record so a later reader knows why a field was thin.

Can we correct a bad parse?

Yes, on the record. Corrections to skill synonyms can be promoted to a tenant-level override so the same mistake is not repeated on the next thousand resumes.

Is the resume text sent to a third-party model?

Text extraction and structuring use named subprocessors that are listed on the trust page. Enterprise accounts can restrict which providers are used. Demographic fields are never part of what is sent.

What happens to resumes for candidates we later delete?

The erasure workflow removes the original file, the extracted text and the parsed record together. Nothing is left in storage, which is the part most systems get wrong.

Does parsing consume AI credits?

Yes, at a published weight per document, with a higher weight for scanned pages that need optical recognition. Bulk imports show the credit cost before the batch runs.

Every vendor claims the best parser. What should we actually test?

Give each of them the same thirty documents from your own inbox, chosen for being awkward rather than clean: a scanned page, a design-tool export, a two-column layout, a fifteen-page consultancy CV and something not in English. Then judge on whether the wrong fields were flagged rather than on how many were right, because a quiet error is the one that costs you a placement.

See it against your own reqs

Bring one live role and three resumes. In twenty minutes you will see the match scores, the shortlist and the placement invoice that comes out the other end.