Equal opportunity
In buildCollect equal-opportunity data without letting it touch the decision
Data collected to measure a hiring decision should never be visible to the people making it, and that is an architecture problem before it is a policy problem.
Surhires is building an equal-opportunity data firewall: demographic fields are collected for reporting, stored in a separate table with its own access control, hidden from recruiters, and excluded from anything sent to a model. Dispositions come from a fixed code list so selection rates can be computed and reviewed.
By Surhires Editorial · Published · Reviewed
In the product: in build for Wave 1: demographic data stored in a separate table with its own access control, and a fixed disposition code list
A firewall is an architecture decision, not a policy statement
Equal-opportunity reporting fails in practice for a boring reason. The demographic answers end up in the same table as the hiring data, on the same screen, in front of the same people. The moment a recruiter can see a candidate's race, sex, veteran status or disability status next to their pipeline stage, data collected to measure the decision has become part of the decision, and no policy document written afterwards undoes that.
The fix is structural rather than procedural. The demographic fields live in their own table with their own access control, keyed to the application through a reporting identifier rather than joined into the candidate view. A recruiter working a shortlist never renders them, because the query that builds the shortlist has no permission to read them. Access is a separate grant held by a reporting role, and every grant is logged.
What this gives you is a set of controls that supports obligations you hold as the employer, or as the agency acting for one. It is not a guarantee and it does not make a screening process lawful. A process is lawful or unlawful because of the criteria you choose and how consistently you apply them. The firewall is what lets you show, afterwards, exactly what you did.
What is collected, and why it is asked for at all
The demographic questions on an application exist because reporting obligations exist. Where they apply to you the categories are prescribed rather than invented, the answers are voluntary, a decline-to-answer option is always present, and the form states plainly that the answers play no part in the hiring decision and are not shown to the people making it.
Because the fields are stored separately, that statement is structurally true rather than aspirational. The application record and the demographic record are linked by an identifier used only for aggregate reporting. There is no view anywhere in the product that puts a candidate's name beside their demographic answers, and there is no export that produces one.
Categories are configurable per market, because different jurisdictions prescribe different sets and asking the wrong questions is its own problem. You configure which set applies to which posting. The product does not decide that an obligation applies to you, and it will not invent a category that nobody asked you to collect.
- Voluntary questions with an explicit decline-to-answer option
- Stored in a separate table with its own access control
- Linked by a reporting identifier, never surfaced on the candidate record
- No recruiter-facing view or export joins a name to an answer
- Category sets configurable per market rather than assumed
Nothing demographic is sent to a model
Surhires scores candidates against requisitions, drafts outreach and parses resumes using models. The rule that governs those features is absolute: demographic data is never included in a prompt, a feature vector, an embedding, a training set or a fine-tune. The payload assembled for any model call is built from an allow-list of fields, and the demographic table is not on the list and cannot be added to it by configuration.
Proxies are the harder problem, and an honest answer beats a confident one. A model handed a resume can infer things from a name, a graduation year, a photograph, a school or a postcode. We reduce the obvious surface by excluding images from the parsing payload and by keeping match explanations at field level, so you see which requirement drove each point rather than being handed a number. We cannot promise that a model infers nothing.
That is precisely why the score is shown with its reasons and why a person makes the decision. A score nobody can inspect cannot be defended, and a scoring system that cannot be explained is worse than no scoring system at all, because it launders a judgement into something that looks like a measurement.
Coded dispositions are what make a decision reconstructable
When a candidate does not go forward, the reason is recorded from a fixed list rather than typed: failed technical screen, salary above budget, better-qualified candidate selected, candidate withdrew, candidate declined the offer, requisition withdrawn, or a stated licence or right-to-work requirement not met. A free-text note can sit alongside it, but the code is mandatory and the drop-out cannot be saved without one.
Free-text reasons cannot be counted, and a reason that cannot be counted cannot be reviewed. Coded dispositions turn a pile of individual decisions into a distribution somebody can look at: this client rejects on salary six times in ten, this requisition rejected everybody from one source, this recruiter's screen rejects at twice the rate of the rest of the desk.
The same codes answer a question about one specific decision. Immutable stage history, plus a coded reason, plus the scorecard behind it, reconstructs what happened at the time it happened, rather than requiring somebody to reconstruct it from memory two years later in front of an audience who was not there.
Selection rates, so an adverse-impact ratio can be computed
Adverse impact is a calculation rather than an opinion: the selection rate of one group set against the selection rate of the group with the highest rate. The calculation needs applicant counts and selection counts by stage, by requisition and by group. Most systems cannot produce them, because the disposition data was free text and the demographic data was never cleanly separated in the first place.
The reporting surface exports exactly those counts, in aggregate, from the separated table. Where a cell is too small to be meaningful it is suppressed rather than published, because a group of three tells you nothing statistically and identifies somebody personally. Exports run as CSV and JSON on a schedule you set, to a destination you control.
The product computes and exports. It does not interpret. What a given ratio means for your organisation, what threshold you treat as a signal worth acting on, and what you do when you see one are decisions for you and your counsel, and a vendor that offers to make them for you is selling something it cannot deliver.
- Applicant and selection counts by stage, requisition, source and group
- Aggregate output only, with small cells suppressed
- A mandatory coded disposition attached to every drop-out
- Scheduled CSV and JSON exports to a destination you control
- Access to the reporting surface granted separately and logged
Where an automated tool is in scope, the audit data is exportable
Where an automated employment decision tool is in scope, the product records which candidates were scored and exports the data an independent bias audit needs. Commissioning the audit and issuing candidate notice remain yours. That division is not us being cautious; it is what the obligation actually looks like, and a vendor cannot hold it for you.
In practice the export carries the score, the requisition, the date, the version of the tool that produced it, and the separated demographic categories joined only in aggregate. Your auditor receives a dataset they can work with rather than a set of screenshots and an assurance.
The notice a candidate may be entitled to before a tool is used, the timing of that notice, and the publication of an audit summary are actions you take. The product can store the text and send it on the schedule you set. It cannot decide that the requirement applies to your process.
Application records run on their own retention clock
Reporting obligations usually come with a record-keeping period, and it is a different clock from the retention window on the candidate relationship. An application, its coded disposition, its scorecards and its stage history are retained for the period configured on the tenant, independently of whether the candidate record itself is still active or has been erased at the candidate's request.
The two clocks are kept apart deliberately, because a candidate exercising a data right and an employer holding an application record for a statutory period are different questions with different answers. Collapsing them into one setting is how organisations end up either deleting records they were required to keep or keeping records they had no reason to hold.
When both apply to the same person, the product says which one governed which part of the record, and the erasure summary shows it. Silence at that moment is the thing that turns a routine request into a complaint.
What you get
Separated storage
Demographic answers live in their own table with their own access control and reporting key.
Invisible to recruiters
No recruiter-facing view or export joins a candidate name to a demographic answer.
Voluntary collection
Prescribed categories, a decline-to-answer option and a plain statement of purpose.
Model allow-list
Every model payload is built from an allow-list; the demographic table is not on it.
Images excluded
Photographs are stripped from the payload assembled for parsing and scoring.
Field-level match reasons
Scores show which requirement drove each point instead of returning a bare number.
Fixed disposition codes
A mandatory coded reason on every drop-out, with optional free text alongside it.
Immutable stage history
Actor, timestamp and reason on every transition; notes editable, history not.
Selection-rate export
Applicant and selection counts by stage, requisition, source and group.
Small-cell suppression
Groups too small to be meaningful are suppressed rather than published.
Scored-candidate register
Which candidates a scoring tool ran against, with the tool version and the date.
Reporting-role access
The reporting surface is a separate grant, and every grant and access is logged.
Per-market categories
Category sets configured per market rather than one assumed list applied everywhere.
Independent record clock
Application records retained on their own period, separate from candidate retention.
Questions recruiters ask
Does this make our screening EEOC compliant?
No. These are controls that support obligations you hold as the employer, or as the agency acting for one. Whether a screening process is lawful depends on the criteria you select, how consistently you apply them and whether they are genuinely job-related. The product separates the data, codes the decisions and produces the counts. The judgement stays with you.
Is the firewall in the product today?
No. The separate demographic table with its own access control and the fixed disposition code list are in build for the Wave 1 release. Coded dispositions already exist in the tracking workflow; the separated storage and the selection-rate reporting are being written now. Ask us for the current build state before you rely on any of it in a procurement answer.
Can a recruiter ever see demographic data?
Not through any recruiter-facing view. Access to the separated table is a distinct grant held by a reporting role, and every grant and every read is logged. Someone who holds both roles sees aggregate output rather than a name beside an answer, because the reporting surface never returns identified rows in the first place.
Is demographic data used in AI matching?
No. Model payloads are assembled from an allow-list of fields and the demographic table is not on it. What we will not claim is that a model infers nothing from a resume, because a name, a school or a graduation year can carry signal. That is why match reasons are shown at field level and a person still makes the call.
What if we operate somewhere with different categories?
Category sets are configurable per market, because different jurisdictions prescribe different questions and asking the wrong ones creates its own exposure. You configure which set applies to which posting. The product will not decide that an obligation applies to you, and it will not add a category nobody asked you to collect.
Can we give an independent auditor what they need?
Where an automated employment decision tool is in scope, the product records which candidates were scored and exports the data an independent bias audit needs, including the tool version and date. Commissioning the audit and issuing candidate notice remain yours. We provide the dataset; we do not provide the audit or the notice.
What happens to application records when a candidate asks to be erased?
The application record runs on its own retention clock, set for the record-keeping period you configure, separately from the candidate relationship. Where both apply, the erasure summary states which part of the record was removed and which part was retained under which reason. You are told, and so is the candidate.
Keep reading
- Know why you hold every candidate record, and for how long
- A blacklisted candidate stays out of reach until someone lifts the flag
- Matching that shows its working
- Applicant tracking built for a billing desk
- Controls that support your equal-opportunity obligations
- Where responsibility sits for every regime that touches recruitment
See it against your own reqs
Bring one live role and three resumes. In twenty minutes you will see the match scores, the shortlist and the placement invoice that comes out the other end.