Fairness & validity

Fairer by construction — and honest about the work that remains

A hiring tool earns trust two ways: by reducing the biases humans bring, and by being transparent about its own limits. Here is how hire.center approaches both.

Design choices that reduce bias

What the model does — and doesn't — see

No résumé, no name, no photo

The score is derived from the simulation transcript against role-relevant constructs. The judge doesn't see the candidate's name, background, school, gender or age.

Same scenario for everyone

Every candidate for a role faces a standardized, structured crisis. Structure is one of the best-evidenced ways to shrink interviewer bias.

Construct-focused, not culture-fit

We score demonstrated behaviors — coordination, judgment, composure — not vague 'fit', which is where a lot of bias hides.

Tone separated from substance

Calibration rules keep charisma, accent and confidence from inflating scores; only evidenced behavior counts.

Evidence you can inspect

Because every score cites a verbatim quote, a reviewer or auditor can check whether the judgment is actually supported.

'Insufficient evidence' over guessing

When a skill wasn't tested, the model says so rather than filling the gap with a biased prior.

Adverse impact & the four-fifths rule

Adverse impact occurs when a selection procedure passes one demographic group at a substantially lower rate than another. The classic screen is the four-fifths (80%) rule: if a group's selection rate is below 80% of the highest group's, that's a flag warranting investigation. Because hire.center retains structured scores and quote-backed evidence per candidate, employers have the data to run this analysis — which is also what NYC Local Law 144 bias audits require.

Validity — what the score is good for

Our approach draws on Situational Judgment Test research, where interpersonal and teamwork constructs — our core lane — show the strongest predictive validity (meta-analytic estimates around r ≈ .26, higher for interpersonal content). A live simulation strengthens this by eliciting a behavior sample rather than a multiple-choice preference.

What we will not do is inflate a single "predicts performance" number the field hasn't established for a new instrument. During beta we are gathering the evidence that actually matters: inter-rater reliability (do two judges agree?), test–retest stability, and adverse-impact monitoring — and we will report it plainly as it matures.

The candidate's experience

Fairness includes the person being assessed. Candidates consent before starting, are told the assessment is AI-assisted and advisory, face a job-relevant scenario rather than a trivia quiz, and are assessed on what they do — not who they are. We are building candidate-facing feedback so the experience gives something back.

How to use scores responsibly

Guidance for deployers

Treat it as one input

Combine the assessment with other job-relevant evidence. Never let a single score be the sole gate.

Keep a human in charge

A competent reviewer should read the evidence and make the call — and be able to explain it.

Monitor your outcomes

Track selection rates across groups over time and audit for adverse impact. We give you the records to do it.

See the evidence behind a score

Run a simulation and inspect the quote-backed scorecard yourself.