The science

Grounded in public occupational science — not a proprietary black box

Our measurement rests on two government-maintained taxonomies, decades of situational-judgment research, and a scoring model built to cite evidence rather than impress. Here is the method, in the open.

Measure with O*NET, speak in ESCO

O*NET (the Occupational Information Network) is the U.S. Department of Labor's free, continuously updated database describing what work requires across 1,000+ occupations — including standardized skills, work styles, work activities and tasks, each with numeric importance ratings. It is the most rigorous public map of the world of work, and it is the measurement spine of hire.center.

We self-host the full O*NET database. When you choose a role, we read that occupation's actual skill-importance ratings to decide which of the 14 assessable transferable skills belong in the blueprint — so the assessment reflects the job, not our guesswork.

ESCO (the European Skills, Competences, Qualifications and Occupations classification) is the EU's parallel framework, spanning ~3,000 occupations in 28 languages. ESCO's competency links are designed for matching, not measurement, so we don't score against them — but we use the official EC/DOL crosswalk to let European occupation titles resolve onto the O*NET data spine, and ESCO's transversal-skills hierarchy to give our constructs official EU identifiers and translations. The result is a dual-taxonomy claim — U.S. DOL and EU — that matters in an AI-Act world.

Why a situational simulation

Decades of selection research converge on a practical point: samples of behavior predict behavior better than self-report about behavior. Situational Judgment Tests (SJTs) — which present realistic work dilemmas and observe how a person responds — show meaningful predictive validity, and they are strongest precisely where we focus: interpersonal effectiveness, teamwork and composure. Peer-reviewed meta-analyses place SJT criterion validity around r ≈ .26 on average, with the interpersonal domain at the higher end.

A live, multi-character simulation is an SJT with the constraints removed: instead of picking A/B/C/D, the candidate must actually act — coordinate people, make a call, hold composure — while the scenario reacts. That is closer to a work sample, and it is far harder to fake than a multiple-choice item or a rehearsed video answer.

The model

A two-axis measurement model

Every assessment scores two independent things, so you can tell a capable-but-brittle candidate from a capable-and-composed one.

Axis 1

Role-relevant transferable skills

2–8 skills selected from the 14 assessable transferable skills by the occupation's O*NET importance ratings. These answer: can they do the work?

Axis 2

The Composure Profile

A fixed five-trait profile — Stress Tolerance, Self-Control, Adaptability, Cooperation, Empathy — drawn from O*NET Work Styles (1.D). This answers: can they do it under pressure?

Evidence-first scoring

A behavioral score is only as trustworthy as the evidence behind it. Three rules make ours auditable:

  • Every score cites a verbatim quote. The scoring model must ground each 1–5 rating in the candidate's own words, and the server independently re-verifies each quote against the transcript. Quotes that don't match are flagged, not trusted.
  • "Insufficient evidence" is a valid outcome. If the conversation never surfaced a skill, the model says so instead of manufacturing a middling score. This is a statement about the evidence, not a verdict on the candidate.
  • Tone is separated from substance. The judge is calibrated to start neutral, ignore charisma, avoid halo effects, and never reward domain name-dropping — only demonstrated behavior counts.

Anchored, calibrated ratings

Scores use a 1–5 scale with behavioral anchors rather than vibes. O*NET provides two complementary scales we lean on conceptually: an importance scale (how much a skill matters to the role, used to build the blueprint) and a level scale (how much of the skill the work demands, used to anchor what "good" looks like). Grounding the rubric in these public anchors keeps scoring consistent from candidate to candidate.

Sources

Where to read more

  • O*NET Resource Center — the U.S. DOL/ETA occupational database (CC BY 4.0).
  • ESCO — the European Commission's multilingual skills/occupations classification.
  • McDaniel et al. and Christian et al., meta-analyses of Situational Judgment Test validity (peer-reviewed I/O psychology literature).
  • Our glossary — plain-language definitions of O*NET, ESCO, SJT, construct validity, adverse impact, and more.

See the method produce a scorecard

Run a simulation and read the quote-backed evidence for yourself.