Last updated 17 August 2026
Pre-employment assessment: what is defensible and what is not
Most guides in this category compare vendors. The question underneath the search is usually different: if this test screens someone out and they ask why, does the answer hold up? This is about that.
What counts as a pre-employment assessment
A pre-employment assessment is any standardised measurement applied to candidates before hiring and used to inform the decision. That includes work samples and structured interviews, not only the psychometric tests the phrase usually evokes.
The breadth matters because obligations attach to the use, not to the label. A scored take-home task that screens people out is a selection procedure in the same way a personality inventory is. Calling it a "practical exercise" does not move it outside the frameworks below.
What is generally not in scope: an unstructured conversation with no scoring, and anything applied after an offer, which sits under different rules again. The line is whether it produces a comparable score used to sort candidates.
The four types, and their legal exposure
Exposure here means the likelihood that a screening decision is challenged and the difficulty of defending it. It rises with how far the measurement sits from the work itself.
| Type | Measures | Evidence it rests on | Exposure |
|---|---|---|---|
| Work sample | Performance on a task drawn from the job. | Content validity — the task visibly resembles the work. | Lowest. The defence is that you asked them to do the job. |
| Structured interview | Reasoning and experience, same questions and rubric for everyone. | Consistency of administration, plus a written rubric. | Low, if it is genuinely structured. Unstructured interviews are the weakest of all. |
| Cognitive ability test | General reasoning, abstracted from any particular job. | Published validation studies and a job-relatedness argument. | High. Well-documented adverse impact in the research literature. |
| Personality inventory | Stable traits, self-reported. | Vendor validation, and an argument that the trait predicts this job. | High, and rises sharply if it strays near health or disability. |
Read the third column as the question you will be asked. For a work sample it is easy to answer. For an off-the-shelf trait inventory the answer has to come from the vendor, and if they cannot produce it, you are the one holding the risk.
Validity: the word that does the work
Validity is the claim that a test measures what it says it measures, for the job it is being used for. It is a property of the use, not of the instrument — the same test can be valid for one role and indefensible for another.
- Content validity — the test samples the actual work. A typing test for a typing job. The easiest to establish and the easiest to explain.
- Criterion validity — scores correlate with later job performance, demonstrated with data. Strongest, and expensive: it needs enough hires and a performance measure worth correlating against.
- Construct validity — the test measures the underlying trait it claims to. The usual basis for cognitive and personality instruments, and the furthest from the work.
For most teams reading this, content validity is both sufficient and free. Build the test out of the job and the argument makes itself. Reaching for a construct-valid instrument buys statistical sophistication and an obligation to justify the leap from trait to role.
Adverse impact, in one section
Adverse impact is what happens when a selection procedure that is neutral on its face passes one group at a substantially lower rate than another. It does not require intent, and intent is not a defence.
The US EEOC Uniform Guidelines on Employee Selection Procedures set out the framework, including the widely-cited convention of comparing selection rates against four-fifths of the highest-scoring group. A jurisdictional layer sits on top: New York City Local Law 144 requires an annual bias audit by an independent party for automated employment decision tools, with results published and candidates notified, and the obligation falls on the employer rather than the vendor.
You cannot check for this without data
Adverse impact is a statistical property of outcomes. If you are not recording who passed each stage, you are not in a position to know, and "we never looked" is not a defence anyone accepts.
The practical floor for a small team: keep pass rates by stage, keep the rubric, and keep the evidence behind each score. That will not satisfy a regulator on its own, and it is the difference between being able to answer a question and not.
What to ask a vendor before you buy
Five questions, in the order that eliminates fastest. Ask for answers in writing.
- What evidence do you have that this predicts performance in a role like mine? "Validated" alone is not an answer. Ask which kind, against what, and how long ago.
- Has adverse impact been analysed for this instrument, and can I see it? A vendor that has done the work will describe the method without being pushed.
- What exactly does the score mean, and what does it not mean? If nobody can tell you what a 62 represents, you cannot defend a rejection caused by one.
- Does anything happen automatically? A product that screens candidates out with no human review is one you are accountable for.
- Who owns the candidate data, how long is it kept, and does it train shared models? In writing, not in a sales call — our answers are the shape to expect.
The answers separate the market quickly. A vendor selling validated psychometrics should answer the first two immediately; one selling work-sample tooling should say plainly that content validity is the basis and not dress it up as something else.
Where NorthAssay fits, and where it does not
NorthAssay is work-sample and interview tooling. It drafts a task from the role and scores submissions and interviews against a rubric it shows you, which puts it in the top two rows of the table above — the low-exposure end, resting on content validity.
It is not a validated psychometric instrument and we do not claim it is. There is no cognitive battery, no personality inventory, and no published criterion-validity study. We have not commissioned a bias audit either; our full position, including what we have not done, is written down rather than implied.
If your requirement is a defensible instrument for a regulated or safety-critical role, buy from a psychometric vendor and ask them the five questions above. If it is a defensible process for an ordinary hire, the work sample plus a written rubric is both cheaper and easier to explain. For the wider comparison, see the assessment tools guide.
A rubric you can show someone
NorthAssay scores every answer with strengths, gaps and a confidence signal, against a rubric written before the first candidate arrives. Nothing is auto-rejected — you make the decision, and you can explain it.
See a scored example