Last updated 17 August 2026
Candidate assessment tools: how to choose, and what each type is for
Most guides in this category rank tools. Ranking assumes they do the same job, and they do not — a psychometric battery and a live interview platform are different products solving different problems. This sorts them by category, says what each is genuinely for, and includes the cases where ours is the wrong answer.
What candidate assessment tools do
A candidate assessment tool produces comparable evidence about applicants before you hire them. It typically does three jobs: it delivers a test or interview to the candidate, it captures what came back, and it scores or organises the results so two candidates can be compared on the same basis.
What it is not is an applicant tracking system. An ATS moves candidates through stages and stores the record; an assessment tool generates the signal you move them on. Most teams need both eventually, they are usually bought separately, and confusing the two is the most common reason a purchase disappoints.
The four categories
Almost every product in this market is one of four things. The categories are not competing implementations of one idea — they measure different things, and a tool is only badly rated when it is judged against a category it was never in.
| Category | What it measures | Suits | Structural limit |
|---|---|---|---|
| Work-sample and skills testing | Whether the candidate can produce the work, via a task or a test library. | Individual-contributor roles with output you can look at. | Unproctored results say less than they used to, because the task is now cheap to generate. |
| Psychometric batteries | Stable traits and general reasoning, measured consistently at volume. | High-volume hiring where the work is genuinely standardised. | Says nothing about role-specific skill. Carries the heaviest validation and compliance burden. |
| Asynchronous video | How a candidate presents an answer to a fixed question, on their own time. | Large funnels, distributed teams, first-round screening at scale. | One-way by design. Nothing can ask the obvious next question. |
| Live interview platforms | How someone reasons and explains when questioned in real time. | Roles where the output is a conversation, and later-stage screening. | Costs real time per candidate, so it does not sit at the top of a large funnel. |
The fourth column is the one to read carefully. Every limit there is structural — a property of the format, not a gap the vendor will close in the next release. Async video cannot follow up on an answer no matter how good the product gets, because the candidate has already left.
A comparison, by category not by rank
Ordered by category, because a ranked list of tools that solve different problems tells you nothing. Each row states what the product is genuinely good at and the limitation that comes with it, including ours.
| Tool | Category | Best for | Comes with |
|---|---|---|---|
| TestGorilla | Work-sample and skills testing | Breadth. A large test library covering many roles without writing anything. | Library tests are generic by construction, and widely-used questions are widely known. |
| Vervoe | Work-sample and skills testing | Role-realistic tasks with automated scoring. | More setup per role than a library, and scoring quality depends on that setup. |
| Criteria | Psychometric | Validated instruments where defensibility is the point. | Measures the person rather than the work. Enterprise procurement, enterprise pricing. |
| HireVue | Psychometric and async video | High-volume enterprise funnels with compliance requirements. | Built for scale most teams do not have, and priced accordingly. |
| Willo | Asynchronous video | Cheap, fast first-round screening at the top of a large funnel. | Recorded answers only. No follow-up, so depth has to come from a later stage. |
| Hireflix | Asynchronous video | One-way interviews done well, with a straightforward setup. | Same structural ceiling: the conversation cannot go anywhere. |
| NorthAssay | Work sample plus live interview | Assessment and a live follow-up conversation as one screen, priced per candidate. | Early-stage product. Not an ATS, not a validated psychometric instrument, and no test library to pick from. |
Two categories are absent on purpose. Live-coding tools such as CoderPad and CodeInterview are an adjacent market — they run the session rather than assess the candidate, and they pair with a rubric rather than replace one. Sourcing platforms are a different product entirely and are frequently sold into the same budget.
How to choose: five questions
Answer these before you look at a demo. Four of them will eliminate most of the market, which is the point — the shortlist is easy once the category is settled.
- How many candidates, and how often? Twenty a year in bursts and two hundred a month are different problems. Per-candidate pricing suits the first; a seat licence suits the second, and buying the wrong one is the most expensive mistake here.
- What does this role fail on? Pick the category that measures the thing most likely to go wrong. If your bad hires could all do the work but could not explain it, no work-sample tool will help you.
- Do you need the result to be defensible to a regulator, or to a candidate? These are different bars. The second is a written rubric. The first is validation evidence, and it narrows you to psychometric vendors quickly.
- Who reads the output? A score with no reasoning attached is not reviewable by a hiring manager, and it is not something you can give a rejected candidate. Ask to see a real result, not a dashboard screenshot.
- What happens to candidate data? Where it is stored, how long it is kept, and whether it trains a shared model. Ask for it in writing — how we answer that is the shape of answer to expect.
One question deliberately not on the list: integrations. It dominates demos because it is easy to demonstrate, and it is close to irrelevant at small scale, where the integration is a person copying a score into a spreadsheet twice a week.
What the category still cannot do
No tool in this market decides who to hire, and every vendor that implies otherwise is selling you a liability. The output is evidence. The decision is a person's, and in several jurisdictions it has to be.
A score is an input, not a verdict
If a tool cannot show you why it produced a number, you cannot defend the rejection it caused — to a hiring manager, to a candidate, or to anyone asking later.
The category also cannot fix a role you have not defined. A tool applied to a vague brief produces a confident ranking against the wrong criteria, and it will produce it faster than the manual process it replaced. That is worse than no tool, because the numbers make it feel settled. For the underlying question of which method suits which role, see what candidate assessment is and how to choose a method; for AI-specific claims, what AI can and cannot screen for.
When NorthAssay is the wrong choice
The cases below are real and we would rather you found them here than in month two:
- You need an ATS. NorthAssay assesses candidates; it does not manage a pipeline, schedule interviews or store your hiring record. If that is the gap, buy that instead and come back later.
- You need validated psychometric instruments for regulated or safety-critical roles. That is a different product with a different evidence base, and we do not have one.
- You are hiring at very high volume for standardised work — hundreds of applicants a month for near-identical roles. Per-candidate pricing is the wrong shape for that, and a battery plus a funnel tool will serve you better.
- You want a test library to pick from. NorthAssay drafts an assessment from the role rather than offering a catalogue. If your preference is choosing a ready-made test off a shelf, that is a genuine preference and we are not it.
- You need a long procurement trail today. We are an early-stage product in open beta, and we say so on the about page rather than in a footnote.
What is left is the case we are built for: a small team hiring people they will never meet in person, who need the assessment and the conversation to be one screen rather than two, and who need to be able to explain the decision afterwards.
The case we are built for
Describe a role in a sentence. NorthAssay drafts the assessment, interviews shortlisted candidates live, and scores both against a rubric it shows you. Free while we are in beta, then priced per candidate.
Try it on a real role