NorthAssay
Product

Product

  • How it worksA sentence to a scored shortlist, in three steps.
  • AI video interviewsLive conversation, not a one-way recording.
  • Verified answersThe interview asks about what they submitted.
Pricing
Guides

Guides

  • Assessment methods

    • Candidate assessment
    • Skills assessments
    • Technical skills assessment
    • Communication assessment
    • Pre-employment assessment
  • Tools & software

    • Assessment tools compared
    • AI candidate screening
    • Hiring software for small teams
  • By role

    • Technical interview questions
    • Assess a DevOps engineer
    • Assess a data engineer
    • Assess an engineering manager
  • Screening practice

    • Five checks
    • When candidates use AI

Trust

  • Responsible AIWhat the score is, what it is not, and who decides.
  • SecurityWhere candidate data lives and how long it stays.
  • AboutWhy NorthAssay exists.
Sign inStart free
  1. Guides
  2. /AI candidate screening

Last updated 17 August 2026

What AI can and cannot screen for

We sell an AI screening tool, which makes this page an odd thing to publish. It is here because every guide in this category is written by a vendor and none of them says where the technology breaks. Ours breaks in the places below.

What AI candidate screening means

AI candidate screening is the use of a model to filter or rank applicants before a human reviews them. In practice the phrase covers two very different things, and most confusion in this market comes from treating them as one.

Résumé parsing reads applications and matches them against a job description. It is the older, more common meaning, and it is what most products labelled "AI screening" actually do. Assessment scoring evaluates something the candidate produced for you — a work sample, a written answer, an interview transcript — against criteria you set.

The distinction matters because the failure modes have nothing in common. Résumé parsing inherits every bias in your historical hiring data and rewards candidates who know how to write for it. Assessment scoring cannot see the résumé at all, and fails in the other direction: it can be confidently wrong about work it has misread.

The three things AI screens today

Almost every product in this category does one of three jobs. They differ in how well they work and in how much damage a mistake does.

What each does, and where it breaks
What it doesHow well it worksWhere it breaks
Résumé parsing and rankingReliable at extraction. Much weaker at judging fit, which is the part being sold.Learns from past hiring, so it reproduces past preferences. Rewards résumé formatting over ability.
Scoring written submissionsGood at consistency — the same rubric applied the same way to every candidate.Reads confident prose as competent prose. Struggles with an unconventional but correct answer.
Analysing interviewsTranscription and structured scoring against a rubric are genuinely solid.Anything inferred from voice, face or manner rather than content. That inference is where the regulatory risk lives too.

The middle row is where most of the value in this category actually sits, and it is the least marketed, because "applies your rubric consistently" is a duller claim than "finds your best candidate". Consistency is the honest version of what these tools deliver.

Where it fails, specifically

These are the failures worth designing around, in rough order of how often they bite:

  • It rewards fluency. A well-structured, confidently-written wrong answer scores above a terse correct one more often than it should. This penalises non-native speakers and people who write tersely because they are busy, and it is the single largest fairness problem in the category.
  • It cannot tell "unusual" from "wrong". A candidate who solves a problem in an unexpected way is the candidate you most want to find, and is exactly the case a rubric-following scorer handles worst.
  • It is confident when it is wrong. Models do not reliably signal their own uncertainty. A score of 4/10 produced from a misreading looks identical to a score of 4/10 produced from a correct reading.
  • It cannot verify authorship. Nothing in an unproctored written submission tells you who wrote it, and detection tools do not close that gap — why detection fails covers the evidence and the check that still works.
  • It inherits whatever the job description assumed. Screening against a brief that quietly encodes a preference will apply that preference at scale and at speed, with a number attached that makes it look objective.

None of these is fixed by a better model. They are consequences of asking a language system to make a judgement about a person from a small sample of text, which is why the correct response is process design rather than waiting for the next release.

Keeping a human decision

The practical mitigations are unglamorous and they work. Score competencies separately rather than producing one number. Show the reasoning next to every score, so a reviewer can catch a misreading. Never auto-reject — rank, and let a person read the bottom of the list occasionally, because that is the only way you find out the ranking is wrong.

If the score has no reasoning, it is not reviewable

A number you cannot interrogate is a decision you cannot defend — to a hiring manager, to a rejected candidate, or to anyone asking a year later why the shortlist looked like that.

Withholding identity from the scorer helps and is cheap: a scorer that never receives a name, a photo or a country cannot act on them. It is a real control and it is not a fairness guarantee, because proxies survive the redaction — where someone studied, how they phrase things, which conventions they follow. Our full position is on responsible AI, including the work we have not done.

What to ask an AI screening vendor

Five questions. The answers separate the market faster than any demo:

  • Does the model see the candidate's identity when it scores? If yes, ask why. If they cannot tell you, that is the answer.
  • Can I see the reasoning behind a score, per answer? Ask to see a real one. A dashboard that shows a number and a confidence bar is not reasoning.
  • What does the system do on its own, and what needs a person? Any product that rejects candidates without a human in the loop is one you are accountable for.
  • What fairness testing has been done, and can I read it? "Our model is fair" is not an answer. A vendor that has done the work will describe the method; a vendor that has not should say so.
  • Is my data used to train shared models? Ask for it in writing, not in a sales call.

A sixth, if you are buying video interview analysis specifically: ask what the system infers from face and voice as opposed to words. Several jurisdictions regulate that inference directly, and some vendors do it without making it obvious.

The regulatory picture, briefly

Not legal advice, and the picture moves. Three instruments matter for most readers:

  • New York City Local Law 144 applies to automated employment decision tools used on candidates in the city. It requires an annual bias audit by an independent party, published results, and notice to candidates. The audit is the employer's obligation, not the vendor's.
  • The EU AI Act classifies AI used in recruitment and candidate selection as high-risk, with obligations covering documentation, human oversight and transparency phasing in over several years.
  • Illinois requires notice and consent before AI is used to analyse video interviews, and several other US states have introduced comparable rules.

The common thread is that regulators are consistently more interested in whether a human made the decision than in which model produced the score. A process where AI ranks and a person decides is the one that survives most of these regimes with the least work.

What we have not solved

NorthAssay is an AI screening tool and the fluency problem above applies to it. We limit identity exposure to the scorer and we score interviews from the transcript as text rather than from voice or face. We have not commissioned a bias audit, and we would rather say that here than imply otherwise.

The honest summary of the category: it makes screening consistent, which is a genuine improvement over five interviewers with five different mental rubrics. It does not make screening correct. For how the methods compare, see what candidate assessment is; for the tools themselves, the buyer's guide.

More in Guides: Tools & software

  • Assessment tools comparedFour categories, and when ours is wrong.
  • Hiring software for small teamsWhether you need an ATS yet.
  • All screening guides

Screening you can read the reasoning on

Every score comes with strengths, gaps and a confidence signal per answer, and the scorer never receives a name, photo or country. You make the decision.

See what it scores
NorthAssay

AI hiring assessments, scored with rationale you can defend.

Guides

All screening guides

Assessment methods

Candidate assessmentSkills assessmentsTechnical skills assessmentCommunication assessmentPre-employment assessment

Tools & software

Assessment tools comparedAI candidate screeningHiring software for small teams

By role

Technical interview questionsAssess a DevOps engineerAssess a data engineerAssess an engineering manager

Screening practice

Five checksWhen candidates use AI

Product

How the product worksHow it worksAI video interviewsVerified answers

Get started

Start freePricingSign in

Trust

Responsible AISecurityAbout

Company

Contact
© 2026 NorthAssay. All rights reserved.
PrivacyTerms