NorthAssay
Product

Product

  • How it worksFrom a job post to the evidence, in three steps.
  • Judgment evaluationMeasure the calls the job actually leaves open.
  • AI video interviewsLive conversation, not a one-way recording.
  • Verified answersThe interview asks about what they submitted.
Pricing
Guides

Guides

  • Assessment methods

    • Candidate assessment
    • Skills assessments
    • Technical skills assessment
    • Communication assessment
    • Pre-employment assessment
  • Tools & software

    • Assessment tools compared
    • AI candidate screening
    • Hiring software for small teams
  • By role

    • Technical interview questions
    • Assess a DevOps engineer
    • Assess a data engineer
    • Assess an engineering manager
  • Screening practice

    • Five checks
    • When candidates use AI

Trust

  • Responsible AIWhat the score is, what it is not, and who decides.
  • SecurityWhere candidate data lives and how long it stays.
  • AboutWhy NorthAssay exists.
Sign inStart free
  1. Product
  2. /Judgment evaluation

Last updated 13 September 2026

Hiring for the job, not for the skill list

Two people can both pass your skills test and be completely different hires. The difference is judgment — the calls your brief leaves open. NorthAssay pulls those out of the job description, shows them to you, and builds the assessment around them.

A skills score answers a question you already know the answer to

By the time someone reaches an assessment you have usually seen their portfolio, their history, and their proposal. "Can they write TypeScript" is rarely the open question. The open question is whether they will make good calls on your job — where the brief is thin, the constraints are real, and nobody is going to be watching them work.

That question is also the one that has got harder, not easier. Work that is mostly execution is being compressed: on Upwork’s own 2026 data, execution-heavy contracts are down 13% in earnings per contract while complex, judgment-led work is up 45%. The part of a job that survives is the part a test of recall never touched.

The thing nobody fixes

Most bad remote hires are not people who could not do the work. They are people who did the wrong work well, because the brief never said which trade-offs mattered — and neither did the assessment.

What an open decision is

An open decision is a real fork in your work where two competent people would reasonably choose differently, and where the choice changes the outcome. Not a gap in the brief to be filled in — a genuine choice the job leaves to whoever takes it.

  • "Rewrite the retry engine, or instrument it for six months first?" — an open decision. Both are defensible; the reasoning is the signal.
  • "Do they know Go?" — a skill. Worth checking, but your CV screen already did.
  • "What database are you using?" — a clarification. Something we need to know to proceed, not something a candidate is judged on. NorthAssay asks you these separately and never confuses the two.

This is the oldest well-evidenced idea in selection research wearing new clothes. Situational judgement tests — put a person inside a realistic dilemma and score the call they make — have been validated for decades, with predictive validity in the r = .32 to .52 range. What has never been practical is writing one per job. That is the part this automates.

How it works, in four steps

1. It reads your job description and pulls out the forks

Alongside the competency model, NorthAssay extracts what the work actually is — responsibilities, deliverables, constraints, business context — and then the judgment calls the brief leaves open. For each one it writes what a weak answer looks like and what a strong one looks like, concretely enough that you could tell two candidates apart by reading them.

2. It tells you where each one came from

Every decision is graded by provenance: in the brief (your posting raises it), inferred (standard for this role and domain, but you did not say it), or speculative (a guess worth confirming or cutting). Thin briefs produce more inferences, which is exactly when you want to know that. Provenance is one of those three words, never a confidence percentage you have to interpret.

3. You confirm or edit them before anything is built

The run stops and shows you the list, least-trustworthy first. You rewrite anything that is wrong in your own words, and confirm. Nothing is generated until you do — because these become the things real candidates are measured on, and you know your job better than any model reading a posting about it.

4. The assessment is built around them

The written and multiple-choice stages put candidates inside the decisions rather than asking about them in the abstract. Distractors are the weak approaches, and each one names why it fails here. Then the live interview probes the same calls in conversation, following up when an answer is shallow — and the criteria it is scored against come from your decisions, so the score is about them rather than about a generic rubric.

What you get back

Evidence per requirement, in your own words. For each decision: what the candidate showed, the verbatim quote it rests on, and — this is the part most tools skip — what was not established, kept visually distinct from what they got wrong.

The score is not the answer

Assessments are scored, and the dossier shows a headline mean. What it does not do is roll this evidence into that number: the requirement verdicts do not sum, and nothing here is weighted into a single figure that says how well someone matches a job. A composite is the thing you would screen on without reading anything underneath it, and the reading is where the value is.

"We never asked them about this" and "they answered this badly" are opposite facts, and most reports render both as an empty row. Collapsing them reports a gap in the assessment as a deficiency in a person, on the screen where you decide whether to hire them. They are kept apart deliberately.

What this does not do

It does not know your business. A model reading your posting can infer a great deal about a role and nothing at all about the politics, the history, or the customer who will not be named in writing. That is why the decisions are graded and why you confirm them — the screen exists to be argued with.

It does not replace judgment about judgment. A strong candidate can reason well and still be wrong for you, and a score here is evidence rather than a verdict. You can override any score, and we ask you to say why.

It also does not work miracles on a two-line brief. It will still find decisions — inferring from the role and the domain is the point, because a thin brief is exactly the case where a client cannot specify their own work — but it will mark them as inferred, and you should read those hardest.

More in Product

  • How it worksFrom a job post to the evidence, in three steps.
  • AI video interviewsLive conversation, not a one-way recording.
  • Verified answersThe interview asks about what they submitted.
  • How the product works

See what it finds in your job description

Paste a real posting and look at the decisions it pulls out. If they are the ones you would have named yourself, that is the whole product working. If they are not, rewrite them — it takes a minute and it is the fastest way to judge this.

Try it on a real posting
NorthAssay

AI hiring assessments, scored with rationale you can defend.

Guides

All screening guides

Assessment methods

Candidate assessmentSkills assessmentsTechnical skills assessmentCommunication assessmentPre-employment assessment

Tools & software

Assessment tools comparedAI candidate screeningHiring software for small teams

By role

Technical interview questionsAssess a DevOps engineerAssess a data engineerAssess an engineering manager

Screening practice

Five checksWhen candidates use AI

Product

How the product worksHow it worksJudgment evaluationAI video interviewsVerified answers

Get started

Start freePricingContactSign in

Trust

Responsible AISecurityAbout
© 2026 NorthAssay. All rights reserved.
PrivacyTerms