NorthAssay
Product

Product

  • How it worksA sentence to a scored shortlist, in three steps.
  • AI video interviewsLive conversation, not record-then-analyse.
  • Verified answersThe interview asks about what they submitted.
Pricing
Guides

Guides

  • Five checksSeparating real skill from a good résumé.
  • When candidates use AIWhy detection fails, and the check that still works.
  • Assess a DevOps engineerJudgement under pressure, not tool familiarity.
  • Assess a data engineerThe stakeholder half that screens never test.
  • Assess an engineering managerAnd what no assessment can tell you.

Trust

  • Responsible AIWhat the score is, what it is not, and who decides.
  • SecurityWhere candidate data lives and how long it stays.
  • AboutWhy NorthAssay exists.
Sign inStart free
  1. Guides
  2. /When candidates use AI

Last updated 15 August 2026

What to do when candidates use AI on your assessment

The take-home test was always a proxy: we accepted that it measured output because output used to be expensive to produce. It is not expensive any more. This is what that changed, what it did not change, and the check that still holds.

What actually changed

A take-home test never measured skill directly. It measured an artefact — code, a document, a plan — and we treated the artefact as evidence of the person because producing a good one used to require being good. That inference is the whole design, and it was sound for as long as the cost of producing a competent artefact stayed high.

That cost collapsed. A task that takes an honest candidate four hours can now be produced in a fraction of it, at a quality that clears most rubrics. Research from Canvas8 and Multiverse found roughly half of job seekers admit to using generative AI to misrepresent their skills — and that was measured before the current generation of models.

The important part is not that some candidates cheat. It is that the test no longer distinguishes. A strong candidate and a weak candidate with a good prompt now submit work you cannot tell apart, which means the assessment has stopped producing a signal in either direction. You are not catching fewer fakers. You are also no longer confirming the real ones.

The failure mode is quiet

A compromised take-home does not look broken. It returns scores, it ranks candidates, and every number in it is wrong in the same direction.

Why detection is a losing position

The instinctive fix is to watch harder. Browser-based platforms flag tab switches, clipboard pastes and keystroke timing; proctoring tools add webcam monitoring, lockdown browsers and identity checks. All of it is worth something, and none of it is worth what a vendor implies it is worth.

The structural problem is that these tools watch the screen and the face, while the assistance arrives somewhere else. A second device sits outside the webcam frame. A virtual machine hides a second session from the proctor entirely. Earbuds with on-device assistants are ordinary consumer hardware now, and no amount of gaze tracking sees a voice in someone's ear.

  • Screen monitoring catches copy-paste from a browser tab, and misses a phone lying beside the keyboard.
  • Keystroke and timing analysis infers effort from rhythm, which is a proxy for a proxy — and penalises fast, fluent candidates alongside dishonest ones.
  • Webcam proctoring establishes who is sitting there, not who is answering. It is genuinely useful against proxy candidates and close to useless against assistance.
  • Output classifiers that claim to detect generated text carry false-positive rates you would not accept if you saw them applied to a candidate you liked.

Every detection method in production has a documented, working bypass, and the bypasses are cheaper to adopt than the detection is to deploy. That is what makes this a losing position rather than a hard one: you are spending on an arms race whose economics run against you, to defend an inference that has already broken.

Worth saying plainly

Proctoring still has real uses — identity assurance and deterrence among them. It just cannot restore the thing the take-home lost.

The check that still works

The most effective defence available is not technical. Ask the candidate to explain their own submission, out loud, in a conversation that can ask a second question.

This works because it inverts what is being measured. A generated artefact is cheap; an explanation of why that artefact is shaped the way it is, delivered live, under follow-up, is not. Someone who did the work remembers the constraint that made it hard, the option they rejected, and the thing they would change. Someone who prompted for it has the output and none of the reasoning underneath it, and the gap shows up within a question or two — not because they are caught out, but because there is nothing below the surface to go to.

The follow-up is the entire mechanism. A recorded answer to a fixed question is the same artefact problem in a different medium — it can be scripted, rehearsed, or read. What cannot be prepared is the second question, the one that depends on what the candidate just said. This is the difference between a conversation and a recording, and it is the reason NorthAssay's video interview is live rather than record-then-analyse.

It also does not need to be long. Ten minutes on work someone actually did is more diagnostic than another hour of unsupervised task.

Redesigning the test itself

The other half of the answer is to stop pretending the tools are not there. Candidates use AI at work — most of the roles you are hiring for will involve it — so a test that forbids it is measuring compliance, and a test that ignores it is measuring nothing.

  • Allow the tools and say so. An honest candidate stops being disadvantaged relative to a dishonest one, which is the actual unfairness in an unproctored test.
  • Ask for the judgement, not the artefact. Trade-offs, rejected options, what would change the decision. These are weak prompts and strong questions.
  • Score the reasoning you can see. A rubric that rewards a defensible decision rewards something generated output is genuinely bad at.
  • Keep the written stage short. Its job now is to decide who is worth a conversation, not to decide the hire.

Read together, these move the written test back to what it is still good at — cheap, fair, parallel filtering — and move the decision to the stage that can carry it.

What this does not fix

A live conversation raises the cost of faking substantially. It does not make it impossible, and anyone telling you otherwise is selling something. A well-prepared candidate with real assistance can hold a shallow conversation, and a nervous strong candidate can interview worse than they work.

It is also not free. A conversation costs candidate time and goodwill, and a hiring process that adds stages without removing any will lose people — particularly freelance and remote candidates, who have the most options and the least obligation to tolerate a long funnel. If you add a live stage, take something out.

The honest position is that assessment produces evidence, not proof, and it did before any of this. What changed is which stage the evidence comes from. The written test used to carry the decision and now cannot; the conversation always could and now has to.

More in Guides

  • Five checksSeparating real skill from a good résumé.
  • Assess a DevOps engineerJudgement under pressure, not tool familiarity.
  • Assess a data engineerThe stakeholder half that screens never test.
  • Assess an engineering managerAnd what no assessment can tell you.
  • All screening guides

A written test that earns a conversation

NorthAssay drafts a role-specific assessment, then interviews candidates live — following up on what they actually say, and scoring both against a rubric it shows you.

Draft your first assessment
NorthAssay

AI hiring assessments, scored with rationale you can defend.

Get started

Start freePricingSign in

Product

How the product worksHow it worksAI video interviewsVerified answers

Guides

All screening guidesFive checksWhen candidates use AIAssess a DevOps engineerAssess a data engineerAssess an engineering manager

Trust & legal

Responsible AISecurityAboutPrivacyTerms

Company

Contact
© 2026 NorthAssay. All rights reserved. Built for teams who hire on proof