Hiring guides · July 29, 2026

How to Tell If a Candidate Can Actually Do the Job

How to Tell If a Candidate Can Actually Do the Job

You know the one. The candidate who talked about the job better than your best employee talks about it. Confident, articulate, said all the right things about "process" and "ownership." Three weeks in, they can't format a simple report without help, or they freeze the first time a customer pushes back. You're now doing their job and yours.

This happens because an interview measures one specific skill: talking about work in a room with a stranger watching. Some people are excellent at that and mediocre at the actual work. Some people are the opposite — quiet in interviews, sharp on the tools. If your only filter is conversation, you're selecting for confidence, not competence. Those overlap sometimes. Not always.

The fix isn't a better interview. It's adding a step where the candidate does something close to the actual job, before you hire them, not after.

Build a work sample, not a trivia round

A work sample is a small, realistic piece of the job. Not a personality quiz, not "tell me about a time you showed leadership." An actual task, done under conditions close to real ones, that produces something you can look at and judge.

The test: could you show two candidates' outputs to your best employee and have them tell you which one is better, without knowing who did which? If yes, it's a good work sample. If the outputs would look identical no matter who did them, you've built a formality, not a filter.

Some examples by role:

For an accounts or bookkeeping role — give them a small, deliberately messy set of transactions (a few miscategorised, one duplicate, one that doesn't balance) and ask them to reconcile it and flag anything odd. You're not testing whether they know what a ledger is. You're testing whether they notice the duplicate.

For a sales role — role-play a real objection your team hears every week, not a generic one. "Your price is 20% higher than the competitor we've used for three years" tells you more than "sell me this pen" ever will, because it's the actual conversation they'll have on day one.

For an ops or admin role — hand them a genuinely tangled inbox (five emails, two contradicting each other, one urgent, one that looks urgent but isn't) and ask them to reply to all five in priority order. Watch what they do first.

For a technical or design role — give them a small piece of real (anonymised) work your team has actually shipped, with a known flaw in it, and ask them to review it. People who can only recite theory struggle here. People who've actually done the work spot the flaw fast.

Keep it to 30–60 minutes. Anything longer and you'll only get candidates who have nothing else going on, which is its own kind of bias.

What to watch while they do it

The output matters, but so does how they got there. Did they ask a clarifying question before diving in, or guess and hope? Did they check their own work, or hand it over the moment it looked finished? Did they say "I don't know, but here's how I'd find out" when they hit something unfamiliar — that answer is often worth more than a lucky guess.

Write down what "good" looks like before you see any candidate's attempt. Otherwise you'll unconsciously grade the first one against a blank page and the fifth one against a rising bar, and you'll have no idea if you're being fair.

Check references for evidence, not opinions

Most reference checks are useless because the questions invite a compliment. "Would you recommend them?" gets you a yes almost every time — people are reluctant to badmouth a former colleague to a stranger on the phone. Ask for specifics instead:

A hesitation, a pause before answering, or a very generic answer tells you something. So does a reference who talks for five minutes without being asked twice.

Where structured testing fits

Work samples and good reference checks solve most of the problem. Where they run thin is comparing candidates fairly at volume — twenty people applying for one warehouse supervisor role, say, and you don't have four hours per candidate to build and grade a custom task each time.

That's the gap tools like AssessFit are built for: standardised, role-relevant tests that are scored the same way for everyone, so a CV and an interview aren't the only inputs left doing the deciding. It's not a replacement for talking to the person or watching them work — it's there so the CV can't carry the whole decision by itself.

What this won't catch

Be honest with yourself about the limits. A 45-minute work sample won't tell you if someone will still be motivated in month six, or how they behave when a project genuinely fails, or whether they'll get along with the specific person they'll sit next to. Those are judgement calls. No test — ours or anyone else's — makes them for you. What a work sample does is remove the biggest, cheapest failure mode: hiring someone who was great at describing the job and never actually good at it.

A checklist for your next hire

  1. Write down what "good work" looks like for this role before you see a single candidate.
  2. Build one small, realistic task — 30 to 60 minutes, real messiness included.
  3. Give every candidate the same task, same instructions, same time limit.
  4. Ask reference-check questions that require a specific story, not a yes or no.
  5. Decide what weight the interview, the task, and the reference each carry — before the scores come in, not after you already like someone.

None of this takes the feeling out of hiring. It just makes sure the feeling has something real underneath it.

Hire on evidence, not gut feel.

Test 5 candidates every month, free forever. No credit card.

Start free