Skip to content
    ← Back to blog

    Defensible Interview Question Design for Hiring Teams (4–6 Competencies)

    Isometric illustration of structured interview design

    Interview question design is the discipline of building role-specific, behaviorally anchored questions and a matching 1-to-5 rubric from a real job analysis, not from a generic template someone found online. The workflow that works: run a short job analysis, pick 4 to 6 core competencies, write one behavioral and one situational question per competency, score against anchored rubrics, then pilot and calibrate before the bank goes live. Follow it and you get interviews that are more valid, fairer to candidates, and comparable across your whole slate.


    TL;DR:

    • Using a short job analysis with top performers and recent reviews helps identify 4 to 6 core competencies that truly differentiate strong candidates.
    • Employing one behavioral and one situational question per competency in a 45 to 60-minute interview yields 8 to 12 targeted questions, maximizing depth over breadth.
    • Developing clear, observable behavioral anchors from excellent and inadequate answers ensures consistent ratings and improves inter-rater agreement.
    • Piloting and calibrating the question bank with multiple raters identify ambiguous anchors and prevent rating inconsistencies before scaling.
    • AI tools can expedite question drafting and anchor suggestions but require SME review and pilot testing to maintain fairness and validity.

    Resyme
    Make Hiring Decisions More Reliable
    Resyme uses tailored, behaviorally anchored AI interviews to compare candidates objectively and reduce time spent on pre-screening.
    Explore Resyme

    Table of Contents

    Running a Job Analysis and Choosing Your Core Competencies

    Skip the job analysis and everything downstream is guesswork. You don’t need a six-week study. You need thirty minutes with the hiring manager, a conversation with two or three top performers already in the role, a look at the last two performance reviews, and the actual task list the job requires day to day.

    From those four sources, pull out what separates a strong performer from an average one. Not tasks, differentiators. “Answers phones” is a task. “Deescalates an angry customer without escalating a manager” is a competency. The OPM structured interview guide builds its entire methodology on this distinction: questions derived from job analysis hold up better under legal scrutiny than questions pulled from a generic list.

    Once you have your list, cut it down hard. Most roles need only:

    • 4 to 6 competencies total, never more
    • 2 to 5 dimensions assigned per interview round so each rater stays focused
    • A two-sentence definition of what “great” looks like for each competency, written before you draft a single question
    • A note on which round owns which competency, so nobody duplicates coverage

    A hiring manager who insists on evaluating ten things in one 45-minute conversation is really evaluating nothing well. Fewer competencies, tested deeper, beats more competencies tested shallow.

    Behavioral vs. Situational Questions: When to Use Each

    Behavioral questions ask about the past: “Tell me about a time you had to deliver bad news to a client.” Situational questions ask about the future or a hypothetical: “If a client called furious about a missed deadline, walk me through what you’d do.” Both formats work. They just work for different candidates and different competencies.

    Behavioral and situational interview comparison

    Behavioral questions are stronger for experienced hires, because you’re measuring what someone actually did, not what they think they’d do. Situational questions earn their keep with entry-level candidates or when you’re testing judgment on a scenario the candidate has never faced. A recent graduate has no five-year track record to draw a behavioral answer from, but they can still reason through a hypothetical.

    The operational rule worth keeping on a sticky note: one behavioral plus one situational question per competency, when your interview length allows it. That template lands you at 8 to 12 total questions for a 45 to 60 minute interview, which is the range most structured-interview guides converge on.

    • Write questions that map directly to a named competency, never a generic “tell me about yourself” filler
    • Avoid stacking two behavioral questions on the same competency. It’s redundant and burns time.
    • If a question could be answered identically by any candidate in any industry, it’s not doing its job

    Building a Behaviorally Anchored 1-to-5 Rubric That Raters Actually Agree On

    A rating scale without behavioral anchors is just an opinion with a number attached. Two raters can watch the same answer and one gives it a 3 while the other gives it a 5, because “good” means something different to each of them. Behaviorally anchored rating scales, or BARS, fix that by tying each number to an observable action or result instead of a vague trait label.

    The fastest way to build one:

    1. Draft the level 5 anchor first. What did the best answer you’ve ever heard for this competency actually include, specifically?
    2. Draft the level 1 anchor next. What does a clearly inadequate answer look like, in concrete terms?
    3. Interpolate levels 2, 3, and 4 between those two poles, describing behavior at each step rather than adjectives like “somewhat effective.”
    4. Require every rater to write an evidence note, quoting or paraphrasing what the candidate actually said, before assigning a score.
    5. Collect scores independently before the panel discusses anything out loud.

    That last step matters more than most teams realize. The moment one rater says “I thought that answer was a 4” out loud before others have scored, you’ve contaminated the data. Levashina, Hartwell, Morgeson, and Campion found that higher degrees of interview structure, independent scoring included, consistently improve validity and inter-rater reliability compared with unstructured formats.

    Pro Tip: Write your anchors around actions and results, never personality traits. “Communicated the delay to the client within 24 hours and offered two alternative timelines” is scoreable. “Was proactive and customer-focused” is not.

    Standardizing Probes, Piloting the Question Bank, and Calibrating Your Raters

    A structured interview stops being structured the moment every interviewer improvises their own follow-up questions. Decide in advance what probes are allowed for each question, write the exact probe text into the guide, and give every interviewer the same script. That’s not bureaucracy, it’s what keeps the comparison fair across candidates interviewed on different days by different people.

    Before any bank goes live company-wide, pilot it:

    • Run it with 5 to 10 candidates and at least two independent raters per candidate
    • Check whether raters land in the same range on the same answers; wide gaps mean an anchor needs rewriting, not that the rater is bad at their job
    • Refine ambiguous anchors based on what actually confused people in the pilot, not what you assumed would confuse them

    Calibration sessions work the same way pilots do, just recurring. Give raters shared exemplar answers, have them score independently first, then bring the group together to talk through any outlier scores and agree on what justified them. Version control your rubrics after every revision, and put adverse-impact re-checks on a calendar instead of hoping someone remembers.

    Pro Tip: If two raters disagree by more than one point on the same answer during calibration, the anchor is the problem almost every time, not the candidate’s answer.

    Standardizing Probes, Piloting the Question Bank, and Calibrating Your Raters — overview diagram

    Scaling Structured Interviews Without Losing the Rigor

    The workflow above works cleanly for one role. It gets harder at volume, when you’re running the same competencies across dozens of requisitions or building question banks for a dozen different roles at once. Scaling means building reusable infrastructure: role templates that map to your competency library, a centralized anchor repository so nobody reinvents a rubric from scratch, and a recurring calibration cadence tracked against KPIs like inter-rater agreement and time-to-screen.

    This is where AI tools earn a real place in the process, not as a replacement for the job analysis, but as an accelerator once the framework exists. A platform can generate a first draft of role-tailored questions, suggest behavioral anchors based on the competency definitions you’ve already written, and surface probe wording worth considering, cutting the drafting time that used to eat a full afternoon per role. AI-assisted rubric generation still requires SME review before anything reaches a live interview. The draft is a starting point, not the final product.

    Resyme.ai builds around exactly this model. Resyme tailors interview questions to deep domain knowledge across industries, which speeds up drafting for specialized or high-volume roles where writing every question by hand isn’t realistic. The platform also validates candidate honesty during automated interviews, which matters given that rejected candidates report 35% higher satisfaction after structured interviews compared with unstructured ones. Fair process, it turns out, is something candidates notice even when they don’t get the job.

    Before you scale anything with AI, run this checklist:

    • Have SMEs review every AI-generated question and anchor before it enters rotation
    • Pilot the AI-assisted bank the same way you’d pilot a hand-written one
    • Track time saved per role and inter-rater reliability before and after
    • Feed results back into your ATS and reporting so the data compounds instead of sitting in a spreadsheet

    What Trips Up Most Teams (and How to Fix It Fast)

    The traps repeat across companies of every size. Teams pick nine competencies when six would do. Anchors describe traits (“collaborative,” “driven”) instead of observable behavior, so two raters can’t agree on what a 4 even means. Pilots and calibration get skipped because “we’re confident the questions are good,” which is exactly the assumption structure exists to test.

    The fixes are boring and that’s the point: cut your dimension count, force yourself to write what a 1, 3, and 5 actually look like in action, and run one small pilot before you roll anything out wide. Track candidate experience and fairness metrics alongside hiring outcomes. A process that produces good hires but leaves rejected candidates feeling jerked around is not a process worth scaling.

    — Raul

    Try a Small Pilot Before You Scale

    If you’re weighing whether to bring AI into your interview design process, the honest answer is to test it on one role first, not roll it out company wide on faith. Some AI platforms offer recruiters tools to draft role-tailored, behaviorally anchored questions faster and to validate candidate honesty during interviews, helping ensure candidates are compared fairly rather than relying on self-reported information.

    Resyme

    Run a pilot on one open requisition. Measure your time-to-screen before and after, check whether inter-rater agreement improves once everyone is scoring against the same anchors, and ask candidates how the process felt. Resyme.ai’s Interview service runs at $4.50 to $7.99 per interview, and Company plans start at €599 per month for teams that want to run this at scale. If you’re not ready to commit, explore the Free plan and see how the workflow fits your hiring calendar before you decide anything bigger.

    Sources

    FAQ

    How Many Interview Questions Should I Ask?

    Aim for 8 to 12 questions covering 4 to 6 competencies, roughly one behavioral and one situational question per competency, fitting a 45 to 60 minute interview. Fewer, deeper competencies consistently beat a long list of shallow ones.

    What’s the Difference Between Behavioral and Situational Questions?

    Behavioral questions ask candidates to describe something they actually did in the past. Situational questions present a hypothetical scenario and ask how the candidate would respond. Experienced hires tend to give richer behavioral answers, while situational questions work better for entry-level roles or judgment-heavy scenarios.

    Why Do I Need a Rubric If I’m Already Asking Good Questions?

    Without a shared rubric, two interviewers can score the same answer completely differently, which undermines the comparison you’re trying to make between candidates. A behaviorally anchored 1-to-5 scale, built from concrete level 5 and level 1 examples, keeps raters aligned and gives you evidence notes to defend a decision later.

    Can AI Actually Write Defensible Interview Questions?

    AI tools can draft role-tailored questions and suggest anchors quickly, which speeds up the heaviest part of the process, but SME review before launch stays necessary. Resyme applies this by generating domain-specific questions and validating candidate honesty during the interview, while still depending on a real job analysis as the foundation.

    How Often Should We Recalibrate Raters?

    Run an initial calibration session before any question bank goes live, then repeat it on a recurring schedule, quarterly is common, alongside a periodic check for adverse impact. Piloting and calibration are what convert a well-written rubric into an instrument raters actually agree on.