Skip to content
    ← Back to blog

    6 Steps to Build Valid Behaviorally Anchored Rating Scales for HR

    Isometric behaviorally anchored rating scale illustration

    A behaviorally anchored rating scale (BARS) defines each numeric rating with a concrete, observable behavior so raters match actions to anchors instead of interpreting vague labels. It works best for roles with repeatable, observable tasks, such as customer service, manufacturing, or technical support. If your organization needs defensible, consistent evaluations and has the staff time to build and maintain the scale, BARS is worth the investment. If you need something fast and cheap, look elsewhere first.


    TL;DR:

    • Building BARS requires collecting specific incidents, translating them into precise behavioral anchors, and validating them through multiple rounds of expert ratings and reliability checks.
    • Use BARS for roles with high turnover of observable, repeatable tasks, especially where legal defensibility and actionable feedback are priorities.
    • Anchors should describe actual actions performed, avoiding vague adjectives or personality-based judgments, and must be updated regularly as job roles evolve.
    • Pilot and calibrate BARS scales with real raters, focusing on inter-rater reliability and discussing discrepancies to improve consistency.
    • Automated interview platforms like Resyme can help standardize behavioral data collection, making it easier to develop reliable, comparable BARS anchors.

    Resyme
    Standardize Behavioral Hiring Data
    Resyme uses AI-driven interviews to collect comparable behavioral insights, helping recruiters make more accurate hiring decisions efficiently.
    Explore Resyme

    Table of Contents

    What Are Behavioral Anchored Rating Scales and How Do They Work?

    Frank Smith and Lorne Kendall introduced the format in 1963 to fix a specific problem: raters interpreting “good communication skills” five different ways on the same team. A BARS replaces that vague label with a specific behavior tied to each point on the scale. Instead of rating someone a 3 out of 5 on “teamwork,” a manager checks whether the employee’s actions match a written anchor like “volunteers to help teammates finish deadline work without being asked.”

    Most BARS instruments run vertically with 5 to 9 points, moving from highly ineffective at the bottom to highly effective at the top. Each point carries a short, specific behavioral statement rather than a number or a generic adjective. A rater does not guess what “exceeds expectations” looks like. They read the anchor and decide which one matches what they actually observed.

    The method behind the anchors is called the critical incident technique, a research approach that collects specific, factual examples of effective and ineffective behavior rather than general impressions. Supervisors and top performers describe real incidents, both good and bad, and those incidents get sorted, rated, and eventually distilled into the anchor language you see on the final scale.

    This is where BARS diverges from its two closest cousins:

    • Graphic rating scales ask raters to mark a point on a line between “poor” and “excellent” with no behavioral definition attached, which leaves interpretation wide open.
    • Behavioral observation scales (BOS) track how frequently a specific behavior occurs, rather than which behavior best matches a performance level.
    • BARS anchors a discrete performance level to a specific behavioral example, combining structure with description.

    The difference matters more than it sounds. Two managers using a graphic scale can rate the same employee a 4 and a 2 on identical performance, purely from differing standards. BARS narrows that gap by giving both managers the same behavioral yardstick.

    What Do Behavioral Anchors Actually Look Like?

    Abstract definitions only get you so far. Here’s how anchors read in practice across four dimensions HR teams evaluate constantly.

    1. Teamwork

    • Level 5: Proactively mentors struggling teammates and redistributes workload before deadlines slip.
    • Level 3: Completes assigned collaborative tasks but rarely offers unsolicited help.
    • Level 1: Misses team commitments and blames others when confronted.

    2. Technical skill

    • Level 5: Diagnoses complex system failures independently and documents the fix for future reference.
    • Level 3: Resolves routine technical issues with occasional supervisor input.
    • Level 1: Requires step-by-step guidance for tasks the role description lists as core competencies.

    3. Initiative

    • Level 5: Identifies a process gap and builds a solution before being asked.
    • Level 3: Responds well to assigned improvement projects.
    • Level 1: Waits for explicit instructions even when problems are visible.

    4. Customer service

    • Level 5: Resolves escalated complaints on the first call and follows up to confirm satisfaction.
    • Level 3: Handles standard requests correctly but escalates most complications.
    • Level 1: Provides incorrect information and shows visible frustration with difficult callers.

    A few rules keep these anchors useful instead of decorative. Write anchors around what someone did, not who they are. “Volunteers to cover shifts during staffing shortages” survives scrutiny; “has a great attitude” does not. Avoid adjectives that require judgment calls, since “handles pressure well” reintroduces the exact ambiguity BARS is built to eliminate.

    Adapting anchors by seniority means changing scope, not swapping in vaguer language. A junior analyst’s Level 5 initiative anchor might be “flags a data inconsistency to the team lead.” A senior analyst’s Level 5 anchor should describe the same behavior at a higher scale, something closer to redesigning the workflow that caused the inconsistency in the first place.

    How Do You Build a Behaviorally Anchored Rating Scale?

    Building a valid BARS instrument is not a weekend project, and treating it like one is the most common reason organizations abandon the format after one cycle. Here’s the sequence practitioners actually use.

    1. Form a governance team and define the purpose. Decide upfront whether the scale will drive developmental feedback, compensation decisions, or both. This choice changes how incidents get worded and how strictly you enforce anchor language later, since a scale feeding into raises needs tighter legal defensibility than one used purely for coaching conversations.

    2. Collect critical incidents from the people closest to the work. Pull incidents from supervisors and, critically, from top performers themselves. Aim for a broad sample across the dimension you’re scaling, not just a handful of stories from one manager’s memory. Write each incident as a specific, observable event rather than a trait or a summary judgment.

    3. Retranslate incidents into draft anchor statements. A second group of subject matter experts (SMEs), ideally people who did not write the original incidents, sorts them into performance dimensions and rewrites vague ones into behavior-specific language.

    4. Have SMEs independently rate each incident’s effectiveness, typically on a 7 to 9-point scale, then compute the mean and standard deviation for each one. Keep incidents where SMEs agree closely (low standard deviation) and discard the ones where ratings scatter widely, since disagreement usually means the incident is ambiguous or dimension-specific interpretation varies too much to be useful.

    5. Pilot the scale with real raters on real (or recent, anonymized) performance data. Measure inter-rater reliability before rolling the scale out organization-wide, and run calibration sessions where raters compare scores on the same incidents and discuss discrepancies openly.

    6. Finalize, document, and set a revision trigger. Lock the anchor language, publish the rationale behind each dimension, and set a review date rather than letting the scale run indefinitely without a check.

    Pro Tip: Recruit your most skeptical manager for the SME panel, not just your most enthusiastic one. Skeptics catch vague anchor language that enthusiastic raters wave through, and their buy-in during development tends to predict adoption better than anyone else’s.

    The AIHR development framework treats the labor involved here as a feature, not a bug. The debate among SMEs about what counts as “exceptional” surfaces disagreements about standards that would otherwise stay hidden until a contested performance review.

    What Validation Checks Keep a BARS Instrument Reliable?

    A BARS scale is only as good as the reliability data behind it, and skipping the psychometric check is how organizations end up defending a “validated” instrument that was never actually tested.

    Start with inter-rater reliability. Have multiple raters independently score the same set of employees or recorded incidents, then calculate agreement using an intraclass correlation coefficient (ICC) or a similar index. A small pilot sample, ideally 15 to 20 rating pairs at minimum, tells you whether raters are interpreting anchors consistently before you scale the instrument organization-wide.

    Watch for two specific bias patterns:

    • Halo error, where a rater’s overall impression bleeds into every dimension score, tends to drop under BARS compared to less structured formats.
    • Leniency error, where raters cluster scores toward the high end regardless of actual performance, can actually increase.

    A comparative psychometric study using 727 raters found exactly this trade-off: BARS reduced halo error relative to summated rating scales but showed greater leniency and somewhat lower inter-rater reliability. That is not an argument against BARS. It is an argument for building leniency checks into your calibration process rather than assuming the format solves bias on its own.

    During pilot analysis, flag any incident where SME effectiveness ratings show high standard deviation and pull it from the final anchor set. Rerun inter-rater checks after your first calibration training round, since reliability typically improves once raters have discussed disagreements together. Revalidate anchors every 18 to 24 months, or sooner if job responsibilities shift meaningfully, a new tool changes how the work gets done, or calibration sessions start surfacing consistent disagreement on the same dimension.

    Illustration of rating calibration convergence

    Where BARS Pays Off and Where It Falls Short

    BARS earns its reputation on a few specific strengths. It gives employees concrete, actionable feedback tied to behavior rather than a number they have to guess the meaning of. It holds up better under legal scrutiny than vague trait-based scales, because a manager can point to the specific behavior an employee did or didn’t demonstrate. Research on the format’s longevity credits its participatory development process with building organizational buy-in that other appraisal formats rarely achieve.

    The limitations are just as real:

    • Development is genuinely time-consuming, often taking weeks of SME involvement per role family.
    • Anchors need updating as job responsibilities evolve, which means ongoing maintenance cost, not a one-time build.
    • The format fits poorly for strategic, creative, or long-horizon work where the “right” behavior varies too much case to case to anchor meaningfully.
    • Small organizations without dedicated HR analytics capacity may struggle to run the statistical validation steps properly.

    The practical decision rule: invest in BARS for roles with high headcount, recurring observable tasks, and evaluations that carry real stakes, whether that’s compensation, promotion, or legal exposure. For a three-person creative team, a simpler method will serve you better.

    How Do You Roll Out BARS Without Losing Rater Buy-In?

    Rollout succeeds or fails on governance and training, not on the anchor wording itself.

    1. Document the governance model. Name who owns anchor updates, who approves revisions, and where the rationale for each dimension lives so a new HR hire isn’t reverse-engineering the logic two years later.
    2. Run frame-of-reference training before go-live. Give raters two or three sample incidents and have them independently score against the anchors, then compare and discuss any mismatches as a group.
    3. Structure calibration meetings around real disagreement. Pull a random audit sample each cycle, review scores together, and treat outliers as a training signal rather than a rater failure.
    4. Set a lightweight revalidation cadence. A short annual anchor review, distinct from the full 18 to 24 month psychometric revalidation, catches language that’s drifted out of date before it causes a dispute.

    Pro Tip: Keep calibration meetings under 45 minutes and focus on three or four disputed ratings, not every score on the sheet. Raters retain the lesson from a focused disagreement far better than from a marathon review session.

    Where Does BARS Fit Across the Employee Lifecycle?

    BARS shows up at more points in the employee lifecycle than most HR teams initially plan for.

    • Hiring interviews use the same anchor logic to score candidate responses against observable behavior rather than gut feel.
    • Onboarding checkpoints at 30, 60, and 90 days can borrow junior-level anchors to track early performance signals.
    • Development plans cite specific anchor gaps as the basis for a coaching conversation, which is more actionable than “needs improvement.”
    • Promotion decisions lean on higher-level anchors to justify the call with documented behavior, not opinion.

    Pair BARS with BOS when you need frequency data (how often does this behavior occur, not just which level it hit) and with MBO (management by objectives) when the role’s success is better measured by outcomes than by observable process. A hybrid approach, BARS for behavior-based competencies and MBO for revenue or output targets, often captures a role more completely than either method alone.

    What Actually Matters When You Build One of These

    Most of the failed BARS rollouts I’ve seen didn’t fail because the anchor wording was bad. They failed because leadership treated the SME debate phase as a bottleneck to rush through, when that argument over what “exceptional” looks like is the actual product. Skip it, and you’ve built a scale that looks rigorous but carries none of the shared understanding that makes it defensible later.

    Start with one role family, not the whole organization. Get the incident collection and calibration right at a small scale before you try to standardize across departments with wildly different definitions of good work.

    — Raul

    How Resyme Standardizes Behavior-Based Interview Data

    Building solid anchors depends on collecting consistent, well-documented behavioral evidence, and that’s exactly where most hiring teams lose quality control. Manual interviews vary wildly by interviewer, mood, and note-taking habits, which makes the incidents you’d use to build or validate a BARS dimension unreliable before you even start scaling them.

    Resyme

    Some platforms run role-tailored, behaviorally anchored interviews automatically, asking candidates structured questions built on deep domain knowledge for the role in question. Because such formats are consistent across candidates, the behavioral data they produce is far easier to scale into anchors than inconsistent interviewer notes. These platforms may validate candidate honesty and produce objective, comparable reports, which can aid internal calibration exercises as well as hiring decisions.

    If your team is piloting a new BARS dimension and needs a faster way to gather standardized behavioral evidence, start with Resyme’s interview platform and see how automated, consistent interview data compares to what your current process produces.

    Sources

    FAQ

    What Is a Behaviorally Anchored Rating Scale?

    A behaviorally anchored rating scale (BARS) is a performance appraisal method that defines each point on a rating scale with a specific, observable behavior instead of a generic label like “excellent” or “poor.” Smith and Kendall introduced the format in 1963, and most versions use a vertical scale of 5 to 9 points.

    What Are the 5 Levels of Performance Rating?

    A common 5-level BARS structure ranges from highly ineffective at the bottom to highly effective at the top, with each level tied to a specific behavioral anchor rather than a numeric score alone. The exact wording varies by organization and by which dimension, such as teamwork or technical skill, the scale is measuring.

    What Is the Difference Between Graphic Rating Scales and BARS?

    A graphic rating scale asks raters to mark a point along a line between “poor” and “excellent” without defining what those points look like in practice, which leaves interpretation open to each rater’s own standard. BARS replaces that open interpretation with a specific behavioral description at each level, which is why BARS shows reduced halo error compared to less structured formats.

    How Much Does Resyme Cost?

    Resyme offers a Company plan and an Agency plan, along with a pay-as-you-go interview option and Free and Enterprise tiers; current pricing details are available on Resyme’s site. Enterprise and Free plan pricing is available on request depending on your hiring volume.

    How Often Should You Update BARS Anchors?

    Most organizations revalidate BARS anchors every 18 to 24 months, or sooner if job responsibilities change significantly or calibration sessions reveal consistent rater disagreement on a specific dimension. A lighter annual language review can catch outdated wording between full revalidation cycles.