Skip to content
SuccessionStack

Is Your 9-Box Grid Secretly Biased?

Research on gender bias in 9-box potential ratings, why the potential axis is exposed to bias, and what actually helps versus what just feels like it does.

By Ryan Grant · Published July 6, 2026

Ask most HR leaders whether their 9-box grid is fair, and they'll say yes, because the categories are the same for everyone and the meeting follows the same process every cycle. Ask the research, and the answer is less comfortable: structured or not, the "potential" axis is where bias slips in, and it slips in quietly enough that nobody in the room notices it happening.

What the research actually found

The clearest evidence comes from a study by Alan Benson, Danielle Li, and Kelly Shue, who analyzed performance and potential ratings for roughly 30,000 management-track employees at a large North American retail chain over several years. Women in the dataset received meaningfully lower ratings on the "potential" axis than men, despite earning higher performance ratings, and the gap in potential scores explained a large share of why they were promoted less often. MIT Sloan's coverage of the research puts it plainly: the same women who were outperforming their peers were still being rated as having less room to grow.

That finding isn't an argument against structure. It's an argument for noticing which part of the structure is doing the damage.

Why potential is where bias hides

Performance is comparatively hard to fake, because it gets checked against something: revenue, deadlines, delivered work. Potential has no equivalent anchor, which is why the high-potential label deserves more scrutiny than it usually gets. It's a forecast, made by a manager, about what someone might become in a role they haven't held yet, and forecasts run on whatever evidence is already lying around, which usually means confidence, visibility, and resemblance to the person making the call.

Gallup has made a version of this same critique of the traditional 9-box model, arguing that the conventional approach leans on a vague, subjective idea of "potential" and proposing a readiness-based alternative instead: sort people into Ready Now, Ready Next, or Expert, categories anchored to an actual timeline rather than a guess about character. It's worth noting how close that proposal already sits to a scored, evidence-based bench: readiness is a claim you can check later, potential is a vibe you can't.

What doesn't fix it

A few common responses to this problem don't actually move the number:

  • More training on the grid. Calibration training helps people use the tool consistently. It does very little to change what the "potential" label is measuring in the first place, because the ambiguity is built into the axis, not into any individual rater's skill.
  • A different set of labels. Renaming the boxes, adding a color legend, or redesigning the grid's layout changes the presentation without touching the underlying problem: an unanchored judgment call still sits behind the label.
  • Waiting for it to average out. Bias in individual ratings compounds across cycles rather than canceling out, because the same names keep landing in the same cells for the same soft reasons every year.

What actually helps

The corrective isn't complicated, but it requires giving up the shortcut of an unexamined "potential" call:

  • Name the criteria before anyone rates anyone. Decide, in writing, what "potential for this role" actually requires: which capabilities, which evidence would demonstrate them, before scores get entered. A criteria list agreed on in January is harder to bend toward a favorite than one improvised in the room during calibration. A fixed, weighted set of leadership dimensions is one way to hold that line, because the weights get argued about once rather than per candidate.
  • Require evidence for every score, not just a number. A rating with no evidence attached is an opinion with a number attached. A rating that has to cite a specific, dated example is a claim someone can actually check.
  • Calibrate against the evidence, not the impression. The calibration meeting should compare what was written down for each person, not re-litigate general impressions of who "feels" ready.
  • Track the pattern, not just the individual placement. A single questionable rating is a judgment call. A pattern of one group consistently landing lower on potential despite comparable or better performance is a signal worth investigating on its own.

Where the audit trail fits, honestly

SuccessionStack logs a reason for every score and every weight change, which is worth being precise about: that's a transparency mechanic, not a bias detector. It doesn't automatically flag a demographic pattern, and nobody should imply it does. What it does is make an ungrounded placement visible in the moment it's made, because "strong presence" isn't an acceptable reason to enter into a required field, and a rater who has to write down the actual evidence is a rater who has already caught a good number of their own weakest calls before anyone else has to.

That's a smaller claim than "software fixes bias," and a more honest one. The fix is the criteria, the evidence requirement, and the willingness to look at the pattern across a whole talent review, not just one placement at a time. Software that makes evidence mandatory and inspectable is what makes doing that actually practical instead of aspirational.

FAQ

Is the 9-box grid biased?

Not inherently, but its potential axis is unusually exposed to bias because it measures a forecast rather than a fact. Research analyzing roughly 30,000 employees found women received lower potential ratings than men despite higher performance ratings, which is the clearest evidence that the axis, left unanchored, tracks something other than actual capability.

Does an audit trail fix bias in talent reviews?

No, and it shouldn't be sold as if it does. A required reason on every score change is a transparency mechanic: it makes an ungrounded rating visible at the moment it's entered instead of hidden inside someone's impression. Fixing bias itself takes named criteria, evidence requirements, and a willingness to examine patterns across a full review, not a single feature.

What should replace the potential axis?

Some organizations keep "potential" but require scored, dated evidence behind it. Others, following Gallup's proposed alternative, replace it with a readiness timeline: Ready Now, Ready Next, or Expert, which anchors the judgment to a checkable prediction instead of an open-ended trait.

Questions buyers actually ask

Not inherently, but its potential axis is unusually exposed to bias because it measures a forecast rather than a fact. Research analyzing roughly 30,000 employees found women received lower potential ratings than men despite higher performance ratings, which is the clearest evidence that the axis, left unanchored, tracks something other than actual capability.

No, and it shouldn't be sold as if it does. A required reason on every score change is a transparency mechanic: it makes an ungrounded rating visible at the moment it's entered instead of hidden inside someone's impression. Fixing bias itself takes named criteria, evidence requirements, and a willingness to examine patterns across a full review, not a single feature.

Some organizations keep "potential" but require scored, dated evidence behind it. Others, following Gallup's proposed alternative, replace it with a readiness timeline: Ready Now, Ready Next, or Expert, which anchors the judgment to a checkable prediction instead of an open-ended trait.

See where your bench breaks before it matters.

Bring your real org chart. We show you the succession gaps, cascade risks, and bench depth in a 30-minute walkthrough. IT security questions answered on the same call.

IT review first? The FAQs answer the security questions honestly →