The 360 Review Alternative: Assess Leadership in Action
Leadership DevelopmentL&DPerformance ManagementWorkplace Learning

The 360 Review Alternative: Assess Leadership in Action

Kontaim

Kontaim

@Argraide

Sep 18, 2026

The report arrives as a 14-page PDF. Three direct reports, two peers, and a manager have rated the leader on communication, strategic thinking, and collaboration. The report says the leader should “communicate more strategically.” The development plan says the leader should “communicate more strategically.” Six months later, the same sentence appears in the next review.

Traditional 360 reviews are not dying because multi-source feedback is useless. They are losing their grip as a standalone annual ritual. A survey can show how a leader is experienced over time, but it rarely shows what that leader does when priorities collide, information is incomplete, or a difficult conversation goes sideways.

An interactive assessment asks a person to make a decision, explain the reasoning, respond to new information, or practice a conversation. It scores observable choices against a defined rubric. That makes it a useful 360 review alternative when the question is what someone does under pressure, rather than how colleagues remember them.

The distinction matters. Smither, London, and Reilly’s meta-analysis of multisource feedback found modest average improvement after feedback, with stronger results when people set goals and received follow-up. The report can be useful. It is not the behavior-change intervention.

Myth 1: “If we get enough raters, the average will be objective.”

An average is a mathematical treatment of disagreement, not proof that everyone observed the same behavior.

Rater count helps when people have enough opportunity to see the behavior and are using a common frame. A manager may see how a leader escalates risk. A peer may see how that leader handles disagreement in a meeting. A direct report may see delegation and follow-through. Those are different slices of the job, not interchangeable camera angles.

Raters also bring different standards. One person’s “decisive” is another person’s “rushed.” A score of 4 may mean “I saw this twice and liked it,” not “this leader consistently demonstrates the behavior.” Anonymity can make criticism safer, but it can also produce vague comments with no accountability for evidence.

The practical fix is to sample behavior, not adjectives. Start with three recurring leadership moments, such as challenging a peer’s recommendation, resetting priorities after a missed deadline, or giving corrective feedback to a strong performer. Turn each moment into a short assessment with an initial set of facts and one new piece of information. Ask for the next action and the reasoning behind it.

Score markers that someone could actually observe: names the trade-off, checks the relevant facts, assigns an owner and date, tests understanding, or communicates the consequence to affected people. Replace “communicates strategically: 3.2” with evidence such as “when priorities changed, named what would stop, assigned ownership, and confirmed the client implication.”

That evidence can sit beside 360 feedback. It should not be swallowed by the average.

Myth 2: “The self–other gap tells you exactly what to fix.”

A gap is a clue, not a diagnosis.

The familiar interpretation is that a leader who rates themselves higher than others lacks self-awareness. Sometimes that is true. Other explanations include limited rater visibility, different expectations for the role, a leader being unusually self-critical, or observers reacting to an outcome rather than the behavior that produced it.

Reviews by Fleenor and colleagues have linked self–other agreement with leadership effectiveness, but the direction and meaning of the gap still require context. A numerical difference cannot tell you whether the issue is judgment, confidence, role ambiguity, or a manager who has failed to explain the standard.

The more useful mismatch is often between intention and enacted choice. A leader may say they value open debate, then choose the first proposal because the meeting is running late. That discrepancy tells a coach more than a self-rating of 3 beside a group rating of 4.

An interactive assessment can expose that mismatch. Before showing feedback, ask the leader to respond to a realistic situation, rate their confidence in the choice, and name the trade-off they are making. Then ask observers to mark which behaviors they saw in the response. Afterward, compare four pieces of evidence: the chosen action, the stated reasoning, the observed behavior, and the likely effect.

The resulting calibration note is far more actionable than “close your self-awareness gap.” It might say: “You recognized the risk but did not make the decision rule explicit. In the next prioritization meeting, state the rule before choosing.”

This method costs more time than sending a survey. It requires a deliberate debrief and a rubric that people understand. That cost is worthwhile when the goal is better judgment rather than a more interesting report.

Myth 3: “Make it interactive and it will be more valid.”

Interactivity can improve an assessment, but it can also add irrelevant difficulty.

Research on situational judgment tests, including meta-analytic work by McDaniel and colleagues, finds that these formats can predict work performance. It also finds variation by job, item design, instructions, and scoring method. A polished branching scenario is not automatically a sound measure.

The format has to resemble the work. A weak item presents a missed deadline and asks the participant to choose between three stock responses. The “correct” answer is obvious, and the score rewards familiarity with corporate language. A stronger item gives the participant a client commitment, limited capacity, and incomplete information. The person chooses a response, explains the trade-off, then receives a new constraint. The rubric looks for clarification, prioritization, ownership, and communication—not a particular personality or tone.

Messick’s validity framework is useful here. Ask four plain questions:

  • Does the task represent an important part of the job?
  • Are the desired behaviors visible in the response?
  • Would trained assessors score the response consistently?
  • Could irrelevant factors such as language fluency, disability, cultural style, or comfort with timed tasks affect the result?

A small panel of three to five experienced role holders can test the item before launch. Have them independently identify the strongest and weakest response, then explain why. If they disagree about the behavior being measured, the item needs work.

This fails when the assessment rewards verbal polish, assumes one culturally correct way to handle conflict, or compresses a six-month strategic decision into a five-minute puzzle. In those cases, use a structured interview, a real work sample, or observation of an actual meeting. Interactive assessment is a tool, not a license to pretend every leadership problem has a hidden right answer.

Myth 4: “Once people receive the report, development begins.”

Report delivery creates awareness at best. Development begins when someone repeats a behavior under pressure and gets useful feedback on the attempt.

This is where many 360 processes quietly stop. The leader receives a debrief, selects a broad goal, and returns to a calendar full of the same situations that produced the original feedback. No one observes the next difficult meeting. No one records whether the behavior changed. The next review then measures memory again.

Use the assessment as rehearsal. A leader completes a baseline scenario, receives feedback on two observable choices, and writes an implementation intention. Gollwitzer’s research on implementation intentions gives the structure: “If X happens, I will do Y.” For example: “If a project owner reports a delay, I will ask for two recovery options before proposing one.”

The manager then looks for that behavior in a real interaction. A ten-minute follow-up is enough to record the situation, the action taken, and the effect. After 30 days, the leader completes a parallel scenario rather than repeating the same question. The point is to see adaptation, not memorization.

This is Kirkpatrick’s Level 3 concern: transfer into behavior. A positive reaction to the assessment or a higher confidence score is not evidence of transfer. If a manager will not make time to observe the behavior, do not claim that the assessment measures workplace change. It measures a response to an assessment.

Myth 5: “An interactive assessment should replace 360 feedback everywhere.”

That is another bad shortcut. Different methods answer different questions.

A 360 review is useful for identifying patterns in how a person is experienced across relationships over time. An interactive assessment is better suited to judgment, prioritization, coaching, escalation, and other behaviors that can be represented in a realistic decision. A work sample or live observation is stronger when execution depends on context that a scenario cannot reproduce.

A credible 360 review alternative is therefore a sequence, not a shinier questionnaire: use multi-source feedback for the pattern, an interactive assessment for the decision, and manager observation for transfer. For promotion, succession, or other high-stakes decisions, do not use a single scenario score—or a single 360 report—as the gate. Use multiple methods, consistent conditions, accommodations, and a documented review of the evidence.

That is performance review innovation worth keeping: changing the evidence and follow-up, rather than dressing the old survey in brighter colors.

Run a five-day pilot before redesigning the whole process

You can test this approach with one leadership behavior and a small group this week.

  1. Choose a behavior with visible consequences, such as reprioritizing work after a missed commitment. Avoid broad labels like “executive presence.” Collect two real examples of the behavior from the past quarter.

  2. Write a seven-minute assessment. Include the situation, the decision required, a request for reasoning, and one new piece of information. Define three or four observable markers before anyone takes it.

  3. Ask three experienced practitioners to score the responses independently. Remove any item where they cannot agree on what good evidence looks like.

  4. Run the assessment, debrief the participant, and write one if–then plan. Keep the pilot developmental; do not attach the score to pay or promotion while the method is being tested.

  5. Schedule a manager observation within 30 days and a parallel assessment afterward. Compare the evidence from both moments, not only the numbers.

By Friday, you should have one scenario, one rubric, and one scheduled observation. That is enough to find out whether the new method is capturing leadership behavior—or merely producing a more entertaining score.