360 Feedback for Executive Coaching: Full Workflow

360 Feedback for Executive Coaching: From Rater Selection to Development Plan

A 360 feedback for executive coaching creates value when it produces a small set of credible behavior hypotheses and a way to test them. The report is an input. The development plan,and what happens in real work after the debrief is the output. The workflow below assumes a development purpose. If results will influence selection, promotion, pay, or formal performance decisions, the method, consent, governance, and assessor qualifications need a different level of scrutiny.

1. Define the purpose and data owner

Write one sentence naming the decision the 360 should support. For example: “Identify two leadership behaviors for the client to practice during a six-month coaching engagement.” Then state who owns the individual report, what the organization will receive, how long data will be retained, and whether comments will be shared verbatim, summarized, or withheld. CCL’s implementation guidance says the purpose, outcomes, confidentiality, and anonymity should be clear and describes a practice in which the participant owns the individual data while the organization may use a group profile. Your programme may differ, but ambiguity should not.

2. Choose the instrument before the raters

Use a validated instrument when norms, reliability, or high-stakes comparison matter. A custom pulse can be appropriate for a narrow coaching outcome, provided it is presented as a practical feedback tool rather than a validated test. Match every item to behavior the intended raters can observe. Avoid a catalogue of leadership virtues. A shorter instrument tied to the client’s role usually yields better examples and a more usable repeat measure.

3. Build the rater map

List the situations in which the target behavior occurs, then identify people with recent exposure. Do not begin with an organizational chart and assume every box has useful evidence.

Rater groupSelection questionCommon risk
Manager or sponsorHas this person observed the behavior in comparable situations?Rating strategic outcomes they did not directly see
PeersDo they share decisions, resources, or customer work with the client?Selecting only allies or frequent collaborators
Direct reportsHave they experienced the leader under routine and pressured conditions?Fear of identification in a small team
Other stakeholdersCan they add a distinct context rather than duplicate another rater?Low exposure and reputation-based scoring

Record role, exposure context, frequency, and any conflict of interest. Let the client propose names, but reserve a three-way check for coverage gaps or a list designed only for comfort.

4. Set anonymity and comment rules

Anonymity needs an operational threshold, not a promise that collapses when only one person responds. Decide the minimum number for each category, what happens below the threshold, and whether groups may be combined without creating a misleading average. Explain that anonymity reduces identification risk; it cannot guarantee that a client will never infer a source from wording or context. Tell raters not to include names, medical information, investigations, or confidential customer details. Review comments before the debrief and remove incidental identifiers without changing the meaning.

5. Invite raters with a precise brief

The invitation should name the development purpose, response deadline, estimated time, confidentiality rule, data recipients, and a contact for questions. Ask raters to use “not enough exposure” rather than guess. Reminders should be proportionate and stop after completion. A practical invitation line is: “Rate only behavior you have observed in the last six months. Use the comment field for a specific work situation, and leave out personal or confidential information.”

Plan for missing data

Set a close date and a decision rule before the first invitation. If a category falls below the anonymity threshold, extend the window once or suppress that category. Do not pressure a lone respondent to recruit colleagues or reveal who has not replied. Completion is voluntary within the programme rules, and non-response should not be scored as a negative rating. When the final rater mix differs materially from the baseline, show the counts and treat the comparison as directional. A higher average from a smaller or friendlier sample is not evidence of improvement. Qualitative examples may still inform the coaching, but the report should state the coverage change.

6. Quality-check before the debrief

Check completion and exposure by rater group. Confirm that anonymity thresholds are met before showing group results. Compare self and other ratings without treating either as objective truth. Look at distribution and comments, not only the mean. Flag contradictory evidence and contextual differences. Separate data-quality problems from genuine development themes. Do not send a dense report without a conversation when the instrument or findings could be emotionally charged. Prepare the client for the process and the limits of interpretation.

7. Debrief in the right order

Begin with the client’s reaction and the conditions needed to stay curious. Then review coverage and limitations before discussing scores. Move from patterns to examples, from examples to hypotheses, and from hypotheses to choices. What feels accurate, surprising, or context-dependent? Where do rater groups see the behavior differently, and what situations differ? Which comment adds an observable example rather than a label? What strength might be overused under pressure? Which one or two changes would matter most in the client’s work? A gap is not automatically a deficit. A manager and direct reports may reasonably want different levels of detail. The coaching task is to find the context in which adaptation would improve outcomes.

8. Convert findings into a development experiment

Each selected behavior needs a situation, action, evidence source, and review point. “Listen more” becomes: “In the weekly portfolio meeting, summarize the opposing view before stating a decision; ask two regular attendees for a two-question pulse after four meetings.” Run a midpoint pulse with the same core items when enough exposure exists. At close, compare like with like: same wording, similar rater groups, clear counts, and recorded context changes. A score delta without comparable coverage should be described, not celebrated.

A minimum-viable 360

When time or sample size is limited, use six to ten behavior items, three to five comment prompts, and a rater set chosen for exposure. Keep the output to two themes and one practice experiment. This is a focused multi-rater pulse, not a comprehensive leadership diagnosis.

Where 360 programmes fail

  • The sponsor wants a performance verdict while the client expects development.
  • Raters are selected for hierarchy or convenience rather than observation.
  • A small category is shown even though anonymity was promised.
  • The coach interprets decimals without examining spread or exposure.
  • Every low score becomes a goal, producing a development plan no one can use.
  • The report is filed after the debrief and never tested in real work.
  • A closing score is compared with a baseline from a different rater mix without qualification.

Operational support without automated judgment

CoachComet’s public guides describe custom and prebuilt assessments, anonymized multi-rater aggregation, reminders, repeated stages, and sponsor-ready reporting. Those controls can reduce spreadsheet work and keep the evidence trail together. They should not choose raters, validate a custom instrument, interpret a result, or decide what the sponsor may see. A strong 360 ends with fewer claims than it began with: two behavior hypotheses, one real-world experiment, named evidence, and a review date. That is enough to make the feedback consequential.

Exceptional Coaching, Made Simple

Deeper insights, effortless practice management, and better outcomes for every client.

Get Started