How to run a team health check that surfaces honest signal: which dimensions to assess, why anonymity matters, questions that work, and how to turn results into real action.

If your team's last retrospective produced the same three action items as the one before it — and nobody is sure whether last quarter's issues were actually resolved — you do not have a retrospective problem. You have a signal problem. Team health checks exist to fix it.
A sprint retrospective looks at what just happened. A health check looks at how the team is doing across time. Those are different questions, and they need different tools.
What a team health check actually is
Workshop Weaver defines a team health check as a structured exercise that helps teams assess how they are working together across multiple dimensions — psychological safety, delivery, process, and morale — rather than just what they shipped. The format was popularised by the Spotify Squad Health Check model in 2014, which remains widely adapted because it separates 'how we feel about our work' from 'what we achieved.'
The Spotify version visualises results as a heat map across multiple squads over time. That visualisation is where the model earns its keep: a single squad with amber 'Fun' scores might have an internal problem. Three squads with amber 'Fun' scores in the same quarter points to something the organisation is doing to those teams. A single retrospective cannot surface that pattern. A health check can.
Atlassian's Team Health Monitor uses a similar approach — eight dimensions, green/yellow/red voting, focused discussion on the lowest scorers — and is freely available for any team to run. One documented pattern in their playbook: teams chronically scoring yellow on 'Delivering Value' often traced the root cause to unclear product ownership, not technical debt. That is the kind of systemic insight that retrospective post-its rarely produce.
The dimensions worth assessing
Effective health checks use 6–10 dimensions. Fewer and you miss important signals. More and you get survey fatigue, which produces garbage data.
The categories that matter most: delivery pace and predictability, work quality, cross-functional collaboration, psychological safety, goal alignment, and personal wellbeing. Every one of those covers territory a sprint retrospective might skim past.
The single dimension teams most often skip — and most need — is psychological safety. Google's Project Aristotle research, a two-year study of 180 teams, found that psychological safety was the number-one factor differentiating high-performing teams. Individual talent, seniority, team size — none of these predicted performance as reliably as whether team members felt safe to take risks and speak up. A health check that omits a safety dimension is ignoring the variable most likely to explain poor outcomes.
Amy Edmondson's research on psychological safety is the academic bedrock here. Her work makes one thing plain: high-performing teams have more visible failures, not fewer, because they are in environments where it is safe to surface them. A health check that shows green on safety while the team is consistently missing commitments should make you suspicious, not reassured.
Why anonymity changes the scores
Without anonymity, health check scores cluster toward the middle. People avoid extremes to protect relationships or because they fear management will draw conclusions about them personally. This is not a character flaw — it is a rational response to a non-anonymous format.
Harvard Business Review research on employee voice consistently shows that individuals withhold negative feedback when they cannot be certain of confidentiality, particularly in teams with power gradients. The more senior people are present, the more scores regress to the mean.
Anonymity matters most for dimensions touching on management quality, interpersonal trust, and psychological safety. For technical dimensions like deployment frequency or code quality, anonymous and attributed scores tend to converge — which itself is a useful signal about where the psychological risk in your team actually lives.
One practical floor: you need at least five respondents for anonymity to hold. Below that, even aggregate scores can be reverse-engineered to individuals. Teams smaller than five are better served by structured facilitated conversation with explicit ground rules than by an anonymous survey.
A pattern worth knowing: when retrospective tools introduced anonymous voting for health dimensions, facilitators reported that 'Fun' and 'Psychological Safety' scores dropped noticeably compared to non-anonymous baselines. The team had not gotten worse. Truthful scoring had become possible. The gap between anonymous and attributed scores is itself a diagnostic.
Questions that produce honest answers
The best health check questions describe observable behaviours or concrete outcomes. 'We know what good looks like for this team and we are moving toward it' is more actionable than 'Do you feel aligned?' because it anchors the response in evidence someone can point to.
Write questions in first-person plural. 'We...' rather than 'Do you...' reinforces collective ownership and makes the follow-up conversation easier — the team owns the result together, which prevents the discussion from collapsing into finger-pointing.
Use a three-point scale: Green / Yellow / Red, or Awesome / OK / Struggling. A 1–10 Likert scale sounds more precise but introduces inter-rater variance without improving the diagnostic value for group assessments. It also makes heat maps unreadable. The simplified scale forces a real judgment and produces visuals you can actually use.
Here is a starter question set adapted from the Spotify and Atlassian frameworks:
- We have a clear, shared goal and we believe in it.
- We deliver value regularly and predictably.
- We learn from mistakes without blame.
- We have the skills and support we need.
- We enjoy working together.
- Our ways of working help us rather than slow us down.
Each rated Green / Yellow / Red, with a one-line comment field. Do not make the comment field optional. Quantitative scores tell you where to look. Comments tell you what you are looking at. The comment field is where the actual diagnostic data lives.
Running the session
The facilitator's job is to hold space for disagreement, not drive consensus. When scores diverge — one person rates 'Collaboration' green, another rates it red — that gap is the most valuable data point in the room. Name it explicitly: 'We have a split here. Can the people at each end tell us what they are seeing?'
Timeboxing matters. Spend roughly five minutes per dimension on initial scoring. Reserve 30–40 minutes of a 60-minute session for the two or three lowest-scoring areas. The most common failure mode is trying to address every dimension at once — teams leave feeling overwhelmed and nothing gets followed up.
The facilitator should not score the dimensions themselves. Even anonymous facilitator scoring changes group dynamics and introduces anchor bias. The facilitator's role is process.
Agile coach Esther Derby describes a principle she calls 'mining the data before moving to solutions'. Teams rush to fixes before fully understanding what the scores are telling them. Spend at least 10 minutes asking 'what do we observe in these results?' before any action planning begins. This single habit consistently improves the quality of commitments that come out of health checks.
Turning results into action
The most common reason teams stop doing health checks is that previous results were collected, acknowledged, and then ignored. Gallup's State of the Global Workplace report consistently finds that a minority of employees feel their feedback leads to action — and that dynamic directly predicts disengagement from feedback processes over time.
Every health check must produce at least one concrete, owned, time-bounded action. Even a small one. If the format does not generate anything that gets done, it loses credibility within two or three cycles and teams start gaming the scores to avoid the conversation.
Track scores over time using a simple heat map: dimensions as rows, quarters or sprints as columns, green/yellow/red as cell values. A dimension that stays yellow for three consecutive periods is a more urgent signal than a new red. It means the team has normalised a structural problem rather than solved it.
ThoughtWorks uses a variant they call a Team Radar. After each quarterly check, each team publishes one 'team commitment' publicly — not to management, but to other teams. Cross-team visibility increases follow-through substantially because teams do not want to appear at the next check with the same red dimension and no visible effort.
Share aggregated results upward selectively and with team consent. The point is to surface systemic blockers the team cannot fix themselves. Persistent amber scores on 'Delivering Value' across five teams is a leadership problem, not a team problem, and leadership needs to see it framed that way.
Templates you can use today
A functional template needs three things: a defined set of dimensions (6–8), a clear rating scale with anchors describing what green and red look like, and a comment field. That is the minimum viable health check.
For digital sessions, tools like Miro, EasyRetro, and Neatro all offer health check templates with anonymous voting built in. TeamRetro publishes a library of over 30 templates covering engineering teams, remote teams, leadership teams, and project teams — their 'Agile Team Health' template maps directly to the Spotify model dimensions and includes facilitator notes for each question, which makes it useful for teams running their first check without an experienced agile coach.
For in-person sessions, a physical whiteboard with sticky dots works well. Anonymity comes from simultaneous dot-placement or private ballot slips collected before the reveal.
Let the template evolve. After three or four cycles, some dimensions will be consistently green and no longer need monitoring. Replace them with dimensions that reflect current challenges. A health check that never changes is one the team has stopped thinking about.
The health check as a team instrument
A team health check is not an HR exercise. It is not a management reporting tool. It is a diagnostic instrument that the team owns, controls, and uses to advocate for what they need — from leadership, from the organisation, from each other.
The goal is not a perfect score. A perfect score in a non-anonymous health check is almost certainly a dishonest one. The goal is an honest score, tracked over time, that produces at least one specific thing the team commits to changing.
Run your first health check this sprint. Pick one of the templates linked above, choose one dimension to act on, and check whether it moves next quarter. That is the whole method. Everything else is refinement.
💡 Tip: See how AI-powered planning changes workshop prep.
Learn More