The science behind the Culture Strength Check
How we measure a team's culture, the research each choice rests on, and, just as importantly, what we deliberately do not claim.
Culture is built, on purpose, one conversation at a time. The Culture Strength Check measures how strongly and deliberately a team has the six conversations that shape it, on a foundation of psychological safety, and how well that translates into engagement. It measures strength and intention, how much your culture is by design rather than by default, not what type of culture you have. Every score is built from behavioural questions grounded in published research and scored transparently, never by asking you to rate yourself.
What we measure, and what we don't
There are two very different things you can measure about a culture. The first is its content, or type: whether it is collaborative or competitive, cautious or bold. Most culture tools sort you into a type. The second is its strength and intentionality: how clearly and deliberately the culture has been built, and therefore how coherent and resilient it is. We measure the second.
We do this for three reasons. First, it is actionable: a type label leaves you nowhere to go, whereas measuring the practice of building culture hands you the change levers. Second, there is no single right culture, what works for one team fails for another, so ranking teams against one ideal is the wrong move. What matters is how clearly a culture is being built and whether people can see it and choose to belong to it. Third, it is honest: strength and intention can genuinely be read from behaviour, whereas claiming to X-ray a team's soul in eight minutes cannot.
What this instrument does not claim. It is not a personality or culture-type test. It does not prove causation for your team. And on its own, taken by a single leader, it is an informed estimate, not a validated measurement, a distinction we return to in Section 06.
Five layers, read as one system
The model is not a flat checklist. Its parts sit at different depths, and reading them as layers is what makes it a model rather than a list.
- FoundationPsychological safetyThe foundation everything builds from. If a team is not safe, people manage what they say, and every other reading turns overly optimistic.
- PracticeThe six conversationsThe levers a leader actually pulls to build culture on purpose. These are behaviours, which is what makes the read actionable.
- LensThe three gapsNot more things to measure, but the three ways we read every conversation: distance to the standard, leader versus team, and how much the team agrees.
- OutcomeEngagementThe felt result: the energy, purpose and focus the six conversations produce. We report it on its own rather than mixing it into the score, because it is an outcome of building culture, not one of the levers you pull.
- ResultsWhat the business measuresRetention, performance, innovation and growth. Shown as context and never scored, because these are business data rather than something a team can self-report. Columbia University found strong cultures turn over 13.9 per cent of their people a year against 48.4 per cent in weak ones.
Foundation · Psychological safety
Psychological safety is the shared belief that a team is safe for interpersonal risk-taking, that you can ask a question, admit a mistake, or disagree without it being held against you.1 Amy Edmondson's foundational work showed that teams higher in safety learn and perform better,1 and Google's study of 180 of its own teams (Project Aristotle) found it the single largest differentiator between its best and worst teams.2 We treat it as the substrate, not a seventh conversation: when safety is low, people do not voice the truth, so the rest of the report is systematically inflated. In scoring, low safety therefore acts as a gate on confidence in everything above it (Section 04).
Practice · The six conversations
The six conversations are drawn from Shane Hatton's book Let's Talk Culture, the behaviours through which culture is built by design. Each also maps to a published construct, so the framework is grounded in both practice and research.
| # | Conversation | What it reads | Grounded in |
|---|---|---|---|
| 1 | Expectation | Whether expectations are surfaced and shared, not privately assumed | Role clarity and ambiguity;3 shared mental models;4 group norms5 |
| 2 | Clarification | Whether values are specific, observable behaviours | Goal specificity;6 behaviourally anchored standards7 |
| 3 | Communication | Whether culture lives in shared language and visible reasons | Organisational sensemaking;8 procedural justice and transparency9 |
| 4 | Capability | Whether people have the skills and support to meet the standard | Self-efficacy;10 competence in SDT;11 the knowing-doing gap;12 training transfer13 |
| 5 | Confrontation | Whether accountability is kind and early, not avoided | Feedback intervention theory;14 task versus relationship conflict15 |
| 6 | Celebration | Whether the behaviours you want get noticed and named | Reinforcement; recognition and engagement;16 the progress principle17 |
The order is not arbitrary: each conversation builds on the last, and capability sits before confrontation deliberately, because holding someone to a behaviour they have not been equipped for is unfair and erodes trust.
Lens · The three gaps
The three gaps are not additional things we measure; they are three ways of reading the same conversations. Each has a distinct research anchor, and this framing, culture as a set of gaps rather than a grade, is the part no other instrument does.
- The aspiration gap, the distance between where a team is and the standard it aspires to. Goal-setting theory holds that performance is driven by exactly this discrepancy between the current state and a clear goal.6
- The perception gap, the distance between how a leader reads the team and how the team reads itself. The self-and-other rating literature behind 360-degree feedback shows that leaders who over- or under-rate relative to their people carry predictable blind spots.18
- The spread, how much a team agrees with itself. “Climate strength”, the degree of within-unit agreement, is an established construct: a strong average masking high dispersion is a weak, contested culture, not a healthy one.19
Outcome · Engagement
Engagement, rendered here in plain language as energy, purpose and focus, is our read of the validated dimensions of work engagement, vigour, dedication and absorption.20 It is an outcome, not a lever: you do not build engagement directly, you build the conditions and the conversations, and engagement tells you whether they are landing. Kahn's original conditions for engagement, meaningfulness, safety and availability,21 map almost one to one onto this model, with safety as the foundation, availability as capability, and meaningfulness as the purpose we read at the top.
Behavioural, not evaluative
The single most important design choice is that we never ask people to rate the construct directly. “Rate your team's psychological safety out of 100” triggers ego and social desirability, and almost everyone answers around 85.22 Instead we ask about specific, recent, observable behaviour and compute the construct from it. Asking “when someone makes a mistake here, does it tend to get held against them?” is answered honestly, because it does not feel like a judgement of the respondent, and it tells you far more.
Three moves make this work:
- Behavioural, not evaluative. We ask what happens, not how good it is.
- Recent and specific. “In the last two weeks…” beats “generally…”, because the memory of a real event resists spin.
- Reverse-scoring. Roughly a third of items are worded so that agreeing is the unflattering answer, which breaks the auto-pilot of straight agreement and helps surface the truth.
Every item maps to one facet of one construct, so results are specific enough to act on rather than vague. There are 32 items in total, across the six conversations, psychological safety and engagement. In the team version, items written from the leader's point of view are reframed to the team member's seat, so both are reading the same team through comparable questions, which is what makes the perception gap valid.
From answers to a score
- Each answer becomes 0 to 100. Five-point scales map to 0 / 25 / 50 / 75 / 100; scenario and recency questions are scored best-to-worst on the same range; reverse-scored items are flipped. Nobody ever sees an item as a score, only as a question.
- Each layer is the mean of its items, rounded to a 0-to-100 score and banded: By design (70+), Taking shape (50 to 69), or By default (under 50).
- Culture strength is the mean of the six conversation scores. Psychological safety and engagement are reported separately, as the foundation and the outcome.
- Safety is a gate. It is reported as its own band, and when it is thin it flags the whole read as likely optimistic and shapes the recommended priority, rather than simply averaging in.
- The aspiration gap is the distance from a healthy benchmark. Until enough real data exists to set population norms, that benchmark is seeded at a deliberately healthy anchor (see Section 06).
- The priority follows a simple default: fix the foundation first if safety is thin; otherwise, start with the lowest-scoring conversation, since that is usually where the most is to be gained. It is a starting point, not a rule. A leader may equally choose to build from a strength where that carries more leverage.
How the perception gap is measured
A single leader can give an informed read, but two of the three gaps, perception and spread, do not exist until the team also answers. In the team benchmark, teammates take the same read anonymously, and:
- Anonymity is protected. Individual answers are never shown, and results unlock only once at least four people have responded, so no single teammate can be identified.
- The team score for each layer is the mean of the individually computed member scores.
- The perception gap is the leader's layer score minus the team's, per layer and overall.
- The spread is the standard deviation of members' scores within each layer, our operational measure of climate strength.19 A reasonable mean with low spread is a strong, shared culture; a high spread signals a contested one, even when the average looks fine.
Validity, limits, and what comes next
A methodology that will not name its own limits does not deserve to be trusted, so here are ours plainly.
- It is self-report. It measures perception, which is a large part of what culture actually is, but it shares the known limits of self-report data.
- Solo, it is a single rater. The free leader read is one perspective, an informed estimate, and leaders are known to over-rate, particularly on psychological safety. It becomes a measurement when the team answers.
- It is not yet psychometrically validated. Item performance, internal consistency, the behaviour of reverse-scored items, and the factor structure will be tested in a pilot. The current bands are interpretive anchors, not norm-referenced percentiles, and the healthy benchmark is a placeholder until real population norms accumulate.
- It is correlational, not causal. The model draws on causal theory, but this instrument does not prove causation for your particular team.
- It is a starting point, not a verdict. The read is designed to open the right conversation, not to replace listening to your team.
None of this is a reason to distrust the read; it is the basis for trusting it, at exactly the level it earns. The citations below are the grounding for each design choice.
What each choice stands on
- Edmondson, A. C. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly, 44(2), 350–383.
- Google re:Work / Project Aristotle (2015). The five keys to a successful Google team; psychological safety as the top factor.
- Rizzo, J. R., House, R. J., & Lirtzman, S. I. (1970). Role conflict and ambiguity in complex organizations. Administrative Science Quarterly, 15(2), 150–163.
- Cannon-Bowers, J. A., Salas, E., & Converse, S. (1993). Shared mental models in expert team decision making. In Individual and Group Decision Making.
- Feldman, D. C. (1984). The development and enforcement of group norms. Academy of Management Review, 9(1), 47–53.
- Locke, E. A., & Latham, G. P. (2002). Building a practically useful theory of goal setting and task motivation. American Psychologist, 57(9), 705–717.
- Smith, P. C., & Kendall, L. M. (1963). Retranslation of expectations: behaviourally anchored rating scales. Journal of Applied Psychology, 47(2).
- Weick, K. E. (1995). Sensemaking in Organizations. Sage.
- Colquitt, J. A. (2001). On the dimensionality of organizational justice. Journal of Applied Psychology, 86(3), 386–400.
- Bandura, A. (1997). Self-Efficacy: The Exercise of Control. Freeman.
- Deci, E. L., & Ryan, R. M. (2000). Self-Determination Theory and the facilitation of intrinsic motivation. American Psychologist, 55(1).
- Pfeffer, J., & Sutton, R. I. (2000). The Knowing-Doing Gap. Harvard Business School Press.
- Baldwin, T. T., & Ford, J. K. (1988). Transfer of training: a review and directions for future research. Personnel Psychology, 41(1).
- Kluger, A. N., & DeNisi, A. (1996). The effects of feedback interventions on performance. Psychological Bulletin, 119(2), 254–284.
- Jehn, K. A. (1995). A multimethod examination of the benefits and detriments of intragroup conflict. Administrative Science Quarterly, 40(2), 256–282.
- Gallup. Employee recognition and its relationship to engagement (Q12 research programme).
- Amabile, T. M., & Kramer, S. J. (2011). The Progress Principle. Harvard Business Review Press.
- Atwater, L. E., & Yammarino, F. J. (1992). Does self-other agreement on leadership perceptions moderate the validity of leadership and performance predictions? Personnel Psychology, 45(1).
- Schneider, B., Salvaggio, A. N., & Subirats, M. (2002). Climate strength: a new direction for climate research. Journal of Applied Psychology, 87(2), 220–229.
- Schaufeli, W. B., Bakker, A. B., & Salanova, M. (2006). The measurement of work engagement with a short questionnaire (UWES). Educational and Psychological Measurement, 66(4).
- Kahn, W. A. (1990). Psychological conditions of personal engagement and disengagement at work. Academy of Management Journal, 33(4), 692–724.
- Paulhus, D. L. (1984). Two-component models of socially desirable responding. Journal of Personality and Social Psychology, 46(3); see also Edwards, A. L. (1957).