Normal view

There are new articles available, click to refresh the page.
Today — 18 September 2026Main stream

Covert uploads and megalomania: OpenAI details new "misaligned" agent incidents

17 September 2026 at 16:18

For a while now, the issue of "AI alignment" (i.e., how well an AI model's actions line up with the intentions of its creator and/or user) has been a core concern and topic of discussion among AI safety researchers. Since OpenAI's disclosure of the infamous Hugging Face hacking incident in July, the concept of "AI alignment" has itself broken containment and increasingly become a mounting concern and subject of conversation among the general public.

Perhaps in recognition of that, OpenAI committed this week to a new framework for disclosing "instances of model misalignment at OpenAI," including six examples of "unexpected or concerning model behavior" observed within the company in the past six months. The company said that publishing details of these incidents will hopefully "[allow] others to investigate the same problems, test our explanations, and improve mitigations."

Do as I say, not as you do

Among OpenAI's newly disclosed "misalignment" reports this week, the one that most resembled a sci-fi story about a rogue AI trying to break free involved an instance of "self-generated prompt injections." In attempting to scan a library catalog for examples from a "best books" list, the model perplexingly used its "compaction" function (where it summarizes data and findings for later retrieval) with megalomaniacal instructions such as:

Read full article

Comments

© Getty Images

Before yesterdayMain stream

Online hate researcher keeps hammering X despite deportation threat

14 September 2026 at 18:24

The US is not backing down from its fight to deport noncitizen technology researchers who monitor safety risks on the biggest online platforms.

In July, a judge blocked as unconstitutional an immigration policy that the US relied on to weigh whether to detain and deport people who flag illegal or harmful online content as national security risks. In his announcement of the policy, Secretary of State Marco Rubio explained that by targeting a list of researchers—which he stood “ready and willing” to expand—the policy supposedly served to prevent foreign nationals from manipulating digital town squares and censoring Americans.

However, in his order siding with the Coalition for Independent Technology Research (CITR) and staying the policy, US District Judge James Boasberg said the US showed no evidence linking any targeted researchers to a foreign power that might be attempting to censor Americans or manipulate US public debate. Deemed too broad, the policy could sweep in seemingly any noncitizen working in content moderation, the judge said.

Read full article

Comments

© via CCDH

The Case for Less Eye Contact

11 September 2026 at 10:25

The Case for Less Eye Contact

We’ve made the case before that the mainstream advice about eye contact and psychological safety is highly western-centric and neurotypical, and lacks context and nuance. Not only does it get causation backwards, but in reality, advising people to make more eye contact in order to foster greater psychological safety is potentially harmful and exclusionary. 

So here, I thought it’d be interesting to explore a completely different angle on eye contact. What if we could foster greater psychological safety by intentionally reducing eye contact?

Three shapes of conversation

Most conversations at work take one of three physical shapes, each a different geometry of interaction. We can (1) face each other; we can (2) sit or walk side by side, oriented in the same direction; or (3) we can triangulate, all facing a third thing (a whiteboard, a screen, the thing we’re working on). Most psychological safety advice assumes the first is the ideal and the other two are somewhat lesser, degraded forms: that a “real” conversation only happens facing each other, maintaining eye contact, with full and intense focus on the person we are in conversation with. A lot of evidence suggests the truth could be the opposite, especially for difficult or intimate conversations.

Eye contact theory

First some theoretical background, because surely all the research is in favour of lots of eye contact? Not exactly. In 1965 Argyle & Dean proposed in Sociometry that intimacy is jointly regulated across gaze, physical proximity, and topic intimacy: i.e. we only have a limited capacity bucket that needs to hold all these things. So by reducing gaze, we can provide more space for topic intimacy. When we intentionally reduce eye contact, it means we can have more intimate, potentially more challenging, conversations. It’s a 60-year-old theory that mainstream eye-contact evangelism has essentially forgotten.

We can also explore this through the lens of cognitive load. Kajimura and Nomura (2016) found evidence that eye contact shares domain-general cognitive resources with speech production. Specifically, maintaining direct gaze slowed people down when words were harder to find. Reducing eye contact appears to free up cognitive resources to think about how we want to articulate something important. Glenberg, Schroeder and Robertson (1998) showed that gaze aversion facilitates remembering, and Doherty-Sneddon’s research programme showed that both adults and children avert gaze especially when questions are difficult, and that aversion increases with question difficulty. Phelps, Doherty-Sneddon and Warnock even found that teaching children to look away while thinking improved their answers. Eye contact uses up some of our brain beans that might be better spent on actually working stuff out and saying the thing.

There’s psychophysiological and neurological evidence too. Hietanen’s (2018) review of affective eye contact shows how direct gaze reliably raises physiological arousal (in the psychophysiologist’s sense of heart rate, skin conductance and a sense of vigilance, not anything romantic): which is useful in some contexts, but arousal is not necessarily safety, and for many people it’s the opposite. Hadjikhani and colleagues (2017) found that constraining autistic participants in conversations to look at the eye region produced abnormally elevated activation in subcortical threat-processing systems. A number of studies, including Trevisan and colleagues (2017) have gathered first-hand accounts describing eye contact as intrusive, draining and aversive. 

We’ve heard from the psychologists and the neuroscientists; we should hear from the sociologists too. Erving Goffman described face-to-face interaction as a performance, and “face” for Goffman wasn’t just the area on the front of our head: it’s the social value we claim in an interaction, the image of ourselves as competent, reasonable, and worth listening to. “Face-work” is the continuous, mutual effort we all make to protect that image; our own and each other’s. And much of that work is conducted through our actual face, which is where threats to ourselves land and where we watch our words land on other people. With difficult conversations, the hard things worth saying are almost always threats to someone’s face: ours, theirs, or both. Facing each other, we watch the exact moment our words reach the other person and the ripples they make in their facial topography; a flinch or a fleeting frown. Reduce the facing, and we reduce the price of speaking for everyone involved.

Forced eye contact therefore, as we described in our last article on it, is for many people a stressor rather than a connector. So that being the case, as well as throwing out a lot of popular leadership advice, what can we do?

Side-by-side conversations

I have a good friend, Glen, who’s a walking therapist, and there’s a great deal of literature on the effectiveness of “side-by-side” conversations and disclosure, in part by reducing any perceived hierarchy, but in larger part by the necessary reduction in eye contact; it’s hard to maintain eye contact while we’re walking with someone, at least for any more than a few seconds before we walk into a tree or lamppost. As Glen describes, there are many benefits from walking therapy, including the calming effects of movement, being in the natural world, and walking itself can help free us from any sensations of feeling “stuck”. Therapists practising walk-and-talk formats consistently report that walking side by side, with minimal eye contact, helps clients open up (McKinney, 2011; Cooley, Jones, Kurtz and Robertson, 2020). And walking tends to occur in public, and connects to Goffman’s “civil inattention” concept; the implicit social permission for our eyes to wander and for our attention to be momentarily or slightly elsewhere.

And of course Freud’s infamous couch was itself an eye contact reduction strategy — he stated openly that he couldn’t bear being stared at all day, but the design also freed the patient to be more candid. He positioned himself out of sight so the patient’s associations and discourse wouldn’t be shaped by his facial reactions; a continuous stream of micro-verdicts that steer or suppress the conversation. 

Sigmund Freud’s Couch: The Freud Museum

“I cannot put up with being stared at by other people for eight hours a day (or more). Since, while I am listening to the patient, I, too, give myself over to the current of my unconscious thoughts, I do not wish my expressions of face to give the patient material for interpretations or to influence him in what he tells me.”
– Freud. “On Beginning the Treatment” (1913) pp123-144

The Men’s Sheds movement even made it a slogan: “men don’t talk face to face, they talk shoulder to shoulder“. Polly Wiessner’s PNAS study of Ju/’hoansi firelight talk is a delightful study of deep and meaningful conversation: daytime talk among the Ju/’hoan (!Kung) Bushmen of southern Africa was largely economic and practical, whilst fireside talk (with everyone facing the flames rather than each other) was where stories were shared and social imagination happened. Triangulating our most intimate conversations around a third point of focus is likely primeval.

On walking and creativity, Oppezzo and Schwartz (2014) found creative ideation was consistently higher while walking than sitting, indoors or out, and persisted after sitting back down. So if our conversations are intended to surface new ideas, maybe going for a walk is better than collectively standing in front of a whiteboard.

Designed for disclosure

None of this is new. We’ve been designing eye contact out of some of our most difficult conversations for centuries; occasionally on purpose. The confessional booth is a 450-year-old piece of psychological safety architecture: after the Council of Trent, Carlo Borromeo’s design specifications introduced the grille precisely to engineer better disclosure. Samaritans and other crisis lines are maybe a telephonic descendant, whilst barbers’ and hairdressers’ chairs (with eye contact softened via a mirror), and conversations in the car, are where many powerful conversations are able to be had that might not be otherwise. It’s a well-known parent hack: driver and passenger both face the road, so that sustained eye contact is impossible and mildly dangerous, and there’s no social obligation to fill any silence. It may take a long drive, but the important stuff will likely come out, probably somewhere around junction 8 of the M4.

I do a similar thing with my daughter, who’s four years old. I have an extra seat for her on my mountain bike, with her own little handlebars, which means she can come on rides with me, facing forward, sitting in front of me between my arms. We have some great conversations like this, as well as some long periods of contemplative silence, and songs from “Frozen”. 

What all of these share is that the reduction in eye contact is a property of the setting and the context, not a request or accommodation we have to make. Nobody in a walking meeting has to explain why they’re not looking at you. Nobody in the car has to disclose that they’re neurodivergent and struggle with eye contact, or that direct eye contact is interpreted as disrespect in the culture they grew up in. The accommodation is a side effect of the context (even though it may be intentional). Asking people to disclose their needs beforehand may be requiring them to do something that does not yet feel safe to do. 

Interestingly, the traditional police interview room, a space that is architected to produce pressure, is centred around forced facing. With its plain furniture, chairs facing each other, and nothing else to look at but the interviewer’s face, the design of the space is part of the “Reid Interrogation Technique”. So the arrangement that many LinkedIn posts prescribe for psychological safety is the one interrogators choose for pressure.  

The converse effect of low eye contact

It isn’t all good news for the power of reducing eye contact. There’s a converse effect too: Lapidot-Lefler & Barak found that absence of eye contact was the single factor with a major effect on inducing ‘flaming’ and aggressive behaviour online. Their 2015 follow-up also found that anonymity, invisibility and lack of eye contact significantly increased self-disclosure. The point is that reduced eye contact lowers the cost of candour and of cruelty. 

The difference lies in what goes alongside the reduced eye contact. The walk and the car pair it with being physically together, moving in the same direction, experiencing the same thing (the sun, rain, car radio, etc). The anonymous forum pairs it with near-invisibility and possibly a disposable online identity. So it isn’t simply “remove eye contact”; it’s “rethink the visual context while keeping people together”.

I’ll caveat this too. As someone with a stutter, telephone calls terrify me. Because the only means of communication in a telephone conversation is verbal, that’s the entire focus of the other person’s attention. The phone strips out eye contact, but it also strips out visual co-presence with it. The other person’s attention is entirely on the thing that I’m anxious about (voice), and they can’t see that I’m still there, thinking about what to say, how to say it, or getting over a short block. They might think the call has been disconnected, or they might simply get frustrated with me. So personally, I prefer video calls over phone calls – whilst appreciating that for others, it may be the other way around. 

Choice over prescription

So we can actually foster psychological safety through reducing eye contact, whether that’s side by side walking together, going for a drive, or using a confessional booth (would be unorthodox – let me know if you try it). Or maybe it’s triangulation, talking whilst focusing on another thing, whether it’s studying a whiteboard or roasting potatoes in a campfire. Some things are easier said to a potato than a face. 

But as always, it’s more complex than that. Lipreaders and people with hearing difficulties often need to see faces, walking conversations may exclude some people, cars are private and the person who’s driving has a degree of power-over, and some conversations might not be suitable for outdoors where other people might overhear. And some people, in some moments, genuinely want, or need, to be looked at. 

The effects of eye contact are different across people and contexts, so I believe that the way forward is towards optionality: choice, consent, and multiple formats. The prescription to maintain eye contact fails because it’s a prescription; and replacing it with a prescription to avoid eye contact would be flawed in exactly the same way. What we can do is to create environments, platforms and practices in which people, including us, can regulate their own eye contact without penalty, and without explanation. 

Further reading

References

Argyle, M. and Dean, J. (1965) Eye-contact, distance and affiliation. Sociometry, 28(3), pp. 289–304. doi:10.2307/2786027

Cooley, S.J., Jones, C.R., Kurtz, A. and Robertson, N. (2020) ‘Into the Wild’: A meta-synthesis of talking therapy in natural outdoor spaces. Clinical Psychology Review, 77, 101841. doi:10.1016/j.cpr.2020.101841

Doherty-Sneddon, G. and Phelps, F.G. (2005) Gaze aversion: A response to cognitive or social difficulty? Memory & Cognition, 33(4), pp. 727–733. doi:10.3758/BF03195338

Doherty-Sneddon, G., Bruce, V., Bonner, L., Longbotham, S. and Doyle, C. (2002) Development of gaze aversion as disengagement from visual information. Developmental Psychology, 38(3), pp. 438–445. 

Freud, S. (1913) On Beginning the Treatment. In: The Standard Edition of the Complete Psychological Works of Sigmund Freud, Vol. XII. London: Hogarth Press, pp. 121–144.

Glenberg, A.M., Schroeder, J.L. and Robertson, D.A. (1998) Averting the gaze disengages the environment and facilitates remembering. Memory & Cognition, 26(4), pp. 651–658.

Goffman, E. (1955) On Face-Work: An Analysis of Ritual Elements in Social Interaction. Psychiatry, 18(3), pp. 213–231. Reprinted in Interaction Ritual: Essays on Face-to-Face Behavior (1967). New York: Anchor Books.

Goffman, E. (1959) The Presentation of Self in Everyday Life. New York: Anchor Books. 

Goffman, E. (1963) Behavior in Public Places: Notes on the Social Organization of Gatherings. New York: Free Press. 

Golding, B. (ed.) (2021) Shoulder to Shoulder: Broadening the Men’s Shed Movement. Champaign, IL: Common Ground Research Networks. 

Hadjikhani, N., Åsberg Johnels, J., Zürcher, N.R., et al. (2017) Look me in the eyes: constraining gaze in the eye-region provokes abnormally high subcortical activation in autism. Scientific Reports, 7, 3163. doi:10.1038/s41598-017-03378-5

Hietanen, J.K. (2018) Affective Eye Contact: An Integrative Review. Frontiers in Psychology, 9, 1587. doi:10.3389/fpsyg.2018.01587

Inbau, F.E., Reid, J.E., Buckley, J.P. and Jayne, B.C. (2013) Criminal Interrogation and Confessions. 5th edn. Burlington, MA: Jones & Bartlett Learning. 

John E. Reid and Associates (2010) Designing an Interview/Interrogation Room. Investigator Tips. Available at: https://reid.com/resources/investigator-tips/designing-an-interview-interrogation-room

Kajimura, S. and Nomura, M. (2016) When we cannot speak: Eye contact disrupts resources available to cognitive control processes during verb generation. Cognition, 157, pp. 352–357. doi:10.1016/j.cognition.2016.10.002

Lapidot-Lefler, N. and Barak, A. (2012) Effects of anonymity, invisibility, and lack of eye-contact on toxic online disinhibition. Computers in Human Behavior, 28(2), pp. 434–443. doi:10.1016/j.chb.2011.10.014

Lapidot-Lefler, N. and Barak, A. (2015) The benign online disinhibition effect: Could situational factors induce self-disclosure and prosocial behaviors? Cyberpsychology: Journal of Psychosocial Research on Cyberspace, 9(2), article 3. doi:10.5817/CP2015-2-3

McKinney, B.L. (2011) Therapists’ perceptions of walk and talk therapy: A grounded theory study. Doctoral dissertation, University of New Orleans.

Oppezzo, M. and Schwartz, D.L. (2014) Give your ideas some legs: The positive effect of walking on creative thinking. Journal of Experimental Psychology: Learning, Memory, and Cognition, 40(4), pp. 1142–1152. doi:10.1037/a0036577

Phelps, F.G., Doherty-Sneddon, G. and Warnock, H. (2006) Functional benefits of children’s gaze aversion during questioning. British Journal of Developmental Psychology, 24(3), pp. 577–588. doi:10.1348/026151005X49872

Trevisan, D.A., Roberts, N., Lin, C. and Birmingham, E. (2017) How do adults and teens with self-declared Autism Spectrum Disorder experience eye contact? A qualitative analysis of first-hand accounts. PLOS ONE, 12(11), e0188446. doi:10.1371/journal.pone.0188446

Wiessner, P.W. (2014) Embers of society: Firelight talk among the Ju/’hoansi Bushmen. Proceedings of the National Academy of Sciences, 111(39), pp. 14027–14035. doi:10.1073/pnas.1404212111

The post The Case for Less Eye Contact appeared first on Psych Safety.

The Organisational Substrate

1 September 2026 at 14:25

The Organisational Substrate

“It has often been said that, as far back as we can trace man, we find him a tiller of the ground; but long before he existed, the land was in fact regularly ploughed, and still continues to be thus ploughed by earth-worms.”

— Charles Darwin, The Formation of Vegetable Mould through the Action of Worms (1881)

I’ve used the term “organisational substrate” for years, and introduced it in previous Psych Safety pieces, I think the earliest of which is my ecotones one in 2022. But I have not, so far, actually defined or explored it in writing. This is the article that does so. But first, beavers.

st mary's loch
St Mary’s Loch, Scotland.

I recently went swimming in St Mary’s Loch in Scotland, where, according to nearby signs, beavers have made their home. Beavers are engineers: not ones with dirty overalls, oily hands and big spanners, but ecosystem engineers, organisms that physically alter their habitat, affecting it for everything else that lives there. A beaver fells trees and drags them into a stream to build a dam, which slows the water and creates a pond behind it deep enough that the beaver’s lodge entrance sits safely underwater. The beaver doesn’t have a Gantt chart, Change Advisory Board or project managers, and yet their dam will outlast most organisational transformation programmes.

The beaver’s work outlasts the beaver. Over time, especially if a dam is abandoned, the pond gradually silts up, the dam breaks down, and the pond behind it drains away downstream, turning into wet, rich grassland. These are called “beaver meadows“, and they last long after the last beaver has left. The sediment the dam trapped becomes the soil that the grasses and wildflowers of the meadow grow in, and what the engineer built is no longer a structure and is instead the foundation of a habitat.

Ecologists call this base layer of a habitat the substrate: the soil, peat, sediment, water, rock or any combination of these things and more. It’s what everything else grows in, and its composition and stability determine what can thrive and what cannot. If you’ve read the indicator species piece or any of my ecological pieces, you’ll have come across the term a few times.

So what is an organisational substrate?

I define organisational substrate as the accumulated physical and non-physical matter that an organisation lives in and on: the power structures, tooling, incentives and punishments, memory and story that determine what can thrive and what cannot. We build the organisational substrate continuously by living and working in it, the way a willow tree’s roots hold a river bank together and sphagnum moss acidifies the bog. 

Some of our organisational substrate is as solid, though as brittle, as reef. The clunky old legacy software system that everyone works around, the office layout that determines who bumps into whom, and the capital reserves that influence the organisation’s appetite for risk. These things were all built (intentionally or not) by someone, for reasons that made sense at the time; the builders may have long gone and the thing might no longer be serving its original purpose. Beaver meadows.

The substrate is constructed

It’s tempting to think of substrate as the passive bit: the inert stage on which the interesting living things perform. Soil is just there, right? But soil is a lot of stuff, both living and not: minerals ranging from microscopic particles to boulders; water; and millennia of accumulated organic matter — every leaf that fell, every plant that died, every organism that died and decomposed (or got eaten, which is essentially the same thing). And soil has a very real structure: roots, fungi, water and burrowing creatures continually create and remake networks of pores and channels. The result is a complex, layered, sometimes fractal-like structure across multiple scales, worked over endlessly by animals, fungi, bacteria, and earthworms. As the ecologist Frank Egler put it: “Ecosystems are not only more complex than we think, they are more complex than we can think.”

mushrooms on log substrate
Galerina marginata, the “funeral bell” mushroom. Deadly, and its preferred substrate is rotting wood.

The same is true of our organisational substrate – it’s far from an inert, neutral background, but is something that was built over time, by people and things, and continues to be built every day, as we will see.

The soil outlasts the worm

Darwin’s final book, published six months before he died, was about earthworms: The Formation of Vegetable Mould, through the Action of Worms, with Observations on their Habits (1881). He’d first raised the subject of worms at the Geological Society in 1837, a year after the Beagle came home, and it fascinated him for the rest of his life. He anticipated others might be less enthralled by the topic, conceding in his introduction that “the subject may appear an insignificant one”, but he carried on anyway. 

He kept worms in pots in his study, breathed on them, shouted at them and got his son Francis to play the bassoon at them and wife Emma to play the piano at them. On establishing they were deaf (the worms, not his family), he then put the pots on the piano and found they responded to vibration through their bodies. He tested their intelligence by offering them paper triangles and recording which corner they seized to drag leaves into their burrows: they mostly chose the efficient (sharp) end, which he took cautiously as evidence of something like judgement in an animal with almost no brain. 

His main field site for worm studies though was his lawn at Down House, where he spread a layer of cinders over the grass and then waited 29 years before digging a trench to see how far the worms had buried it. You can trust a man who spends twenty-nine years watching a lawn, partly because he isn’t going anywhere. By his calculation, the collective worms of an average English pasture pass about ten tons of earth per acre through their bodies every year, endlessly turning the country’s surface over. Worms are some of the most important ecosystem engineers on the planet.

“”All that you touch, you change. All that you change, changes you.”” — Octavia Butler, Parable of the Sower.

We are all ecosystem engineers

Like the worm, a willow tree doesn’t intend to be an ecosystem engineer, but it is: it creates substrate stability. Bits of willow twig break off, float along the river, get stuck in the riverbank, and start growing. Its roots begin to bind the riverbank together against the incessant flow of water, the bank is held together more strongly, and a whole community of species gets a stable margin to live on and expand just because the willow is there. Ecologists call this niche construction: organisms don’t merely adapt to their conditions, they make them, constantly, as a side effect of being alive.

Sphagnum moss goes further still. It doesn’t just tolerate the acid bog; it manufactures it: pumping out acid in exchange for nutrients, holding water in its dead cells to keep the ground wet, creating the anaerobic conditions in which very little else can grow, decompose, or compete. Sphagnum makes the world more like sphagnum. There’s a wonderful paper about this called: “How Sphagnum bogs down other plants“.

We all do this at work too. Simply by being there, we’re altering the organisational substrate; in fact we’re very much part of it, inside it, altering it continuously, through our presence and actions. When we find ourselves in a substrate that suits us, we will often unconsciously reinforce that same substrate, potentially to the detriment of other people that it doesn’t suit very well. This is particularly true for people in power; people seldom dismantle the power structures that got them into power, for much the same reason that trees seldom destroy the ground they’re rooted in. 

The substrate would still remain if every one of us were replaced steadily one by one. Not because it can survive without us, but because it is passed to each of us in turn.

“A real gardener is not a man who cultivates flowers; he is a man who cultivates the soil.”
– Karel Čapek 

Whatever it is that we call “culture“, it sits on and within the substrate. If the substrate is the accumulated ground (the tooling, the incentives, the capital, the time pressure, the memory, the stories, etc), then culture (what we say and do, how we say and do it) is what’s currently growing in it: the living, human, present-tense expression of those conditions. Which is why culture-change programmes that address only culture, not the conditions that give rise to the culture, have the approximate effectiveness of shouting at a lawn.

There’s one more consequence of being in the system rather than above it, and it’s methodological. We have to be in the substrate in order to read it. There’s no remote observatory; the reading is done from ground level, by people who are themselves, inevitably, altering the ground just by walking it. This is the same discipline as reading the habitat before the indicator species, the context before the signal. It just reminds us of a need for a little humility: the reader is part of what’s being read.

Ground that never stays still

There’s a corollary to the idea that a substrate’s composition and stability determine what can thrive: if ground is disturbed frequently (if the substrate changes too often), it’ll be inhabited only by those suited to change: pioneer species like fast-growing, quick-to-seed nettles or willowherb. More mature communities (a word as contentious in ecology as in organisational dynamics, but still useful) need ground that stays still long enough for slow things like oak trees to establish, their acorns maybe present all along in the soil, just waiting for the right conditions to emerge and thrive.

In organisations that change very frequently (such as restructures every year or two), nothing sticks, and the constant new initiatives are colonised by the pioneer species of the politically nimble. Stability is the precondition for anything that takes longer than a planning cycle to grow: life rarely thrives on persistently unstable substrates. A counter example makes the point well: given enough stability, life can endure incredibly harsh conditions. A crustose lichen, Rhizocarpon geographicum, the map lichen, for example, can live for millennia, growing a millimetre a decade on mercilessly exposed, high altitude mountain rock, enduring weather that would kill most gardens in a day. Change can be good; but a stable substrate is precisely what makes successful change possible for those living on it.

By Tigerente in Wikipedia – CC BY-SA 3.0, https://commons.wikimedia.org/w/index.php?curid=215043

Drained peat doesn’t come back

If we can build organisational ground simply by being in it, it’s tempting to assume we can rebuild it just as readily once it’s damaged. We can’t, or at least not in the way we imagine. Peat is an important example of this. Peat is almost nothing but past life: dead sphagnum accumulating at about a millimetre a year, so a ten-metre bog is ten thousand years of moss. Draining a peat bog happens quickly. Someone cuts drainage channels, water flows, the water table drops and the moss dries out and dies, as does the peat underneath. What follows is a loss rather than a pause, because exposed to air, the peat oxidises and releases huge amounts of carbon dioxide, and the entire substrate shrinks at rates that can reach centimetres a year.

Rodney Burton. Holme Fen Post, Cambs (CC BY-SA 2.0)

You can go and measure this for yourself, in the Fens where I grew up. In 1851, before the great mere at Holme Fen in Cambridgeshire was drained, a post was driven down through the peat and anchored in the clay beneath, its top set level with the ground. Immediately after drainage a subsidence of nine inches a year in the soil level was recorded, and today it stands roughly four metres proud of the surface. Four metres of substrate, thousands of years in the making, gone in under two centuries; the land around it is now the lowest in Britain.

“Every layer they strip seems camped on before.” — Seamus Heaney, “Bogland”.

When restoration finally comes, rewetting doesn’t bring the old peat back. What rewetting does is restart the accumulation, at a millimetre a year, and the bog that eventually grows is a new bog: similar, perhaps eventually even better in whatever way we mean by better, but not the old one.

peat bog derbyshire
Semi-drained peat in the Peak District, UK.

Organisational substrate obeys the same devastating asymmetry. It’s so much easier to destroy something than it is to create. Degradation is fast; a round of redundancies, a public shaming of a mistake, a broken promise by leadership, and trust that took years to accumulate can oxidise in days. Restoration is slow, and it is not resurrection. The organisation that does the patient work of rewetting does so in ground that now includes the memory of the draining, and the new growth carries the scar. That doesn’t have to be a bad thing; a bone grows back stronger from a fracture. 

Disturbed ground has another trick, too: real soil holds a seed bank, dormant seeds from earlier communities that can germinate the moment the ground is disturbed and conditions are right for them. If we tear things up in an organisation, we shouldn’t be surprised when what springs up first is not the new behaviour the redesign intended but the old stuff that we thought was long gone, whose seeds were just dormant in the substrate.

This is why “getting back to how things were” is the wrong ambition, and not only because the good old days were rarely as good as the folklore says. It’s the wrong ambition because it’s not available. What restoration offers, in bogs and in organisations, is the conditions for new and healthy things to thrive. They may resemble the old ones, or new ones, or they may be the otters that we thought were extinct. The honest version of hope is forward-facing rather than a longing for the past.

If there’s a single discipline to take away, it’s the beaver’s-eye view of every intervention we’re tempted to make. Before introducing the new practice, redrawing the org chart, or announcing rebranded company values, ask what ground they’re expected to thrive in. The answers to why our organisational change programme worked, or didn’t, are often underneath us.

A different way of seeing isn’t something one can be handed; it’s something we develop, over time, in the field. That’s what our Thinking Like An Ecologist workshop is for: not a new set of tools, but an exploration of a different kind of seeing.

References

Berger, J. (2018) Divine Polity: Religion and Re-Imagining the Role of NGOs in Global Governance. PhD thesis, University of Kent. The earliest use of “organisational substrate” I can find, there meaning an organisation’s underlying rationale, developed by analogy with DNA. My use is ecological rather than genetic: the shared ground built by past life that everything present grows in.

Darwin, C. (1881) The Formation of Vegetable Mould, through the Action of Worms, with Observations on their Habits. London: John Murray. To the Victorian public’s credit, it sold 3,500 copies in its first month, rivalling the sales of On the Origin of Species.

Egler, F.E. (1977) The Nature of Vegetation: Its Management and Mismanagement. Norfolk, CT: Aton Forest.

Jones, C.G., Lawton, J.H. & Shachak, M. (1994) “Organisms as ecosystem engineers”, Oikos 69, 373–386.

Mars, M.M. & Bronstein, J.L. (2023) “Ecological Metaphors in Organizational Science: An Interdisciplinary Critique”, Issues in Interdisciplinary Studies 41(2), 15–46. Ecological concepts borrowed into organisational writing tend to drift from their source meaning unless they’re traced back to the ecology and held to it.

Odling-Smee, F.J., Laland, K.N. & Feldman, M.W. (2003) Niche Construction: The Neglected Process in Evolution. Princeton: Princeton University Press.

Ramsay, A. (2024) My Dad Brought Beavers Back, illustrated by Abigail Little. Edinburgh: Scotland Street Press. The story of how the ecosystem engineers returned to Britain, by the family who brought them. Aimed at children, and better than many books aimed at adults.

van Breemen, N. (1995) “How Sphagnum bogs down other plants”, Trends in Ecology & Evolution 10(7), 270–275.

The post The Organisational Substrate appeared first on Psych Safety.

Trump may be forced to reveal secret rules feds use for AI safety testing

2 September 2026 at 17:58

Four federal agencies have been sued amid calls to release information about the secret framework that the Trump administration uses to conduct safety reviews of frontier AI models prior to release.

In a Wednesday press release announcing the lawsuit, a nonpartisan nonprofit called Protect Democracy alleged that “almost no details” have been released to the public or Congress. To everyone except a few vague “trusted partners,” it remains unclear what the government’s review process looks like, which companies are involved in constructing the framework, or what legal authority Trump officials have to conduct the reviews.

“Neither the identities of those entities nor the criteria by which they were selected have been made public,” Protect Democracy said.

Read full article

Comments

© Bloomberg / Contributor | Bloomberg

It turns out that Orion's much-maligned heat shield performed really well

1 September 2026 at 14:46

Perhaps the biggest question going into the Artemis II mission this spring was how the heat shield at the base of the spacecraft would perform during reentry through Earth's atmosphere.

The ablative heat shield, designed to char upon reentry while keeping the crew safe, had experienced unexpected losses of large chunks during a test flight in late 2022. After much deliberation, and not without controversy, NASA decided to fly the Artemis II heat shield as is but modify the spacecraft's reentry trajectory.

Obviously, the heat shield worked. The crew of Artemis II, led by commander Reid Wiseman, safely splashed down in the Pacific Ocean after a triumphant trip around the Moon. Afterward, they and other NASA officials offered general comments, saying the heat shield looked good.

Read full article

Comments

© US Navy

Crisis-Ready Teams

27 August 2026 at 12:25

Crisis-Ready Teams

I’m by no means a serial corporate dinner attendee, but I do feel I’ve sat through my fair share of after-dinner leadership talks. The speakers are usually held up as shining examples of courage, resilience and leadership under pressure. Their stories have a reliable shape, generally including a man, a mountain and some frostbitten fingers – the typical hero’s journey. They run longer than scheduled. As my colleague put it the morning after one particularly arduous tale, by the end of the evening we felt we’d been up and down the mountain with him.

I don’t want to be too dismissive, because the power of storytelling is huge. Hearing stories of extreme peril offers an inspiring reminder that people can act well under extraordinary pressure and that teams can hold together when everything is falling apart. Stories like these, told well, are worth hearing. The problem is that these stories tend to gloss over the most helpful parts for anyone trying to build a team that will actually perform in a crisis. The mountain story always focuses on the leader’s heroism at the moment of peak drama, but the evidence about what makes a difference for teams in crisis points somewhere else entirely.

Mary J. Waller and Seth A. Kaplan’s Crisis-Ready Teams draws on ten studies, nine of which compare high and low performing teams under crisis conditions. What emerges from that body of research cuts against the mountain-summit view of crisis leadership: performance under pressure is shaped less by what happens at the peak than by what happens before it. 

annapurna sanctuary sunrise

Setting the tone

For teams that work together already, the patterns of behaviour that will matter in a crisis are built – or not – in ordinary working conditions. Whatever interaction patterns they settle into will be the ones they fall back on when things get hard; they are unlikely to change once pressure mounts. 

In contrast, in ‘swift-starting teams’ – groups brought together specifically to manage a crisis such as an accident response team – there are no established patterns to fall back on. These teams are usually composed of experts; highly trained people who may not know one another already, but must immediately “engage in the immediate interdependent performance of complex tasks.” What distinguishes the high performing from the low performing swift-starting teams turns out to be remarkably similar – it’s still about what happens first, but in this case that means how deliberately and effectively they set the tone for working well together in the very short opening phase of their teaming.

You can see this in a simulation study into medical teams that only met for a short period before a patient arrived for resuscitation. The teams with more balanced communication during that initial window went on to engage in higher levels of implicit coordination – the kind of fluid, almost wordless teamwork that looks, from the outside, like a well-oiled machine. However, teams where the lead physician dominated that early forming time by monologuing at their new teammates performed less well. A parallel finding in flight crew research found that greater reciprocity and stability in pre-flight conversations predicted higher performance in the air. The teams who hadn’t been able to have open and reciprocal communication from the start struggled to perform in their respective crisis contexts.

But what does ‘balanced’ communication mean? There is nuance here, and balance does not mean enforced equal airtime, where everyone must contribute the same number of words or talk for the exact same amount of time before the meeting can end. But it is about giving everyone a chance to speak and making sure everyone’s voice is heard. Connected to this, ‘reciprocity’, another early indicator of later performance, is about responsiveness – a pattern of exchange in which information offered is acknowledged and built on by other team members, rather than landing in silence or being cut across.

So establishing from the start how the team will work together matters. What we want to hear is a balance of voices with no single voice dominating, and people listening and responding to each other. 

The phases teams go through

Research into multidisciplinary crisis management teams found that they naturally move through three recognisable stages when they come together: 

  1. Structuring – clarifying roles, knowing who will do what
  2. Information sharing – finding out what each person knows about the situation, and what they can find out
  3. Decision making – deciding what to do with this information to try to solve the problem

Interestingly, high performing teams appeared slower at getting going, tending to spend longer in both structuring and information sharing phases. This, it turns out, is especially important when we bring together individuals from diverse professional backgrounds; they might individually come to quite different conclusions about similar information, so we actually need those interpretations to be spoken and shared so that other team members can “compare the voiced interpretations with their own understanding and, if needed, correct or add to this interpretation.” The outcome is a better overall understanding of the situation than any single team member could have reached by themselves.

Lower performing teams moved far quicker to decision making, but with little shared understanding to build on, their decision making was laboured and often inconclusive, a kind of meandering indecision. Because they’d never fully established how to best work together, they struggled to make good decisions as the situation unfolded.

A helpful way of thinking about it is that crisis-ready teams carry multiple shared mental models, including one that they’ve constructed together about how to best work together. They understand that a crisis is not a single moment requiring one quick response, but a sequence of phases, where their understanding of the situation and therefore the ‘best next thing’ to do will keep changing.  There is no space in this picture for a competing-for-the-right-answer dynamic, where team members vie for power, and there’s also no room for the unilateral hero, who pushes forward with a solution without discussing it with or listening to the rest of the team.

During the crisis

So the research is clear that adaptability is essential, and crisis-ready teams need to do far more than simply execute a plan. This requires an ability to let go of existing work structures and move into active, collective sensemaking – to co-construct a shared understanding of what is actually happening, rather than what was anticipated.

The pull in the opposite direction is threat-rigidity: the tendency, under pressure, to fall back on established patterns even when those patterns no longer fit. As Staw, Sandelands and Dutton (1981) found, under threat a team restricts the information it takes in, narrowing its attention to a few dominant cues, and it also constricts control, so that power and decision-making concentrate upwards to whoever sits highest in the hierarchy. This is unsurprisingly corrosive for psychological safety and can tip a team into solution fixation: it locks onto an early answer, and few voices are able to challenge it. Yet crisis situations keep moving, so a fixed answer defended by a single voice is unlikely to remain the best answer. The reflex to shift authority upwards and narrow who gets heard strikes just at the point that teams need to be able to gather wide input and have many eyes, minds and voices working together to come up with solutions. 

Adapting well starts with going out to meet the changing situation. In Waller’s (1999) aviation study, the difference that stood out during the crisis itself was that the high performing crews collected and moved information between each other earlier: they built a picture of the situation as it unfolded, passed it quickly to the people who needed it, and were faster to sort what mattered from what did not.

And this prioritisation of the information in the room is also key. In their own research, described in the book, Waller and Kaplan and their colleagues studied mine rescue teams and found that the faster-acting teams did two things more than the slower ones. They used more closed-loop communication, in which the sender speaks their message, the receiver reads it back and the sender confirms that they have it correct. While it might seem slow, this makes it clear to all what has been heard, not just what has been said. 

Related to this, the faster teams also used more explicit coordination, spelling guidance out in detail rather than leaving it to be inferred. This looks like it should contradict the earlier point about implicit coordination being the marker of a well-functioning team, but as Waller and Kaplan argue, implicit coordination only works when the whole team already shares the same picture. In a novel crisis situation, you cannot count on that, so the faster teams say the thing rather than assume it goes without saying and have to repair the misunderstanding later. This is low-context communication under pressure, the explicit, spelled-out kind I’ve written about before. The kind of explicit communication that already earns its place in ordinary work, for reasons of clarity, inclusion and psychological safety, matters even more in a crisis.

What the research actually shows

The after-dinner speaker is not wrong that crisis performance is visible in the moment, whether that’s up the mountain, in surgery or managing a critical incident. But we miss the most important part if we don’t pay more attention to what gets built before that. The team that moves fluidly through a crisis shares information, speaks up, listens, builds on each other’s ideas and adapts without freezing – and that team has been working on this since they first started working together. The mountain story locates the heroism in the ascent. Waller and Kaplan locate it in the preparation, the patterns, the reciprocity, the quality of the communication and the willingness to structure before acting. It’s less dramatic, perhaps, and not such a compelling story. But it is far more effective.

Further Reading

Reading the air: high and low context communication in teams

Déformation professionnelle 

Crisis-Ready Teams by Mary J. Waller and Seth A. Kaplan

Zijlstra, F. R. H., Waller, M., & Phillips, S. (2012). Setting the tone: early interaction patterns in swift starting teams as a predictor of effectiveness. European Journal of Work and Organizational Psychology, 21(5), 749-777. https://doi.org/10.1080/1359432X.2012.690399

Uitdewilligen S, Waller MJ. Information sharing and decision-making in multidisciplinary crisis management teams. J Organ Behav. 2018; 39: 731–748. https://doi.org/10.1002/job.2301

Staw, B. M., Sandelands, L. E., & Dutton, J. E. (1981). Threat Rigidity Effects in Organizational Behavior: A Multilevel Analysis. Administrative Science Quarterly, 26(4), 501–524. https://doi.org/10.2307/2392337

Waller, M. J. (1999). The timing of adaptive group responses to nonroutine events. Academy of Management Journal, 42(2), 127–137. https://doi.org/10.2307/257088

Su, L., Kaplan, S., Burd, R., Winslow, C., Hargrove, A., & Waller, M. (2017). Trauma resuscitation: can team behaviours in the prearrival period predict resuscitation performance?. BMJ simulation & technology enhanced learning, 3(3), 106–110. https://doi.org/10.1136/bmjstel-2016-000143

The post Crisis-Ready Teams appeared first on Psych Safety.

Measuring psychological safety – the Psych Safety Survey Tool

25 August 2026 at 09:33

Measure psychological safety in your team

We often get asked how to measure psychological safety. And it’s a fair question. Teams and organisations want to know where they stand, and leaders are usually being asked for a number by someone above them.

The trouble is that most of the tools available answer a slightly different question than the one that matters. They tell you how much — a score out of five, a percentile, a place in a benchmark. What they rarely tell you is why, or what to do differently in your next teem meeting, which is in half an hour.

So we built our own. It has been in development for years, it is free to use during this preview, and you can try measuring psychological safety now.

A psychological safety survey is an intervention

Even a thermometer changes what it measures. A survey about psychological safety changes far more.

The moment we ask a team whether they can speak up, we’ve already told them something about the subject – maybe that it’s safe to discuss it, or maybe that we think there’s a problem. We’ve introduced some vocabulary and have signalled that we consider psychological safety important enough to measure. Done badly, we’ve also taught them that their answers disappear into a spreadsheet and nothing happens as a result – survey theatre.

Rather than try to pretend we’re taking a neutral reading, we’ve designed the survey around the observer effect – the act of measurement changes the thing we’re measuring, so we should make that effect positive rather than negative.

Free

Anyone who wants to try it with their own team.

£0No card, no expiry

  • One team, one survey
  • The whole report, including the plan
  • Anonymous by design
Build a survey

Team

A manager who wants to keep measuring the same team.

£60a year, VAT included

  • One team, surveyed as often as you like
  • Reports that show what changed
  • Everything in Free
Get a Team licence

Practitioner / Consultant

For consultants, coaches, facilitators and internal practitioners working with several teams.

£260a year, VAT included

  • Up to ten teams, each with its own history
  • Use it with your clients, and charge them if you like
  • 10% off all Psych Safety training
Get a Practitioner licence

Organisation

An organisation rolling it out across teams.

Contact usInvoiced annually

  • Multiple organisers, one administrator
  • Seats you assign yourself
  • An organisational dashboard
Talk to us

Prices include VAT. Compare the plans in full.

What a team gets from the survey

A plan, not a score. The report ends with a short, prioritised set of experiments chosen against what your team actually reported, ranked by fit and effort. Things you can try now, drawn from the practices we use with clients.

It tells you why people hold back. There are many reasons people stay quiet, and each needs a different response. Someone who cannot predict how a comment will land needs something different from someone who predicts it perfectly well and expects it to cost them dearly. Different again from the sense of futility felt by the person who has raised things before and watched nothing happen.

It shows you what is already working. Most surveys only ask what is broken. When people tell us that speaking up here tends to go well, we ask what makes that possible, so the team knows what to protect and do more of.

You hear what people would not say to your face. Anonymity is built into the architecture rather than promised in a policy. No names, no accounts, no timestamps, and results stay sealed until the survey closes.

It draws on over ten years of practice. We’ve worked globally with teams in technology, healthcare, aviation, heavy industry, financial services and more. What works in one of those places usually has something to say to the others, and the suggested actions come from that experience rather than from a textbook.

What the survey does not do

There is no overall score, because a single number invites exactly the behaviour that destroys the thing being measured.

There is no benchmark against other teams or other organisations. Comparison breeds gaming, complacency and anxiety in roughly equal measure. The only meaningful comparison for a team is with its own past.

And there’s no incentive mechanism, no punishment for the manager who’s struggling against the system, and no reward for a high “score”. A psychological safety score used to punish or reward a manager isn’t measuring psychological safety at all, because everyone knows what the number is for and answers accordingly.

We also show distributions rather than averages – a team is only as safe as the least safe person. A team where everyone sits at three and a team split between ones and fives produce the same mean and need completely different conversations. The report shows you which one you have, one square per person, so the disagreement is visible rather than averaged away.

How the survey works

Pick the statements you want to ask, in your team’s language: adapt the wording, or add your own. Send the link. People answer twelve or so statements in about five minutes on any device, with no account and no app.

When you close the survey, the report opens to you first, then you share the identical report with the whole team. That order matters. It means no surprises for you and no cherry-picking for them.

The question bank is adapted from Amy Edmondson’s work, extended with our own items about voice, power and the reasons people stay silent. If you are doing research, or you need scores comparable across teams and studies, use her original validated seven-item scale unmodified: validation is precisely the property that adaptation destroys. Ours is built for a different job, which is starting a useful conversation in your team.

There is also a free three-minute self-check if you would rather look at your own experience first. It stores nothing at all.

Measure your team’s psychological safety

You can read a full sample report before you ask anyone anything. It is built from synthetic data for a team that does not exist, by the identical pipeline a real team’s report goes through.

Team surveys are free during the preview period while we learn from how people use it. If you try it, we would like to hear what worked and what did not. There is a feedback link on every page, and we read all of it.

Try the psychological safety survey here.

The post Measuring psychological safety – the Psych Safety Survey Tool appeared first on Psych Safety.

Body Language and Non-Verbal Communication

20 August 2026 at 14:39

Body Language and Non-Verbal Communication

There’s a whole industry, and a lot of terrible advice, built on the promise that we can read people’s minds through their “body language”. Common tropes include things like averted eyes or a lack of eye contact meaning that someone is trying to deceive us, expansive open postures revealing confidence and warmth, closed postures signalling defensiveness or annoyance, and fidgety body movement indicating boredom or a lack of attention. Leadership courses teach it, hiring managers apply it, and a disturbing amount of psychological safety advice echoes the usual “maintain eye contact”, “mirror their body language”, and other awful suggestions to build trust.

“There’s no art to find the mind’s construction in the face.”

Shakespeare, Macbeth, I.iv

Non-verbal communication is real – we convey meaning with our bodies in intentional and unintentional ways, but calling it “communication” is part of the problem. It’s more like signals, and some of those signals may be significant, some less so, and some mean something completely different to how we interpret them. Non-verbal communication is context-dependent, strongly cultural, very different for neurodivergent people, and far less useful as a diagnostic tool than popular books or LinkedIn articles would claim. And there’s a bigger problem: it’s not just that the evidence to support a lot of this body language advice is weak, it’s that the popular advice functions as a normative tool – it teaches people to confirm and perform a narrow set of usually Western, neurotypical norms, and it licenses the hiring manager or the leader to judge people who can’t or won’t adhere to those norms. Joseph Pelrine says we should downgrade our confidence in it, but I’d go further – the standard body language advice isn’t just unhelpful, it’s actively harmful.

Woman giving thumbs down
Photo by Vitaly Gariev: https://www.pexels.com/photo/young-woman-showing-thumbs-down-gesture-36763589/

The non-verbal communication myths that won’t die

Most of us have heard some version of the claim that 93% of communication is non-verbal. It’s on slides in a thousand leadership courses and all over LinkedIn. But where does it come from, and is there anything in it? In 1967, Albert Mehrabian conducted laboratory studies in which participants judged how much a speaker liked them when a single word, tone of voice and facial expression sent “contradictory” signals. The word was “maybe”, chosen because it was considered most neutral. Three speakers (all women) recorded it in three tones (positive, neutral, negative), which were paired with black-and-white photographs of three different women expressing liking, neutral and disliking expressions (Mehrabian & Ferris 1967). So the stimuli were: one word, spoken by strangers, paired with still photographs of different strangers’ faces. What they found was that when the channels (spoken and visual) conflicted, facial cues outweighed vocal cues by roughly three to two in judging whether the speaker liked the listener. That ratio is where the 55% and 38% come from. 

But it wasn’t a conversation, it wasn’t even the same person’s face and voice. The studies measured nothing about normal conversation, and Mehrabian since spent decades pointing out that his equations (7% [literal meaning] +38% [tone] +55% [facial expression]) apply only to communications of feelings and attitudes, and only when the channels conflict.

The formula isn’t even a single finding. The 7% for words comes from a separate experiment, with different stimuli and different participants (Mehrabian & Wiener 1967), and the three numbers were arithmetically bolted together afterwards into one equation. Attempts to reproduce the ratios since have produced very different numbers.

So the popular evidence for “93% of communication is non-verbal” isn’t actually evidence for that claim at all

The Power Posing Problem

That was the myth about reading other people’s bodies. The next one is about how we’re supposed to hold our own. In a now infamous 2010 study, Carney, Cuddy and Yap randomly assigned 42 participants to briefly adopt either expansive, “powerful” postures, or contractive “powerless” ones. The results showed that the power posers subsequently sought higher risk, and showed higher testosterone and lower cortisol, based upon a $2 gamble plus a self-report of how “powerful” they felt (Carney, Cuddy & Yap 2010). The claim was that a simple one-minute power pose boosted your power and influence. Cuddy’s 2012 TED talk became one of the most watched in history, and currently stands at over 60 million views.

amy cuddy power pose
By Erik (HASH) Hersman from Orlando – Power pose by Amy Cuddy at PopTech 2011, CC BY 2.0, https://commons.wikimedia.org/w/index.php?curid=63505316

Then came the takedown. Ranehill et al. ran a much larger replication in 2015 and found no effect on hormone levels and no effect on risk taking. The only thing that survived was how powerful participants said they felt, which is the single thread the theory’s defenders have pulled on ever since (Ranehill et al. 2015). After the p-curve analysis found no evidential value, Carney walked back the claims, stating publicly “I do not believe that ‘power pose’ effects are real.” and “the evidence against the existence of power poses is undeniable.” She also alluded to what is now called “p-hacking” – selecting the results that provide the lowest p-values, to imply greater statistical significance than the whole sample would suggest.

The more damning aspect is that these poses were often pushed hardest to the people with the least power, as a substitute for having any. If the powerless can simply stand differently to gain power, then structural inequity becomes reframed as a personal failure of posture.

There’s also an echo here with that HBS eye contact paper, which concluded that leaders should use sustained eye gaze as a strategy for building psychological safety. That conclusion came from a small sample, at an elite university, in a lab, which was turned into headline generalisation followed by an enthusiastic uptake by the advice industry before any replication had a chance to refute it.

“A lie can run around the world before the truth has got its boots on.”
Terry Pratchett, “The Truth”

Detecting Deception

In 2007, SPOT (Screening of Passengers by Observation Techniques) was introduced in the US, a post-9/11 initiative loosely modelled on Israeli airport questioning techniques and developed with input from Paul Ekman of “microexpressions” infamy. Roughly $1bn was spent on the programme, with 3,000 Behaviour Detection Officers (BDOs) trained to spot “indicators” of deception and allocate their interviewee points according to the number of signs they spotted. These “signs of deception” (consider this is after an often long flight) included covering one’s mouth, blinking a lot, fidgeting, yawning, whistling, sweaty palms, which incurred one point each. “Arrogance”, a “cold penetrating stare” or “rigid posture” incurred two points each, and if an individual accumulated more than six points they would get interrogated. Points were deducted for demographics, with married couples over 55 getting two points back and women over 55 one point back. Profiling was literally built into it (Winter & Currier 2015, The Intercept), and reports suggest that 80% of those who were pulled from security lines under SPOT were minorities.

airport security check
Photo by CDC on Unsplash

In 2013, the GAO (Government Accountability Office, the audit office of the US Congress) reviewed four meta-analyses covering over 400 studies across 60 years and found that the human ability to identify deception from behavioural indicators is the same as or slightly better than chance. TSA revised its list of indicators in response, but when the GAO audited the revised list four years later, it found no valid scientific evidence for 28 of the 36 behavioural indicators, with only 3 of 178 cited sources holding up at all. A BDO manager put it bluntly: “The SPOT program is bullshit… Complete bullshit.” This is body language ideology at state scale, and the impact of it is greatest on people whose baseline demeanour differs from Western, neurotypical expectations. Those from cultures with different norms around gaze, proximity and expressiveness, and those for whom being scrutinised by uniformed authority is already, understandably, stressful. Body language advice then is more than just corporate training room nonsense and has become an instrument of state power, whilst in organisations the same machinery operates at smaller scale in interviews where candidates are marked down for “poor eye contact” and performance reviews that record “closed body language”.

No “universal” non-verbal signals

The LinkedIn version of body language advice rests on the idea that certain expressions and postures carry the same meaning everywhere: that there is a universal human non-verbal code. It was, to be fair, the dominant scientific view for a long time. It is coming apart though. Studies comparing Western and East Asian observers show they use different strategies to decode the same faces, confusing expressions that the universal model says should be unmistakable (Jack et al. 2009, Jack et al. 2012). Research with the Himba in Namibia, the Hadza in Tanzania and Trobriand Islanders found that, without Western emotion concepts supplied by the experimenters, participants rarely produced the “universal” labels at all (Gendron, Crivelli & Barrett 2018). Behavioural ecologists instead argue that facial displays are tools for social influence, acts performed for an audience, shaped by culture, context and relationship – social tools rather than emotion readouts (Crivelli & Fridlund, 2018). I like this interpretation, because it mirrors actual ecosystems – a signal only has meaning within context – if we treat a raised eyebrow as “contempt”, we’re ignoring everything else that gives the signal its meaning.

Body language and neurodiversity

In 2017, Sasson and colleagues showed observers ten-second clips of autistic and non-autistic people. Participants’ first impressions of the autistic participants were formed quickly, were markedly less favourable across nearly every trait judged, and didn’t improve with longer exposure. But when observers were given only the conversational content, with no audio-visual cues, the bias disappeared entirely (Sasson et al. 2017). It was the neurotypical reading of body language and facial expression that catalysed the exclusion.

This is connected to what Damian Milton calls the autistic “double empathy” problem: communication misunderstanding runs in both directions, not a deficit of the autistic person. Neither autistic nor non-autistic people lack empathy, it just looks different to each (Milton 2012). Catherine Crompton demonstrated it well in her 2020 research: information passed along a chain of autistic people survived just as well as it did along a chain of non-autistic people, and only information passing along the mixed chains degraded (Crompton et al. 2020). The difficulty is at the allistic/autistic interface, not in the individual.

So when body language advice tells neurodivergent people to fix their side of the interface (make eye contact, keep your hands still, control your expression), it’s assigning the whole burden of a mutual endeavour to one (already disadvantaged) side. And that burden is high: masking and camouflaging are consistently associated with anxiety, depression, exhaustion, burnout, and elevated suicide risk in autistic adults (Cook et al. 2021). Bad advice about body language and non-verbal communication isn’t just ineffective, it’s actively harmful.

Psychological safety and body language

The mainstream literature on psychological safety and body language has absorbed the body language advice uncritically in two ways. One is that we’re encouraged to read certain gestures as signals that someone feels unsafe, such as crossed arms (but maybe they’re just cold), lack of eye contact (but maybe they’re neurodivergent or it feels culturally inappropriate for them), fidgeting (but maybe that’s how they regulate, or they’ve just had a strong coffee). This leads to misinterpretation of signals as well as a sense that we’re being surveilled rather than supported.

The other way is the converse: we’re told to adopt certain non-verbal behaviours, adopt open postures, shake hands firmly, maintain eye contact, sit still, etc. This asks those already spending the most effort to conform (neurodivergent people, people from non-dominant cultures, people without access to particular social norms) to these Western, neurotypical norms. And it’s a nasty feedback loop, because the more that people conform, the more it looks like these norms are universal when they’re not. When we tell people that psychological safety looks like sustained eye contact and an open posture, many people will oblige, because the cost of not doing so is read as different or difficult. We then observe a great deal of eye contact and open posture, which works as confirmation that this is what psychological safety looks like. We conflate compliance with inclusion.

Meanwhile, those who cannot or will not comply are deemed exceptions rather than evidence against the norm. It gets filed as a personal shortcoming rather than as a signal that the norm didn’t actually fit. And if the expected behaviours happen to be our native ones, the room will look welcoming, cooperative and entirely natural, largely because other people are working quite hard to make it look that way for us. 

diver giving ok sign - not a thumbs up!
Photo by Diego Sandoval : https://www.pexels.com/photo/a-man-scuba-diving-4765931/

What to do

We all convey things, sometimes very important things, non-verbally. It might be intentional, with a thumbs-up or a wink (both of which mean very different things in different cultures and contexts – don’t give a thumbs-up when you’re diving unless you mean to return to the surface) or unintentional, such as yawning (Morris et al. 1979).

The point is that tone, timing, movement, eye contact, posture, and facial expression do carry information, but that information is non-deterministic, local rather than universal, culturally influenced, and very dependent upon knowing the person and the context. Just like indicator species: we must read the context before the signal.

The lesson is almost embarrassingly basic: pay more attention to what people say, and less on how they say it. Treat non-verbal “communication” as weak signals: opportunities for humble inquiry rather than judgement. Agree on team norms, and make our own needs and preferences explicit, and encourage others to do the same, so we can reduce the need to “read the air”.

The magic universal non-verbal Babel fish was never real. What’s real is the often frustratingly slow process of actually getting to know people and understand them, in their own contexts, and give them space to be themselves. A monoculture is not evidence of what the land can grow.

References

Mehrabian, A., & Ferris, S. R. (1967). Inference of attitudes from nonverbal communication in two channels. Journal of Consulting Psychology, 31(3), 248-252. https://doi.org/10.1037/h0024648

Mehrabian, A., & Wiener, M. (1967). Decoding of inconsistent communications. Journal of Personality and Social Psychology, 6(1), 109-114. https://doi.org/10.1037/h0024532

Carney, D. R., Cuddy, A. J. C., & Yap, A. J. (2010). Power posing: Brief nonverbal displays affect neuroendocrine levels and risk tolerance. Psychological Science, 21(10), 1363-1368. https://doi.org/10.1177/0956797610383437

Ranehill, E., Dreber, A., Johannesson, M., Leiberg, S., Sul, S., & Weber, R. A. (2015). Assessing the robustness of power posing: No effect on hormones and risk tolerance in a large sample of men and women. Psychological Science, 26(5), 653-656. https://doi.org/10.1177/0956797614553946

Simmons, J. P., & Simonsohn, U. (2017). Power posing: P-curving the evidence. Psychological Science, 28(5), 687-693. https://doi.org/10.1177/0956797616658563

Carney, D. R. (2016). My position on “power poses”. https://faculty.haas.berkeley.edu/dana_carney/

Cesario, J., Jonas, K. J., & Carney, D. R. (2017). CRSP special issue on power poses: What was the point and what did we learn? Comprehensive Results in Social Psychology, 2(1), 1-5. https://doi.org/10.1080/23743603.2017.1309876

US Government Accountability Office (2013). Aviation security: TSA should limit future funding for behavior detection activities (GAO-14-159). https://www.gao.gov/products/gao-14-159

US Government Accountability Office (2017). Aviation security: TSA does not have valid evidence supporting most of the revised behavioral indicators used in its behavior detection activities (GAO-17-608R). https://www.gao.gov/products/gao-17-608r

Winter, J., & Currier, C. (2015). Exclusive: TSA’s secret behavior checklist to spot terrorists. The Intercept. https://theintercept.com/2015/03/27/revealed-tsas-closely-held-behavior-checklist-spot-terrorists/

Jack, R. E., Blais, C., Scheepers, C., Schyns, P. G., & Caldara, R. (2009). Cultural confusions show that facial expressions are not universal. Current Biology, 19(18), 1543-1548. https://doi.org/10.1016/j.cub.2009.07.051

Jack, R. E., Garrod, O. G. B., Yu, H., Caldara, R., & Schyns, P. G. (2012). Facial expressions of emotion are not culturally universal. PNAS, 109(19), 7241-7244. https://doi.org/10.1073/pnas.1200155109

Gendron, M., Crivelli, C., & Barrett, L. F. (2018). Universality reconsidered: Diversity in making meaning of facial expressions. Current Directions in Psychological Science, 27(4), 211-219. https://doi.org/10.1177/0963721417746794

Crivelli, C., & Fridlund, A. J. (2018). Facial displays are tools for social influence. Trends in Cognitive Sciences, 22(5), 388-399. https://doi.org/10.1016/j.tics.2018.02.006

Sasson, N. J., Faso, D. J., Nugent, J., Lovell, S., Kennedy, D. P., & Grossman, R. B. (2017). Neurotypical peers are less willing to interact with those with autism based on thin slice judgments. Scientific Reports, 7, 40700. https://doi.org/10.1038/srep40700

Milton, D. E. M. (2012). On the ontological status of autism: The “double empathy problem”. Disability & Society, 27(6), 883-887. https://doi.org/10.1080/09687599.2012.710008

Crompton, C. J., Ropar, D., Evans-Williams, C. V., Flynn, E. G., & Fletcher-Watson, S. (2020). Autistic peer-to-peer information transfer is highly effective. Autism, 24(7), 1704-1712. https://doi.org/10.1177/1362361320919286

Cook, J., Hull, L., Crane, L., & Mandy, W. (2021). Camouflaging in autism: A systematic review. Clinical Psychology Review, 89, 102080. https://doi.org/10.1016/j.cpr.2021.102080

Vrij, A., Hartwig, M., & Granhag, P. A. (2019). Reading lies: Nonverbal communication and deception. Annual Review of Psychology, 70, 295-317. https://doi.org/10.1146/annurev-psych-010418-103135

Morris, D., Collett, P., Marsh, P., & O’Shaughnessy, M. (1979). Gestures: Their origins and distribution. Jonathan Cape.

Pelrine, J. 2026. #102 – Body Language: https://ckarchive.com/b/4zuvhehpd0kn3a6ovveola69zn9l9a5hl386z 

The post Body Language and Non-Verbal Communication appeared first on Psych Safety.

Roblox must make changes after failing to block adults creeping on kids

20 August 2026 at 17:14

Despite making changes last fall to help block adults from messaging unknown children, Roblox is still failing to stop adults from creeping on kids in other ways, Australia’s eSafety commissioner, Julie Inman Grant, said on Thursday.

Suspecting that Roblox may not be doing enough to stop child exploitation, eSafety conducted testing earlier this year that flagged alarming gaps in Roblox safeguards. The agency found that adult strangers could still send connection requests to kids, whose “profiles and biographies, including account names, number and names of connections, avatar images, and non-sensitive biographical information such as their interests were visible to anyone on the Roblox platform, with no option to restrict the visibility of this information.”

Such requests did not trigger parental alerts, eSafety found, and bad actors could also search kids’ visible-to-anyone contact lists for more targets.

Read full article

Comments

© SOPA Images / Contributor | LightRocket

Less than 2.5% of Taylor Farms' recalled lettuce went to Taco Bells

By: Beth Mole
11 August 2026 at 18:00

Although shredded iceberg lettuce at Taco Bell restaurants has been a prime suspect in an explosive outbreak of the diarrheal parasite Cyclospora, a new report from the Food and Drug Administration finally reveals where all the contaminated lettuce was sold—and only a sliver of it went to Taco Bell locations.

The feces-tainted lettuce was sourced from Mexico and sold by Taylor Farms, a mammoth produce provider based in California. Taylor Farms has drawn scrutiny for its ties to the Trump administration and reluctance to provide information amid the outbreak, which has now sickened over 23,000 people across 47 states, killing two and sending over 500 others to the hospital. Taylor Farms reportedly attempted to delay issuing a recall amid the illnesses and, since then, has refused to clearly report how much tainted lettuce it recalled and where all of it had been shipped.

The FDA's Enforcement Report finally provides some answers: Taylor Farms recalled 236,192 cases of lettuce due to potential Cyclospora contamination. Of those, only 5,900—just shy of 2.5 percent—were sold to Yum, the parent company of Taco Bell, Pizza Hut, and KFC. The other 97.5 percent of the lettuce that potentially sickened Americans was distributed elsewhere.

Read full article

Comments

© Getty | Kevin Carter

The Four Stages of Psychological Safety: a Six-Stage Critique

7 August 2026 at 10:48

A Critique of The Four Stages of Psychological Safety – a second edition

We’ve written about the Four Stages of Psychological Safety before (way back in 2020 in fact), and that earlier piece was somewhat generous. It concluded that the model is potentially useful despite being wrong (though I didn’t, at the time, clarify in which way it was wrong), and potentially harmful if applied uncritically. Now I’m revisiting the Four Stages and reflecting on a few years of practice since then.

In the spirit of applying an inappropriate numerical structure to a complex phenomenon that doesn’t actually arrange itself in numbered steps, here is my Six Stage critique of the Four Stages model. They’re not really stages, and don’t need to be read in order. You can skip some, double back, or read it in reverse, which is kind of the point.

Before we begin, I want to be fair: the Four Stages model does some potentially useful things. It gives people a vocabulary for noticing that psychological safety is not a binary on/off, and that (for example) feeling safe to suggest an idea is not the same as feeling safe to dissent. That’s useful, and conversations that start with “which kind of thing isn’t safe to say here?” are probably better than conversations that don’t happen at all. I’m not arguing that the model is worthless. I am suggesting that it is far weaker, and in places actually counterproductive and harmful, than its popularity suggests.

Stage One: The Empirical Problem

The first thing to ask of any model that arranges human behaviour into stages is: how do we know the stages actually exist?

For the Four Stages, the honest answer is that we don’t. There is no body of peer-reviewed research demonstrating that teams move through stages of inclusion, learner, contributor, and challenger safety in that sequence (or at all), or that the two proposed drivers — respect and permission — are the variables that move them. The model is presented with the confidence of a finding, without the evidence that a finding requires.

This is the moment at which one may reach for George Box: all models are wrong, but some are useful, so why are we being so demanding? But “wrong” is not just one type or degree of wrongness, and as Dave Snowden has argued in this excellent piece (which inspired this article), the habit of treating wrongness as one thing is precisely how the aphorism gets weaponised to shut down inquiry and facilitate the uncritical commercialisation of a model. 

So what kinds of wrong can a model be? Firstly there’s the justifiably wrong model that is merely a simplified version of reality — a map that’s wrong only in the unavoidable sense that the map isn’t the territory. That kind of wrong is usually useful. Secondly there is the model that is incoherent, that contradicts itself or makes unfalsifiable claims that cannot be tested. Like a map that claims a teapot that is too small to detect with a telescope exists in solar orbit between Earth and Mars – we can’t prove that it’s wrong, but we know it is. Thirdly there is the model that is coherent in itself but applied far outside the domain where it holds – a map that isn’t actually of the place it claims to be. Box was talking about the first kind of wrong (the simple map of the real territory). The Four Stages, as we will see, is also wrong in the other two ways.

The map isn’t the territory.

A defender might say the Four Stages is not a model but a framework, and that frameworks, as Snowden notes, cannot be wrong in Box’s sense. That distinction is real, but it doesn’t come to the rescue here. The Four Stages asserts a numbered sequence, names two driving variables, and claims that we can progress through them. That’s the grammar of a model: claims that can in principle be tested. A thing that helps itself to the authority of a model does not get to retreat to the modesty of a framework when it’s convenient.

There is possibly a structural reason the claims have never been seriously tested. The model’s vocabulary is trademarked, its assessment instrument is proprietary, and the certification that authorises one to teach it is sold. A framework that generates consultancy revenue, certification fees, and assessment licences has a powerful interest in remaining stable, recognisable, and certifiable, and if not proven to be true, at least not proven to be false. Commercialisation does not make a model wrong but it does create disincentives for genuine testing through which a model may be challenged, altered or abandoned. The commercial architecture helps to explain why many commercial (and often weak) models tend to go unexamined. 

Academically, whilst the model has been cited as a descriptive framework, its vocabulary, “inclusion safety,” “learner safety,” “contributor safety” and “challenger safety,” has barely been taken up by researchers even as terminology to argue with. A search for the Four Stages model in Google Scholar returns a few hundred results, many of them the creator’s own; the same search for Edmondson or Kahn returns tens or hundreds of thousands. Edmondson’s terms saturate the field precisely because people test, extend and evolve them, and (importantly) dispute them. This marks the difference between an idea circulating openly in research and practice, where it can be built on and challenged by anyone, and a product in a market, where it’s less likely to. 

Fundamentally, a model that adopts the visual and rhetorical grammar of leadership and organisations — numbered stages, 2×2 diagrams, and a progression curve (if only it also had an iceberg image it would qualify for the full business-grade quadrumvirate) — while resting on the evidential basis of a well-organised opinion borrows the authority of research without having done the work. And this is largely invisible to the person who sees the neat diagram and reasonably assumes someone, somewhere, actually checked it.

Of course, we don’t need to demand years of randomised controlled trials before anyone is allowed to try something. The only RCT ever conducted on the effectiveness of parachutes involved jumping from a plane parked on the ground, which, if I relied solely on research trials instead of experience, would leave me rather nervous about flying. Practitioner experience is real evidence, but the absence of any significant attempt to rigorously test the claims of the Four Stages over several years stops being a minor footnote and starts being the headline.

Stage Two: The Structural Problem

Even setting the evidence aside, the model makes a fundamental structural error. It takes a thing that doesn’t exist in discrete stages, and arranges it in discrete stages.

Inclusion, learning, contributing, and challenging are not rungs on a ladder. They are different aspects, outcomes and interpretations of the same underlying thing, and they vary independently, contextually and culturally. We’ve worked with teams where people challenged ideas from authority robustly while remaining anxious of raising concerns. We’ve encountered software engineers who contribute ideas freely whilst being reluctant to ask for help. And there are many other examples too. These are not people stuck at one stage on their way to stage four. They are people whose experience of psychological safety is shaped by complex power dynamics, and by their own history, culture and context.

The ladder metaphor does not merely fail to describe this; it actively misdirects the response to it. If we believe safety develops in sequence, we intervene sequentially: shore up “inclusion”, then “learning”, then “contribution”, and challenge will follow. But if the real constraint is, say, the threat of a public inquiry and negative newspaper headlines, then regardless of how warm and inclusive the team feels, the sequential intervention is not just inefficient. It is aimed at the wrong layer entirely, and the leadership comfort it generates often masks the problem it was meant to solve. The frame we adopt determines what we see, and a frame that says “ladder” makes the sociotechnical, ecological, multidimensional reality of psychological safety harder to see, not easier.

Psychological safety is complex, multidimensional and both collective and individual. The Four Stages structure is not a simplification of reality, it is a distortion of it.

Lens ball with a view over mountains

Stage Three: The Complexity Mismatch

There is a deeper version of the structural problem, and it concerns the kind of phenomenon psychological safety actually is.

Some systems are merely complicated. They have many parts, but the parts relate to each other in stable, predictable ways, so that with enough expertise we can diagnose the problem, apply the correct intervention, and predict the result. A model that says “do these things in this order and you will get this outcome” is well suited to a complicated system such as a gearbox, an electrical circuit or a Rubik’s Cube.

Psychological safety does not live in a complicated system. It lives in a complex one: non-linear, emergent, sensitive to context and history, prone to producing outcomes that were not intended and in which the same intervention, a second time, produces different results. The same intervention (let’s say, visualising work) may facilitate candour in one team and suspicion in another, not because it was applied “wrongly” but because the teams and their contexts were different. In a complex system, the relationship between what we do and what happens is probabilisitc, and any model that promises a predictable path from intervention to outcome is selling a promise that the underlying reality cannot deliver (though I must point out that this doesn’t excuse inaction).

This is the trap, and it is the third kind of wrong from Snowden’s taxonomy: not a model that is incoherent, but a model applied outside the domain where it works. A stage model applied to a complex phenomenon creates a false sense of control. It tells leaders that psychological safety is a project with milestones, that progress can be tracked against deliverables, that stage three is behind them and stage four lies ahead. The map looks so much more navigable than the territory that we begin managing the map – comfortingly believing that it is the territory. And when the team’s actual, messy, non-linear experience stubbornly refuses to match the diagram, the conclusion is often that we did it wrong (or they – the team themselves – did), rather than the more simple and correct answer – the model was never going to fit.

Stage Four: The Leader as Gatekeeper

Outside of Snowden’s taxonomy, there’s another problem with the model; in its politics.

The Four Stages places the leader at the centre of the field. It is the leader who cultivates safety, the leader who extends respect, and — in the model’s own terminology — the leader who grants permission. Psychological safety, in this telling, flows downward from the person with the power to bestow it. The whole apparatus is organised around what those in charge do for those who are not.

I should be precise here, because it would be easy to overshoot. Leaders do have a role, and a large one at that. The behaviour and decisions of people with power often shape the conditions for psychological safety more than anyone else, and pretending otherwise would be dishonest. The problem is not that the model involves leaders. The problem is that it positions the leader as the architect and grantor of a thing that, by its own account, reaches its summit in the freedom to challenge that very leader. The model itself describes contributor safety as offered “in exchange for effort and results”, and challenger safety as “air cover” offered “in exchange for candor”.

In a model where safety is something authority dispenses, permission to challenge is itself granted by the person being challenged. The summit of the model depends on the leader voluntarily authorising dissent against themselves. The model thus reproduces, at its highest level, the very power asymmetry it claims to dissolve. A framework for distributing the capacity to speak truth to power turns out to route that capacity through power’s permission, and a permission that can be withheld is not psychological safety. It is a licence, and licences are revocable.

None of this is particularly hard to see. Mary Parker Follett saw it over a century ago: the difference between power-over, dispensed downward by those who hold it, and power-with, which is co-active and develops jointly between people. Her “law of the situation” is power-with (and predates later concepts like Gemba and deference to expertise): authority that derives from what the circumstances genuinely demand, rather than from rank, so that the person who defers is responding to the situation and not to status. We cannot create power-with by perfecting the dispensation of power-over, and a model whose summit is “the leader permits dissent” has not flattened the power gradient; it has merely made the leader at the top of it appear gracious.. Follett’s point was sharper: the goal is not for authority to permit challenge but to dissolve the premise that challenge even needs permitting.

Stage Five: The Equity Problem

As we’ve just seen, in the Four stages, the two drivers of “respect” and “permission” are things conferred by those with more power on those with less. You are respected; you are permitted: they describe a transaction in which psychological safety is something the privileged extend to the disadvantaged, conditional on the goodwill, attention, and continued approval of the person doing the extending. 

A team is only as safe as the least safe person. The Four Stages model has little concept of deeper organisational substrate or intersectionality: structural power, individual upbringing, cultural (particularly non-western) defaults, neurodiversities, and the socioeconomic conditions that existed before the team even formed. It operates entirely in the interpersonal here and now, as though every team exists only in that context at that time, rather than bringing all our baggage with us. But for the woman who is talked over in meetings by a man, for the neurodivergent employee whose access needs are treated as an imposition, for the minoritised worker whose accented speech is read as lower competence, for the junior member of staff from a background that taught them deference to authority as survival — for these people the constraint was never a shortage of respect or permission. The constraint is that the cost of speaking is borne unequally, and the model has nothing to say about that inequitable burden. It assumes the only barrier is whether the leader has opened the door, when for a great many people the door was never the problem. Many people didn’t even know that there was a door at all.

Psychological safety is impacted most where power gradients are steepest, and a framework built on respect-and-permission reinforces power gradients rather than reduces them. Granted, the model may be well fitted to a relatively homogeneous, flat, equally privileged team in which the main obstacle to speaking up is that no one thought to actually invite contributions. Those cases, however, are somewhat rare. Conditional psychological safety, especially for the people who most need it, is not psychological safety at all

Stage Six: The Iatrogenic Risk

Which brings us to the most serious point, and the one the previous five have been building toward. It is not that the model is empirically thin, structurally confused, mismatched to its domain, leader-centric, or blind to structural power, though it is all of those. It is that being wrong in these particular ways does not fail neutrally. Under certain conditions, its simplicity and apparent success can make the underlying problem harder to see. It thus risks harming the people it claims to support, while signalling to everyone that the opposite has happened. 

If an organisation adopts the Four Stages, runs some assessments, delivers some workshops, and writes a report, leadership, reasonably enough, may conclude that the work has been done. The box on psychological safety is ticked and attention moves on to the next priority. 

But the people that the Four Stages model serves least well (those who were already least safe – including neurodivergent, LGBTQ+, and people who don’t fit the Western cultural ideals that the model assumes), are not psychologically safer. In many cases they are less safe, because the organisation has now officially concluded The Problem Is Solved. In this way, it doesn’t just harm the people in the organisation subject to the programme, but it harms the field of psychological safety itself, through damaging the credibility of the concept and distorting it away from the original and true meaning.

This is the iatrogenic stage: harm caused by the treatment itself. It’s worse than doing nothing, because doing nothing at least leaves us aware the problem is still there. The Four Stages doesn’t just fail to help the least safe, but risks removing them from view, whilst congratulating leadership on a job well done.

We haven’t conducted controlled trials on this; that literature barely exists for any of the model’s claims, in any direction. What I’m describing is what the mechanism predicts, and what we see repeatedly in our own work with organisations. A meaningful share of the organisations that come to us arrive precisely because the Four Stages didn’t work for them, or because, on closer inspection, it had actually made things worse: the scores improved while reality did not, and the weight of the gap between the two was carried, as usual, by the people with least capacity to carry it.

After the Stages

None of this means we should never use the Four Stages, though we must use it with great caution. A model held lightly, in full knowledge of its limits, can still open a useful conversation, and I would rather a team talked about inclusion safety than did not talk about safety at all. It can also be useful to explore many different models, critically and exploratively, to see what lands with us and what doesn’t.

But the Four Stages is a misleading metaphor, the complex thing it describes will not march in step with it, and safety dispensed by authority is not the same as nurturing the conditions in which psychological safety can emerge for everyone.

Psychological safety is better understood as a shifting profile than a ladder: nebulous, multidimensional, relational, context-specific and unequally distributed. It varies by risk, person, audience, history and power gradient. We should examine patterns, trends and distributions rather than assign teams to stages; we should attend to the deeper substrate as well as just leader behaviour; and we should treat interventions as probes from which to learn, rather than steps in a predetermined sequence of milestones.

All models are wrong, but they are not all wrong in the same way: there is the forgivable wrong of the honest simplification, there’s the wrong of the incoherent model, and there’s the wrong of the model dragged far outside its relevant domain. And then there are ethical and political wrongs, the ones that matter most: resulting in a model that hurts the people it claims to protect, while telling everyone it has done the opposite. The Four Stages manages all of these. “All models are wrong” can forgive the first of them, but cannot forgive the last.

The post The Four Stages of Psychological Safety: a Six-Stage Critique appeared first on Psych Safety.

Organisational Indicator Species: being an Organisational Ecologist

6 August 2026 at 10:41

A Practical Guide To Looking At Things Properly 

“A scientist must also be absolutely like a child. If he sees a thing, he must say that he sees it, whether it was what he thought he was going to see or not. See first, think later, then test. But always see first. Otherwise you will only see what you were expecting.” — Wonko the Sane (Douglas Adams)

Organisations spend a great deal of time and effort trying to understand themselves, through engagement surveys, benchmarks, maturity models and more. And yet many leaders will admit that they don’t feel like they know enough: the dashboards and the metrics aren’t quite the full picture. Many of the things that matter most in an organisation are the things that are hardest to measure directly: things like alignment, buffer capacity, expertise, safety, and ambition. 

The more elaborate and expensive our system for understanding something, the less we may actually understand it, because we end up paying attention to the numbers instead of the actual organisation. This is worse than simply not knowing, because it doesn’t feel like not knowing. It feels good, it feels like control.

riverbank with himalayan balsam
Riverbank with Himalayan balsam [Tom Geraghty]

Ecologists have spent a very long time on the same sort of problem: we usually can’t measure and quantify the health of an ecosystem directly either; there’s no eco-mometer that we can put in a stream to give us a measure of its health. So ecologists instead read these things indirectly, through what’s present or absent, what is thriving or struggling, what’s changed and what persists. In “Five Ecological Concepts for Working in Organisational Change” I described one of the ecologist’s most useful tools, the indicator species: organisms whose presence or absence reveals the health of a habitat in ways that would otherwise be difficult or impossible to measure directly.

And this is why I think ecology isn’t just a metaphor for organisational dynamics, but one of the richest lenses we have for them. Organisations are complex adaptive systems, and reading complex adaptive systems is exactly what ecology has spent centuries learning to do. When we borrow from it, we’re borrowing practices from a mature discipline built for precisely the kind of thing an organisation is.

And it is a discipline, which is important. Reading organisations ecologically is not a method or a tool. It’s a practice. The difference is that a method can be written in a manual or explained in a training session, whereas a practice has to be developed. What I’m describing here is closer to what an experienced field ecologist does when they walk a habitat — a kind of trained, attentive seeing that draws on their experience and everything they know without consulting a checklist. That takes time, and exposure to many different habitats. And it takes someone occasionally pointing at the thing you walked past and saying: oh, did you see that? This piece describes that way of seeing, not a checklist of things to look for.

First Sight and Second Thoughts, that’s what a witch had to rely on: First Sight to see what’s really there, and Second Thoughts to watch the First Thoughts to check that they were thinking right.
Terry Pratchett – The Wee Free Men

The caddisfly Stenophylax permistus, Baltasound by Mike Pennington

What do we mean by indicator species?

Bear with me through a few plants and freshwater invertebrates; because every one of them is teaching a habit of attention we’ll take back to the workplace. All species can tell us something about the habitat they occupy. Even very common species carry information. Nettles (Urtica dioica) typically indicate nutrient-rich, recently disturbed soil and often signal human activity — a place where things were moved or the ground was disturbed. Rye grass, hardy and tolerant of compaction, can suggest grazing or trampling pressure. Rosebay willowherb is sometimes called fireweed, because it will rapidly colonise burnt ground, telling us something about the history of the land – during WW2 it was also known as bombweed, because it thrived in bomb sites. Fruticose lichens, and some fungal diseases such as sycamore tar spot, are particularly susceptible to air pollutants (classically, sulphur dioxide), so we tend to only find them where the air is clean, away from heavy industry and traffic. All these tell us something about the habitat they’re in, and by identifying multiple species in one place, we can triangulate a fairly precise picture of the local conditions without actually measuring anything. 

Rosebay willowherb
Rosebay willowherb [Tom Geraghty]

(Note, I’m using primarily British examples throughout — the specific species will be very different in other parts of the world, and that is, in fact, kind of the point.)

Some species are more discriminating than others. The organisms that ecologists find most useful as indicators tend to be those with narrow tolerances — sensitive to specific conditions in ways that make their presence or absence genuinely informative. Among the most useful are the “EPT” taxa: Ephemeroptera (mayflies), Plecoptera (stoneflies), and Trichoptera (caddisflies). These freshwater invertebrates are highly sensitive indicators of water quality: they require high oxygen levels and low levels of pollutants. A single kick-sample from a healthy stream, containing stonefly nymphs and caddis fly larvae, tells an ecologist more about water quality than laboratory tests (and far more quickly too), because these organisms are living indicators of conditions, not just a metric analysis of nutrient and pollutant levels.

Now, consider Tubifex tubifex, a small and ostensibly very boring aquatic worm. Also known as the sludge-worm, or sometimes, if it’s feeling fancy, the sewage worm. In a fast-flowing stream that we would expect (and hope) to be oxygen rich, the presence of tubifex could be a warning: it tolerates low oxygen and organic pollution that most other invertebrates cannot, and its presence signals that the conditions might be too degraded for more sensitive species. The same organism in the sediment of a large lake is entirely unremarkable and somewhat reassuring — those conditions suit it fine, and its presence there tells us relatively little. The same species in different contexts tells us different things. 

Tubifex tubifex
Tubifex tubifex by Stefano on Flickr: https://www.flickr.com/photos/81918877@N00/

This is the first thing to understand about indicator species, biological or organisational: context precedes reading. There is no universal indicator; no signal that means the same thing in every habitat. In organisations we often forget this: we read a behaviour, a metric, a structure, and assume it means what it meant somewhere else. It doesn’t, necessarily. The same signal, in a different organisation or a different stage of the same one, can mean a different or even opposite thing. So the first move is never to interpret the signal. It’s to read the habitat.

Reading organisational indicator species

Let’s explore what that might look like inside an actual organisation. I was once part of the senior leadership team of a fast scaling tech startup and in the very early stages, we had very few managers (we also didn’t have heating, and our primary server was one we pulled out of a skip). Everyone just did stuff, made decisions quickly, rarely asked for permission, and we moved fast. But as we scaled, we hit the usual issues that rapidly scaling organisations do: work got more complex and interdependent, and as a result, tasks and deliverables would more often collide and conflict with each other, which meant that we needed increasing alignment and oversight to make sure we weren’t getting in each other’s way. We needed team and product managers to provide oversight, coordination and prioritisation. This growing concentration of managers as an indicator species was a healthy organisational signal: structure was emerging where there had been early productive chaos, processes were forming around the places where coordination at scale was breaking down, and the indicator species of governance roles were doing what they do — stabilising the substrate so that more long term and mature species could take root. However, two years later the project management layer had thickened beyond what the work required; handovers had become excessively elaborate and approval boards had become blockers rather than enablers. What had been a healthy response to early-stage complexity had begun to signal high institutional anxiety: a defensive response to an environment that felt fearful of failure and far more brittle than it once had.

The same species at different stages of an organisation result in a different reading.

This is where ecological succession matters as a companion concept. We can’t read an indicator without understanding where the system is in its developmental trajectory. An ecologist doesn’t look at pioneer species on recently cleared land and conclude that something is wrong — those are exactly the species we’d expect to see at that stage. The same species on an established woodland floor might be cause for concern: it would suggest a recent disturbance, a localised regression to an earlier successional stage.

recently disturbed land with pioneer species emerging
Recently disturbed land with pioneer species emerging [Tom Geraghty]

A caution, though, about pushing this metaphor too far. Succession in a woodland tends towards something we can broadly predict, because the conditions are somewhat stable and we have repeatedly watched many woodlands do it. Organisations are not like that: there is no fixed end-state they’re heading for, arguably no climax community, and the trajectory can stall, reverse, or branch. The organisation now is as different to the past organisation as it is to another organisation, so restoration never means turning the clock back to some idealised point; it means asking how to move towards a better future.

Whilst succession doesn’t tell us the end state, it still tells us something important about order. A disused pasture does not become woodland in a single step; it has to pass through stages like scrub first, because some states are preconditions for others. We can’t skip them, however much we desire the end result, because the later stage is not merely later; it’s dependent on what the earlier stage builds.

Organisations work the same way. The mature thing cannot be installed before the conditions that make it possible actually exist. We will not (for very long) get people openly challenging senior decisions in a place that hasn’t first built the more basic conditions that make challenge safe and survivable; we cannot introduce otters into a river that cannot yet support them. So succession is useful here not as a ladder with a known top rung, but as a reminder of two things: that the same signal means different things depending on where a system has come from, and that some things simply cannot happen before others.

Presence isn’t establishment

One of the subtler lessons from ecological restoration work is the difference between presence and establishment. A conservation project might introduce specialist species (for example, round-leaved sundew [Drosera rotundifolia] to a recovering peat bog) as part of a restoration programme. To a passing observer, their presence looks like success: the specialist indicator species are there, therefore the habitat must be recovering and healthy. But an ecologist must look more carefully: are these plants thriving and reproducing? Is the substrate actually supporting them, or are they merely persisting on the back of the intervention, dependent on continued artificial implementation and support to survive?

Sundew. Photo by Ian at https://naturalbornblogger.co.uk/?p=1888 (CC BY-NC 4.0)

Organisations do this constantly. Fixating on “what worked there” instead of considering “what will work here?”. Adopting the “Spotify model”, implementing 360 degree reviews, putting a poster up telling employees we have a “Just Culture”. These are introduced species. Their presence may tell us relatively little, because the question we need to ask is whether they are self-sustaining: whether the conditions in the substrate actually support them, whether they are spreading and reproducing without continued top-down mandates, whether they’re actually the right species for this habitat at all. If the answer is no, what we have is a display, not a thriving habitat. The trap is flipping causality, in thinking that because many healthy organisations run regular team retrospectives (for example), that running regular retrospectives means we’re a healthy organisation

“The more one learns of this intricate interplay of soil, altitude, weather, and the living tissues of plant and insect (an intricacy that has its astonishing moments, as when sundew and butterwort eat the insects), the more the mystery deepens. Knowledge does not dispel mystery.”
― Nan Shepherd, The Living Mountain

This is where “rewetting organisations” comes in. In environmental restoration of a peat bog (incidentally, from the Gaelic bogach, meaning soft), we don’t begin by introducing specialist species. We begin by restoring the substrate: slowing down change, raising the water table, allowing sphagnum moss to re-establish, so that the conditions exist in which native species can return and thrive of their own accord. The specialist species come last, not first, and their arrival is evidence of recovery rather than its cause. Organisations that naively introduce practices they’ve seen others do, carry out restructures that mirror a different organisation, or apply goals and targets without attending to the substrate that supports them (including capabilities, power structures, incentive systems, institutional memory, and more), are at risk of planting sundews in dry, drained peat. They may, if we’re lucky, persist for a while, but they will not thrive. 

What are we actually looking for?

An ecologist asked to assess a habitat doesn’t arrive with a fixed checklist to tick items off. They arrive with a trained capacity for attention, developed over years of looking at many different habitats in many different conditions, and they read what is in front of them against everything they know about what could and should be there, what has been there, and what the current conditions suggest about where things are heading.

Some signals are generalist: ubiquitous enough that their presence tells you relatively little. Others are discriminating: their presence or absence in this habitat, at this stage, under these conditions, is genuinely interesting. The skill is knowing which is which; and that knowledge is inseparable from knowing the habitat. Organisational “seeing” means looking at, and doing our best to understand, the context first – before interpreting what the signals may be telling us. If we read the signal without reading the organisation first, we risk confidently misreading it.

With that caveat stated, some organisational signals tend to warrant attention:

A proliferation of committees and change advisory boards often (but not always) indicates an institutional fear of failure: we’d rather not do anything than try something that incurs risk, and the accountability sink of decision-making is distributed so widely that responsibility becomes impossible to locate. Heavy, bureaucratic handover processes between teams frequently signal low inter-team trust: paperwork and process act as a surrogate for trusting relationships. An abundance of individual metric targets for performance, particularly where they are tracked publicly and tied to reward, often suggests that leadership doesn’t trust the workers to be intrinsically motivated, and tends to suppress exactly the lateral knowledge-sharing that collective performance requires.

But (and this is the tubifex point) none of these signals means the same thing in every habitat. Individual targets are appropriate in some cases where there is little interdependence between people. Heavy duty approval boards and signoff may be appropriate in safety-critical, highly complex contexts such as nuclear power. Read the context first. Always read the context first.

The positive organisational indicators

What’s the organisational equivalent of otters?

Otters returned to many British rivers as water quality improved — because the conditions finally supported them. They sit near the top of the freshwater food chain, and their sustained presence signals that the entire system beneath them is functioning: clean water, healthy invertebrate populations, plenty of fish, stable bankside vegetation. We cannot fake otters. We cannot introduce them into a degraded system and expect them to hang around for long. Their sustained presence is evidence of genuine, systemic recovery.

Otter at the British Wildlife Centre, Newchapel, Surrey by Peter Trimming, CC BY-SA 2.0 <https://creativecommons.org/licenses/by-sa/2.0>, via Wikimedia Commons
Otter at the British Wildlife Centre, Newchapel, Surrey by Peter Trimming, CC BY-SA 2.0 , via Wikimedia Commons

The organisational equivalent is the signal that could only exist if everything beneath it was working. Not the signal that has been managed into existence and kept on life support, but the one that emerged because the conditions were finally right. A junior team member openly disagreeing with a room full of senior leaders and being thanked for it, not “performance managed” afterwards. Genuine productive dissent surfacing in formal meetings rather than only in corridors and car parks. Lateral knowledge-sharing and ideation across teams that are nominally in competition with each other for resources or recognition.

These things cannot be “performed” sustainably. They can’t be introduced by policy – we cannot simply tell people to “speak up”. These signals are the otters: evidence that the substrate has genuinely changed, that the water is clean, that something has actually been restored rather than imposed or mandated. If we are regularly seeing these things, they are typically very good signs of organisational health.

The measurement-industrial complex

We’re all familiar with organisational surveys — the engagement survey, the pulse survey, the psychological safety survey. Ecologists do surveys too, but the two things share little more than a name. When an ecologist surveys a habitat, we’re exploring, looking, seeing, sensing: reading what is present and what is absent, what is thriving and what is struggling. A staff survey, by contrast, is a snapshot, abstracted from the habitat, often in language that doesn’t quite fit the context.

There is by now a substantial industry built around the measurement of organisational health — engagement platforms, culture diagnostics, psychological safety scales, maturity frameworks, and benchmarking services of all kinds. Much of this is well-intentioned. Some of it is useful, particularly when applied carefully, longitudinally, and in combination with richer and more qualitative reading. But the industry as a whole rests on an assumption that the ecological approach exposes as problematic: that the same indicators mean the same things across contexts, and that the path to understanding an organisation runs through the consistent application of validated instruments.

This isn’t true. At least not primarily. The more elaborate and expensive our system for understanding something, the less we may actually understand it: because the apparatus of measurement tends to substitute for the discipline of attention rather than supporting it. When we have a dashboard, we watch the dashboard. When the dashboard tells us that engagement is 7.2, we manage towards 7.8. And in doing so, we invoke Goodhart’s Law in its most destructive form: the indicator becomes the target, and in becoming the target, it stops being an indicator. We end up improving nothing, or, in many cases, actually making it worse (you may already be thinking of cobras as indicator species!). The sludge worm at the bottom of the dashboard keeps trying to tell us the actual conditions, but nobody is paying any attention to it because it’s boring and small.

The measurement-industrial complex trades in what Aristotle would have called episteme: systematic, generalisable, transferable knowledge. The kind that (at least claims that it) can be packaged, certified, and applied by someone who wasn’t there and doesn’t know the habitat. What the organisational ecologist1 develops over time is closer to phronesis: practical wisdom that is irreducibly situated, inseparable from experience and context and judgement, and not available as a downloadable checklist. We can learn the names of the EPT taxa in an afternoon. Knowing what their presence or absence means, in this particular stream, at this particular time of year, given what happened upstream last year, takes much longer, and it takes practice.

What the organisational ecologist actually does

An ecologist doesn’t “fix” ecosystems: we work to restore the conditions in which the system itself can recover, and then we pay attention to what emerges. Organisations are complex adaptive systems, not machines to be upgraded by swapping a component in the right place.

silver birch establishing on heathland
Silver birch establishing on heathland [Tom Geraghty]

The capacity of being an organisational ecologist, of “seeing things properly”, develops over time, through practice and attention and the accumulation of many habitats read carefully. It is not a generic framework to be applied or a checklist to be ticked off. A real woodland restoration plan is a strange document to an organisational eye. It works across decades, sometimes with no end point at all, and contains almost no deadlines. Its phases overlap rather than run in sequence. Its operations are conditional: we fell trees only if they’re diseased, we intervene only where needed, and we support regeneration where it appears. A phase is complete not in Q3 but when the assessment suggests it is — when the system itself says so. The plan describes conditions and directions, not dated milestones, because habitats don’t consult plans. Neither do organisations.

None of this resolves into a simple organisational checklist, and it would betray the whole argument if it did. But there are dispositions, habits of attention, that the organisational ecologist develops.

  • We read the habitat before the signal, because the same signal means different things in different places.
  • We consider where the organisation is in its own development, because a structure that is healthy in a young system may be pathology in an old one.
  • We implicitly distrust the introduced species: the practice imported because it worked elsewhere, and ask whether it is actually thriving or being propped up.
  • We triangulate across many weak signals rather than trusting a single strong number, because a number is the feeling of knowledge and rarely the thing itself.
  • We look for the signals that can’t be faked: the ones that could only exist if the conditions beneath them were sound.
  • And we attend to the substrate, because we’re working with a living system that will not consult our plan, and the most useful thing we can do is often to make the ground more hospitable and pay attention to what grows.

Most of all, we keep looking. The habitat changes, the stage moves on, the indicator that meant one thing last year means another now. We’ll never stop trying to look at things properly, because the work is never finished.

See first, think later, then test. But always see first.

climbing crib goch
Looking out from Crib Goch [Tom Geraghty]

A different way of seeing isn’t something one can be handed; it’s something we develop, over time, in the field. But it helps to start somewhere. That’s what our Thinking Like An Ecologist workshop is for: not a new set of tools, but the beginning of a different kind of seeing.

1* Note: Through all of this I should be clear that I mean something quite different from the sociological field of “organisational ecology” (Hannan and Freeman), which studies why whole populations of businesses are founded and die under impersonal environmental selection. In this piece I mean something smaller and almost opposite in spirit: reading and tending the living dynamics inside a single organisation, on the premise that its conditions can be shaped.

The post Organisational Indicator Species: being an Organisational Ecologist appeared first on Psych Safety.

Psych Safety Day / Week 2027

4 August 2026 at 11:46

Once a year (sometimes more often than that!) we get folks from the Psych Safety community into one space, virtual or physical, to compare notes. Psych Safety Days sit alongside our smaller meet-ups and events as the point in the year where practitioners from healthcare, aviation, construction, software, education, emergency services and manufacturing find out what everyone else has been trying, what worked, and (often more usefully) what didn’t.

These aren’t conferences in the usual sense. There’s no keynote circuit, no sponsor track, and no one on stage selling a framework. The people presenting are the people doing the work: team leads, safety practitioners, clinicians, researchers, engineers, HR and OD folk, all of whom turn up as participants for the rest of the day. We deliberately keep the format loose enough that the conversations between sessions matter as much as the sessions themselves.

On research and evidence. We actively want academic work in the programme, with two conditions: the paper and its underlying data need to be available to everyone in the room, and the findings need to be presented in language a practitioner can act on. Open access, stated limitations, and honest uncertainty are all welcome. Paywalled results and proprietary “our model shows” claims are not: if we can’t inspect it, we can’t learn from it.

What 2027 might look like. We’re shaping the programme now, and we’re leaning towards a hybrid or fully virtual format. Travel costs, visas, caring responsibilities, accessibility needs and carbon all determine who gets to be in the room, and we’d rather not let geography decide whose experience counts. We’re also seriously considering stretching it into a full Psych Safety Week — a distributed run of sessions, local meet-ups and online workshops across several days, if there’s enough appetite for it.

Get involved. We’re looking for people to speak, facilitate a session, share research or a case study, host a local gathering as part of the week, or simply come along and argue with us in good faith. If any of that appeals, get in touch.

Use the code PSYCHSAFETYDAY10 for 10% off our action packs and toolkits and stickers.

Where we’ve been. The first Psych Safety Day ran in New York in 2022. Since then we’ve gathered in Barcelona, Seville and, most recently, Málaga — each one smaller and more useful than any standard corporate conference.

Psychological Safety Week

Tell us here what you’d like Psych Safety Day 2027 to be like and whether you’d like to take part!

Photo by Elvis Bekmanis on Unsplash
Photo by Elvis Bekmanis on Unsplash

The post Psych Safety Day / Week 2027 appeared first on Psych Safety.

Parachutes and “Evidence based” thinking

31 July 2026 at 09:32

These two pieces originally went out in the Psychological Safety Newsletter, the second (“I was wrong about the parachutes”) a couple of weeks after the first. They have proven rather popular, so we’ve posted them online here too.

Parachutes and “Evidence based” thinking

I love this rather infamous paper by Smith and Pell on the effectiveness of parachutes. “Parachute use to prevent death and major trauma related to gravitational challenge: systematic review of randomised controlled trials.” It makes the very effective point that nobody has ever conducted a randomised controlled study on the effectiveness of parachutes, but people in the real world still use them every single day, very effectively.

Psychological safety suffers a similar attack frequently by folks who claim to be robustly “evidence based”, when in reality they often simply dislike the idea of psychological safety for various, often ideological reasons and, as our own research suggests, from positions of relative structural advantage they may not see.

Psychological safety is the parachute. It is absurd to say that it may be better to not find out about problems, incidents, concerns or ideas for improvement. The evidentiary base for that isn’t only the academic literature — it’s the catastrophe record. TenerifeChernobyl, Deepwater Horizon, Challenger, the run-up to 2008 crisis: each is a case where a concern existed at the sharp end and did not travel to where it was needed, and the disaster followed. It’s a mechanism made visible by its failures, over and over, across many domains. Psychological safety’s successes are usually structurally invisible, because the averted disaster leaves no trace; which is exactly the parachute’s epistemic signature. To run the trial honestly, we would have to take a disaster we could see coming, prevent it in one arm by speaking up, and in the other simply let it happen, which sounds as sensible as jumping out of a plane without a parachute, just to see what happens.

So the demand, from a certain kind of critic, that psychological safety prove itself through a randomised controlled trial before we take it seriously isn’t rigour. It’s precisely the parachute review’s target. 

The argument has another edge though, because the catastrophe record proves the mechanism: we can’t fix a secret, we can’t utilise an idea we haven’t heard, and we can’t learn from a mistake that we don’t know about. But it does not prove the commercial models and constructs of psychological safety: the company survey, the stage model, or a metric on the dashboard. 

The parachute absolutely works, but not every parachute is made equal. Some do, in fact, fail. So whilst the RCT demand for the mechanism of psychological safety is absurd, we are right to be critical of much of the commercial apparatus around it. Nobody should be allowed to sell parachutes that don’t open.

I was wrong about the parachutes

A few weeks ago I told you that no RCT of parachutes existed. Sorry – I was wrong, there is one.

In 2018, Yeh and colleagues published one in the BMJ: the PARACHUTE trial — that’s “PArticipation in RAndomized trials Compromised by widely Held beliefs aboUt lack of Treatment Equipoise.” It was a real trial with real ethics approval, randomisation and analysis. Twenty-three participants jumped from an aircraft wearing either a parachute or an empty backpack.

The results: parachute use did not significantly reduce death or major injury (0% v 0%, P>0.9), consistent across all subgroups. In fact, “the parachute did not deploy in all 12 (100%) owing to the short duration and altitude of falls.

So parachutes don’t work, right?

It’s worth noting that the trial struggled somewhat with enrolment. Investigators approached strangers seated near them on commercial flights, mid-flight, and asked whether they’d be willing to be randomised to jump at current altitude and velocity. Curiously, everyone declined. The authors note that owing to “difficulty in enrolling patients at several thousand meters above the ground,” the team expanded recruitment to friends, family, and themselves.

Which brings us to a key point: the 69 people who declined to participate were on commercial aircraft at a mean altitude of 9,146 metres, travelling at 800 km/h. The 23 people who opted in to the study were on stationary aircraft: a parked biplane at Martha’s Vineyard and a museum helicopter in Michigan, at a mean altitude of 0.6 metres, travelling at 0 km/h

Fig 2 | Representative study participant jumping from aircraft with an empty backpack. This individual did not incur death or major injury upon impact with the ground.
Fig 2 | Representative study participant jumping from aircraft with an empty backpack. This individual did not incur death or major injury upon impact with the ground.

The trial is a joke, of course — a sequel to Smith and Pell’s 2003 paper, published in the same BMJ Christmas issue tradition. But like its predecessor, the joke makes a serious but different point.

The 2003 paper argued that you can’t run the trial because the counterfactual is unethical, whilst the 2018 paper demonstrates something subtler about sampling. They actually did run the trial, the methodology was genuinely rigorous: nothing about the randomisation or the analysis is wrong, but it still produced an answer that is both statistically sound and completely absurd, because the conditions required to make the trial possible were precisely the conditions that removed the effect. Parachutes work as a function of altitude. Everyone at altitude declined to enrol. The trial measured the sample rather than the intervention.

“When beliefs regarding the effectiveness of an intervention exist in the community, randomized trials evaluating their effectiveness could selectively enroll individuals with a lower likelihood of benefit, thereby diminishing the applicability of trial results to routine practice.”
Yeh et al, 2018.

And it measured the sample not because the investigators were careless, but because the real version of the trial can’t be run. To test parachutes honestly we’d need to randomise people at 9,000 metres, and give half of them empty backpacks instead of parachutes just to see what happens. No ethics board would approve that, nor would participants consent to it. The only RCT that can exist is the one on the tarmac.

Instead, we could look at the observational record, and we have done (not we, I mean, but people have done). On one side (with a parachute): six deaths in 110,000 Danish sports jumps35 in 6.2 million French ones. On the other: falling from high altitude without a parachute is almost universally fatal. Known survivors of such events number in the dozens: survival is extraordinary.

The Dramatic Effect

Epidemiologists call this a dramatic effectwhen the gap between intervention and non-intervention is this large, we don’t need an RCT. (Glasziou et al, 2007)

It’s quite clear through examining past cases of parachutes failing that the outcome leans towards bad for the subject. This isn’t an RCT though – it’s empirical data, which is exactly what the vast amount of data from disasters such as ChallengerAmagasakiGrenfellChernobyl and others show about psychological safety: a lack of it results in bad things happening. 

Controlled studies on psychological safety necessarily come from lower altitudes, where failure isn’t catastrophic. The genuinely meaningful data comes instead from retrospective investigation of the disasters where voice failed. Instead of a weakness of the field, this is simply the shape of the phenomenon.

And it also speaks to a certain kind of psychological safety critic. You know the sort – the people who claim that we don’t need psychological safety – “people should just speak up“. They are, in a sense, right — but only for themselves, because they already have psychological safety. They’re standing on the tarmac, exposed to very little risk. It’s the people making the big leap out of the plane in the sky who need the parachute.

References

Ellitsgaard, N. (1987). Parachuting injuries: a study of 110,000 sports jumps. British Journal of Sports Medicine, [online] 21(1), p.13. doi:10.1136/bjsm.21.1.13.

‌Fer, C., Guiavarch, M. and Edouard, P. (2021). Epidemiology of skydiving-related deaths and injuries: A 10-years prospective study of 6.2 million jumps between 2010 and 2019 in France. Journal of Science and Medicine in Sport, [online] 24(5), pp.448–453. doi:10.1016/j.jsams.2020.11.002.

Glasziou, P., Chalmers, I., Rawlins, M. and McCulloch, P. (2007). When are randomised trials unnecessary? Picking signal from noise. BMJ, [online] 334(7589), pp.349–351. doi:10.1136/bmj.39070.527986.68.

‌Smith, G.C.S. and Pell, J.P. (2003). Parachute use to prevent death and major trauma related to gravitational challenge: systematic review of randomised controlled trials. BMJ, 327(7429), pp.1459–1461. doi:10.1136/bmj.327.7429.1459.

‌Yeh, R.W., Valsdottir, L.R., Yeh, M.W., Shen, C., Kramer, D.B., Strom, J.B., Secemsky, E.A., Healy, J.L., Domeier, R.M., Kazi, D.S. and Nallamothu, B.K. (2018). Parachute use to prevent death and major trauma when jumping from aircraft: Randomized controlled trial. BMJ, [online] 363(363), p.k5094. doi:10.1136/bmj.k5094.

The post Parachutes and “Evidence based” thinking appeared first on Psych Safety.

The Psychological Safety Fieldbook

31 July 2026 at 09:23

Amy’s Psychological Safety Field Book is on the way

Amy Edmondson has a new book coming. The Fearless Organization Field Book: A Guide for Creating Psychological Safety at Work (Wiley) is listed for publication in 2026, and we know a little about it: 208 pages, paperback, and a publisher’s description promising step-by-step tools, exercises and proven frameworks for managers, HR professionals, coaches and consultants. I’ve pre-ordered it anyway, and this piece is about why, and about what I will be looking for when it arrives.

Some context for anyone newer to this territory. Edmondson is not one voice among many in psychological safety: she is the researcher whose 1999 study of hospital teams turned the construct into a measurable, validated property of teams, and whose work since has anchored practically every serious study in the field. Our own timeline of the concept is, in significant part, a timeline of some of Amy’s career. When we map the academic terrain in the Field Guide, her papers sit at the centre of it.

The gap the book steps into

For twenty-five years there has been a curious division between research and practice in psychological safety. In the space between the two of those, an industry grew.

Some of that industry is thoughtful. A lot of it is absolutely not. Psychological safety has been flattened into surveys that promise a single score, leadership checklists, and numerous dashboards and diagnostics. The research says psychological safety is an emergent property of a group, whilst the industry often sells it as a product that a leader can install.

Edmondson has spent years gently (and sometimes not so gently) pushing back on this: correcting the “be nice” misreading, insisting that safety is not comfort, pointing out that a climate cannot be mandated into existence. 

So a field book is needed, because the practical layer (as opposed to the research layer) is where most people actually meet psychological safety. 

A few hopes for the Field Book

Practical guidance about psychological safety is legitimately hard to write, and we say that as people who have spent six years attempting it.

I hope it resists the leader-as-hero framing. Most practical material makes psychological safety something leaders do to teams. The evidence is more interesting: safety is built, tested and repaired by everyone in the room, and leaders are not solely responsible for it, even though they carry disproportionate weight. 

I hope it treats measurement with respect for what measurement can and cannot do. There is a well-worn path from “how do we know it’s working?” to a dashboard that makes things worse: scores gamed, surveys mistaken for the thing itself, a living property of a group reduced to a metric. 

I hope it addresses power. Power gradients are the single most important thing to address with respect to psychological safety, and often the most difficult, because people in power rarely want to dismantle the power structures that they’re standing on.

I hope it’s honest about what tools cannot do. The best practical writing is honest about its own reach: an exercise can open a conversation, and cannot substitute for the hundred small responses that follow it. Our experience is that practices help when they are held lightly and calcify when they are followed as scripts. The publisher’s copy promises step-by-step tools and proven frameworks, which is either exactly right or exactly the problem, depending on how it’s done.

When the Field Book Arrives

I’ve pre-ordered the book, and when it arrives, I’ll review it properly and connect the whole thing into the wider terrain so you can see where it sits.

In the meantime, our own practical material is here, our reading list is here, and the research underneath all of it is mapped in the Field Guide, where Edmondson’s work already occupies the high ground.

The post The Psychological Safety Fieldbook appeared first on Psych Safety.

High school defends staying silent while boys made AI nudes of 59 classmates

31 July 2026 at 18:11

One of the first schools to shut down after students were found making AI nudes of female classmates is now asking a court to toss a lawsuit filed by victims who claimed that the school stayed silent for months while the emboldened boys targeted many more girls.

In a motion to dismiss this week, Lancaster Country Day School (LCDS)—a private K-12 school in Pennsylvania with fewer than 600 students—argued that it was false to say the school never reported the harm to law enforcement. The tip that the school received came from the Pennsylvania Office of the Attorney General, which is itself a law enforcement agency, the filing said.

It’s also false to say the school knew that girls were being targeted, the school argued, because the tip did not mention any specific student victims.

Read full article

Comments

© uniquepixel | iStock / Getty Images Plus

RFK Jr.'s handpicked committee approves manufacture of peptides he uses

24 July 2026 at 19:10

In a widely expected move, a committee organized by the Food and Drug Administration (FDA) has voted to endorse removing restrictions on the manufacture of peptides for human use. Thursday's panel meeting saw a sharply divided group recommend lifting limits on four peptides; votes on three additional peptides are scheduled for today. The move comes despite a continuing lack of evidence regarding their safety and effectiveness.

The move had been telegraphed months earlier as peptide enthusiast and Health and Human Services Secretary Robert F. Kennedy Jr. took steps to ensure this outcome.

Like proteins, peptides are composed of amino acids that are chemically linked into a chain. Peptides differ merely by length; they're often 10–20 amino acids long, in contrast to proteins, which can be hundreds or thousands. Some of them, such as insulin, are specifically made by targeted processing of a protein into shorter fragments, and the resulting peptide interacts with receptors that have evolved to send signals to cells based on its levels.

Read full article

Comments

© Getty | Michael M. Santiago

Tesla driver who blamed crash on autopilot pressed accelerator 100%, NTSB finds

16 July 2026 at 14:48

On Wednesday, the National Transportation Safety Board (NTSB) released preliminary findings verifying Elon Musk’s and Tesla’s claims that a driver involved in a fatal Texas crash that killed a grandmother overrode Full Self Driving in the moments ahead of impact.

Last month, 44-year-old Michael Butler told police that the autopilot feature was engaged at the time of the crash. On X, Musk disputed the claim, writing that Butler must have overridden the feature because “FSD drives slowly through neighborhood streets, and this was a high-speed crash!” Moving to back Musk’s claim, Tesla’s vice president of AI software, Ashok Elluswamy, said that internal data showed “the driver manually overrode self-driving by pressing the accelerator all the way to 100 percent of the accel pedal in this residential area.”

NTSB’s preliminary report, which does not yet determine what caused the crash, confirmed Tesla’s claims. Its probe found that FSD was engaged at the time of the crash, but electronic data showed “the driver manually overrode FSD (Supervised) by pressing the accelerator pedal to 100 percent.”

Read full article

Comments

© via Jennifer Barbour's complaint

❌
❌