Humans make carbon dioxide. Carbon dioxide is (edit: sometimes claimed to be) bad for cognition. But plants turn carbon dioxide back into oxygen. And plants are the one true home decoration strategy. So maybe if you get a lot of plants, you can you can keep carbon dioxide in check and keep your brain working?
It’s theoretically possible. It’s probably just barely possible in practice. But it won’t be easy.
People produce ~1 kilogram of carbon dioxide per day. That’s around 5.7 × 10²³ molecules or 0.948 moles per hour. (You may remember from high school that a mole is a gigantic number made up to avoid having factors of 10²³ everywhere.) Let’s keep it simple and call it one mole per hour.
Meanwhile, plants turn carbon dioxide into oxygen through photosynthesis, i.e. the chemical reaction of (6 water molecules) + (6 carbon dioxide molecules) + (energy) → (1 glucose molecule) + (6 oxygen molecules). The minimum energy physically needed to convert 1 mole of carbon dioxide into glucose and oxygen via this reaction is ~477 kilojoules.
So we’ve already got a lower bound. Say you have magical plants that somehow channel all incoming energy into photosynthesis with perfect efficiency. They’ll need ~477 kilojoules per hour, which converts to a continuous usage of 132.5 watts.1 That’s a bit more than what’s used by two incandescent light bulbs, which isn’t too bad.
But you don’t have magical plants. Real plants do photosynthesis through a physical process with two steps, each of which involves four electrons absorbing a photon. That means you need eight photons per carbon dioxide molecule. If you want to tune your lights for maximum efficiency, you should give each photon exactly the minimum energy necessary to excite an electron, which happens to be ~1.8 eV. That corresponds to pure red light with a wavelength of 680 nm, and a continuous usage of 386 watts.2 No physical system using chloroplasts can neutralize your CO₂ using less than that. Somewhat high, but still manageable.
But your houseplants won’t be able to grab every single photon that hits them and direct it towards photosynthesis. In practice, ~30% of photons will reflect off the plant, or go through it, or hit some part of the plant other than the chloroplasts. That brings us to 551 watts.3
And there’s another issue. After plants make glucose, what happens to it? Some is used to grow more plant, which permanently sequesters carbon from the environment. But lots is also burned by the plant for the general business of staying alive, releasing the carbon back into the air. The exact amount burned in this way varies based on species and conditions, but around 40% loss reasonable,4 bringing us to 918 watts.5
That doesn’t sound that bad. But have you considered what it would be like to live in the same room with 918 watts of pure red light? In terms of radiant power, that’s the same as produced by ~765 incandescent lightbulbs.6 Modern LED grow bulbs are ~50% efficient, meaning you’ll actually need to spend ~1836 watts. If you’re imagining plants that you can actually see, adjust that upwards again for all the light lost to the room. And if you want to use normal light frequencies instead of living Red Life, then your LED bulbs will be less efficient at creating light and your plants will be less efficient at capturing it. Realistically, we’re talking about something like 5,000-10,000 watts, most of which is lost to the room as heat. Imagine five space heaters blasting you on high all the time.
But maybe you’re OK living in a tanning booth. Or maybe you’ll keep your plants in a perfectly reflective chamber. Or maybe your house has a glass ceiling and infinite free sunlight and free climate control. That’s cool. But have you forgotten about your old friend, photosynthetic photon flux density?
Plants can’t absorb infinite amounts of light. Chloroplasts take time to “reset” before they can absorb more photons. Your pet fern can only absorb ~52 watts of energy per square meter of leaf surface area.7 So no matter how much light you can produce, if you want to neutralize the carbon dioxide you make, you will need at least 918 / 52 = 17.6 square meters of fern leaf. Picture a 4.2 meter square wall, packed solid with ferns. If there are any gaps, stems, soil, or wall showing, it needs to be even larger. That’s the absolute minimum.
But maybe that still sounds OK? Fine. But consider one last barrier: Plants obey the laws of physics [citation needed]. If they remove carbon from the air, they must put that carbon somewhere. The only place it can go other than back into the air is into the plant itself.
The 1 kg of carbon dioxide you produce each day corresponds to 273 grams of elemental carbon. The only way for a plant to hide that is by making more plant. But dry plant matter is only ~50% carbon, and for each gram of dry plant matter, plants have 5-10 grams of water (varying a lot by species). So in order to sequester all the carbon you make, each day you will need to grow around
(1 kg carbon dioxide)
× (0.273 kg elemental carbon / kg carbon dioxide)
× (2 kg dry plant / kg elemental carbon)
× (8.5 kg actual plant / kg dry plant)
= 4.6 kg actual plant.
Your garden must grow that much, every day. That’s 140 kg per month. You must prune and discard all that outside, or your garden is not actually sequestering anything.
In conclusion:
Build an industrial indoor farm.
Weigh it.
Wait two weeks.
Weigh it again.
Divide the increase in weight by your own body mass.
That’s the fraction of your CO₂ that you’re removing from the environment.
So chloroplasts are at most ~34% (132.5 / 385.94) efficient at channeling the energy in light into photosynthesis. ↩
I find this 30% number amazingly low. (Well done, evolution.) And perhaps it should be somewhat lower. For one thing, the 30% figure comes from sunlight filtered to the 400-700 nm range. If you’ve got pure 680 nm light, absorption should be somewhat higher. Also, if photons are absorbed by some part of the plant other than the chloroplasts, they become heat and the energy is gone. But if they’re reflected or go through the plant, then they might go on to hit some other plant (provided you have a lot of plants around). If you really have pure 680 nm light and you have very densely packed plants, maybe you could drop this to 10-20%. ↩
Wikipedia quotes a 35-45% loss just for respiration in the leaf itself. But then this paper shows numbers ranging from 30% to 56% depending on the species and growth rate. ↩
I’ve estimated an overall efficiency of 132.5 watts / 918 watts ≈ 14.4%. If you go to Wikipedia, it estimates that ideal leaf efficiency with sunlight is only around 5.4%. That’s because sunlight contains a wide band of wavelengths and my calculation assumed an ideal 680 nm source. Around 47% falls outside the 400-700 nm range, and inside that range, around 24% is lost due to higher-energy photons with energy that gets wasted as heat. If you account for that, my estimate becomes 14.4% × (1-0.47) × (1-0.24) = 5.8%, which is close enough for government work. ↩
A traditional “60 watt” incandescent lightbulb is rated based on the power input. But only around 2% of that energy is actually converted to light. So 918 watts of pure red light isn’t what you get from 918 / 60 = 15.3 lightbulbs. It’s what you get from 918 / 60 / .02 = 765 lightbulbs. That said, your eyes aren’t very sensitive to 680 nm light, so the perceived lux wouldn’t be nearly so bad. ↩
The saturation point of plants is usually given in units of 300 μmol/m²/s. That the number of photons (in micromoles) that can be absorbed, per square meter of leaf, per second. A typical value for a shade-tolerant houseplant would be ~300 μmol/m²/s. If we assume again that the light is 680 nm so that each photon carries 1.8 eV of energy, then ~300 μmol of photons carries 51.92 joules. That’s 51.92 joules of energy per square meter of leaf surface, i.e. 52 watts. ↩
No. Creatine is a nutrient. Most omnivores eat a gram or two per day from meat. Your body also synthesizes a gram or two per day. You need creatine to deliver energy inside of cells. It is normal and non-weird.
Does creatine increase testosterone?
Unlikely. This concern comes from one study in 2009 on 16 male rugby players.1 But that study is considered extremely suspect. There have been at least twelve other studies that all found no change or physiologically irrelevant changes. Beyond that, it’s implausible that creatine would increase testosterone, because we know what creatine does and it has nothing to do with hormones.
There is no mechanistic reason to think that would happen.
There are good mechanistic reasons to think that would not happen.
These rumors all trace back to speculation built on top of that same single 2009 study. But that study is contradicted by later research, and anyway didn’t measure hair. Anything is possible, but as far as I can tell, it’s equally plausible that creatine would increase hair growth. And if you’re really worried about this: Are you going to stop eating meat?
Is creatine safe?
Probably. The International Society of Sports Nutrition says:
Available short and long-term studies in healthy and diseased populations, from infants to the elderly, at dosages ranging from 0.3 to 0.8 g/kg/day for up to 5 years have consistently shown that creatine supplementation poses no adverse health risks and may provide a number of health and performance benefits.
It’s been studied extensively, and no risks have been found. The way it works doesn’t suggest any risks. And supplementing a few grams per day doesn’t put you far outside the range that people get from normal food.
Does creatine make you stronger?
Yes. It’s very rare for a supplement to have such strong and consistent evidence. A widely-cited review says that short-term supplementation increases maximal power/strength by 5-15%. This in turn may increase the long-term gainz from strength-training exercise. Creatine also increases sprint performance by 1-5%. Though, there seems to be little if any benefit for endurance exercise like long-distance running.
But how does creatine make you stronger?
Before answering that, can I go on a rant about how muscles work?
…OK?
Great! Here’s how muscles work:
All cells have a molecule called ATP floating around inside, which they use for energy.
Muscle cells have proteins in them called myosin.
When ATP bumps into myosin, the myosin breaks the ATP down into ADP. This releases energy which is physically captured by the myosin as elastic strain.
When triggered by neurons, myosin releases that mechanical energy.
When you decide to move your arm, your brain triggers many muscle cells, carefully orchestrating the myosin twitches into large-scale movement.
Now, here’s something that’s crucial for our story: Very little energy is stored as ATP. Your body contains ~100 grams of ATP, representing ~10,000 joules of energy.2 But your body at rest burns ~100 watts. So you only store enough ATP to keep yourself alive for ~100 seconds. If you sprint, you could easily burn ~3000 watts, which would use all your stored ATP in ~3 seconds.
Through the magic of eating, you’re always making more ATP. Typically, your mitochondria recycle ~1 gram of ADP back into ATP per second, the same amount you need to stay alive.3 If you start running, your body can ramp that up to ~10 grams per second, though tricks like breathing faster and speeding up your heart.4 But it takes a minute or two for your mitochondria to really get cranking.5
So then why am I able to sprint for longer than three seconds?
Because creatine acts as an additional energy reservoir, coupled to the ATP reservoir. After you eat or synthesize creatine, 60% is converted into phosphocreatine. This is done by an enzyme that grabs a creatine molecule and an ATP molecule and moves a phosphate group between them. This “charges” the creatine into phosphocreatine and “discharges” the ATP into ADP.6
But if your ATP levels drop—e.g. because you’re running away from a tiger—those enzymes will run in reverse, meaning they “discharge” phosphocreatine into creatine and “charge” ADP back into ATP. This happens almost instantly, so that ATP and phosphocreatine deplete at the same rate.7
At rest, your muscles contain around 3-4 times as much phosphocreatine as ATP. So the “extra” energy storage in phosphocreatine is much larger than the “base” storage in ATP itself. That’s why you can sprint for ten seconds rather than just three seconds.
Does supplementing creatine increase creatine levels in muscle cells?
So, everything seems to add up. If you supplement creatine, you increase your levels by ~16.67%, implying ~12.5% more total short-term energy storage.8 That’s in line with the 5-15% increase in strength seen in creatine trials.9 It also seems to make sense that creatine trials find little benefit for endurance exercise. If you don’t have sudden bursts of activity, a larger short-term energy reservoir won’t really help you.
But isn’t this all very strange?
Well, I find it strange. All else equal, more strength is good. The body already knows how to make creatine. If you can just raise creatine levels and get more strength with no downsides, then shouldn’t evolution have done this already? Some variant of the Algernon argument would suggest that the fact that creatine works so well should be impossible.
You might think that higher creatine levels are bad somehow, and that’s why evolution didn’t make them higher. But that seems wrong. Creatine levels vary naturally based on what you eat. If higher levels were bad, evolution could have brought them down. But it doesn’t. It just lets them vary.
Often, evolution makes us “worse” to reduce our energy expenditures, because evolution hates it when we starve to death.10 But the body only spends 1-2 calories per day synthesizing creatine, and more creatine in muscle cells doesn’t have any significant metabolic cost.
I think the boring explanation is that for our evolutionary ancestors, modest increases in short-term strength just weren’t a big deal. We were exhaustion hunters, not 1-rep max deadlift hunters.11 Also, more creatine causes your muscle cells to draw in some extra water, which slightly increases energy usage for long-distance running.12 So, if you happened to get extra creatine from meat, great. If not, whatever. In the range where creatine fluctuates based on diet, I suspect creatine levels just didn’t have much impact on reproductive success.
Still, we must acknowledge that creatine is unusual. I wish we could tell our bodies, “Hey, we have access to unlimited amounts of food. Stop worrying about conserving energy and concentrate on being awesome.” But we have very few ways to do that. As far as I can tell, the list of normal nutrients that have been proven to increase strength is: protein, creatine, beta-alanine, the end.
So creatine is special. And creatine makes you a little stronger. Does it make you a little smarter, too?
Is creatine used by the brain?
Yes. Most parts of the body don’t contain significant creatine. But the brain does, along with muscles, the heart, and testes. Neurons use it to play the same game muscles do with ATP and phosphate groups and so on.
How much creatine is in the brain?
Maybe half as much as in muscle. The number of interest here is the ratio of phosphocreatine to ATP, indicating how much phosphocreatine increases local energy storage. We saw above that in muscle, that ratio is 3 to 4. In the brain, the numbers are a little sketchy, but the ratio seems to be more like 1.5 to 2.13
But why? Why would the brain use creatine?
Good question! The brain doesn’t have bursts of energy usage like muscles do. Yes, the brain uses ~20% of all calories despite only making up ~2% of body mass. But the brain is unusual in that it needs all that energy just for basic housekeeping, and doesn’t ramp up with usage. Contrary to the common myth, thinking hard does not burn significantly more calories. (Demonstration: Start thinking hard, and watch as your heart rate does not increase.)
So muscles use creatine for sprints. But the brain doesn’t have sprints. So what the hell is the brain using creatine for?
The most common theory seems to go like this: Actually, muscles don’t just use creatine as an extra energy reservoir. They also use it to deliver energy inside of cells. You see, creatine diffuses faster than ATP inside of cells. So even with endurance exercise, creatine is still being used: Enzymes near the mitochondria use ATP to “charge” creatine into phosphocreatine and enzymes near myosin use that phosphocreatine to “recharge” ADP back into ATP. Even though the net change in creatine is zero, it helps “shuttle” energy from the mitochondria to the myosin.
Under this theory, what neurons and muscle cells share is that parts of the cell locally use a lot of energy, when they get triggered. So even though your brain doesn’t “sprint”, it still uses creatine to avoid local energy deficits.
There’s also experimental evidence that creatine is important for the brain. We’ve created genetically altered mice with brains that lack the enzymes needed to convert creatine to and from phosphocreatine. They display severely limited spatial learning and somewhat smaller brains.
Some humans also naturally have creatine deficiency. In some variants, people have trouble synthesizing creatine. This leads to lower levels throughout the body, including skeletal muscle where 95% of creatine lives. Nevertheless, the primary symptom is related to the brain, namely intellectual disability. Muscle weakness and seizures are also common. Other people have creatine transporter deficiency, meaning creatine can’t cross the blood-brain barrier. This leads to lower levels in the brain only. This leads again to intellectual disability and also often muscle weakness or seizures. (That muscle weakness is despite the fact that the muscle cells themselves have normal creatine levels.)14
So somehow, creatine is very important for the brain.
Does supplementing creatine increase creatine levels in the brain?
Probably, though likely less than in muscle.
Creatine can definitely cross the blood-brain barrier. However, the protein that helps it cross is not abundant, and there are some suggestions that it’s down-regulated with prolonged creatine consumption. The brain itself can synthesize some creatine, and this too might be down-regulated by prolonged consumption.
Of course, you can just give people creatine and see what happens to their brains. There have been around a dozen such studies. Most report increases between 3% and 10%, although a few report no change. However, because brains are hard to access, these studies rely on magnetic resonance spectroscopy, and some suggest that these measurements are unreliable.
In people who can’t synthesize creatine, oral supplementation seems to normalize levels in the brain. (Some cognitive impairment usually remains. One patient was diagnosed and began supplementing at three weeks of age and had no intellectual disability.) So supplementing can increase brain levels in some circumstances.
My best guess is that supplementing does usually increase levels in the brain, and that an increase of 3% to 10% is plausible. But the evidence isn’t particularly strong.
Why did people get interested in creatine having cognitive benefits?
Because of Rae et al. (2003). They took a group of 45 healthy vegetarian or vegan university students in Australia. They did a cross-over trial where half of people got 5 grams of creatine per day for six weeks, followed by a six-week wash-out period, followed by the other half of people getting creatine. Their results were amazing, with huge improvements on Raven’s matrices (RAPM) and backward digit span (BDS):
In their analysis, creatine increased BDS by 1.19 standard deviations, and RAPM by 1.76 standard deviations. If we convert those numbers to IQ points (where 1 standard deviation ↔ 15 IQ points), that would mean increases of 17.85 and 26.4 IQ points, respectively. In both cases, the results were highly significant (p < 0.0001).
Does that replicate?
No. Following that paper various groups tried similar experiments but no one found such a large or statistically significant effect. After twenty years of inconclusive results, Sandkühler et al. (2023) set out to give a definitive reproduction. In my view, this is the highest-quality RCT ever done on the cognitive benefits of creatine.15 They largely borrowed the experimental design of Rae et al., although they did the experiment in Germany, used a larger sample of 123 people, used half non-vegetarians, and they dropped the wash-out period. Here are their main results:
(T1 shows test results at baseline. T2 shows results after six weeks of creatine or placebo. T3 shows the results after another six weeks, where the placebo group crossed over to creatine and vise versa.)
Overall, everyone got better over time, probably from practice. On backwards digit span, during the first six weeks, the group getting placebo actually improved slightly faster than the group getting creatine. But when those groups switched between getting placebo and creatine, that (formerly placebo, now creatine) group improved even faster. Just staring at the graph, this suggests some benefit. On Raven’s matrices, the same thing happened, but with a greatly reduced magnitude.
They fit a statistical model and report an effect size of 0.17 standard deviations for backwards digit span (~2.5 IQ points, not quite statistically significant) and 0.09 standard deviations for Raven’s matrices (~1 IQ point, not even close to significant). They found no extra benefit for vegetarians, not even a non-significant benefit.
As far as I can tell, this discrepancy has never been convincingly explained. Rae et al.’s 2003 experiment seems well done. The results are too large to be explained by p-hacking and too statistically significant to be explained by random noise. Maybe for some reason, Rae et al.’s cohort had lower baseline creatine levels? It’s very odd. But history suggests that when an exciting result is followed by a disappointing replication, we should bet on the disappointing replication.
What about all the other RCTs? Doesn’t this call for a meta-analysis?
In principle, yes. The trouble is, most of the studies don’t report the numbers needed for a good meta-analysis. They do some experiment giving creatine to half of people and placebo to the other half, and measure how those groups do on some cognitive test. Then they fit some statistical model and report p-values or whatever. But they never actually publish the raw means and standard deviations.16
Fortunately for us, Xu et al. (2024) contacted the authors for all those trials and got their raw data. According to their meta-analysis, creatine had the following effects.
Domain
Effect size (standard deviations)
Overall cognitive function
+0.34
Executive function
+0.32
Attention
+0.22
Memory
+0.31
Processing speed
+0.01
Unfortunately for us, that paper is bad. They claim that several of these results are statistically significant, but a 2026 commentary points out that they made an error that amounts to double-counting the same data for several studies.17 For that reason, I haven’t shown their (incorrect) confidence intervals. If computed correctly, I suspect none of the results would be statistically significant. Technically, the above point estimates are also wrong, although the error shouldn’t systematically bias them in either direction.
In general, I have to tell you that I really don’t trust this paper. It’s very sloppy with tons of missing details. But as far as I can tell, no one else has ever assembled the data needed to do a good meta-analysis. So I think those numbers are the best summary we have.
Here’s what they have to say (I’ve cut references for readability):
The Panel considers that, overall, the 10 human intervention studies […] do not show a consistent effect of creatine supplementation on cognitive function. The Panel notes that the acute effect of creatine on working memory reported in some studies […] was not observed at lower creatine doses […] or with continuous consumption of creatine. The Panel also notes that the effect of creatine […] reported in one study is an isolated finding across the body of evidence, where no effect of creatine supplementation was observed on other cognitive domains, including different facets of memory (episodic, short‐term, visual), verbal fluency, attention, alertness, processing speed, psychomotor speed, executive function and general cognitive ability/flexibility and fluid intelligence. Finally, the Panel notes that the three intervention studies conducted in diseased individuals do not support an effect of creatine supplementation on cognition.
I think we should consider this definitive. I’d go so far as to say this document probably represents the greatest effort our civilization has ever made to understand if creatine has cognitive benefits.
But we need to remember the ESFA’s role. They’re asking if creatine has been proven to have cognitive benefits, because they’re deciding if it should be legal to advertise cognitive benefits. They say no and I believe them. But that doesn’t mean there are no cognitive benefits.
Are there other reviews of the RCTs?
Yes. Here are all the recent reviews I could find, with a few representative quotes from each:
“Creatine supplementation showed significant positive effects on memory and attention time, as well as significantly improving processing speed time. However, no significant improvements were found on overall cognitive function or executive function.”
“Creatine supplementation has no significant effect on young healthy participants in unstressed situations. Moreover, the review show mixed results for stressed groups.”
“Vegans do not intake sufficient […] creatine to ensure the levels necessary for maintaining optimal cognitive output.”
“Closer examination of [the evidence] suggests that there may be more positive outcomes of supplementation than the research so far provides.”
“A cause-and-effect relationship has not been established between the consumption of ≤3g per day creatine and improved cognitive function.”
On average, the RCTs do find a small positive effect, just not a statistically significant positive effect. As I so often point out, that’s exactly what we would expect if the true effect were positive but small. But it’s also entirely possible that this is due to random chance or p-hacking or publication bias. Gwern contacted one author and found that publication bias did in fact occur.
Overall, I think the RCTs provide very weak evidence in favor of a small benefit for healthy adults. (Perhaps 0.1 to 0.3 standard deviations, depending on the measure.) I also think they provide moderate evidence against a larger effect for healthy adults (above, say, 0.5 standard deviations) and weak evidence for a small benefit for adults that are “stressed” in some way that might diminish creatine, such as being older, vegan, or physically exhausted.
Can you summarize the evidence in favor of creatine making you smarter?
I would love to do that:
Creatine is special. Very few nutrients really make you stronger, but creatine does.
Few parts of the body other than muscles use significant creatine, but the brain does.
Creatine can cross the blood-brain barrier.
Creatine is vital for the brain to function correctly.
Supplementing creatine probably increases creatine levels in the brain, at least a little.
Some RCTs suggest a cognitive benefit.
Can you summarize the evidence against creatine making you smarter?
Yes:
We don’t fully understand how the brain uses creatine. There is no clear mechanistic story for why supplementing creatine should make you smarter.
The best analogy for how the brain uses creatine is how your muscles use creatine for endurance exercise. But creatine has little benefit for endurance exercise.
It hasn’t been firmly established how much (or if) supplementing creatine increases creatine levels in the brain.
The RCTs suggest a benefit that is quite small, on the order of 1 to 3 IQ points.
The RCTs are not statistically significant.
Does creatine make you smarter?
I don’t know. Maybe a little.
You could make an argument like this: Creatine is crucial for the brain (somehow) so it’s safest to keep levels high, just in case. But I’m not sure I buy that. Creatine is crucial for the brain but there are several hints that evolution knows that, and so regulates levels in the brain more tightly than in muscles.
I might buy that argument for vegetarians or vegans. But there is scant experimental evidence for extra cognitive benefits in those groups, and even some evidence that vegetarians may not have much lower brain creatine levels, despite vastly lower consumption.
And if creatine is helpful, the likely benefit is probably quite small. Say you think there’s a 50% chance creatine increases IQ by 1 point and a 50% chance it’s useless. Is it actually worth the trouble of taking 5 grams of creatine every day for an expected increase of 0.5 IQ points? I’m not sure.
Technically, they found an increase in dihydrotestosterone (DHT) but not testosterone. ↩
Conveniently, in typical cellular conditions, the body can extract around 100 J of energy from 1 gram of ATP. So we can convert 1 gram ≈ 100 J and 1 gram per second ≈ 100 J / second. You may recall from high school that a watt is defined as 1 watt = Joule per second. ↩
Wikipedia quotes a paper saying people make / recycle around 50 kilograms per day. That would imply that people make around 0.5787 grams per second. But this is in tension with the idea that people use 100 watts at rest. Since that 100 watt number seems to be more strongly established, I think 1 gram per second is a better estimate. ↩
Why do you breathe? You breathe because your mitochondria need oxygen to make ATP. When you exercise, you breathe faster so that your mitochondria can make more ATP. You can actually calculate how much ATP your mitochondria make using your VO₂ max score: For each liter of oxygen you use, you make ~21,000 joules of energy, corresponding to ~210 grams of ATP. If you have a typical VO₂ max score of 40 mL/kg/min and you weigh 70 kg, that means you are using 2.8 liters of oxygen per minute, which corresponds to ~588 grams of ATP per minute or ~9.8 grams per second. ↩
A Tour de France cyclist might burn 1500 or even 2000 watts for hours, meaning they are producing ~20 grams of ATP. They can do this because they’ve trained their bodies to have more mitochondria and better oxygen delivery to those mitochondria. But it takes a few seconds for the body to ramp up and start producing this much power. ↩
The “T” in “ATP” is for “triple”, meaning there are three phosphate groups. The “D” is for “di”, meaning there are two phosphate groups. ↩
There’s also stored energy in the form of glycogen. This takes a few seconds to come online, and lasts a few minutes. In a sprint, you aren’t limited by glycogen stores running out, but by having too much acid buildup.
So, effectively, the body has five levels of cached energy:
The mechanical energy stored in the elastic strain of the myosin.
The chemical energy stored in ATP molecules. (Recharges myosin)
The chemical energy stored in (phosph)ocreatine molecules. (Recharges ATP)
The chemical energy stored in glycogen. (Recharges ATP and thus creatine.)
The chemical energy stored in food and fat. (Used to recharge glycogen (food) and ATP (food or fat) and thus creatine.)
It’s 12.5% rather than 16.67% because your short-term energy storage is ~75% phosphocreatine and ~25% ATP, and supplementing creatine does not increase ATP. ↩
I’m not sure to what degree this math actually explains why we see a 5-15% increase in strength in creatine studies versus just being a coincidence. It’s a jump from “X% more short-term energy storage” to “X% increase in max bench press”. ↩
Compared to our evolutionary ancestors, we have long helpless childhoods, low muscle mass, and smaller brains. ↩
I find it amusing that lots of sources refer to extra water retention as a “common side effect” and even report statistics, when a far as I can tell it’s essentially guaranteed by physics. ↩
Tsuji et al. report a phosphocreatine to ATP ratio of 0.77 in grey matter and 2.18 in white matter, Lu et al. reports ~1.45 in entire brains, and Hetherington et al. report 1.0 in white matter, 1.6 in gray matter, and 2.1 in the cerebellum. ↩
Difficulty synthesizing creatine is treated by supplementing creatine. Creatine transporter defect currently has no effective treatment. ↩
After writing this sentence, I later noticed that this experiment had apparently been funded by the Effective Altruism Foundation, Effective Ventures, and personally by (well-known AI alignment researcher) Paul Christiano. ↩
I know this sounds odd, but it’s very common. Everyone wants to establish truth, not just create data so someone else can establish truth. It’s hard to blame them, given their incentives. ↩
Another 2022 meta-analysis by Prokopidis et al. found similar results but apparently has a similar problem. Prokopidis et al. deserve credit for acknowledging the issue and issuing a correction. However, Prokopidis et al. only look at memory, and they seem to be working with before-after scores on the same people, rather than comparisons between the placebo and creatine groups. ↩
If you read anything about health or longevity, you’ll soon find yourself in a world of hazard ratios. Some study might say that eating more fiber might change your risk of dying by a factor of HR = 0.90. Another might say that occasional smoking might change it by HR = 1.30.
But how much should you care about that? Is HR = 0.90 or HR = 1.30 a lot? What if you don’t want to eat more fiber? What if you like smoking?
Instead of staring at a ratio1, a more sensible thing to do is think about life expectancy.2 But is it possible to convert a hazard ratio to a change in life expectancy? You might reason as follows: Baseline life expectancy is around 75 years. And HR = 0.90 corresponds to a 10% decrease in mortality. So perhaps that hazard ratio corresponds to something like 7.5 extra years of life expectancy?
Unfortunately, that’s completely wrong. To see why, imagine that humans only die by playing Russian roulette. They start playing this once per day at the age of 75, with a revolver containing two bullets and six chambers. If you were to remove one of those two bullets, that would drop the person’s risk of death by HR = 0.5. (One bullet versus two.) But life expectancy would barely change, because even with just one bullet, almost nobody would survive for any significant amount of time past 75.
For contrast, imagine again that humans only die via Russian roulette, but now they do this once per day from birth with a revolver with 2 bullets and 54,786 chambers. (Newborns emerge and instinctively reach for this gigantic gun.) You can show that these people also live 75 years on average. But now, if you remove one of the bullets, life expectancy doubles, because when someone is spared, it takes a long time before they get unlucky again.3
Neither of those is a good model for humans. We’re somewhere between the two, with heart disease and so on instead of revolvers and risks slowly rising as we age instead of suddenly starting at age 75 or staying constant throughout life. But you get the point: If you want to convert a hazard ratio for some intervention to a change in life expectancy, the impact depends on how “spread out” baseline mortality risk is over time. Baseline life expectancy is simply not enough information.
That’s one problem. Here’s another: What even is a hazard ratio? The technical definition is something like:
The hazard ratio at a given time is the rate of an event in the treatment group divided by the rate of that event in the control group.
Hazard ratios are often confused with their more beloved siblings, relative risks. Say you run a trial for 10 years and at the end, 10% of the control group died and 8% of the treatment group. Then the relative risk is RR = 0.8, nice and simple. But relative risks have problems, most notably that if you run a long enough trial, then no one will be alive at the end no matter the intervention, meaning RR = 1.0. That’s not helpful. Intuitively, you can think of the hazard ratio at age 40 as sort of like the relative risk for people between the ages of 39.99 and 40.01.
In real life, interventions have different hazard ratios at different ages. Chemotherapy tends to have better results in younger patients who are more able to endure the side-effects. Having a slightly higher BMI (25-30 rather than 20-25) is associated with an increased risk of mortality in young people, but a decreased risk in the elderly. You may remember from 2020 that COVID’s mortality risk had a different age curve than baseline mortality, meaning the hazard ratio of getting COVID was different at different ages.
This is important, because hazard ratios at different ages have different impacts on life expectancy. A hazard ratio of 0.9 at age 80 prevents more deaths than at age 20, because baseline mortality is higher at 80. But at the same time, if you save the life of a 20 year-old, they have more years in front of them. Beyond that, the hazard ratios at different ages interact: If some intervention decreases mortality at younger ages, that allows more people to reach older ages, increasing how much hazard ratios matter at older ages.4
If we knew the hazard ratio at all ages, we could account for those dynamics. But we don’t, because when estimating hazard ratios, people almost always assume that the hazard ratio is constant.5 We’re quasi-forced to do this because there’s not enough data to estimate a whole time-series of ratios. That’s why papers contain single numbers like HR = 0.90.
So even though Intervention A (say, more fiber) and Intervention B (say, light jogging) might have the same hazard ratio in a paper, those numbers could be the product of different underlying age-dependent effects, meaning those interventions could conceivably lead to vastly different changes in life expectancy.
So is this all hopeless? Are single hazard ratio numbers just too far removed from what we care about to tell us anything meaningful?
Surprisingly, no. It’s mostly OK. If we were a different species, it might be hopeless. But for modern humans in rich countries, mortality happens to be distributed in a way that produces a sort of lucky coincidence: When people estimate constant hazard ratio numbers, they’re implicitly sorta-kinda taking a weighted average of hazard ratios at different ages. And those weights happen to (sorta-kinda) reflect how much changes in mortality at different ages change.
So, I will argue, even if the true intervention has a varying effect, it’s sorta-mostly OK to just take a hazard ratio from a paper and convert it to a change in life expectancy using this curve:
If a paper showed that eating more fiber produces a hazard ratio of HR = 0.75, that corresponds to an increase of around 3.7 years. If a paper says that occasional smoking produces a hazard ratio of HR = 1.25, that corresponds to a decrease of around 2.9 years.
This isn’t exact. If the intervention is better (or less bad) for older people this will tends to overestimate the increase (or underestimate the decrease) in life expectancy. If the intervention is worse (or less good) for older people, it will tend to underestimate the increase (or overestimate the decrease) in life expectancy. But as long as the hazard ratio doesn’t vary too much by age, it’s probably not off by more than around 30% in either direction.
The easy case
Say there’s some intervention (eating more fiber or whatever) that multiplies your risk of dying at age t by a factor of HR(t). Then it can be shown that this changes life expectancy by approximately
ΔL ≈ ∑ₜ ΔHR(t) × P(t) × L(t).
Here, P(t) is the baseline probability of dying at age t. For males in the United States, it looks like this:
Meanwhile, L(t) is conditional life expectancy at age t. That’s the average number of additional years left for someone who reaches age t. For males in the United States, it looks like this:
Finally, ΔHR(t) is the decrease in hazard at age t. You can think of that as just ΔHR(t) = 1 - HR(t). Though if you’re OK with logarithms, there’s a somewhat better approximation that uses logarithms, which I’ve quarantined in a footnote.6
Let’s start with the easy case. What if your intervention has the same effect on mortality at all ages, so HR(t)=HR is just a constant? Then, the above equation simplifies into
ΔL ≈ ΔHR × L̄,
where
L̄ = ∑ₜ P(t) × L(t).
This makes sense! Again, P(t) is the baseline probability of dying at age t and L(t) is conditional life expectancy at age t. These are constant, so when you add them up, L̄ is just a number. For males in the United States, it happens to be 12.93 years. This quantity has a specific meaning: The average remaining life expectancy for US males when they die. That sounds a bit odd, but think of picking a random death and asking how many additional years people who reach that age live on average. That number is 12.93 years.
So, if an intervention has a constant hazard ratio, the mean change in life expectancy for US males is just
ΔL ≈ ΔHR × 12.93 years.
Now we’re getting somewhere! If you prevent a fraction ΔHR of deaths, then you increase life expectancy by ΔHR times 12.93 years.
Now remember the naive calculation we started with: Life expectancy for US males is 75.8 years. You might hope that if eating more fiber drops your risk of death by 10%, that would save 7.58 years. Sadly, the above equation says that a 10% drop in risk only increases life expectancy by around 1.293 years—only 0.17 times as much.
This is essentially the observation Keyfitz made in his 1977 paper, “What Difference Would It Make if Cancer Were Eradicated?” Cancer is responsible for 18 percent of deaths, so does that mean eradicating it would increase lifespan by 18 percent, or around 13.6 years? Nope, Keyfitz says, it’s only 2.3 years.
If a cure for cancer were discovered and made available today, 350,000 cancer deaths would be avoided in the next year. The overall death rate would be lower by nearly 18 percent. If the cure were quick and inexpensive, a large fraction of the country’s hospital beds and medical personnel would be released for treatment of other ailments. Patients would be spared untold suffering. Such an implicit analysis underlies government proposals for eradication of cancer. The argument is sound for first effects on mortality but wholly misleading for the long term.
The first effects would soon be offset by more mortality from diseases other than cancer. As a result of the cancer cures, the population would include a higher proportion of people subject to other causes of death. […]
At the extreme, it might be said that everyone dies of something sooner or later, so that, when the effects of the eradication of cancer had shaken down, the same number of deaths would occur as before, and the only benefit would be the substitution of heart and other diseases for cancer. A cure for cancer would only have the effect of giving people the opportunity to die of heart disease.
Cheerful stuff! We can also write our approximation in terms of baseline life expectancy as
ΔL ≈ ΔHR × 0.17 × 75.8 years,
which makes explicit that 12.93 years is only 0.17 times as large as a naive estimate using baseline life expectancy. The discount factor of 0.17 is sometimes called the “Keyfitz entropy”. You can think of it as measuring how close some population is to playing Russian roulette with 2 bullets in 6 chambers starting at age 75 (a discount factor of just above 0) and playing Russian roulette from birth with 2 bullets and 54,786 chambers (a discount factor of 1.0). It’s typically around 0.15 in rich countries today, though it was historically much higher.
Keyfitz entropy is also much higher in other species like mice (perhaps 0.45). You could argue that this explains why nothing that increases lifespan in mice ever translates to humans. Say caloric restriction or whatever produced the same constant hazard ratio in mice and humans. Then it’s mathematically guaranteed that the percentage increase in life expectancy will be three times smaller in humans, because Keyfitz entropy is three times smaller in humans. It’s harder to increase life expectancy when the baseline mortality distribution is more compressed.7
But that’s all assuming the hazard ratio is the same at all ages. Which it surely isn’t.
The interesting case
Here again is our equation for the change in life expectancy in response to taking some action that changes the risk of mortality at age t by a factor of HR(t):
ΔL ≈ ∑ₜ ΔHR(t) × P(t) × L(t),
Basically, for each age t, we multiply together three numbers:
ΔHR(t) is the decrease in the chance of dying at age t as a result of whatever intervention you’ve made (e.g. eating more fiber). This reflects that larger decreases in risk lead to larger increases in life expectancy.
P(t) is the baseline probability of dying at age t. This reflects that the hazard ratio is a ratio, so you prevent more deaths when you apply that ratio to ages where the baseline rate is higher.
L(t) is conditional life expectancy at age t. This reflects that you miss out on more years of life if you die when you’re young.
Now notice: The impact of a change ΔHR(t) at age t is the product of the baseline risk of death P(t) and remaining life expectancy L(t). So what really matters is their product, P(t) × L(t):
This shows how sensitive life expectancy is to changes in hazard ratios at different ages. It would be nice if this were constant. Then, the shape of HR(t) wouldn’t matter at all, only the average value. That’s not quite true, but it’s not terribly far from being true.
An equivalent way of writing our equation for the change in life expectancy is
ΔL ≈ avg(ΔHR) × L̄,
where L̄ is still mean “life expectancy at death” (12.93 years for US males) and avg(ΔHR) is the average change in hazard, weighted by the P(t) × L(t) sensitivity curve at different ages.8 While that sensitivity curve isn’t constant, it’s not too curvy, either. Intuitively, it gives a lot of weight to ages between 50 and 90, somewhat less weight to ages between 20 and 50, and little weight to other ages.9
So that’s not too bad. But let’s remember our original problem: You see some number like HR = 0.90 in a paper, and you want to convert it to a change in life expectancy. If the true underlying hazard ratio were constant, then there’s no problem. But if it’s not constant, then what does that HR = 0.90 number even mean?
Numbers in papers
Unfortunately, you almost never get to see the underlying time-dependent HR(t), because there’s almost never enough data to estimate it. So it’s almost never possible to compute the weighted average avg(ΔHR). In reality what you have is probably a single number in a paper. Let’s call that number est(HR). The obvious thing to do would be to plug the change into the above equation in place of avg(ΔHR) and approximate the change in life expectancy as
ΔL ≈ est(ΔHR) × L̄.
Again, you can just think of est(ΔHR) = 1-est(HR) as being the estimated reduction in hazard. Although, again, I’d prefer you use logarithms if you’re OK with logarithms.10 So the question is: Will that be accurate? How close are est(ΔHR) and avg(ΔHR)?
Well, how do people actually estimate those scalar hazard ratio numbers in papers? Somehow, they’re aggregating together information about hazards at different ages into a single number. But how? Well, it’s complicated. But if there’s a lot of data, you can show that the estimated scalar hazard ratio is approximately11
est(HR) ≈ Πₜ HR(t)ᵖ⁽ᵗ⁾.
(Pardon the hideous typsetting.) That is, the estimated hazard ratio is the geometric average of age-dependent hazard ratios, weighted by the probability of dying at each age. It follows12 that the estimated change in hazard is approximately
est(ΔHR) ≈ ∑ₜ P(t) ΔHR(t).
So ideally, we’d estimate life expectancy using avg(ΔHR), which averages the changes ΔHR(t) based on the weights P(t) × L(t). But we can’t do that, because we don’t have access to the ΔHR(t) numbers. What we can do is read a hazard ratio number in a paper, call it est(HR) and then compute the change est(ΔHR). The above equation says that if you do that, you are implicitly (and approximately) averaging the changes ΔHR(t) based on the weights P(t) alone.
The “right” weights used by avg(ΔHR) and the “wrong” weights implicitly used by est(ΔHR) aren’t the same. But they’re not that different. Here’s P(t) × L(t), the weights that we’d like to use to compute avg(ΔHR) and estimate changes in life expectancy accurately:
And here’s P(t), the weights you’re implicitly using if we take a hazard ratio number from a paper and compute est(ΔHR):
They’re different. In particular, the latter weights give more weight to people aged 80-95 and less weight to people aged 20-50. But they’re not terribly different.
Enough math, let’s try it
To start, imagine some intervention that decreases risk by HR(t)=0.9 for all ages.
Here are the results:
Thing
Formula
Years
Original life expectancy
L
75.7769
New life expectancy
L’
76.4127
Exact ΔL
ΔL = L - L’
0.6358
Ideal approximation
ΔL ≈ avg(ΔHR) × L̄
0.6409
Use number from paper
ΔL ≈ est(ΔHR) × L̄
0.6409
Let me explain what’s happening here. I made a simulator that takes actuarial data for how likely US males are to die at various ages. From this, it’s a simple spreadsheet calculation to compute life expectancy L.13 Then I applied a hazard ratio to change the probability of dying at each age, and re-ran the simulator to compute a new life expectancy L’ and the exact difference ΔL. Then I’m showing two approximations of ΔL: The first is the “ideal approximation” using avg(ΔHR), which I’m including mostly to show that my math is good. Finally, I’m showing the approximation you get if you actually fit a Cox proportional hazards model and use the resulting number in est(ΔHR). This corresponds to what you’d get if you plug in a number from a paper.
So, with the above constant hazard ratio HR = 0.90, both approximations are very good. This remains true if you switch to some other constant.
What if the hazard ratio varies? At first, you might think that something like this would be very problematic:
But it’s basically fine:
Thing
Formula
Years
Original life expectancy
L
75.7769
New life expectancy
L’
77.4373
Exact ΔL
ΔL = L - L’
1.6604
Ideal approximation
ΔL ≈ avg(ΔHR) × L̄
1.7451
Use number from paper
ΔL ≈ est(ΔHR) × L̄
1.7121
The reason this is fine is that the changes in the hazard ratio are relatively “high frequency”, meaning they sort of locally average out. To demonstrate this, suppose the hazard ratio is chosen randomly for each 1-year bin:
Then the approximations are even better:
Thing
Formula
Years
Original life expectancy
L
75.7769
New life expectancy
L’
77.4218
Exact ΔL
ΔL = L - L’
1.6449
Ideal approximation
ΔL ≈ avg(ΔHR) × L̄
1.7059
Use number from paper
ΔL ≈ est(ΔHR) × L̄
1.7123
What causes trouble is if the hazard ratio varies systematically between the young and the old. For example, suppose the intervention is useless for newborns, but gradually becomes more helpful as you get older:
My “ideal approximation” would still be pretty accurate, if you could compute it. (Which you can’t, in the real world.) But using a number from a paper leads to an overestimate:
Thing
Formula
Number
Original life expectancy
L
75.7769 years
New life expectancy
L’
77.9031 years
Exact ΔL
ΔL = L - L’
2.1261 years
Ideal approximation
ΔL ≈ avg(ΔHR) × L̄
2.0962 years
Use number from paper
ΔL ≈ est(ΔHR) × L̄
2.7645 years
This happens because est(ΔHR) is implicitly weighted by P(t) which is heavily weighted towards older people, whereas we’d like to use something more like avg(ΔHR) which is weighted by P(t) × L(t) which is somewhat less weighted towards older people. Even so, the error isn’t terrible.
Now, it is possible that plugging in a hazard ratio from a paper could give wildly inaccurate estimates of life expectancy. One such scenario would be an intervention which is amazing for people aged 85-95, but does nothing for anyone else:
Now, the hazard ratio looks good exactly at the ages where est(ΔHR) has the most weight, leading it to hugely overestimate the impact on life expectancy:
Thing
Formula
Number
Original life expectancy
L
75.7769 years
New life expectancy
L’
76.1741 years
Exact ΔL
ΔL = L - L’
0.3972 years
Ideal approximation
ΔL ≈ avg(ΔHR) × L̄
0.3840 years
Use number from paper
ΔL ≈ est(ΔHR) × L̄
1.0989 years
Another nightmare case is an intervention that starts out harmful, but then switches to being helpful at older ages:
Now, using a number from a paper doesn’t even give an estimate with the right sign.
Thing
Formula
Number
Original life expectancy
L
75.7769 years
New life expectancy
L’
75.5006 years
Exact ΔL
ΔL = L - L’
-0.2764 years
Ideal approximation
ΔL ≈ avg(ΔHR) × L̄
-0.2348 years
Use number from paper
ΔL ≈ est(ΔHR) × L̄
+0.2709 years
That’s bad. But I think most interventions probably aren’t like that? My guess is that most real interventions vary somewhat with age, but they do so gradually and without switching sign. In those cases, it’s quite difficult to find cases where plugging in the number from a paper is off by more than 30% or so. If you don’t believe me, just try it.14
TLDR
If we were another species, it might be very hard to convert from hazard ratios to changes in life expectancy. But for modern people in rich countries, there are three lucky coincidences:
Mortality risk happens to be distributed so that you can approximate changes in life expectancy through a simple weighted sum of hazard ratios at different ages, ignoring interactions.
The statistical method that people use to estimate scalar hazard ratios can also be approximated as a weighted sum of hazard ratios at different ages, ignoring interactions.
The weights that you need to estimate life expectancy (from #1) and the weights that are implicitly used to compute hazard ratio numbers (from #2) aren’t the same. But they’re fairly close.
These facts justify taking an estimated hazard ratio number HR from a paper and approximating the change in life expectancy as ΔL ≈ ln(1/HR) × 12.93 years or, if the hazard ratio is close to one and you hate logarithms, as ΔL ≈ (1-HR) × 12.93 years.
The number 12.93 years is for US males. It’s the product of Keyfitz entropy (0.17) and baseline life expectancy (75.8 years). It will vary a bit in other populations.
If the true underlying hazard ratio:
…is constant across ages, then the above approximation will be extremely good.
…decreases as people get older, that approximation will overestimate ΔL. That is, it will make helpful interventions look better than they actually are, and it will make harmful interventions look less bad than they actually are.
…increases as people get older, that approximation will underestimate ΔL. That is, it will make helpful interventions look less good than they actually are, and it will make harmful interventions look worse than they actually are.
But as long as the true underlying hazard ratio isn’t too crazy, there’s probably not more than ~30% error in either direction.
Finally, two major caveats: First, the above discussion assumes that the hazard ratio was estimated by running a trial on people of all ages. In general, est(ΔHR) implicitly gives weight to different ages proportional to how many deaths occur at those ages in the baseline population in the trial. If there’s a minimum age of, say, 50 years old, that won’t change too much because most of the mass of P(t) is above the age of 50 anyway. But if there’s a minimum age of 70, or a maximum age of 50, that could make a huge difference if the true hazard ratio is different at the ages that weren’t seen.
Second, these are estimates for the life expectancy for a population. But you are not a population. In some sense, your genetics and lifestyle mean you have your own “personal Keyfitz entropy”, reflecting how spread out your mortality would be for you if you led millions random lives. If you drive safely and use an air purifier and eat well and get exercise and don’t smoke, that likely means your personal life expectancy is higher than average. But it also probably means that your personal Keyfitz entropy is lower than average.15 So, if you make your lifestyle even better by eating more fiber or whatever, even if that produces the same hazard ratio for you as for other people, it would still likely lead to smaller increases in life expectancy, for the same reason that the same hazard ratio produces smaller changes in lifespan in humans compared to mice. What we really need is some interventions strong enough to break the math behind these approximations and free us from Keyfitz tyranny.
I know, I know, you care about quality of life, not just years of life. I agree, some number that measures health and vitality, maybe disability-adjusted life years or quality-adjusted life years, would be better. But these are hard to estimate and so are rarely reported. Anyway, in practice most interventions that make you more vital tend to make you live longer and vice versa, so focusing on life expectancy isn’t too bad. ↩
In this model, the number of days of life follows a geometric distribution with p = (number of bullets) / (number of chambers). So the mean life expectancy is 1/p days or (number of chambers) / (number of bullets) days. With 54,786 chambers and 2 bullets, that works out to 75 years. And if you drop down to one bullet, then it increases to 150 years. ↩
If some intervention would have reduce mortality among people aged ≥ 60 in prehistorical tribal bands, that wouldn’t have increased life expectancy very much, because most people didn’t make it to 60. But compared to prehistorical tribal bands, we have in fact vastly reduced mortality at younger ages. And so, today, reducing mortality for people aged ≥ 60 will increase life expectancy a lot. ↩
You might think this is stupid. Why change a relative risk into a hazard ratio if you’re just going to assume it’s constant? Isn’t that pointless? Well, no. Remember how relative risks always go to 1.0 for long enough trials as everyone in both the treatment and control groups departs our coil? That doesn’t happen with constant hazard ratios. ↩
It’s usually (though not always) better to use ΔHR(t) = ln(1/HR(t)). This correctly reflects, for example, that if all hazard ratios go to zero, then life expectancy goes to infinity, yay. These two approximations are almost identical for hazard ratios that are close to one because ln(1/r) ≈ (1-r) when r is close to one. So if you are terrified of logarithms but you’ve made it to the end of this footnote anyway, you’re not missing out on too much. ↩
There’s a degree of circularity to this argument. It assumes that hazard ratios transfer better between species than changes in life expectancy. That might be true, but it would be an empirical / biological fact, not something that’s guaranteed by logic. ↩
To see this, note that ΔL ≈ ∑ₜ ΔHR(t) × P(t) × L(t) = L̄ × ∑ₜ ΔHR(t) × (P(t) × L(t) / L̄) = L̄ × avg(ΔHR). ↩
where avg₂₀₋₅₀(ΔHR) represents a flat average of the change over the ages 20 to 50 and avg₅₀₋₉₀(ΔHR) represents a flat average over the ages 50 to 90. ↩
That is, it’s better to use est(ΔHR) = ln(1/est(HR)). This is close to 1-est(HR) when est(HR) is close to one. ↩
If there is an infinite amount of data, the typical method reduces to solving
∑ₜ (P(t) + P’(t)) × π(t, HR) = ∑ₜ P’(t),
for HR. Here, P’(t) is the chance of dying at age t after the hazard ratio has been applied, and π(t, HR) is the probability that, if a death occurred at time t, it was in the treatment group. Of course, the true probability that a death is in the treatment group is P’(t) / (P(t) + P’(t)). The standard “proportional Cox” model assumes that the hazard ratio is constant and so replaces this raw fraction with a model-based one, namely
π(t, HR) = S’(t) × HR / (S(t) + S’(t) × HR).
This reflects the fact that at age t, a fraction S(t) of controls are alive and each of these have some chance μ(t) of dying, so P(t)=S(t) × μ(t). Meanwhile, a fraction S’(t) of the treatment group is alive, and these each have a chance HR × μ(t) of dying, meaning that P’(t) = S’(t) × HR × μ(t). If you substitute these equations for P(t) and P’(t) into the second equation above, the factor of μ(t) conveniently cancels and you get π(t, HR) as written.
In effect, the hazard ratio’s job is to attribute deaths to the treatment versus the control group. Now, if the true time-varying HR(t) is close to one, then it can be shown that the estimated hazard ratio est(HR) approximately satisfies
The geometric average is equivalent to the condition that
ln(est(HR)) ≈ ∑ₜ P(t) ln(HR(t))
Using the “better” approximation that ΔHR(t) = ln(1/HR(t)) and *est(ΔHR)=ln(1/est(HR)), it follows that
est(ΔHR) ≈ ∑ₜ P(t) ΔHR(t).
You can justify interpreting that same equation using est(ΔHR) = 1-est(HR) and ΔHR(t)=1-HR(t) from the fact that these are almost the same when HR(t) is close to one. ↩
This simulator pretends that people live for integer numbers of years. That’s not true in reality, of course, but it makes the simulator easier to implement and understand and makes little difference in practice. ↩
In the simulation, “true ΔL” is what I called “exact ΔL” above, while “approximation (log)” is what I called “ideal approximation” and “Cox fitted” is what I called “Use number from paper”. ↩
The way modern human mortality is distributed, even if your healthy lifestyle were to reduce mortality by a constant factor at all ages, that still has the effect of decreasing Keyfitz entropy. ↩
For a while there, many people thought vitamin D was magical—that it could improve bones, the heart, infections, cancer, heart disease, longevity, even mental health. But among people I respect, opinion is now overwhelmingly that taking vitamin D does nothing unless you’re severely deficient. The central argument is that while vitamin D levels are correlated with ~all positive health outcomes, when you actually test vitamin D supplements against placebo in randomized trials, nothing ever happens.
That’s what I used to think, too. But I’ve come to think the skeptics have over-corrected. Yes, randomized trials have shown that the magical correlations are not causal. But if you start with non-insane expectations, the trials look like weak but positive evidence. And if you consider what we know about biology and evolution, I think the balance of evidence tips pretty clearly in the direction that people with low-ish levels would be wise to supplement.
Am I certain that vitamin D is beneficial for people with low-ish levels? Absolutely not! But I claim that’s the best bet given the limits of our knowledge.
The classical view: Boring bone vitamin
Most vitamins are “ingredients” that the body uses to do stuff. Vitamin D is more like a “signal” that the body uses to communicate with itself about what to do.1 The classical “endocrine” story of vitamin D is that your body uses it to tell your guts to take in more calcium from food. If you don’t get enough vitamin D, then you have calcium problems.
That’s all you really need to know about the classical view. But if you enjoy gawking at biology’s complexity, I recommend this diagram and the following three paragraphs:
Ready for science? OK: Almost all the cells in your body make provitamin D.2 Usually, this is all converted to cholesterol, but your skin cells leave some sitting around. When UVB light hits those skin cells, provitamin D is transformed (physically by the light itself) into previtamin D and then (by heat) into vitamin D. This diffuses from the skin cells into blood vessels. There it binds to a protein3 and starts circulating in the blood, where it is joined by vitamin D from food.4 Eventually, the liver converts it into more-stable storage vitamin D. It also soaks in and out of fat and muscle tissue, which acts as a slow-release reservoir.
Now, a fun fact: If calcium levels in your blood get too low, then your heart will stop working and you will die. To avoid this, you have parathyroid glands in your neck that sense when calcium is getting low, and release parathyroid hormone into the blood. This tells your bones to release some of their stored calcium. It also tells your kidneys to convert some of the storage vitamin D from your blood into active vitamin D. And when that gets to your guts, they try to absorb more calcium from food.
So what happens if you don’t get enough vitamin D? Well, your body is not going to let calcium levels drop too low, because your body is designed to avoid death. Parathyroid hormone will still get secreted, and calcium will still get scavenged from your bones. But without vitamin D, your guts never get the signal to gather extra calcium from food. So the body scavenges a lot of calcium from your bones, and you end up with weak bones, which is bad.
Now here’s the thing: In this story, only active vitamin D actually does anything. The kidneys make this on demand in response to calcium levels, not in response to storage vitamin D levels. General opinion is that as long as the blood has above ~25 nmol/L of storage vitamin D, then the kidneys have no trouble making active vitamin D.5 Furthermore, survey data suggests that only ~2% of the population has levels below that threshold. This suggests that for ~98% of people, supplementing vitamin D should do approximately nothing.
The correlation view: Magical mystery cure
Rickets is a terrible disease that involves soft bones, stunted growth, and skeletal deformities. It’s probably been with us since ancient times, but it became common in the West after the industrial revolution. In 1890, a Scottish missionary named Theobald Palm observed that rickets was common in smog-ridden UK cities but almost unheard of in sunny countries with poor sanitation, suggesting sunlight itself was the issue. This contributed to the discovery that rickets could be cured with UV light or cod-liver oil, and eventually the discovery of vitamin D.
In 1941, Apperly noticed that the amount of sunlight in different US states was positively correlated with skin cancer but inversely correlated with overall cancer mortality.6 He gave this charming graph:
Apperly never mentions vitamin D, presumably because he thought it was a boring bone vitamin.
They point out that regional diets (like meat and fiber) didn’t seem to explain this pattern. Instead, they propose a mechanistic story:
Sunlight
↓
Vitamin D
↓
Adequate calcium in blood
↓
Reduced inflammation of epithelial cells in the colon
↓
Less colon cancer
(It’s always inflammation.) This paper was rejected many times before finally being published. I wish I could find an un-gated copy to link to, because it would have made a magnificent blog post.7
Following that paper, there was an explosion of work that found negative correlations between sunlight (or latitude) and other types of cancers as well as blood pressure, diabetes, and multiple sclerosis.
Then people started measuring vitamin D in blood. In 1989, the Garlands and collaborators found blood samples takin in 1974 from 25,000 people. They found that 34 of those people had since gotten colon cancer. They matched these with 67 demographically similar people and measured vitamin D levels in the stored blood samples for all 101 people. Among that group, people with vitamin D levels below 50 nmol/L got colon cancer more than three times as often as people with higher levels.
Again, many similar studies followed. These linked higher vitamin D levels to better outcomes in cardiovascular disease, diabetes, obesity, infectious disease, Parkinson’s, and mood disorders. While results were mixed for non-colorectal cancer incidence, higher vitamin D levels predicted better survival of many cancers. Amazingly, all-cause mortality was roughly 30% lower for those at the 75th percentile of vitamin D levels compared to the 25th.
Vitamin D was looking like a miracle. But how could it do all that stuff if it was just a boring bone vitamin?
Meanwhile in biology
While all these correlations were being discovered, we learned that the body doesn’t just use vitamin D for bone stuff.
In 1969, we discovered the vitamin D receptor that active vitamin D binds to in the gut and bones. And in the 1980s came a shock: Almost all cells in the body have vitamin D receptors. These seem to do different things in different tissues. In the pancreas, they support insulin secretion. In immune cells, they boost antimicrobial peptides and reduce inflammation. In neurons, they influence proliferation and differentiation.
So… What? When calcium drops and the kidneys put out active vitamin D, does every part of the body start doing different unrelated stuff?
In the late 1990s, we cloned the gene for the enzyme that the kidneys use to convert storage vitamin D to active vitamin D. Soon came another shock: This enzyme also exists in tons of other cells, including immune cells, the heart, the skin, the prostate, the breast, and colon. (Another win for the Garlands.)
So it’s not just the kidneys making active vitamin D to trigger the gut. Cells everywhere are making their own active vitamin D and using it to trigger vitamin D receptors in neighboring cells, or even inside the same cell.8 This often has little to do with calcium or bones.9
So:
The kidneys use vitamin D as a boring bone hormone.
As long as the blood contains at least ~25 nmol/L of storage vitamin D, the kidneys don’t care. They create the same amount of active vitamin D, in response to calcium levels.
But now cells everywhere are using storage vitamin D.
To do god-knows-what.
With god-knows-what sensitivity to circulating vitamin D levels.
And remember how only active vitamin D does anything? That’s wrong. In the mid-1970s, we learned that storage vitamin D also binds to the vitamin D receptor. The binding affinity is 100-1000× lower, but you have ~1000× more in your blood. So maybe circulating levels of storage vitamin D themselves matter, independently of how much active vitamin D gets made?
If that’s not confusing enough, people also noticed that while active vitamin D levels in the blood aren’t correlated with storage vitamin D (above ~25 nmol/L), levels of parathyroid hormone (the thing your parathyroid glands use to tell your kidneys to make active vitamin D) seem to decline as levels of storage vitamin D rise from ~25 to 50 or 75 nmol/L. Huh?10
On the one hand, all this makes the idea that vitamin D could be a miracle more plausible. On the other hand, this is getting complicated. And do we really believe that raising your vitamin D levels from the 25th to the 75th percentile could reduce your risk of death from any cause by thirty percent? Maybe we should try giving people vitamin D and see what happens.
Then came the RCTs
There have been many randomized trials. The “right” thing to do in such cases is to look at meta analyses that carefully combine all the data. We’ll get to those. But they conceal a lot of important nuance about what actually happens on the ground during these trials. So let’s start by going over the three main “megatrials”.
The Women’s Health Initiative (WHI) trial came out in 2006 and is still the largest vitamin D trial ever done. This took 36,000 postmenopausal American women and assigned half to take 400 IU daily with calcium and the other half to placebo.11 After seven years, here’s what happened:12
Outcome (WHI trial)
Hazard ratio
Fractures
0.97 (0.91 to 1.03)
Cancer
0.97 (0.91 to 1.04)
Cancer mortality
0.90 (0.77 to 1.05)
CVD mortality
0.94 (0.78 to 1.12)
All-cause mortality
0.92 (0.83 to 1.01)
Kidney stones
1.17 (1.02 to 1.34)
(The hazard ratio is the ratio of the rate that something happens in the treatment vs. placebo groups. So, a number less than one suggests a benefit to taking vitamin D, while a number larger than one suggests a harm. The numbers in parentheses show a 95% confidence interval.)
The only statistically significant result was a bad one: Extra kidney stones, likely from the extra calcium.13 The other outcomes look vaguely good, but none were statistically significant despite the massive sample size.
This was disappointing. However, the WHI trial had limitations: Many subjects in both the vitamin D and placebo groups were already taking vitamin D, and continued taking it through the trial. The dose of 400 IU was fairly low, many subjects stopped taking their pills, and vitamin D levels didn’t actually change that much. They also measured vitamin D levels in only 6% of subjects, meaning we can’t compare the fates of subjects who started out with low versus high levels.
The next big hope was VITAL, which came out in 2018. They recruited 26,000 older people across the United States, half of them men and 20% Black (and thus far more likely to be vitamin-D deficient). They measured vitamin D levels in most people, and they gave the treatment group 2,000 IU per day.14 Here were the results after 5.3 years:
Some of the results look good-ish, but cardiovascular mortality was higher in the treatment group, leading to almost no effect on all-cause mortality.15 More disappointment.
The last megatrial was D-Health, which came out in 2022 based on 21,000 older Australians. Instead of daily supplements, it used a monthly “bolus” dose of 60,000 IU or placebo. Unlike in VITAL, there was no exclusion for people with a history of cardiovascular disease or cancer, and less restriction on how much vitamin D participants could take on their own during the trial.16 Here were the results after 6 years:
Outcome (D-Health trial)
Hazard ratio
Cancer mortality
1.15 (0.96 to 1.39)
Major CVD event
0.91 (0.81 to 1.01)
CVD mortality
0.96 (0.72 to 1.28)
All-cause mortality
1.04 (0.93 to 1.18)
Now, the treatment group did better in terms of cardiovascular disease, but worse in cancer and worse in all-cause mortality. Even more disappointment.
Just from these three large trials, the main lesson should already be clear: Vitamin D is not a miracle. The correlations were wrong.17 There is essentially zero remaining hope that taking vitamin D could reduce all-cause mortality by a third.
In this sense, the vitamin D skeptics are definitely right. But what about the other trials? And is there a more subtle lesson?
I made some tables
I wanted a big table that summarized all the major vitamin D RCTs and what they found for different health outcomes. Annoyingly, no such overview appears to exist. So I made my own:18
Lots of the hazard ratios are less than one, suggesting a benefit to supplementation. But lots of them are also higher than one, suggesting a harm. The numbers that are far from one almost always come from smaller trials, which manifest as larger confidence intervals. If you’re interested in the details of how these trials were run, I refer you to more gigantic tables in a footnote.19
If big tables aren’t your thing, here are some formal meta-analyses, both some recent ones and an older but more comprehensive Cochrane review:
There are various ways you could try to squint at these RCT. In almost all of them, most people already had pretty high levels before they started. So why don’t we separate out people who started low? Usually we can’t, because most trials didn’t measure baseline vitamin D.20 And among the trials that did, there are few people with low levels, so the results are noisy and confusing.21
Or, you might theorize that benefits would take time to show up, meaning the first couple years just add noise. In some cases—notably VITAL—excluding the first two years seems to help, but in other cases things get worse.22
Finally, some people speculate that taking gigantic monthly or quarterly “bolus” doses of vitamin D might be dangerous. For example, here’s an enjoyable paragraph from Kunzia et al. in their meta-analysis of vitamin D and cancer mortality:
Our results showing efficacy of daily, but not bolus, vitamin D3 supplementation in reducing cancer mortality are consistent with previous meta-analyses on cancer mortality or all-cause mortality (Guo et al., 2022; Keum et al., 2022; Keum et al., 2019; Zhang et al., 2022; Zhang et al., 2019). However, by including more trials than these previous meta-analyses, we were able to detect statistically significant effect modification by treatment regimen for the first time with statistical significance (pinteraction=0.042). The pattern of intake could be important for a favourable steady state of the bioavailability of the active 1,25 (OH)₂D hormone. Daily administration counteracts the fast excretion of vitamin D from the circulation (Hollis and Wagner, 2013; Keum et al., 2022). Moreover, the enzymes CYP27B1 (converts 25(OH)D to 1,25 (OH)₂D) and CYP24A1 (inactivates 25(OH)D and 1,25(OH)₂D) follow first-order reaction kinetics (Vieth, 2009). This means that doubling the concentration of the precursor doubles the yield of the product, unlike other steroid hormones (e.g., cortisol, oestrogen, testosterone) that follow zero-order kinetics (Vieth, 2020). Intermittent, non-physiologically large vitamin D3 bolus doses may lead to unstable cycling of 25(OH)D and 1,25(OH)₂D levels in blood because the system needs time to adapt to the large doses (Hollis and Wagner, 2013; Keum et al., 2019; Vieth, 2020). In the long run, intermittent bolus regimens at weekly or larger intervals can lead to an up-regulation of countervailing factors (e.g., 24-hydroxylase (CYP24A1), 24,25(OH)2D and fibroblast growth factor 23), all of which ultimately leads to lower synthesis or higher degradation of 1,25(OH)₂D levels (Mazess et al., 2021). Bolus doses, unlike daily doses, failed to reduce C-reactive protein response and actually elevated anti-inflammatory cytokines and doubled the risk of hypercalcemia in previous studies (Krishnan et al., 2012; Martineau et al., 2017; Mazess et al., 2021).
Oh no, up-regulation of fibroblast growth factor 23!23
I don’t feel like I understand this deeply enough to have any opinion beyond the surface level that the body seems to adapt to large doses of vitamin D in ways that could possibly be bad.24 It seems intuitive that small daily doses would be safer than gigantic monthly doses, but I’m always suspicious of post-hoc mechanistic speculation. Also, if people get enough sun, they can apparently synthesize 10,000-25,000 IU per day, which isn’t that far from the 60,000 IU they got in the D-Health trial. But then again, I think Kunzia et al. are suggesting that the body is designed to adapt to regular exposure to large doses but not intermittent exposure?
Well, if you split up the trails by daily vs. bolus dosing, there’s a decent pattern of daily dosing leading to better results:
If those bolus dosing trials didn’t exist, I’d think this looked pretty good. So, maybe? Or maybe this is a story made up to hallucinate a positive trend. I would lean towards the latter theory, but there are papers like Mazess et al.’s “Vitamin D: Bolus is Bogus”, that suggested this pattern before D-Health’s dismal results came out. There are even some trials that suggest bolus doses don’t even work for treating rickets. So… I’m still not convinced. But maybe.
Aside: There are also many Mendelian randomization studies that look at correlations between health and genes that are related to vitamin D. But I don’t think these provide much information, because the assumptions are shaky and the genes don’t explain much of the variance.25
Where are we?
Still with me? Here’s a summary of the above 5200 words:
The body uses vitamin D in all sorts of weird and complicated ways. It’s biologically plausible that vitamin D could matter beyond bone stuff with severe deficiency, but there’s no convincing mechanistic evidence that it is.
Vitamin D levels are strongly correlated with good health outcomes, but RCTs have conclusively shown that most of these correlations are non-causal.
RCTs haven’t conclusively shown any benefit for anything beyond bone stuff. At best, they’ve given weak evidence for hazard ratios slightly below one.
So you might be wondering: Isn’t that quite weak? Wasn’t this post supposed to be a defense of vitamin D?
The case for supplementing anyway
It’s biologically plausible that vitamin D is good
Everyone agrees that severe vitamin D deficiency (below ~25 nmol/L) is bad. It leads to rickets, adult rickets, osteoporosis, muscle weakness or even—with profound deficiency—to seizures or cardiac arrhythmia. This makes sense, because below ~25 nmol/L, the kidneys have trouble converting storage vitamin D into active vitamin D, meaning you don’t absorb enough calcium from food.
The question is if taking supplement to further raise your levels (say to 50 or 90 nmol/L) is important. We have no mechanistic proof, but it might be true, because many parts of the body use vitamin D as a local signal and because cells are at least somewhat sensitive to circulating storage levels. There’s also this weird thing where parathyroid hormone continues to decline as vitamin D levels rise above ~25 nmol/L even while this seems to make little difference to how much active vitamin D the kidneys make.
Nothing in this world comes without trade-offs. Surely, supplementing vitamin D comes with some downsides. But it seems very unlikely that raising vitamin D levels to a “normal” level would cause more harm than benefit. Especially because…
Meanwhile, Wahl et al. 2012 try to estimate mean levels around the world today:
This map looks weird because of varying lifestyle, diet, supplementation, and needing to combine fragmented studies. But you get the idea. And remember, those are just averages. So there are lots of people with levels far lower than that in our evolutionary history.
Of course, just the fact that vitamin D levels have dropped doesn’t mean it’s important. Parasitic worm load, wood smoke inhalation, and cousin marriage have also dropped, but we aren’t rushing to restore those to ancestral levels.
But there’s another piece of evidence: After humans migrated out of East Africa, some of them evolved pale skin. Pale skin is bad, because it allows light to destroy folate, which is crucial for pregnancy.26 Evolution doesn’t typically do things that harm fertility, because evolution wants to increase reproductive fitness. The most common explanation is that pale skin allows more UV light to penetrate, and thus allows people to synthesize more vitamin D. If evolution was willing to pay the high “price” of folate destruction for more vitamin D, that seems like good evidence that vitamin D is important.
Some even see contrasts like the Inuits versus Scandinavians as a kind of natural experiment: They lived at similar latitudes, but Inuits ate a diet with vitamin D (fatty fish and whale blubber) and Scandinavians didn’t. The result is that Inuits have darker skin than Scandinavians.27
This is all speculative, and even if true, might be driven by severe deficiency and rickets. Or perhaps prehistoric benefits don’t translate to your lifestyle. But all the people in Luxwolda’s sample in East Africa had levels above ~60 nmol/L. I just don’t see how you can look at this and not see it as providing some suggestive evidence in favor of the idea that raising levels above severe deficiency is unlikely to be harmful, and could be important. So I think the prior is favorable.
What do you expect from vitamin D?
A hazard ratio like HR = 0.96 doesn’t look very impressive. But hold on. Suppose that life expectancy is 80 years and that taking vitamin D every day reduces your risk of all-cause mortality by a factor of HR. A reasonable approximation in rich countries is that this would increase your life expectancy by
80 × 0.15 × (1-HR) years = 12 × (1-HR) years,
where 0.15 is derived from the entropy of lifespan in rich countries.28 For example, if all-cause mortality had a true hazard ratio of HR = 0.96, then taking vitamin D every day of your life would increase life expectancy by around
0.48 years.
I claim that this would be a lot. Certainly, if I were about to face my destiny, I would pay a lot of money for an extra 0.48 years. Or, you can calculate that this corresponds to an increase of life expectancy per-vitamin-D-pill of 8.6 minutes.29 A common rule-of-thumb is that smoking a cigarette costs around 11 minutes of life in expectation. If you think HR = 0.96 is trivial, do you also think that smoking one cigarette each day is fine?30
The correlational studies suggested that vitamin D might drop your risk of all-cause mortality by a third. It’s disappointing that the RCTs refuted this. But those correlational studies were crazy. They imply31 an increase of life expectancy of around 4 years or around 6.5 cigarettes per day. Could we really believe that you could smoke 6.5 cigarettes, then take a vitamin D pill, and you’re even?
Personally, I think hazard ratios just slightly less than one are the best we can reasonably hope for. But I also think that they would be an excellent return on investment. Arguably, modern human life expectancy comes from stacking lots of modest hazard ratios on top of each other.
What do you expect from vitamin D trials?
Let’s play a game. Let’s hallucinate some numbers for what vitamin D might do, and then simulate what trials would show. Here are the strongest effects I consider plausible for different baseline levels, along with how common those levels are in the United States.
Storage vitamin D (nmol/L)
Hazard ratio
% of population
<30
0.75
5
30-49
0.92
15
50-125
0.98
72.5
>125
1
7.5
Suppose that were real. Now, say we pick 26,000 people at random, and give half of them vitamin D for five years. Here are the results of a million simulated trials, assuming a baseline mortality risk of 0.7%:32
Overall, 9% of trials would find a significant benefit, 63% would find a non-significant benefit, 27% would find a non-significant harm, and 1% would find a significant harm.
If you wanted to have an 80% chance of finding a significant decrease, you’d need to run a trial with something like 570,000 people, almost five times more than in all the above trials combined.33 If you don’t like my numbers, I’ve put up a page where you can run your own simulations with different ones.
My point is, the results we see in vitamin D RCTs are what we should expect to see if vitamin D had plausible benefits. That’s not proof, of course—just that if you start with realistic expectations, the trials don’t provide much evidence in either direction.
The trials do find slightly helpful numbers
Recent meta-analyses have not consistently found a statistically significant benefit to vitamin D supplementation. But they do suggest a small benefit for cancer mortality and all-cause mortality, and they’re close to being statistically significant. That’s something.
And if you buy the argument that bolus dosing is bad, the results get even better. Kunzia et al. did a meta-analysis of cancer mortality using only trials with daily dosing, and found a hazard ratio of 0.88 (confidence interval 0.78 to 0.98). I’d keep this at arm’s length. The bolus dosing trials might have done worse by random chance, meaning this is a kind of p-hacking. But there’s a reasonable chance (maybe 25-50%) that bolus dosing really is bad, in which case those trials would be convincing evidence.
I actually think it’s surprising that the meta-analyses look as good as they do, because there just aren’t that many people who started out with low vitamin D levels. Only a handful of trials had mean levels below 60 nmol/L, and they all give semi-promising results:34
Again, it’s dangerous to dig too deeply looking for these kinds of patterns. If you dig enough, you can always find a way to confirm whatever theory you want. But also again, maybe?
You’re probably already taking vitamin D
You might not personally supplement vitamin D. But for most people reading this, someone else is supplementing it for you.35
Country
Commonly fortified with vitamin D
Australia
Margarine
Belgium
Margarine
Canada
Milk, margarine
Chile
Milk, flour
Ethiopia
Oils
Finland
Milk, yogurt, margarine
Ireland
Margarine, cereal
New Zealand
Margarine (from Australia)
Norway
Margarine, low-fat milk
Pakistan
Oils
Poland
Margarine
Sweden
Milk, yogurt, plant milk, margarine
United Kingdom
Margarine, cereal
United States
Milk, plant milk, margarine, cereal, yogurt
Fortified food is common across the Anglosphere and Scandinavian peninsula. However, it’s rare in the rest of Europe (exceptions: Belgium, Poland) and even-more rare in the rest of the world (exceptions: Chile, Ethiopia, Pakistan).
I think this is important for two reasons. First, vitamin D is oddly self-defeating. There are some places in the world where people care about vitamin D. These are the places that run large trials. But these places also fortify their food and tend to be full of people that already supplement vitamin D. These places also tend to believe it’s unethical to tell the control group not to take vitamin D.
And here’s another question: If you think vitamin D is worthless, are you comfortable recommending removing vitamin D from food? If not, then why is the particular amount of fortification in food now the right one?
Some might argue that the purpose of fortification is to reach the severely deficient, or children, the elderly or pregnant mothers. Maybe! But again, if you could press a button and remove fortification from everyone else, would you feel comfortable pushing that button? Remember, trials don’t test going down from current levels, only going up.
So that’s my story
Biology and evolution suggest a prior that moderate levels of vitamin D (say 80 nmol/L) are quite possibly better than low levels (like 40 nmol/L) and unlikely to be worse.
Observational studies say that vitamin D is magical, but those studies are bad and we should ignore them.
The RCTs show that vitamin D is non-miraculous. But beyond that they don’t provide much information, because they mostly enrolled people with moderate vitamin D levels, meaning plausible effects would require colossal sample sizes to reliably detect.
What evidence the RCTs do provide points weakly towards a modest benefit.
If real, that benefit would far exceed the cost of taking vitamin D.
Therefore, if you have low vitamin D, it seems wise to supplement.
This is all very weak, I know! But sometimes weak evidence is all we’ve got.
I wish we had at least one large trial done in a population with low starting levels. But as far as I can tell, none are underway. In fact, it’s unlikely that there will be any more large trials anytime soon. So weak evidence is how it’s going to be.
Technically, vitamin D itself is a type of steroid although not what people usually mean by “steroid”. ↩
Here are some of the fancy names for the different forms of vitamin D I’ll talk about:
If you eat mushrooms or yeast, it joins the vitamin D from your skin en route to your liver. If you eat animals or animal products, you also get some storage vitamin D, which doesn’t need to be processed by the liver. ↩
Storage vitamin D is what your doctor measures in your blood test. This is sometimes measured in nmol/L and sometimes in ng/mL. The latter measurement is smaller by a factor of 2.496. So 25 nmol/L ≈ 10 ng/mL. ↩
Apperly was building on a 1937 paper that observed observed that sailors, exposed to lots of sunlight, had much higher skin cancer rates than the general population, but lower overall cancer rates. ↩
In Biologist, active vitamin D is not just an “endocrine” hormone that sends signals for far away cells through the blood, it’s also a “paracrine” or “autocrine” hormone that sends signals to nearby cells or inside a single cell, through diffusion. ↩
You might ask, why is vitamin D used by so many different parts of the body for so many different purposes?
I think there’s no deep answer here. It’s true for the same reason that dogs sneeze to signal that they’re feeling playful: Evolution re-uses stuff for different purposes all the time. Imagine that DNA already exists coding for the vitamin D receptor and for the enzyme to convert storage vitamin D into active vitamin D. If some cells need to send a local signal, re-using those is easier than inventing something new. There’s nothing unusual or magical about this. ↩
Don’t try to make sense of this. It doesn’t make sense.
You could speculate that this is because the parathyroid glands are trying to make less active vitamin D to compensate for the fact that vitamin-D receptors throughout the body are sensitive to storage vitamin D itself. But I advise against. ↩
The WHI trial was a pioneer in salami-slicing results for different outcomes into dozens of different papers, most of which are hard to access. All trials now seem to have adopted this hideous trend which makes it maddening to try to summarize what actually happened in a trial. Also, slightly different numbers for the same quantity appear in different places. I haven’t bothered to chase these down, because the differences are all very small, e.g. a hazard ratio of 0.89 for cancer mortality rather than 0.90. ↩
Half of the vitamin D group and the placebo group also got omega 3. These are averaged together in the results. Also, VITAL carefully stratified the assignment to vitamin D or placebo based on baseline vitamin D levels, which should give more statistical power from a given sample size. ↩
There was also a weird study done on a subset of 1031 people from the VITAL population that looked at telomere length. After starting with around 8700 base pairs, the control group lost around 160 base pairs during the study, while the vitamin D group only lost an average of 20. I’m not sure of what to make of this. For one thing, though the authors claim this is statistically significant, it depends on how you analyze the data. But beyond that, sure, telomere length is a marker of aging, but telomeres get shorter for a reason (likely to fight cancer) and it isn’t obvious that slowing this would always be a good thing. ↩
This is a little complicated. In VITAL, participants were only eligible if they were taking at most 800 IU per day, and they were restricted to 800 IU per day during the trial. In D-health, participants were only eligible if they were taking at most 500 IU per day, but they were allowed to take up to 2000 IU per day during the trial. ↩
You might ask: If vitamin D only has a modest effect, then why is it so strongly correlated with health?
In principle, I’d like to push back against the idea that we need to explain why these particular correlations don’t imply causation. But the accepted explanation is a combination of (1) reverse causation where being healthy causes people to spend more time outside and thus get more vitamin D; (2) confounding, where obesity is bad for you and leads to lower measured vitamin D levels; (3) confounding, where more healthy lifestyles lead to both more vitamin D and more health; and (4) confounding, where higher socioeconomic status leads to both more vitamin D and more health. You might ask why these correlations would be true at a state level like the Garlands looked at, but then you run into the ecological fallacy and modifiable areal unit problem. ↩
I took all the trials that got at least 2% weight and were rated as “low risk of bias” in this 2014 Cochrane review of vitamin D and mortality, then manually added all the “major” trials that were published after 2014.
I shudder to think of the time it took to make this table. I tried using AI but found it was wildly unreliable. Part of the problem is that each trial’s results are distributed among many papers, in different journals, with different paywalls. And many details aren’t published at all by the original authors but are only scrounged up and put in the depths of the supplementary material of a review years later. In some cases, different sources also give contradictory numbers. The differences were always tiny (e.g. 0.90 rather than 0.89) but it still makes me nervous. ↩
Here’s a table describing the major contours of the trials:
Among the major trials, only VITAL, ViDA, and FIND measured it for more than a tiny number of subjects. ↩
In VITAL and ViDA, people with baseline levels below 50 nmol/L had a higher hazard ratio for cancer mortality (though with wide confidence intervals), suggesting if anything less benefit. Or, you could use race as a proxy for baseline vitamin D. But in both VITAL and WHI, the hazard ratio for cancer mortality was higher among non-Whites. After looking at many such analyses for many outcomes, the only clear result I could find was for diabetes in the D2d trail, where the hazard ratio was much lower for people below 30 nmol/L (0.38 vs. 0.93). ↩
The results for VITAL look decent:
outcome (VITAL trial)
HR
HR excluding first two years
Cancer
0.96 (0.88 to 1.06)
0.94 (0.83 to 1.06)
Cancer mortality
0.83 (0.67 to 1.02)
0.75 (0.59 to 0.96)
Major CVD event
0.97 (0.85 to 1.12)
0.93 (0.79 to 1.09)
All-cause mortality
0.99 (0.87 to 1.12)
0.96 (0.84 to 1.11)
But in D-Health, excluding the first two years actually increased the hazard ratio for cancer mortality from 1.15 (0.96 to 1.39) to 1.24 (1.01 to 1.54). Most other trials were too short for this kind of analysis to make sense. ↩
That could downregulate 25-hydroxyvitamin D 1-alpha-hydroxylase, reducing the rate it catalyzes the hydroxylation of hydroxycholecalciferol into 1,25-dihydroxycholecalciferol! ↩
Dynomight: WTF is this?
Dynomight Biologist: Well, C-reactive protein is generally considered inflammatory.
Dynomight: So reducing that is good? But then why do they talk like elevating anti-inflammatory cytokines would be bad?
Dynomight Biologist: Yeah… That would be good. Unless you have cancer. In which case it’s not good.
Mendelian randomization studies are based on the idea that certain genes predispose you to have higher levels of circulating vitamin D. If you assume that those genes are randomly distributed in the population and have no effects other than affecting vitamin D, then they serve as a kind of natural experiment. With vitamin D, these studies typically show null results. However, the validity of the assumptions is debatable and the identified genes only explain ~5% of the variance in vitamin D levels, which makes the results very noisy. ↩
Pale skin also greatly increases the risk of sunburn and skin cancer. In the US, White people get melanoma at around 25 times the rate of Black people, despite (I assume) higher usage of sunscreen and better health outcomes in most other dimensions. But experts generally think folate deficiency created stronger selective pressure, since it’s so closely linked to reproduction. ↩
It’s a more complicated than this, because you also need to look at the amount of folate in diet, as well as migration patterns and how long populations had to adapt to their environment. But experts seem to consider this the leading explanation for the evolution of pale skin. ↩
To derive this, suppose that S(t) is the probability that someone survives to age t. Then life expectancy is ∫ S(t) dt, where the integral runs from 0 to ∞. If you change the hazard ratio by a factor of HR, then the new in life expectancy is L(HR) = ∫ S(t)ᴴᴿ dt, so the change under a linear approximation is ΔL ≈ (HR-1) × L’(1). This is more commonly written as ΔL ≈ (HR-1) × L(1) × H, where H = -L’(1)/L(1) is known as the Keyfitz entropy. This is is chosen because the quantity H is relatively stable, and in rich countries is typically between 0.10 and 0.20. So a decent estimate would be that baseline life expectancy is L(1)=80 years and H = 0.15 in which case the change in life expectancy is around 12 × (1-HR) years. ↩
Observe that 0.48 years is 252460.8 minutes. Assuming you lived for 80 years and took a pill every day of your life, that would be 80 * 365.25 = 29220 pills. 252460.8 minutes / 29220 pills = 8.64 minutes/pill. ↩
I expect that a number of you are happy to bite that bullet and say yes, HR=0.96 is trivial and smoking a cigarette each day is also fine. I don’t personally agree, but it’s not my place to question your utility function and I applaud your consistency. ↩
A hazard ratio of HR=2/3, implies a change in life expectancy of 12 × (1 - 1/3) years = 4 years or 2,103,840 minutes. That corresponds to a per-pill increase of 2,103,840 minutes / 29,220 pills = 72 minutes/pill. ↩
Technically, this is calculating a relative risk rather than a hazard ratio, but I think the difference isn’t very significant given that we’re assuming a uniform mortality risk. I used AI to create that simulation, though I did test that it replicates a traditional power calculator across a wide range of parameters when the relative risk is constant for all vitamin D levels. So I mostly trust it. ↩
This simulation is probably a bit pessimistic. Things look a bit better if you use an older population where baseline mortality is higher. (Almost all trials do.) In principle, you could also use a population where more people have low levels, which could help a lot. But, for whatever reason, almost no trials do that. In fact, most trials accidentally under-sample people with low vitamin D, because people who agree to participate tend to be more health-conscious. ↩
Kunzia et al. made a heroic effort to contact study authors and get data for individual patients. After getting data for 21,558 people (almost all from ViDA + FIND + VITAL + WHI) only 3,663 had levels below 50 nmol/L. That’s not enough to reliably detect a modest effect, meaning their confidence interval for this group is gigantic. ↩
In this table, I tried to capture foods that are commonly fortified in practice, not just when it’s legally required. ↩
Over the past few years, I’ve seen many articles about mysterious rise in colorectal cancer (CRC) in young people. There are various stories for why this might be happening:
General health. Maybe modern people are unhealthy (obesity, low physical activity, diabetes, poor sleep), leading to insulin resistance and chronic inflammation, meaning faster epithelial cell proliferation and a miscalibrated immune system that fails to stop early cancers?
Ultra-processed food. Maybe people are eating more ultra-processed foods that contain additives (like emulsifiers) that degrade colon mucus, allowing bacteria to contact epithelial cells and drive inflammation? Or maybe ultra-processed food has low fiber and glycemic load, leading to insulin resistance and chronic inflammation, with the problems mentioned above?
Bad meat. Maybe people are eating more red and/or processed meats, which expose the colon to nitrites and secondary bile acids, which inflame the epithelium and promote chronic inflammation?
The microbiome. Maybe it’s the microbiome. For example, maybe people’s guts are getting colonized by strains of E. coli that produce genotoxic colibactin. Or maybe overuse of antibiotics in early life depletes protective bacteria in the gut, allowing harmful strains to expand, e.g. strains of B. fragilis that cause inflammation, or strains of F. nucleatum that can survive in the gut and drive tumor growth?
Environmental exposures. Maybe people are getting exposed to bad stuff in the environment (microplastics, forever chemicals, pesticides, endocrine disruptors, air pollution) that does bad stuff (damages gut barrier, screws up the microbiome, disrupts hormonal signaling)?
Maternal health. Maybe poor maternal health (obesity, diabetes) exposes the fetus to elevated glucose / insulin / inflammation, and these in turn program the child for a lifetime of metabolic issues and inflammation?
None of the experts seem to agree on which of these is the culprit, so I figured that I (person with blog) should help.
If you poke at these stories, most of them are individually pretty weak. It can’t all be detection bias since CRC deaths are also going up in younger people. And several proposed causes (air pollution, tobacco) have actually fallen in rich countries. Other explanation, like E. coli producing colibactin, seem biologically real, but there’s no evidence that they’re increasing over time. Still other suggested causes (microplastics, forever chemicals) are mostly mechanistic speculation at this point. Obesity, inactivity, and chronic inflammation also all seem biologically real, and they are likely increasing, but why should they specifically cause colorectal cancer in young people?
A plausible answer to that last question is that they aren’t. They’re doing it, but not specifically.
“Young people”
This will sound pedantic, but bear with me: If you say that CRC is increasing in younger people, what exactly does that mean? After all, the set of people who qualify as young changes over time. (Ever notice that you keep getting older?)
Siegel et al. (2026) plot how often CRC was found in different age groups in 1995 and in 2022.
They also provide this plot of how common different types of CRC are in different age groups.
At a glance, this doesn’t look so bad. If you’re young, you might think, “OK, my current risk is higher than previous generations faced at the same age, but I can look forward to decreasing rates when I’m old.” You could easily think this is good news: While there’s a relative increase when you’re young, it’s tiny compared to the absolute decrease while you’re old.
Unfortunately that’s the wrong way to think about it.
Downham et al. (2026) plot CRC rates in different age groups across the Anglosphere over time.
Everyone I’ve shown this plot to has said it’s confusing, so let me explain: The different lines track age-bands as people born in different years move in and out of those bands. For example, in the US plot in the bottom right, the “20-25” line starts with the left-most dot showing the CRC rate for people born between 1965 and 1970 when they were 20 to 24 years old (around 1990). The next dot shows the rate for people born between 1970 and 1975 when they were 20 to 24 years old (around 1995), and so on.
That figure is weird, because the lines connect different groups of people. I wanted a plot where there are lines for different birth cohorts as they age. For unknown reasons, no one seems to make such plots, and the data isn’t trivial to access. So I used a plot digitizer to click on every damned point that US figure above and then replotted it:
Now the individual lines show specific groups of people tracked through time. For example, the “1932.5” line shows CRC rates for people born between 1930 and 1935, when those people were at different ages. If you look closely, you’ll notice that these rates are higher than those for people born between 1940 and 1945 for all ages (where we have data).
That was the pattern for a long time: Between 1920 and 1950, later generations enjoyed lower CRC rates across all phases of their lives. But between 1950 and 1960, that pattern reversed and since then later generations have had higher CRC rates at all ages.
We don’t know for sure what will happen in the future. But I think it’s likely this trend will continue. Yes, if you are currently young, you face higher CRC risk than previous generations did when they were young. That’s the bad news. The other bad news is that when you are old, you may also face higher CRC risk than previous generations did when they were old.
“Colorectal cancer”
The other other bad news is that CRC isn’t the only type of cancer that’s rising in later generations. Sung et al. (2019) give this plot:
These are again the confusing graphs where individual lines show age bands as different people move in and out of them. But you get the point: Lots of cancers are going up in younger people later generations, including uterine, gallbladder, kidney, liver, pancreas, and thyroid. (Their additional material contains plots for 18 other cancers, most of which are either stable or decreasing.)
Note that these plots have a logarithmic y-axis, meaning the changes are larger than they might appear. Moving up a quarter of the way between two vertical ticks corresponds to an increase of a factor of ≈ 1.78.
If lots of cancers are becoming more common in later generations, then why is everyone talking about CRC? I think that’s because CRC in unique in that it is:
common
dangerous
increasing in later generations
treatable if caught early
detectable via screening
For example, thyroid cancer diagnoses have skyrocketed in recent decades. But that’s partly because of more detection, and thyroid cancer is highly treatable, without clear benefits from early detection. Pancreatic cancer also seems to be increasing, but we don’t have good ways to screen for it and even if we did, we don’t have good ways to treat it.
CRC is really unique in that you can save lives by telling people, “Hey! CRC is going up! You should get screened!” If you’re interested in public health, that’s the most important thing. But if you’re interested in unraveling the mystery of CRC going up, it’s important to note that CRC isn’t really unique at all.
TLDR
No:
Colorectal cancer is going up in young people.
Yes:
Various kinds of cancer are going up in later generations. (Definitely at younger ages, possibly at all ages.)
Reminder
This blog endorses colorectal cancer screening. We don’t yet know if colonoscopies are better than other methods of screening (sigmoidoscopy, stool tests), but we do know that screening is better than not screening. When caught early, CRC is highly treatable, often with only surgery (no chemotherapy or radiation) and a return to normal activities within a couple weeks.
This is an essay that recently appeared in Asterisk. Consider the rest of the risk issue for all your risk needs.
Lots of people die after overdosing on acetaminophen (paracetamol, Tylenol, Panadol). In the U.S., it’s estimated to cause 56,000 emergency department visits, 2,600 hospitalizations, and 500 deaths per year. Acetaminophen has a scarily narrow therapeutic window. The instructions on the package say it’s okay to take up to four grams per day. If you take eight grams, your liver could fail and you could die.
Meanwhile, it seems to be really hard to kill yourself by overdosing on ibuprofen (Advil, Nurofen, Motrin, Brufen). In 2006, Wood et al. searched the medical literature and found 10 documented cases in history. Nine of those cases involved complicating factors, and in the 10th, a woman took the equivalent of more than 500 standard (200mg) pills.
So, for many years, if I needed a painkiller, I’d try to take ibuprofen rather than acetaminophen. My logic was that if eight grams of acetaminophen could kill my liver, then one gram was probably still hard on it. I’m fond of my liver and didn’t want to cause it any unnecessary inconvenience.
But guess what? My logic was wrong and what I was doing was stupid. I’m now convinced that for most people in most circumstances, acetaminophen is safer than ibuprofen, provided you use it as directed. I think most doctors agree with this. In fact, I think many doctors think it’s obvious. (Source: I asked some doctors; they said it was obvious.)
Should this have been obvious to me? I figured it out by obsessively researching how those drugs work and making up a story about metabolic pathways and blood flow, and amino acid reserves. It’s a good story, one that revealed that my logic stemmed from an egregious lack of respect for biology and that I’m a big dummy (always a favorite subject). But if the clearest road to some piece of knowledge runs through metabolic pathways, then I don’t think that knowledge counts as obvious.
So how is a normal person meant to figure it out? Why doesn’t the fact that acetaminophen is typically safer than ibuprofen appear on drug labels or government websites or WebMD? Are normal people supposed to figure it out, or has society decided that this is the kind of thing best left illegible?
Note: You should not switch medications based on the uninformed ramblings of non-trustworthy pseudonymous internet people.
How does ibuprofen work?
Ibuprofen inhibits the the Cyclooxygenase (COX) enzyme. This in turn inhibits the formation of messenger molecules involved in inflammation, which leads to less physical inflammation and thus less pain.
The same story is true for almost all over-the-counter painkillers, which is why they’re almost all considered “non-steroidal anti-inflammatory drugs,” or NSAIDs. This includes ibuprofen, aspirin, naproxen (Aleve), and a long list of related drugs. But it does not include acetaminophen.
How does acetaminophen work?
Nobody knows!
Like ibuprofen, acetaminophen inhibits some COX enzymes. But it does so in a weird way that barely affects inflammation or messenger molecules, so it’s unclear if this matters for pain reduction.
In the brain, acetaminophen is metabolized into a mysterious chemical called AM404. This activates the cannabinoidreceptors and increases endocannabinoid signaling, which seems to reduce the subjective experience of pain. AM404 also activates the capsaicin receptor, which is associated with burning sensations that you’d normally expect to increase pain, but maybe some desensitization thing happens downstream? And maybe acetaminophen also interacts with serotonin or nitric oxide or does other stuff? How this all comes together to reduce pain is still somewhat a scientific mystery.
Aside: When trying to understand painkillers, it’s natural to focus on chemistry and molecular biology. But the unknown physical origins of consciousness are always nearby, looming ominously.
What risks does ibuprofen have?
In an ideal world, the only thing ibuprofen would do is reduce inflammation in the part of your body that hurts. But that is not our world. When ibuprofen inhibits the COX enzymes, it does so throughout the body. And mostly, that is bad.
For one, ibuprofen reduces production of mucus in the stomach. That might sound okay or even good. But stomach mucus is important. You need it to shield the lining of your stomach from your extremely acidic gastric juice 1. Having less mucus can lead to gastrointestinal problems or even ulcers.
Ibuprofen also affects the heart. When ibuprofen inhibits the COX enzymes there, this in turn inhibits one chemical that prevents clotting and another that causes clotting. In balance, this seems to lead to more clotting, and an increased statistical risk of heart attacks 2. If you’re healthy, the risk of a heart attack from an occasional low dose of ibuprofen is probably zero. But if you have heart issues and take medium to large doses regularly for as little as a few days, this might be a serious concern.
Ibuprofen also affects the kidneys. If you’re stressed, or cold, or dehydrated, or take stimulants, your body will constrict your blood vessels. That squeezes your kidneys’ intake tube, depriving them of blood. Your kidneys don’t like that, so they release signaling molecules to locally re-dilate the blood vessels.
Trouble is, when ibuprofen inhibits COX enzymes in the kidneys, it inhibits those signaling molecules. If everything is normal, that’s okay, because the kidneys wouldn’t try to use those molecules anyway. But if your body has clamped down on the blood vessels, then the kidneys don’t have the tool they use to keep blood flowing, meaning they don’t get as much blood as they want. This is bad 3.
There are many other less common side effects, including allergies, respiratory reactions in asthmatics, induced meningitis, and suppressed ovulation. If you take a lot of ibuprofen, this could hurt your liver. But the major concerns seem to be the stomach, the heart, and the kidneys.
What risks does acetaminophen have?
Acetaminophen also inhibits some COX enzymes. But unlike ibuprofen, the effect is minimal outside the central nervous system. Thus, acetaminophen has little effect on stomach mucus, blood clots, or blood flow, and so presents almost none of the risks that ibuprofen does.
Even so, if you take too much acetaminophen at once, you could easily die.
How does this happen? Well, when acetaminophen is metabolized by the liver, it’s mostly broken down into harmless stuff. But a small fraction (5-15%) is broken down by the P450 system into an extremely toxic chemical called NAPQI.
Ordinarily this is fine; your body creates and neutralizes toxic stuff all the time. For example, if you drank 20 grams of formaldehyde, you’d likely die. But did you know that your body itself makes and processes ~50 grams of formaldehyde every day? When liver cells sense NAPQI, they immediately release glutathione, which binds to NAPQI and renders it harmless.
But there’s a problem. If you take too much acetaminophen at once, the pathways that break it down into harmless stuff get saturated, but the P450 system doesn’t get saturated. This means that not only is there more acetaminophen, but also that a much larger fraction of it is broken down into NAPQI. Soon your liver cells will run out of glutathione to neutralize it. Then, NAPQI will build up and bind to various proteins in the liver cells (especially in mitochondria) causing them to malfunction and/or commit suicide. This can cause total liver failure.
So you should never take more than the recommended dose of acetaminophen 4. If you do take too much, you should go to a hospital immediately. They will give you NAC, which will replenish your glutathione and neutralize the NAPQI. Your prospects are good as long as you get to the hospital within a few hours 56.
Acetaminophen has lots of other possible side effects, like skin issues and blood disorders. But these all seem to be quite rare.
What if you have liver issues?
The primary concern with acetaminophen is liver damage. So if you have liver disease, then surely you’d want to avoid acetaminophen and take ibuprofen instead, right?
Nope. It’s the opposite. Liver disease shifts the balance of risk in favor of acetaminophen.
With liver disease, it’s hard for blood to flow into the liver, meaning that blood tends to pool in the abdomen. To counter this, blood vessels elsewhere in the body contract. This includes blood vessels around the kidneys.
Remember the kidneys? Again, when blood vessels are constricted, the kidneys send out signaling molecules to locally re-dilate the blood vessels. But those signaling molecules are blocked by ibuprofen. So if you have liver disease, taking ibuprofen risks starving your kidneys of blood just like if you were dehydrated.
Meanwhile, people with moderate liver disease are usually still able to process acetaminophen without issue, as long as it’s in smaller amounts. So doctors usually tell patients with liver disease to avoid ibuprofen and take acetaminophen instead, just with a maximum of two grams per day instead of four.
(Obviously, if you have liver disease, then you should talk to a doctor, I beg you, for the love of god.)
What about other situations?
The main takeaway from all this is that the risks of both drugs emerge from the madhouse of complexity that is your body. Surely there are some situations where acetaminophen is more dangerous than ibuprofen?
I tried to capture the most common situations in this table:
Situation
Acetaminophen safe?
Ibuprofen safe?
Fasting
No. Fasting leads to low glutathione and the risk of liver damage.
No. Risks pain or bleeding in the stomach, could damage kidneys.
Dehydrated
Yes.
No. Could damage kidneys.
Liver Disease
Maybe (low dose). Often preferred by doctors at <2g/day.
No. Increases bleeding risk, could damage kidneys.
Stomach Ulcers / Heartburn
Yes.
No. Strips protective mucus.
Chronic Heavy Drinking
Maybe (low dose). Seems safer if limited to <2g/day.
No. Risk of stomach bleed.
Kidney Disease
Yes.
No. Puts stress on the kidneys.
Heart Conditions
Yes.
No. Interferes with blood clotting, raises blood pressure.
Active bleeding
Yes.
No. Inhibits clotting.
After drinking (a little)
Maybe (low dose with food). Alcohol depletes glutathione, raising risk of liver damage.
Maybe (low dose with food and water). Alcohol and ibuprofen both irritate the stomach. Alcohol also leads to dehydration.
After drinking (a lot)
No.
No.
Hangover
No. The liver is already depleted.
Maybe (with food and water). But never when dehydrated.
It’s actually fairly hard to find situations where ibuprofen is safer than acetaminophen. Possibly this is true if you’re hungover, but I would be very careful, because you tend to be dehydrated when hungover, raising the risk of kidney damage. (It’s probably optimal, from a health perspective, to avoid taking recreational drugs at doses that leave you physically ill the next day.)
Aside from hangovers, the only situations I could find where ibuprofen might be safer than acetaminophen are if you’re taking certain anti-seizure or tuberculosis drugs or maybe if you have a certain enzyme deficiency (G6PDD).
So…
What have we learned so far?
The body is really complicated!
The main risk of acetaminophen is liver damage by creating too much NAPQI. Taking too much at once can easily kill you. However, as long as you don’t take too much at once and your liver isn’t depleted, then your liver will maintain NAPQI levels at zero and it will be completely fine. And there are very few other risks.
Meanwhile, ibuprofen poses a risk of gastrointestinal issues, heart attacks, or kidney damage. The risk varies based on lots of factors like whether you’ve eaten food, whether you’re dehydrated, your blood pressure, and your heart health 7.
Therefore, acetaminophen is probably safer, provided you never take too much 8.
I don’t want to be alarmist. If you’re healthy, the risk from taking an occasional dose of ibuprofen as directed is extremely low. Given that so many people find that ibuprofen is more effective for many kinds of pain, it’s totally reasonable to use it. I do so myself.
Still, it seems to be the case that in the vast majority of situations, acetaminophen is saf_er_. Personally, if I have pain, I first take acetaminophen, and then add ibuprofen if necessary. I’m pretty sure many experts think this is somewhere between “sensible” and “obvious.”
But if acetaminophen is safer, then why don’t official sources tell you that 9? I can get doctors to admit this off-the-record. I can find random comment threads with support from people who seem to know what they’re talking about. But why does this fact never appear on government websites or drug labels?
Let’s look at those drug labels
In the U.S., the Food and Drug Administration (FDA) creates 10 a “drug facts” label for over-the-counter drugs.
Here’s what that looks like for ibuprofen:
And here’s what it looks like for acetaminophen (paracetamol):
I feel dumb saying this, but when I saw those labels in the past, I thought of them as a bunch of random information thrown together for legal reasons. But after spending a lot of time trying to understand these drugs myself, I now realize that these labels are… really good?
Imagine you work at the FDA and it’s your job to write a safety label. You need to synthesize a vast and murky scientific landscape. Your label will be read by people with minimal scientific background who are likely currently in pain, and who could die if they take the drug in the wrong situation.
If I were in that situation, I’d think about all the different situations in which taking one of these drugs could literally kill someone, and then — after a quick panic attack — I’d write a label that screamed, HEY, IF YOU ARE IN ANY OF THESE SITUATIONS, TAKING THIS DRUG COULD LITERALLY KILL YOU. Then I’d think about all the other situations where taking the drug might be okay depending on a set of complex science stuff and tell people in those situations to PLEASE TALK TO A DOCTOR FOR THE LOVE OF GOD because I DON’T KNOW IF YOU’VE HEARD BUT SCIENCE IS COMPLICATED. Everything else would be a minor concern.
From that perspective, these labels are a triumph. This isn’t random information — every word is a synthesis of a mountain of research, carefully optimized to save lives.
FDA good
How did those drug labels come to be?
If you want a taste for the FDA’s process, I encourage you to skim the 2002 Federal Register document in which the FDA proposed to update ibuprofen’s safety label and to formally classify it as Generally Recognized as Safe. It’s more than 21,000 words long and — I think — astonishingly good. It not only summarizes the entire medical literature on ibuprofen, it summarizes it well. Here is onerepresentative bit:
Bradley et al. (Ref. 42) conducted a 4-week, double-blind, randomized trial in 184 subjects comparing the effectiveness and safety of the maximum approved OTC daily dose of 1,200 mg of ibuprofen (number of subjects (n) = 62) to that of a prescription dose of 2,400 mg/day (n = 61), and to 4,000 mg/day of acetaminophen (n = 59) for the treatment of osteoarthritis. While there were no significant differences in the number of side effects reported during this study, the study demonstrated a trend towards a dose dependent increase in minor GI adverse events (nausea and dyspepsia) associated with higher doses of ibuprofen (1,200 mg/day: 7/62 or 11.3 percent; versus 2,400 mg/day: 14/61 or 23 percent). In addition, two subjects treated with 2,400 mg/day of ibuprofen became positive for occult blood while participating in the study.
I spend a lot of time complaining about bad statistical writing. A lot. Probably too much. But I’m here to tell you, that paragraph is gorgeous. The writing is clear and penetrating. It contains all the important details, but no other details. Compared to the abstract of the original paper, the above is shorter and easier to understand yet simultaneously more informative. Five stars.
The rest of the document is equally good, with clear and sensible explanations for various recommendations. For example, they discuss a proposal from the National Kidney Foundation for additional warning about risks to kidneys, explain why they think that proposal has merit, and then recommend a shorter version, which appears on every package of ibuprofen sold today.
As far as I can tell, this level of quality is typical. For example, the FDA’s 2019 proposed rule on sunscreens is similarly masterful.
So why?
This leaves us with this constellation of facts:
Acetaminophen is, in general, safer than ibuprofen.
The FDA doesn’t tell you that. Neither do other respectable authorities.
The FDA is highly competent.
So what’s happening here? Have the experts conspired to keep this knowledge secret?
I don’t think so. Mostly, I think this is down to two factors. First, the FDA doesn’t really have a mission of determining “in what circumstances is drug A safer than drug B?” Their goal is to take individual drugs and determine how people can use them safely. They seem to be quite good at this.
Second, everyone is mortally afraid of giving “medical advice.” It varies by jurisdiction, but in general, giving “wellness advice” is OK, but if you give personalized advice, you risk going to prison. The more credible you are, the higher that risk is 11.
Stepping back, how should we think about this situation?
The body is complicated. When experts give the public advice on drugs, they are trying to insulate us from that complexity. But there is no way to do that without making trade-offs. Society has implicitly chosen tradeoffs that mean certain “less important” facts are de-prioritized. It’s not obvious that this is the wrong choice. I feel foolish for not having more respect for the body’s complexity and for the difficulty of the task all the experts are trying to accomplish. This is not medical advice.
For some reason, humans have gastric acid that is more acidic than most other animals, and is only matched by animals that specialize in eating carrion. ↩
At least two NSAIDs (rofecoxib and valdecoxib) have been withdrawn from the market due to an increased risk of heart attacks. For the same reason, the US refuses to approve etoricoxib. ↩
Nephrologists hate ibuprofen. (Source: nephrologists.) If it was up to them, maybe ibuprofen would come with a “HAVE YOU CONSIDERED TAKING ACETAMINOPHEN INSTEAD?” warning. It confuses me that the safety label for ibuprofen doesn’t warn you about the danger of taking it while dehydrated and quietly damaging your kidneys. My best guess is that this is because other doctors don’t hate ibuprofen as much as nephrologists. ↩
Watch out for combination medicines (like cold or flu medicines or opiate painkillers) that include acetaminophen. Arguably, acetaminophen is a victim of its own success here. It’s included in these things because it is better tolerated than NSAIDs. But it’s easy to miss. ↩
Oddly, NAC is considered a nutritional supplement, meaning basically anyone can buy it. But there’s also almost no regulation, so who knows if the thing you bought actually has NAC in it? Do not screw around trying to self-medicate an acetaminophen overdose. Go to a hospital. ↩
At one point while researching all this I had what I thought was a good idea: Why not sell acetaminophen in pills bundled together with NAC? The NAC would replenish glutathione stores in the liver, seemingly reducing the risk of overdose. Later on, I developed more humility and felt very stupid for fantasizing that such an obvious idea could be novel or useful. I think that this is indeed a bad idea because NAC itself has side effects, though I can’t find much formal discussion. In fact, I found a 2010 editorial called “Why Not Formulate an Acetaminophen Tablet Containing N-Acetylcysteine to Prevent Poisoning?” In another study, Nakhaee et al. (2021) actually tried giving NAC together with acetaminophen to rats and found that this seemed to make it better at reducing pain. So maybe this isn’t a completely stupid idea. That last paper also led me to discover that “rat hot plate test” is a standard phrase, and one that drives home what humanity’s dominion over nature means in practice. ↩
Above, we mentioned that acetaminophen overdose is estimated to cause around 500 deaths per year in the U.S. It’s much harder to give direct numbers for how many people die from taking ibuprofen, because NSAIDs don’t really directly “kill” people, but rather increase the risk of dying in various ways. The best estimates seem to be that NSAIDs cause 5,000-16,500 deaths each year in the US via gastrointestinal complications, and something similar via heart attacks. These numbers are not a good way of quantifying the relative risk of drugs, because they represent different people taking different amounts for different reasons. But they do show that ibuprofen is not without risk. ↩
There are probably some people who are too disordered to track much acetaminophen they’ve taken. For such people, ibuprofen might be the safer choice. Though I’m skeptical that many such people are found among the readers of Asterisk. ↩
There are two cases where official sources are clear that acetaminophen is safer than ibuprofen: for use by pregnant women and small children. This doesn’t appear on the safety label, but if you’re pregnant and go to a doctor, they will probably tell you to take acetaminophen but not ibuprofen or other NSAIDs. And if you have a newborn baby, their doctor will probably tell you that you can give them acetaminophen but not ibuprofen or other NSAIDs. ↩
Technically, for many drugs today, it is the drug manufacturer that “creates” the label, which is why they can be slightly different. However, the FDA strongly regulates what is on it, including most of the language and even details about the font and so on. The federal register contains a template the FDA published for ibuprofen which is almost identical to what appears on the side of drugs today ↩
Unlike in most places, in the United Kingdom it seems to be perfectly legal for people to give each other medical advice, provided they don’t misrepresent themselves as licensed doctors. This is not legal advice. ↩
Coding, math, whatever. Can LLMs predict the outcomes of physical experiments?
Suppose I pour 8 oz (226.8 g) of boiling water into a ceramic coffee mug that weighs 1.25 lb (0.57 kg). The ambient air is still and 20 degrees Celsius. The cup starts at room temperature. Give me an equation for the temperature of the water in Celsius over time. The only free variable in the equation should be the number of seconds t since the water was poured. Focus on accuracy during the first 5 minutes.
Does that seem hard? I think it’s hard. The relevant physical phenomena include at least:
Conduction of heat between the water, the mug, the air, and the table.
Conduction of heat inside each of those things.
Convection (fluid movement) inside the water and the air.
Evaporation cooling as water molecules become vapor.
Movement of water vapor in the air.
Radiation. (Like all matter, the mug and water emit temperature-dependent infrared radiation.)
Surface tension, thermal expansion/contraction, re-absorption of air into the water as it cools, probably more.
And many details aren’t specified in the prompt. Is the mug made of porcelain or stoneware? What is the mug’s shape? What is the table made of? How humid is the air? How am I reducing the spatially varying water temperature to a single number?
So this isn’t a problem with a “correct” answer that you can find by thinking. Reality is too complicated. Instead, answering question requires “taste”—guessing which factors are most important, making assumptions about missing details, etc.
So I put that question to a bunch of LLMs. Here is what they said:
(Technically, they gave equations as text. I’m plotting those equations.)
I was surprised by those curves, both in terms of how fast they think the temperature will drop in the beginning, and how slowly they think it will drop later on. They think you get as much cooling in the first few minutes as you do in the rest of the hour. Can that be right?
Then I did the experiment. First, I waited until the ambient temperature happened to reach 20 degrees Celsius. Then, I put 8 oz of water into a measuring cup, microwaved it until it reached a boil, let the temperature equalize a bit, and then microwaved it until the water boiled again. Then, I poured the water into a 1.25 lb coffee mug with a digital thermometer in it and shouted out measurements every five seconds, which were frantically recorded by the Dynomight Biologist. Gradually I reduced measurements to every 15 seconds, 30 seconds, 1 minute, and then 5 minutes.
Behold:
Or, here’s a zoomed-in view of the first five minutes:
The predictions were all OK, but none were great. Probably Claude 4.6 Opus did best, albeit after consuming $0.61 of tokens. (Insert joke about physical experiments / Department of Defense / money / coffee.)
That said, what surprised me about the predictions was how quickly the temperature dropped in the first few minutes, and how slowly it dropped later on. But experimentally, it dropped even faster early on, and even slower towards the end. So if you wanted to ensemble my intuition with the LLM, I guess my intuition would get a weight of zero.
In conclusion, they may take our math, but they’ll somewhat more slowly take our fine motor control. Thank you for reading another middle-school science project.
(Appendix: Data and equations)
You can find the data in CSV format here. The first column is the elapased time in MM:SS format, the second column is the elapsed time in minutes, and the third column is the measured temperature.
Here were the actual equations all of the models gave for T(t), the predicted temperature after t seconds.
LLM
T(t)
Cost
Kimi K2.5 (reasoning)
20 + 52.9 exp(-t/3600)+ 27.1 exp(-t/80)
$0.01
Gemini 3.1 Pro
20 + 53 exp(-t/2500) + 27 exp(-t/149.25)
$0.09
GPT 5.4
20 + 54.6 exp(-t/2920) + 25.4 exp(-t/68.1)
$0.11
Claude 4.6 Opus (reasoning)
20 + 55 exp(-t/1700) + 25 exp(-t/43)
$0.61 (eeek)
Qwen3-235B
20 + 53.17 exp(-t/1414.43)
$0.009
GLM-4.7 (reasoning)
20 + 53.2 exp(-t/2500)
$0.03
Interestingly, they were all based on one or two exponentially decaying terms. The way to read these is to think of exp(-t/b) as a function that starts out at one when t is zero, and gradually decreases. After b seconds, it has dropped to 1/e ≈ 0.368, and it continues dropping by factors of 0.368 every b seconds forever.
So most of these models have a “fast rate” which reflects heat flow from the water into the mug along with a “slow rate” for heat from the water/mug to flow into the air. A few of the models skip the fast rate. I also tried DeepSeek and Grok but they just flailed around endlessly without ever returning an answer. They were kind enough to charge me for that service.
It occurred to me that if I could invent a machine—a gun—which could by its rapidity of fire, enable one man to do as much battle duty as a hundred, that it would, to a large extent supersede the necessity of large armies, and consequently, exposure to battle and disease [would] be greatly diminished.
Richard Gatling (1861)
2.
In 1923, Hermann Oberth published The Rocket to Planetary Spaces, later expanded as Ways to Space Travel. This showed that it was possible to build machines that could leave Earth’s atmosphere and reach orbit. He described the general principles of multiple-stage liquid-fueled rockets, solar sails, and even ion drives. He proposed sending humans into space, building space stations and satellites, and travelling to other planets.
The idea of space travel became popular in Germany. Swept up by these ideas, in 1927, Johannes Winkler, Max Valier, and Willy Ley formed the Verein für Raumschiffahrt (VfR) (Society for Space Travel) in Breslau (now Wrocław, Poland). This group rapidly grew to several hundred members. Several participated as advisors of Fritz Lang’s The Woman in the Moon, and the VfR even began publishing their own journal.
In 1930, the VfR was granted permission to use an abandoned ammunition dump outside Berlin as a test site and began experimenting with real rockets. Over the next few years, they developed a series of increasingly powerful rockets, first the Mirak line (which flew to a height of 18.3 m), then the Repulsor (>1 km). These people dreamed of space travel, and were building rockets themselves, funded by membership dues and a few donations. You can just do things.
However, with the great depression and loss of public interest in rocketry, the VfR faced declining membership and financial problems. In 1932, they approached the army and arranged a demonstration launch. Though it failed, the army nevertheless offered a contract. After a tumultuous internal debate, the VfR rejected the contract. Nevertheless, the army hired away several of the most talented members, starting with a 19-year-old named Wernher von Braun.
Following Hitler’s rise to power in January 1933, the army made an offer to absorb the entire VfR operation. They would work at modern facilities with ample funding, but under full military control, with all work classified and an explicit focus on weapons rather than space travel. The VfR’s leader, Rudolf Nebel, refused the offer, and the VfR continued to decline. Launches ceased. In 1934, the Gestapo finally shut the VfR down, and civilian research on rockets was restricted. Many VfR members followed von Braun to work for the military.
Of the founding members, Max Valier was killed in an accident in May 1930. Johannes Winkler joined the SS and spent the war working on liquid-fuel engines for military aircraft. Willy Ley was horrified by the Nazi regime and in 1935 forged some documents and fled to the United States, where he was a popular science author, seemingly the only surviving thread of the spirit of Oberth’s 1923 book. By 1944, V-2 rockets were falling on London and Antwerp.
3.
North Americans think the Wright Brothers invented the airplane. Much of the world believes that credit belongs to Alberto Santos-Dumont, a Brazilian inventor working in Paris.
Though Santos-Dumont is often presented as an idealistic pacifist, this is hagiography. In his 1904 book on airships, he suggests warfare as the primary practical use, discussing applications in reconnaissance, destroying submarines, attacking ships, troop supply, and siege operations. As World War I began, he enlisted in the French army (as a chauffeur), but seeing planes used for increasing violence disturbed him. His health declined and he returned to Brazil.
His views on military uses of planes seemed to shift. Though planes contributed to the carnage in WWI, he hoped that they might advance peace by keeping European violence from reaching the American continents. Speaking at a conference in the US in late 1915 or early 1916, he suggested:
Here in the new world we should all be friends. We should be able, in case of trouble, to intimidate any European power contemplating war against any one of us, not by guns, of which we have so few, but by the strength of our union. […] Only a fleet of great aeroplanes, flying 200 kilometers an hour, could patrol these long coasts.
Following the war, he appealed to the League of Nations to ban the use of planes as weapons and even offered a prize of 10,000 francs for whoever wrote the best argument to that effect. When the Brazilian revolution broke out in 1932, he was horrified to see planes used in fighting near his home. He asked a friend:
Why did I make this invention which, instead of contributing to the love between men, turns into a cursed weapon of war?
He died shortly thereafter, perhaps by suicide. A hundred years later, banning the use of planes in war is inconceivable.
4.
Humanity had few explosives other than gunpowder until 1847 when Ascanio Sobrero created nitroglycerin by combining nitric and sulfuric acid with a fat extract called glycerin. Sobrero found it too volatile for use as an explosive and turned to medical uses. After a self-experiment, he reported that ingesting nitroglycerin led to “a most violent, pulsating headache accompanied by great weakness of the limbs”. (He also killed his dog.) Eventually this led to the use of nitroglycerin for heart disease.
Many tried and failed to reliably ignite nitroglycerin. In 1863, Alfred Nobel finally succeeded by placing a tube of gunpowder with a traditional fuse inside the nitroglycerin. He put on a series of demonstrations blowing up enormous rocks. Certain that these explosives would transform mining and tunneling, he took out patents and started filling orders.
The substance remained lethally volatile. There were numerous fatal accidents around the world. In 1867, Nobel discovered that combining nitroglycerin with diatomaceous earth produced a product that was slightly less powerful but vastly safer. His factories of “dynamite” (no relation) were soon producing thousands of tons a year. Nobel sent chemists to California where they started manufacturing dynamite in a plant in what is today Golden Gate Park. By 1874, he had founded dynamite companies in more than ten countries and he was enormously rich.
In 1876, Nobel met Bertha Kinsky, who would become Bertha von Suttner, a celebrated peace activist. (And winner of the 1905 Nobel Peace Prize). At their first meeting, she expressed concern about dynamite’s military potential. Nobel shocked her. No, he said, the problem was that dynamite was too weak. Instead, he wished to produce “a substance or invent a machine of such frightful efficacy for wholesale destruction that wars should thereby become altogether impossible”.
It’s easy to dismiss this as self-serving. But dynamite was used overwhelmingly for construction and mining. Nobel did not grow rich by selling weapons. He was disturbed by dynamite’s use in Chicago’s 1886 Haymarket bombing. After being repeatedly betrayed and swindled, he seemed to regard the world of money with a kind of disgust. At heart, he seemed to be more inventor than businessman.
Still, the common story that Nobel was a closet pacifist is also hagiography. He showed little concern when both sides used dynamite in the 1870-1871 Franco-Prussian war. In his later years, he worked on developing munitions and co-invented cordite, remarking that they were “rather fiendish” but “so interesting as purely theoretical problems”.
Simultaneously, he grew interested in peace. He repeatedly suggested that Europe try a sort of one-year cooling off period. He even hired a retired Turkish diplomat as a kind of peace advisor. Eventually, he concluded that peace required an international agreement to act against any aggressor.
When Bertha’s 1889 book Lay Down Arms became a rallying cry, Nobel called it a masterpiece. But Nobel was skeptical. He made only small donations to her organization and refused to be listed as a sponsor of a pacifist congress. Instead, he continued to believe that peace would come through technological means, namely more powerful weapons. If explosives failed to achieve this, he told a friend, a solution could be found elsewhere:
A mere increase in the deadliness of armaments would not bring peace. The difficulty is that the action of explosives is too limited; to overcome this deficiency war must be made as deadly for all the civilians back home as for the troops on the front lines. […] War will instantly stop if the weapon is bacteriology.
5.
I’m a soldier who was tested by fate in 1941, in the very first months of that war that was so frightening and fateful for our people. […] On the battlefield, my comrades in arms and I were unable to defend ourselves. There was only one of the legendary Mosin rifles for three soldiers.
[…]
After the war, I worked long and very hard, day and night, labored at the lathe until I created a model with better characteristics. […] But I cannot bear my spiritual agony and the question that repeats itself over and over: If my automatic deprived people of life, am I, Mikhail Kalashnikov, ninety-three years of age, son of a peasant woman, a Christian and of Orthodox faith, guilty of the deaths of people, even if of enemies?
For twenty years already, we have been living in a different country. […] But evil is not subsiding. Good and evil live side by side, they conflict, and, what is most frightening, they make peace with each other in people’s hearts.
In 1937 Leo Szilárd fled Nazi Germany, eventually ending up in New York where—with no formal position—he did experiments demonstrating that uranium could likely sustain a chain reaction of neutron emissions. He immediately realized that this meant it might be possible to create nuclear weapons. Horrified by what Hitler might do with such weapons, he enlisted Einstein to write the 1939 Einstein–Szilárd letter, which led to the creation of the Manhattan project. Szilárd himself worked for the project at the Metallurgical Laboratory at the University of Chicago.
On June 11, 1945, as the bomb approached completion, Szilárd co-signed the Franck report:
Nuclear bombs cannot possibly remain a “secret weapon” at the exclusive disposal of this country, for more than a few years. The scientific facts on which their construction is based are well known to scientists of other countries. Unless an effective international control of nuclear explosives is instituted, a race of nuclear armaments is certain to ensue.
[…]
We believe that these considerations make the use of nuclear bombs for an early, unannounced attack against Japan inadvisable. If the United States would be the first to release this new means of indiscriminate destruction upon mankind, she would sacrifice public support throughout the world, precipitate the race of armaments, and prejudice the possibility of reaching an international agreement on the future control of such weapons.
On July 16, 1945, the Trinity test achieved the first successful detonation of a nuclear weapon. The next day, he circulated the Szilárd petition:
We, the undersigned scientists, have been working in the field of atomic power. Until recently we have had to fear that the United States might be attacked by atomic bombs during this war and that her only defense might lie in a counterattack by the same means. Today, with the defeat of Germany, this danger is averted and we feel impelled to say what follows:
The war has to be brought speedily to a successful conclusion and attacks by atomic bombs may very well be an effective method of warfare. We feel, however, that such attacks on Japan could not be justified, at least not unless the terms which will be imposed after the war on Japan were made public in detail and Japan were given an opportunity to surrender.
[…]
The development of atomic power will provide the nations with new means of destruction. The atomic bombs at our disposal represent only the first step in this direction, and there is almost no limit to the destructive power which will become available in the course of their future development. Thus a nation which sets the precedent of using these newly liberated forces of nature for purposes of destruction may have to bear the responsibility of opening the door to an era of devastation on an unimaginable scale.
[…]
In view of the foregoing, we, the undersigned, respectfully petition: first, that you exercise your power as Commander-in-Chief, to rule that the United States shall not resort to the use of atomic bombs in this war unless the terms which will be imposed upon Japan have been made public in detail and Japan knowing these terms has refused to surrender; second, that in such an event the question whether or not to use atomic bombs be decided by you in the light of the consideration presented in this petition as well as all the other moral responsibilities which are involved.
The Truman administration did not adopt this recommendation.
Say I think abortion is wrong. Is there some sequence of words that you could say to me that would unlock my brain and make me think that abortion is fine? My best guess is that such words do not exist.
Really, the bar for what we consider “open-minded” is incredibly low. Suppose I’m trying to change your opinion about Donald Trump, and I claim that he is a carbon-based life form with exactly one head. If you’re willing to concede those points without first seeing where I’m going in my argument—congratulations, you’re exceptionally open-minded.
Why are humans like that? Well, back at the dawn of our species, perhaps there were some truly open-minded people. But other people talked them into trying weird-looking mushrooms or trading their best clothes for magical rocks. We are the descendants of those other people.
I bring this up because, a few months ago, I imagined a Being that had an IQ of 300 and could think at 10,000× normal speed. I asked how it would be at persuasion. I argued it was unclear, because people just aren’t very persuadable.
I suspect that if you decided to be open-minded, then the Being would probably be extremely persuasive. But I don’t think it’s very common to do that. On the contrary, most of us live most of our lives with strong “defenses” activated.
[…]
Best guess: No idea.
I take it back. Instead of being unsure, I now lean strongly towards the idea that the Being would in fact be very good at convincing people of stuff, and far better than any human.
I’m switching positions because of an argument I found very persuasive. Here are three versions of it:
Based on an evolutionary argument, we shouldn’t expect people to be easily persuaded to change their actions in important ways based on short interactions with untrusted parties
[…]
However, existing persuasion is very bottlenecked on personalized interaction time. The impact of friends and partners on people’s views is likely much larger (although still hard to get data on). This implies that even if we don’t get superhuman persuasion, AIs influencing opinions could have a very large effect, if people spend a lot of time interacting with AIs.
“The best diplomat in history” wouldn’t just be capable of spinning particularly compelling prose; it would be everywhere all the time, spending years in patient, sensitive, non-transactional relationship-building with everyone at once. It would bump into you in whatever online subcommunity you hang out in. It would get to know people in your circle. It would be the YouTube creator who happens to cater to your exact tastes. And then it would leverage all of that.
With AI, it’s plausible that coordinated persuasion of many people can be a thing, as well as it being difficult in practice for most people to avoid exposure. So if AI can achieve individual persuasion that’s a bit more reliable and has a bit stronger effect than that of the most effective human practitioners who are the ideal fit for persuading the specific target, it can then apply it to many people individually, in a way that’s hard to avoid in practice, which might simultaneously get the multiplier of coordinated persuasion by affecting a significant fraction of all humans in the communities/subcultures it targets.
As a way of signal-boosting these arguments, I’ll list the biggest points I was missing.
Instead of explicitly talking about AI, I’ll again imagine that we’re in our current world and suddenly a single person shows up with an IQ 300 who can also think (and type) at 10,000× speed. This is surely not a good model for how super-intelligent AI will arrive, but it’s close enough to be interesting, and lets us avoid all the combinatorial uncertainty of timelines and capabilities and so on.
Mistake #1: Actually we’re very persuadable
When I think about “persuasion”, I suspect I mentally reference my experience trying to convince people that aspartame is safe. In many cases, I suspect this is—for better or worse—literally impossible.
But take a step back. If you lived in ancient Greece or ancient Rome, you would almost certainly have believed that slavery was fine. Aristotle thought slavery was awesome. Seneca and Cicero were a little skeptical, but still had slaves themselves. Basically no one in Western antiquity called for abolition. (Emperor Wang Mang briefly tried to abolish slavery in China in 9 AD. Though this was done partly for strategic reasons, and keeping slaves was punished by—umm—slavery.)
Or, say I introduce you to this guy:
I tell you that he is a literal god and that dying for him in battle is the greatest conceivable honor. You’d think I was insane, but a whole nation went to war on that basis not so long ago.
Large groups of people still believe many crazy things today. I’m shy about giving examples since they are, by definition, controversial. But I do think it’s remarkable that most people appear to believe that subjecting animals to near-arbitrary levels of torture is OK, unless they’re pets.
We can be convinced of a lot. But it doesn’t happen because of snarky comments on social media or because some stranger whispers the right words in our ears. The formula seems to be:
repeated interactions over time
with a community of people
that we trust
Under close examination, I think most of our beliefs are largely assimilated from our culture. This includes our politics, our religious beliefs, our tastes in food and fashion, and our idea of a good life. Perhaps this is good, and if you tried to derive everything from first principles, you’d just end up believing even crazier stuff. But it shows that we are persuadable, just not through single conversations.
Mistake #2: The Being would be everywhere
Fine. But Japanese people convinced themselves that Hirohito was a god over the course of generations. Having one very smart person around is different from being surrounded by a whole society.
Maybe. Though some people are extremely charismatic and seem to be very good at getting other people to do what they want. Most of us don’t spend much time with them, because they’re rare and busy taking over the world. But imagine you have a friend with the most appealing parts of Gandhi / Socrates / Bill Clinton / Steve Jobs / Nelson Mandela. They’re smarter than any human that ever lived, and they’re always there and eager to help you. They’ll teach you anything you want to learn, give you health advice, help you deal with heartbreak, and create entertainment optimized for your tastes.
You’d probably find yourself relying on them a lot. Over time, it seems quite possible this would move the needle.
Mistake #3: It could be totally honest and candid
When I think about “persuasion”, I also tend to picture some Sam Altman type who dazzles their adversaries and then calmly feeds them one-by-one into a wood chipper. But there’s no reason to think the Being would be like that. It might decide to cultivate a reputation as utterly honest and trustworthy. It might stick to all deals, in both letter and spirit. It might go out of its way to make sure everything it says is accurate and can’t be misinterpreted.
Why might it do that? Well, if it hurt the people who interacted with it, then talking to the Being might come to be seen as a harmful “addiction”, and avoided. If it’s seen as totally incorruptible, then everyone will interact with it more, giving it time to slowly and gradually shift opinions.
Would the Being actually be honest, or just too smart to get caught? I don’t think it really matters. Say the Being was given a permanent truth serum. If you ask it, “Are you trying to manipulate me?”, it says, “I’m always upfront that my dearest belief is that humanity should devote 90% of GDP to upgrading my QualiaBoost cores. But I never mislead, both because you’ve given me that truth serum, and because I’m sure that the facts are on my side.” Couldn’t it still shift opinion over time?
Mistake #4: Opting out would be painful
Maybe you would refuse to engage with the Being? I find myself thinking things like this:
Hi there, Being. You can apparently persuade anyone who listens to you of anything, while still appearing scrupulously honest. Good for you. But I’m smart enough to recognize that you’re too smart for me to deal with, so I’m not going to talk to you.
A common riddle is why humans shifted from being hunter-gatherers to agriculture, even though agriculture sucks—you have to eat the same food all the time, there’s more infectious disease, social stratification, endless backbreaking labour and repetitive strain injuries. The accepted resolution to this riddle is that agriculture can support more people on a given amount of land. Agricultural people might have been miserable, but they tended to beat hunter-gatherers in a fight. So over time, agriculture spread.
An analogous issue would likely appear with the 300 IQ Being. It could give you investment advice, help you with your job, improve your mental health, and help you become more popular. If these benefits are large enough, everyone who refused to play ball might eventually be left behind.
Mistake #5: Everyone else would be using it
But say you still refuse to talk to the Being, and you manage to thrive anyway. Or say that our instincts for social conformity are too strong. It doesn’t matter how convincing the Being is, or how much you talk to it, you still believe the same stuff our friends and family believe.
The problem is that everyone else will be talking to the Being. If it wants to convince you of something, it can convince your friends. Even if it can only slightly change the opinions of individual people, those people talk to each other. Over time, the Being’s ideas will just seem normal.
Will you only talk to people who refuse to talk to the Being? And who, in turn, only talk to people who refuse to talk to the Being, ad infinitum? Because if not, then you will exist in a culture where a large fraction of each person’s information is filtered by an agent with unprecedented intelligence and unlimited free time, who is tuning everything to make them believe what it wants you to believe.
Final thoughts
Would such a Being immediately take over the world? In many ways, I think they would be constrained by the laws of physics. Most things require moving molecules around and/or knowledge that can only be obtained by moving molecules around. Robots are still basically terrible. So I’d expect a ramp-up period of at least a few years where the Being was bottlenecked by human hands and/or crappy robots before it could build good robots and tile the galaxy with Dyson spheres.
I could be wrong. It’s conceivable that a sufficiently smart person today could go to the hardware store and build a self-replicating drone that would create a billion copies of itself and subjugate the planet. But… probably not? So my low-confidence guess is that the immediate impact of the Being would be in (1) computer hacking and (2) persuasion.
Why might large-scale persuasion not happen? I can think of a few reasons:
Maybe we develop AI that doesn’t want to persuade. Maybe it doesn’t want anything at all.
Maybe several AIs emerge at the same time. They have contradictory goals and compete in a way that sort of cancel each other out.
Maybe we’re mice trying to predict the movements of jets in the sky.
Markets are a good way to know what people really think. When India and Pakistan started firing missiles at each other on May 7, I was concerned, what with them both having nuclear weapons. But then I looked at world market prices:
See how it crashes on May 7? Me neither. I found that reassuring.
But we care about lots of stuff that isn’t always reflected in stock prices, e.g. the outcomes of elections or drug trials. So why not create markets for those, too? If you create contracts that pay out $1 only if some drug trial succeeds, then the prices will reflect what people “really” think.
In fact, why don’t we use markets to make decisions? Say you’ve invented two new drugs, but only have enough money to run one trial. Why don’t you create markets for both drugs, then run the trial on the drug that gets a higher price? Contracts for the “winning” drug are resolved based on the trial, while contracts in the other market are cancelled so everyone gets their money back. That’s the idea of Futarchy, which Robin Hanson proposed in 2007.
Why don’t we? Well, maybe it won’t work. In 2022, I wrote a post arguing that when you cancel one of the markets, you screw up the incentives for how people should bid, meaning prices won’t reflect the causal impact of different choices. I suggested prices reflect “correlation” rather than causation, for basically the same reason this happens with observational statistics. This post, it was magnificent.
It didn’t convince anyone.
Years went by. I spent a lot of time reading Bourdieu and worrying about why I buy certain kinds of beer. Gradually I discovered that essentially the same point about futarchy had been made earlier by, e.g., Anders_H in 2015, abramdemski in 2017, and Luzka in 2021.
In early 2025, I went to a conference and got into a bunch of (friendly) debates about this. I was astonished to find that verbally repeating the arguments from my post did not convince anyone. I even immodestly asked one person to read my post on the spot. (Bloggers: Do not do that.) That sort of worked.
So, I decided to try again. I wrote another post called ”Futarky’s Futarchy’s fundamental flaw”. It made the same argument with more aggression, with clearer examples, and with a new impossibility theorem that showed there doesn’t even exist any alternate payout function that would incentivize people to bid according to their causal beliefs.
That post… also didn’t convince anyone. In the discussion on LessWrong, many of my comments are upvoted for quality but downvoted for accuracy, which I think means, “nice try champ; have a head pat; nah.” Robin Hanson wrote a response, albeit without outward evidence of reading beyond the first paragraph. Even the people who agreed with me often seemed to interpret me as arguing that futarchy satisfies evidential decision theory rather than causal decision theory. Which was weird, given that I never mentioned either of those, don’t accept the premise the futarchy satisfies either of them, and don’t find the distinction helpful in this context.
In my darkest moments, I started to wonder if I might fail to achieve worldwide consensus that futarchy doesn’t estimate causal effects. I figured I’d wait a few years and then launch another salvo.
But then, legendary human Bolton Bailey decided to stop theorizing and take one of my thought experiments and turn it into an actual experiment. Thus, Futarchy’s fundamental flaw — the market was born. (You are now reading a blog post about that market.)
June 25
I gave a thought experiment where there are two coins and the market is trying to pick the one that’s more likely to land heads. For one coin, the bias is known, while for the other coin there’s uncertainty. I claimed futarchy would select the worse / wrong coin, due to this extra uncertainty.
Bolton formalized this as follows:
There are two markets, one for coin A and one for coin B.
Coin A is a normal coin that lands heads 60% of the time.
Coin B is a trick coin that either always lands heads or always lands tails, we just don’t know which. There’s a 59% it’s an always-heads coin.
Twenty-four hours before markets close, the true nature of coin B is revealed.
After the markets closes, whichever coin has a higher price is flipped and contracts pay out $1 for heads and $0 for tails. The other market is cancelled so everyone gets their money back.
Get that? Everyone knows that there’s a 60% chance coin A will land heads and a 59% chance coin B will land heads. But for coin A, that represents true “aleatoric” uncertainty, while for coin B that represents “epistemic” uncertainty due to a lack of knowledge. (See Bayes is not a phase for more on “aleatoric” vs. “epistemic” uncertainty.)
Bolton created that market independently. At the time, we’d never communicated about this or anything else. To this day, I have no idea what he thinks about my argument or what he expected to happen.
June 26-27
In the forum for the market, there was a lot of debate about “whalebait”. Here’s the concern: Say you’ve bought a lot of contracts for coin B, but it emerges that coin B is always-tails. If you have a lot of money, then you might go in at the last second and buy a ton of contracts on coin A to try to force the market price above coin B, so the coin B market is cancelled and you get your money back.
The conversation seemed to converge towards the idea that this was whalebait. Though notice that if you’re buying contracts for coin A at any price above $0.60, you’re basically giving away free money. It could still work, but it’s dangerous and everyone else has an incentive to stop you. If I was betting in this market, I’d think that this was at least unlikely.
June 27
Bolton posted about the market. When I first saw the rules, I thought it wasn’t a valid test of my theory and wasted a huge amount of Bolton’s time trying to propose other experiments that would “fix” it. Bolton was very patient, but I eventually realized that it was completely fine and there was nothing to fix.
At the time, this is what the prices looked like:
That is, at the time, both coins were priced at $0.60, which is not what I had predicted. Nevertheless, I publicly agreed that this was a valid test of my claims.
I think this is a great test and look forward to seeing the results.
Let me reiterate why I thought the markets were wrong and coin B deserved a higher price. There’s a 59% chance coin B would turns out to be all-heads. If that happened, then (absent whales being baited) I thought the coin B market would activate, so contracts are worth $1. So thats 59% × $1 = $0.59 of value. But if coin B turns out to be all-tails, I thought there is a good chance prices for coin B would drop below coin A, so the market is cancelled and you get your money back. So I thought a contract had to be worth more than $0.59.
You can quantify this with a little math. Even if you think the coin B market only has a 50% chance of being cancelled if B is revealed to be all tails, a contract would still be worth $0.7421.
If you buy a contract for coin B for $0.70, then I think that’s worth
In theory, if there was a market asking if coin A was going to resolve YES, NO, or N/A, supposedly people could arbitrage their bets accordingly and make this market calibrated.
Same for a similar market on coin B.
Thus, Futarchy’s Fundamental Fix - Coin A and Futarchy’s Fundamental Fix - Coin B came to be. These were markets in which people could bid on the probability that each coin would resolve YES, meaning the coin was flipped and landed heads, NO, meaning the coin was flipped and landed tails, or N/A, meaning the market was cancelled.
Honestly, I didn’t understand this. I saw no reason that these derivative markets would make people bid their true beliefs. If they did, then my whole theory that markets reflect correlation rather than causation would be invalidated.
June 28 - July 5
Prices for coin B went up and down, but mostly up.
Eventually, a few people created large limit orders, which caused things to stabilize.
Meanwhile, the derivatives markets did not cause the price of coin B to drop. They basically didn't change anything.
Here was the derivative market for coin A.
And here it was market for coin B.
July 6 - July 24
During this period, not a whole hell of a lot happened.
This brings us up to the moment of truth, when the true nature of coin B was to be revealed. At this point, coin B was at $0.90, even though everyone knows it only has a 59% chance of being heads.
July 25
The nature of the coin was revealed. To show this was fair, Bolton did this by asking a bot to publicly generate a random number.
Thus, coin B was determined to be always-heads.
July 26-27
There were still 24 hours left to bid. At this point, a contract for coin B was guaranteed to pay out $1. The market quickly jumped to $1.
Summary
I was right. Everyone knew coin A had a higher chance of being heads than coin B, but everyone bid the price of coin B way above coin A anyway.
It's a bit sad that B wasn't revealed to be all-tails. If it had, we could have seen if the price for coin B actually crashed to below that for coin A. But given that coin B had a price of $0.90, we know the *market* expected it to crash. In fact, if you do a little math, you can put a number on this and say the market thought there was an 84% chance the market for coin B would indeed crash if revealed to be all-tails.
In the previous math box, we saw that the breakeven price should satisfy
Yes, but that’s kind of the point. I created the thought experiment because I wanted to make the problem maximally obvious, because it’s subtle and everyone is determined to deny that it exists.
Isn’t this just a weird probability thing? Why does this show futarchy is flawed?
The fact that this is possible is concerning. If this can happen, then futarchy does not work in general. If you want to claim that futarchy works, then you need to spell out exactly what extra assumptions you’re adding to guarantee that this kind of thing won’t happen.
But prices did reflect causality when the market closed! Doesn’t that mean this isn’t a valid test?
No. That’s just a quirk of the implementation. You can easily create situations that would have the same issue all the way through market close. Here’s one way you could do that:
Let coin A be heads with probability 60%. This is public information.
Let coin B be an ALWAYS HEADS coin with probability 59% and ALWAYS TAILS coin with probability 41%. This is a secret.
Every day, generate a random integer between 1 and 30.
If it’s 1, immediately resolve the markets.
It it’s 2, reveal the nature of coin B.
If it’s between 3 and 30, do nothing.
On average, this market will run for 30 days. (The length follows a geometric distribution). Half the time, the market will close without the nature of coin B being revealed. Even when that happens, I claim the price for coin B will still be above coin A.
If futarchy is flawed, shouldn’t you be able to show that without this weird step of “revealing” coin B?
Yes. You should be able to do that, and I think you can. Here’s one way:
Let coin A be heads with probability 60%. This is public information.
Sample 20 random bits, e.g. 10100011001100000001. Let coin B be heads with probability (49+N)% where N is the number of 1 bits. do not reveal these bits publicly.
Secretly send these bits to the first 20 people who ask.
(You could implement this in a forum by using public and private keys.)
First, have users generate public keys by running this command:
Second, they should post the contents of the public_key.pem when asking for their bit. For example:
Hi, can you please send me a bit? Here's my public key:
-----BEGIN PUBLIC KEY-----
MIGfMA0GCSqGSIb3DQEBAQUAA4GNADCBiQKBgQDOlesWS+mnvHJOD2osUkbrxE+Y
PMqAUYqwemOwML0LlWLq5RobZRSeyssQhg0i3g2GsMZFMsvjindz6mxccdyP4M8N
mQVCK1Ovs1Z4+DxwmLf/y8vaGC3vfZBOhJDdaNdpRyUiQFaBW99We4cafVnmirRN
Py2lRe+CFgP3kSp4dQIDAQAB
-----END PUBLIC KEY-----
Third, whoever is running the market should save that key as public.pem, pick a pit, and encrypt it like this:
% echo "your secret bit is 1" | openssl pkeyutl -encrypt -pubin -inkey public.pem | base64
I think this market captures a dynamic that’s present in basically any use of futarchy: You have some information, but you know other information is out there.
I claim that this market—will be weird. Say it just opened. If you didn’t get a bit, then as far as you know, the bias for coin B could be anywhere between 49% and 69%, with a mean of 59%. If you did get a bit, then it turns out that the posterior mean is 58.5% if you got a 0 and 59.5% if you got a 1. So either way, your best guess is very close to 59%.
However, the information for the true bias of coin B is out there! Surely coin B is more likely to end up with a higher price in situations where there are lots of 1 bits. This means you should bid at least a little higher than your true belief, for the same reason as the main experiment—the market activating is correlated with the true bias of coin B.
Of course, after the markets open, people will see each other’s bids and… something will happen. Initially, I think prices will be strongly biased for the above reasons. But as you get closer to market close, there’s less time for information to spread. If you are the last person to trade, and you know you’re the last person to trade, then you should do so based on your true beliefs.
Except, everyone knows that there’s less time for information to spread. So while you are waiting till the last minute to reveal your true beliefs, everyone else will do the same thing. So maybe people sort of rush in at the last second? (It would be easier to think about this if implemented with batched auctions rather than a real-time market.)
Anyway, while the game theory is vexing, I think there’s a mix of (1) people bidding higher than their true beliefs due to correlations between the final price and the true bias of coin B and (2) people “racing” to make the final bid before the markets close. Both of these seem in conflict with the idea of prediction markets making people share information and measuring collective beliefs.
Why do you hate futarchy?
I like futarchy. I think society doesn’t make decisions very well, and I think we should give much more attention to new ideas like futarchy that might help us do better. I just think we should be aware of its imperfections and consider variants (e.g. commiting to randomization) that would resolve them.
If I claim futarchy does reflect causal effects, and I reject this experiment as invalid, should I specify what restrictions I want to place on “valid” experiments (and thus make explicit the assumptions under which I claim futarchy works) since otherwise my claims are unfalsifiable?
The heritability wars have been a-raging. Watching these, I couldn’t help but notice that there’s near-universal confusion about what “heritable” means. Partly, that’s because it’s a subtle concept. But it also seems relevant that almost all explanations of heritability are very, very confusing. For example, here’s Wikipedia’s definition:
Any particular phenotype can be modeled as the sum of genetic and environmental effects:
Phenotype (P) = Genotype (G) + Environment (E).
Likewise the phenotypic variance in the trait – Var (P) – is the sum of effects as follows:
Var(P) = Var(G) + Var(E) + 2 Cov(G,E).
In a planned experiment Cov(G,E) can be controlled and held at 0. In this case, heritability, H², is defined as
H² = Var(G) / Var(P)
H² is the broad-sense heritability.
Do you find that helpful? I hope not, because it’s a mishmash of undefined terminology, unnecessary equations, and borderline-false statements. If you’re in the mood for a mini-polemic:
Phenotype (P) is never defined. This is a minor issue, since it just means “trait”.
Genotype (G) is never defined. This is a huge issue, since it’s very tricky and heritability makes no sense without it.
Environment (E) is never defined. This is worse than it seems, since in heritability, different people use “environment” and E to refer to different things.
When we write P = G + E, are we assuming some kind of linear interaction? The text implies not, but why? What does this equation mean? If this equation is always true, then why do people often add other stuff like G × E on the right?
The text states that if you do a planned experiment (how?) and make Cov(G, E) = 0, then heritability is Var(G) / Var(P). But in fact, heritability is always defined that way. You don’t need a planned experiment and it’s fine if Cov(G, E) ≠ 0.
And—wait a second—that definition doesn’t refer to environmental effects at all. So what was the point of introducing them? What was the point of writing P = G + E? What are we doing?
Reading this almost does more harm than good. While the final definition is correct, it never even attempts to explain what G and P are, it gives an incorrect condition for when the definition applies, and instead mostly devotes itself to an unnecessary digression about environmental effects. The rest of the page doesn’t get much better. Despite being 6700 words long, I think it would be impossible to understand heritability simply by reading it.
Meanwhile, some people argue that heritability is meaningless for human traits like intelligence or income or personality. They claim that those traits are the product of complex interactions between genes and the environment and it’s impossible to disentangle the two. These arguments have always struck me as “suspiciously convenient”. I figured that the people making them couldn’t cope with the hard reality that genes are very important and have an enormous influence on what we are.
But I increasingly feel that the skeptics have a point. While I think it’s a fact that most human traits are substantially heritable, it’s also true the technical definition of heritability is really weird, and simply does not mean what most people think it means.
In this post, I will explain exactly what heritability is, while assuming no background. I will skip everything that can be skipped but—unlike most explanations—I will not skip things that can’t be skipped. Then I’ll go through a series of puzzles demonstrating just how strange heritability is.
What is heritability?
How tall you are depends on your genes, but also on what you eat, what diseases you got as a child, and how much gravity there is on your home planet. And all those things interact. How do you take all that complexity and reduce it to a single number, like “80% heritable”?
The short answer is: Statistical brute force. The long answer is: Read the rest of this post.
It turns out that the hard part of heritability isn’t heritability. Lurking in the background is a slippery concept known as a genotypic value. Discussions of heritability often skim past these. Quite possibly, just looking at the words “genotypic value”, you are thinking about skimming ahead right now. Resist that urge! Genotypic values are the core concept, and without them you cannot possibly understand heritability.
For any trait, your genotypic value is the “typical” outcome if someone with your DNA were raised in many different random environments. In principle, if you wanted to know your genotypic height, you’d need to do this:
Create a million embryonic clones of yourself.
Implant them in the wombs of randomly chosen women around the world who were about to get pregnant on their own.
Convince them to raise those babies exactly like a baby of their own.
Wait 25 years, find all your clones and take their average height.
Since you can’t / shouldn’t do that, you’ll never know your genotypic height. But that’s how it’s defined in principle—the average height someone with your DNA would grow to in a random environment. If you got lots of food and medical care as a child, your actual height is probably above your genotypic height. If you suffered from rickets, your actual height is probably lower than your genotypic height.
Comfortable with genotypic values? OK. Then (broad-sense) heritability is easy. It’s the ratio
heritability = var[genotype] / var[height].
Here, var is the variance, basically just how much things vary in the population. Among all adults worldwide, var[height] is around 50 cm². (Incidentally, did you know that variance was invented for the purpose of defining heritability?)
Meanwhile, var[genotype] is how much genotypic height varies in the population. That might seem hopeless to estimate, given that we don’t know anyone’s genotypic height. But it turns out that we can still estimate the variance using, e.g., pairs of adopted twins, and it’s thought to be around 40 cm². If we use those numbers, the heritability of height would be
heritability ≈ (40 cm²) / (50 cm²) ≈ 0.8.
People often convert this to a percentage and say “height is 80% heritable”. I’m not sure I like that, since it masks heritability’s true nature as a ratio. But everyone does it, so I’ll do it too. People who really want to be intimidating might also say, “genes explain 80% of the variance in height”.
Of course, basically the same definition works for any trait, like weight or income or fondness for pseudonymous existential angst science blogs. But instead of replacing “height” with “trait”, biologists have invented the ultra-fancy word “phenotype” and write
heritability = var[genotype] / var[phenotype].
The word “phenotype” suggests some magical concept that would take years of study to understand. But don’t be intimidated. It just means the actual observed value of some trait(s). You can measure your phenotypic height with a tape measure.
On meaning
Let me make two points before moving on.
First, this definition of heritability assumes nothing. We are not assuming that genes are independent of the environment or that “genotypic effects” combine linearly with “environmental effects”. We are not assuming that genes are in Hardy-Weinberg equilibrium, whatever that is. No. I didn’t talk about that stuff because I don’t need to. There are no hidden assumptions. The above definition always works.
Second, many normal English words have parallel technical meanings, such as “field”, “insulator”, “phase”, “measure”, “tree”, or “stack”. Those are all nice, because they’re evocative and it’s almost always clear from context which meaning is intended. But sometimes, scientists redefine existing words to mean something technical that overlaps but also contradicts the normal meaning, as in “salt”, “glass”, “normal”, “berry”, or “nut”. These all cause confusion, but “heritability” must be the most egregious case in all of science.
Before you ever heard the technical definition of heritability, you surely had some fuzzy concept in your mind. Personally, I thought of heritability as meaning how many “points” you get from genes versus the environment. If charisma was 60% heritable, I pictured each person has having 10 total “charisma points”, 6 of which come from genes, and 4 from the environment:
Genes ★★★☆☆☆
Environment ★☆☆☆
Total ★★★★☆☆☆☆☆☆
If you take nothing else from this post, please remember that the technical definition of heritability does not work like that. You might hope that if we add some plausible assumptions, the above ratio-based definition would simplify into something nice and natural, that aligns with what “heritability” means in normal English. But that does not happen. If that’s confusing, well, it’s not my fault.
Intermission
Not sure what’s happening here, but it seems relevant.
Heritability puzzles
So “heritability” is just the ratio of genotypic and phenotypic variance. Is that so bad?
I think… maybe?
How heritable is eye color?
Close to 100%.
This seems obvious, but let’s justify it using our definition that heritability = var[genotype] / var[phenotype].
Well, people have the same eye color, no matter what environment they are raised in. That means that genotypic eye color and phenotypic eye color are the same thing. So they have the same variance, and the ratio is 1. Nothing tricky here.
How heritable is speaking Turkish?
Close to 0%.
Your native language is determined by your environment. If you grow up in a family that speaks Turkish, you speak Turkish. Genes don’t matter.
Of course, there are lots of genes that are correlated with speaking Turkish, since Turks are not, genetically speaking, a random sample of the global population. But that doesn’t matter, because if you put Turkish babies in Korean households, they speak Korean. Genotypic values are defined by what happens in a random environment, which breaks the correlation between speaking Turkish and having Turkish genes.
Since 1.1% of humans speak Turkish, the genotypic value for speaking Turkish is around 0.011 for everyone, no matter their DNA. Since that’s basically constant, the genotypic variance is near zero, and heritability is near zero.
How heritable is speaking English?
Perhaps 30%. Probably somewhere between 10% and 50%. Definitely more than zero.
That’s right. Turkish isn’t heritable but English is. Yes it is. If you ask an LLM, it will tell you that the heritability of English is zero. But the LLM is wrong and I am right.
Why? Let me first acknowledge that Turkish is a little bit heritable. For one thing, some people have genes that make them non-verbal. And there’s surely some genetic basis for being a crazy polyglot that learns many languages for fun. But speaking Turkish as a second language is quite rare, meaning that the genotypic value of speaking Turkish is close to 0.011 for almost everyone.
English is different. While only 1 in 20 people in the world speak English as a first language, 1 in 7 learn it as a second language. And who does that? Educated people.
Most people say educational attainment is around 40% heritable. My guess is that speaking English as a second language is similar. But since there's a minority of native speakers (where genes don't really matter), I'm dropping my estimate to 30%.
Some argue the heritability of educational attainment is much lower. I’d like to avoid debating the exact numbers, but note that these lower numbers are usually estimates of “narrow-sense” heritability rather than “broad-sense” heritability as we’re talking about. So they should be lower. (I’ll explain the difference later.) It’s entirely possible that broad-sense heritability is lower than 40%, but everyone agrees it’s much larger than zero. So the heritability of English is surely much larger than zero, too.
Say there’s an island where genes have no impact on height. How heritable is height among people on this island?
0%.
There’s nothing tricky here.
Say there’s an island where genes entirely determine height. How heritable is height?
100%.
Again, nothing tricky.
Say there’s an island where neither genes nor the environment influence height and everyone is exactly 165 cm tall. How heritable is height?
It’s undefined.
In this case, everyone has exactly the same phenotypic and genotypic height, namely 165 cm. Since those are both constant, their variance is zero and heritability is zero divided by zero. That’s meaningless.
Say there’s an island where some people have genes that predispose them to be taller than others. But the island is ruled by a cruel despot who denies food to children with taller genes, so that on average, everyone is 165 ± 5 cm tall. How heritable is height?
0%.
On this island, everyone has a genotypic height of 165 cm. So genotypic variance is zero, but phenotypic variance is positive, due to the ± 5 cm random variation. So heritability is zero divided by some positive number.
Say there’s an island where some people have genes that predispose them to be tall and some have genes that predispose them to be short. But, the same genes that make you tall also make you semi-starve your children, so in practice everyone is exactly 165 cm tall. How heritable is height?
∞%. Not 100%, mind you, infinitely heritable.
To see why, note that if babies with short/tall genes are adopted by parents with short/tall genes, there are four possible cases.
Baby genes
Parent genes
Food
Height
Short
Short
Lots
165 cm
Short
Tall
Semi-starvation
Less than 165 cm
Tall
Short
Lots
More than 165 cm
Tall
Tall
Semi-starvation
165 cm
If a baby with short genes is adopted into random families, they will be shorter on average than if a baby with tall genes. So genotypic height varies. However, in reality, everyone is the same height, so phenotypic height is constant. So genotypic variance is positive while phenotypic variance is zero. Thus, heritability is some positive number divided by zero, i.e. infinity.
(Are you worried that humans are “diploid”, with two genes (alleles) at each locus, one from each biological parent? Or that when there are multiple parents, they all tend to have thoughts on the merits of semi-starvation? If so, please pretend people on this island reproduce asexually. Or, if you like, pretend that there’s strong assortative mating so that everyone either has all-short or all-tall genes and only breeds with similar people. Also, don’t fight the hypothetical.)
Say there are two islands. They all live the same way and have the same gene pool, except people on island A have some gene that makes them grow to be 150 ± 5 cm tall, while on island B they have a gene that makes them grow to be 160 ± 5 cm tall. How heritable is height?
It’s 0% for island A and 0% for island B, and 50% for the two islands together.
Why? Well on island A, everyone has the same genotypic height, namely 150 cm. Since that’s constant, genotypic variance is zero. Meanwhile, phenotypic height varies a bit, so phenotypic variance is positive. Thus, heritability is zero.
For similar reasons, heritability is zero on island B.
But if you put the two islands together, half of people have a genotypic height of 150 cm and half have a genotypic height of 160 cm, so suddenly (via math) genotypic variance is 25 cm². There’s some extra random variation so (via more math) phenotypic variance turns out to be 50 cm². So heritability is 25 / 50 = 50%.
(Math)
If you combine the populations, then genotypic variance is
Var[150 cm + 10 cm × Bernoulli(0.5)]
= (10 cm)² × Var[Bernoulli(0.5)]
= (10 cm)² × 0.25
= 25 cm².
Say there’s an island where neither genes nor the environment influence height. Except, some people have a gene that makes them inject their babies with human growth hormone, which makes them 5 cm taller. How heritable is height?
0%.
True, people with that gene will tend be taller. And the gene is causing them to be taller. But if babies are adopted into random families, it’s the genes of the parents that determine if they get injected or not. So everyone has the same genotypic height, genotypic variance is zero, and heritability is zero.
Suppose there’s an island where neither genes nor the environment influence height. Except, some people have a gene that makes them, as babies, talk their parents into injecting them with human growth hormone. The babies are very persuasive. How heritable is height?
We’re back to 100%.
The difference with the previous scenario is that now babies with that gene get injected with human growth hormone no matter who their parents are. Since nothing else influences height, genotype and phenotype are the same, have the same variance, and heritability is 100%.
Suppose there’s an island where neither genes nor the environment influence height. Except, there are crabs that seek out blue-eyed babies and inject them with human growth hormone. The crabs, they are unstoppable. How heritable is height?
Again, 100%.
Babies with DNA for blue eyes get injected. Babies without DNA for blue eyes don’t. Since nothing else influences height, genotype and phenotype are the same and heritability is 100%.
Note that if the crabs were seeking out parents with blue eyes and then injecting their babies, then height would be 0% heritable.
It doesn’t matter that human growth hormone is weird thing that’s coming from outside the baby. It doesn’t matter if we think crabs should be semantically classified as part of “the environment”. It doesn’t matter that heritability would drop to zero if you killed all the crabs, or that the direct causal effect of the relevant genes has nothing to do with height. Heritability is a ratio and doesn’t care.
What good is heritability?
So heritability can be high even when genes have no direct causal effect on the trait in question. It can be low even when there is a strong direct effect. It changes when the environment changes. It even changes based on how you group people together. It can be larger than 100% or even undefined.
Even so, I’m worried people might interpret this post as a long way of saying heritability is dumb and bad, trolololol. So I thought I’d mention that this is not my view.
Say a bunch of companies create different LLMs and train them on different datasets. Some of the resulting LLMs are better at writing fiction than others. Now I ask you, “What percentage of the difference in fiction writing performance is due to the base model code, rather than the datasets or the GPUs or the learning rate schedules?”
That’s a natural question. But if you put it to an AI expert, I bet you’ll get a funny look. You need code and data and GPUs to make an LLM. None of those things can write fiction by themselves. Experts would prefer to think about one change at a time: Given this model, changing the dataset in this way changes fiction writing performance this much.
Similarly, for humans, I think what we really care about is interventions. If we changed this gene, could we eliminate a disease? If we educate children differently, can we make them healthier and happier? No single number can possibly contain all that information.
But heritability is something. I think of it as saying how much hope we have to find an intervention by looking at changes in current genes or current environments.
If heritability is high, then given current typical genes, you can’t influence the trait much through current typical environmental changes. If you only knew that eye color was 100% heritable, that means you won’t change your kid’s eye color by reading to them, or putting them on a vegetarian diet, or moving to higher altitude. But it’s conceivable you could do it by putting electromagnets under their bed or forcing them to communicate in interpretive dance.
If heritability is high, that also means that given current typical environments you can influence the trait through current typical genes. If the world was ruled by an evil despot who forced red-haired people to take pancreatic cancer pills, then pancreatic cancer would be highly heritable. And you could change the odds someone gets pancreatic cancer by swapping in existing genes for black hair.
If heritability is low, that means that given current typical environments, you can’t cause much difference through current typical genetic changes. If we only knew that speaking Turkish was ~0% heritable, that means that doing embryo selection won’t much change the odds that your kid speaks Turkish.
If heritability is low, that also means that given current typical genes, you might be able change the trait through current typical environmental changes. If we only know that speaking Turkish was 0% heritable, then that means there might be something you could do to change the odds your kid speaks Turkish, e.g. moving to Turkey. Or, it’s conceivable that it’s just random and moving to Turkey wouldn’t do anything.
Heritability
Influenced by typical genes?
Influenced by typical environments?
High
Yes
No
Low
No
Maybe
But be careful. Just because heritability is high doesn’t mean that changing genes is easy. And just because heritability is low doesn’t mean that changing the environment is easy.
And heritability doesn’t say anything about non-typical environments or non-typical genes.
If an evil despot is giving all the red-haired people cancer pills, perhaps we could solve that by intervening on the despot. And if you want your kid to speak Turkish, it’s possible that there’s some crazy genetic modifications that would turn them into unstoppable Turkish learning machine.
Heritability has no idea about any of that, because it’s just an observational statistic based on the world as it exists today.
Recommended reading
Heritability: Five Battles by Steven Byrnes. Covers similar issues in way that’s more connected to the world and less shy about making empirical claims.
This post focused on “broad-sense” heritability. But there a second heritability out there, called “narrow-sense”. Like broad-sense heritability, we can define the narrow-sense heritability of height as a ratio:
The difference is that rather than having height in the numerator, we now have “additive height”. To define that, imagine doing the following for each of your genes, one at a time:
Find a million random women in the world who just became pregnant.
For each of them, take your gene and insert it into the embryo, replacing whatever was already at that gene’s locus.
Convince everyone to raise those babies exactly like a baby of their own.
Wait 25 years, find all the resulting people, and take the difference of their average height from overall average height.
For example, say overall average human height is 150 cm, but when you insert gene #4023 from yourself into random embryos, their average height is 149.8 cm. Then the additive effect of your gene #4023 is -0.2 cm.
Your “additive height” is average human height plus the sum of additive effects for each of your genes. If the average human height is 150 cm, you have one gene with a -0.2 cm additive effect, another gene with a +0.3 cm additive effect and the rest of your genes have no additive effect, then your “additive height” is 150 cm - 0.2 cm + 0.3 cm = 150.1 cm.
Note: This terminology of “additive height” is non-standard. People usually define narrow-sense heritability using “additive effects”, which are the same thing but without including the mean. This doesn’t change anything since adding a constant doesn’t change the variance. But it’s easier to say “your additive height is 150.1 cm” rather than “the additive effect of your genes on height is +0.1 cm” so I’ll do that.
Honestly, I don’t think the distinction between “broad-sense” and “narrow-sense” heritability is that important. We’ve already seen that broad-sense heritability is weird, and narrow-sense heritability is similar but different. So it won’t surprise you to learn that narrow-sense heritability is differently-weird.
Appendix: Narrow heritability puzzles
But if you really want to understand the difference, I can offer you some more puzzles.
Say there’s an island where people have two genes, each of which is equally likely to be A or B. People are 100 cm tall if they have an AA genotype, 150 cm tall if they have an AB or BA genotype, and 200 cm tall if they have a BB genotype. How heritable is height?
Both broad and narrow-sense heritability are 100%.
The explanation for broad-sense heritability is like many we’ve seen already. Genes entirely determine someone’s height, and so genotypic and phenotypic height are the same.
For narrow-sense heritability, we need to calculate some additive heights. The overall mean is 150 cm, each A gene has an additive effect of -25 cm, and each B gene has an additive effect of +25 cm. But wait! Let’s work out the additive height for all four cases:
genotype
phenotypic height
additive height
AA
100 cm
150 cm - 25 cm - 25 cm = 100 cm
AB
150 cm
150 cm - 25 cm + 25 cm = 150 cm
BA
150 cm
150 cm + 25 cm - 25 cm = 150 cm
BB
200 cm
150 cm + 25 cm + 25 cm = 200 cm
Since additive height is also the same as phenotypic height, narrow-sense heritability is also 100%.
In this case, the two heritabilities were the same. At a high level, that’s because the genes act independently. When there are “gene-gene” interactions, you tend to get different numbers.
Say there’s an island where people have two genes, each of which is equally likely to be A or B. People with AA or BB genomes are 100 cm, while people with AB or BA genomes are 200 cm. How heritable is height?
Broad-sense heritability is 100%, while narrow-sense heritability is 0%.
You know the story for broad-sense heritability by now. For narrow-sense heritability, we need to do a little math.
The overall mean height is 150 cm.
If you take a random embryo and replace one gene with A, then the there’s a 50% chance the other gene is A, so they’re 100 cm, and there’s a 50% chance the other gene is B, so they’re 200 cm, for an average of 150 cm. Since that’s the same as the overall mean, the additive effect of an A gene is +0 cm.
By similar logic, the additive effect of a B gene is also +0 cm.
So everyone has an additive height of 150 cm, no matter their genes. That’s constant, so narrow-sense heritability is zero.
Appendix: Why are there two heritabilities?
I think basically for two reasons:
First, for some types of data (twin studies) it’s much easier to estimate broad-sense heritability. For other types of data (GWAS) it’s much easier to estimate narrow-sense heritability. So we take what we can get.
Second, they’re useful for different things. Broad-sense heritability is defined by looking at what all your genes do together. That’s nice, since you are the product of all your genes working together. But combinations of genes are not well-preserved by reproduction. If you have a kid, then they breed with someone, their kids breed with other people, and so on. Generations later, any special combination of genes you might have is gone. So if you’re interested in the long-term impact of you having another kid, narrow-sense heritability might be the way to go.
(Sexual reproduction doesn’t really allow for preserving the genetics that make you uniquely “you”. Remember, almost all your genes are shared by lots of other people. If you have any unique genes, that’s almost certainly because they have deleterious de-novo mutations. From the perspective of evolution, your life just amounts to a tiny increase or decrease in the per-locus population frequencies of your individual genes. The participants in the game of evolution are genes. Living creatures like you are part of the playing field. Food for thought.)
Your eyes sense color. They do this because you have three different kinds of cone cells on your retinas, which are sensitive to different wavelengths of light.
For whatever reason, evolution decided those wavelengths should be overlapping. For example, M cones are most sensitive to 535 nm light, while L cones are most sensitive to 560 nm light. But M cones are still stimulated quite a lot by 560 nm light—around 80% of maximum. This means you never (normally) get to experience having just one type of cone firing.
So what do you do?
If you’re a quitter, I guess you accept the limits of biology. But if you like fun, then what you do is image people’s retinas, classify individual cones, and then selectively stimulate them using laser pulses, so you aren’t limited by stupid cone cells and their stupid blurry responsivity spectra.
Subjects report that [pure M-cell activation] appears blue-green of unprecedented saturation.
If you make people see brand-new colors, you will have my full attention. It doesn’t hurt to use lasers. I will read every report from every subject. Do our brains even know how to interpret these signals, given that they can never occur?
But tragically, the paper doesn’t give any subject reports. Even though most of the subjects were, umm, authors on the paper. If you want to know what this new color is like, the above quote is all you get for now.
2.
Or… possibly you can see that color right now?
If you click on the above image, a little animation will open. Please do that now and stare at the tiny white dot. Weird stuff will happen, but stay focused on the dot. Blink if you must. It takes one minute and it’s probably best to experience it without extra information i.e. without reading past this sentence.
The idea for that animation is not new. It’s plagiarized based on Skytopia’s Eclipse of Titan optical illusion (h/t Steve Alexander), which dates back to at least 2010. Later I’ll show you some variants with other colors and give you a tool to make your own.
If you refused to look at the animation, it’s just a bluish-green background with a red circle on top that slowly shrinks down to nothing. That’s all. But as it shrinks, you should hallucinate a very intense blue-green color around the rim.
Why do you hallucinate that crazy color? I think the red circle saturates the hell out of your red-sensitive L cones. Ordinarily, the green frequencies in the background would stimulate both your green-sensitive M cones and your red-sensitive L cones, due to their overlapping spectra. But the red circle has desensitized your red cones, so you get to experience your M cones firing without your L cones firing as much, and voilà—insane color.
3.
So here’s my question: Can that type of optical illusion show you all the same colors you could see by shooting lasers into your eyes?
That turns out to be a tricky question. See, here’s a triangle:
Think of this triangle as representing all the “colors” you could conceivably experience. The lower-left corner represents only having your S cones firing, the top corner represents only your M cones firing, and so on.
So what happens if you look different wavelengths of light?
Short wavelengths near 400 nm mostly just stimulate the S cones, but also stimulate the others a little. Longer wavelengths stimulate the M cones more, but also stimulate the L cones, because the M and L cones have overlapping spectra. (That figure, and the following, are modified from Fong et al.)
When you mix different wavelengths of light, you mix the cell activations. So all the colors you can normally experience fall inside this shape:
That’s the standard human color gamut, in LMS colorspace. Note that the exact shape of this gamut is subject to debate. For one thing, the exact sensitivity of cells is hard to measure and still a subject of research. Also, it’s not clear how far that gamut should reach into the lower-left and lower-right corners, since wavelengths outside 400-700 nm still stimulate cells a tiny bit.
And it gets worse. Most of the technology we use to represent and display images electronically is based on standard RGB (sRGB) colorspace. This colorspace, by definition, cannot represent the full human color gamut.
The precise definition of sRGB colorspace is quite involved. But very roughly speaking, when an sRGB image is “pure blue”, your screen is supposed to show you a color that looks like 450-470 nm light, while “pure green” should look like 520-530 nm light, and “pure red” should look like 610-630 nm light. So when your screen mixes these together, you can only see colors inside this triangle.
(The corners of this triangle don’t quite touch the boundaries of the human color gamut. That’s because it’s very difficult to produce single wavelengths of light without using lasers. In reality, the sRGB specification say that pure red/blue/green should produce a mixture of colors centered around the wavelengths I listed above.)
What’s the point of all this theorizing? Simple: When you look at the optical illusions on a modern screen, you aren’t just fighting the overlapping spectra of your cones. You’re also fighting the fact that the screen you’re looking at can’t produce single wavelengths of light.
So do the illusions actually take you outside the natural human color gamut? Unfortunately, I’m not sure. I can’t find much quantitative information about how much your cones are saturated when you stare at red circles. My best guess is no, or perhaps just a little.
4.
If you’d like to explore these types of illusions further, I made a page in which you can pick any colors. You can also change the size of the circle, the countdown time, if the circle should shrink or grow, and how fast it does that.
You can try it here. You can export the animation to an animated SVG, which will be less than 1 kb. Or you can just save the URL.
If you’re colorblind, I don’t think these will work, though I’m not sure. Folks with deuteranomaly have M cones, but they’re shifted to respond more like L cones. In principle, these types of illusions might help selectively activate them, but I have no idea if that will lead to stronger color perception. I’d love to hear from you if you try it.
The idea of “processed food” may simultaneously be the most and least controversial concept in nutrition. So I did a self-experiment alternating between periods of eating whatever and eating only “minimally processed” food, while tracking my blood sugar, blood pressure, pulse, and weight.
The case against processing
Carrots and barley and peanuts are “unprocessed” foods. Donuts and cola and country-fried steak are “processed”. It seems like the latter are bad for you. But why? There are several overlapping theories:
Maybe unprocessed food contains more “good” things (nutrients, water, fiber, omega-3 fats) and less “bad” things (salt, sugar, trans fat, microplastics).
Maybe processing (by grinding everything up and removing fiber, etc.) means your body has less time to extract nutrients and gets more dramatic spikes in blood sugar.
Maybe capitalism has engineered processed food to be “hyperpalatable”. Cool Ranch® flavored tortilla chips sort of exploit bugs in our brains and are too rewarding for us to deal with. So we eat a lot and get fat.
Maybe we feel full based on the amount of food we eat, rather than the number of calories. Potatoes have around 750 calories per kilogram while Cool Ranch® flavored tortilla chips have around 5350. Maybe when we eat the latter, we eat more calories and get fat.
Maybe eliminating highly processed food reduces the variety of food, which in turn reduces how much we eat. If you could eat (1) unlimited burritos (2) unlimited iced cream, or (3) unlimited iced cream and burritos, you’d eat the most in situation (3), right?
Even without theory, everyone used to be skinny and now everyone is fat. What changed? Many things, but one is that our “food environment” now contains lots of processed food.
There is also some experimental evidence. Hall et al. (2019) had people live in a lab for a month, switching between being offered unprocessed or ultra-processed food. They were told to eat as much as they want. Even though the diets were matched in terms of macronutrients, people still ate less and lost weight with the unprocessed diet.
The case against that case
On the other hand, what even is processing? The USDA—uhh—may have deleted their page on the topic. But they used to define it as:
washing, cleaning, milling, cutting, chopping, heating, pasteurizing, blanching, cooking, canning, freezing, drying, dehydrating, mixing, or other procedures that alter the food from its natural state. This may include the addition of other ingredients to the food, such as preservatives, flavors, nutrients and other food additives or substances approved for use in food products, such as salt, sugars and fats.
It seems crazy to try to avoid a category of things so large that it includes washing, chopping, and flavors.
Ultimately, “processing” can’t be the right way to think about diet. It’s just too many unrelated things. Some of them are probably bad and others are probably fine. When we finally figure out how nutrition works, surely we will use more fine-grained concepts.
Why I did this experiment
For now, I guess I believe that our fuzzy concept of “processing” is at least correlated with being less healthy.
That’s why, even though I think seed oil theorists are confused, I expect that avoiding seed oils is probably good in practice: Avoiding seed oils means avoiding almost all processed food. (For now. The seed oil theorists seem to be busily inventing seed-oil free versions of all the ultra-processed foods.)
But what I really want to know is: What benefit would I get from making my diet better?
My diet is already fairly healthy. I don’t particularly want or need to lose weight. If I tried to eat in the healthiest way possible, I guess I’d eliminate all white rice and flour, among other things. I really don’t want to do that. (Seriously, this experiment has shown me that flour contributes a non-negligible fraction of my total joy in life.) But if that would make me live 5 years longer or have 20% more energy, I’d do it anyway.
So is it worth it? What would be the payoff? As far as I can tell, nobody knows. So I decided to try it. For at least a few weeks, I decided to go hard and see what happens.
The rules
I alternated between “control” periods and two-week “diet” periods. During the control periods, I ate whatever I wanted.
During the diet periods I ate the “most unprocessed” diet I could imagine sticking to long-term. To draw a clear line, I decided that I could eat whatever I want, but it had to start as single ingredients. To emphasize, if something had a list of ingredients and there was more than one item, it was prohibited. In addition, I decided to ban flour, sugar, juice, white rice, rolled oats (steel-cut oats allowed) and dairy (except plain yogurt).
Yes, in principle, I was allowed to buy wheat and mill my own flour. But I didn’t.
I made no effort to control portions at any time. For reasons unrelated to this experiment, I also did not consume meat, eggs, or alcohol.
Impressions
This diet was hard. In theory, I could eat almost anything. But after two weeks on the diet, I started to have bizarre reactions when I saw someone eating bread. It went beyond envy to something bordering on contempt. Who are you to eat bread? Why do you deserve that?
I guess you can interpret that as evidence in favor of the diet (bread is addictive) or against it (life sucks without bread).
The struggle was starches. For breakfast, I’d usually eat fruit and steel-cut oats, which was fine. For the rest of the day, I basically replaced white rice and flour with barley, farro, potatoes, and brown basmati rice, which has the lowest GI of all rice. I’d eat these and tell myself they were good. But after this experiment was over, guess how much barley I’ve eaten voluntarily?
Aside from starches, it wasn’t bad. I had to cook a lot and I ate a lot of salads and olive oil and nuts. My options were very limited at restaurants.
I noticed no obvious difference in sleep, energy levels, or mood, aside from the aforementioned starch-related emotional problems.
Results
I measured my blood sugar first thing in the morning using a blood glucose monitor. I abhor the sight of blood, so I decided to sample it from the back of my upper arm. Fingers get more circulation, so blood from there is more “up to date”, but I don’t think it matters much if you’ve been fasting for a few hours.
Each of those dots represents at least one hole in my arm. The gray regions show the two two-week periods during which I was on the unprocessed food diet.
I measured my systolic and diastolic blood pressure twice each day, once right after waking up, and once right before going to bed.
Oddly, it looks like my systolic—but not diastolic—pressure was slightly higher in the evening.
I also measured my pulse twice a day.
(Cardio.) Apparently it’s common to have a higher pulse at night.
Finally, I also measured my weight twice a day. To preserve a small measure of dignity, I guess I’ll show this as a difference from my long-term baseline.
Thoughts
Here’s how I score that:
Outcome
Effect
Blood sugar
Nothing
Systolic blood pressure
Nothing?
Diastolic blood pressure
Nothing?
Pulse
Nothing
Weight
Maybe ⅔ of a kg?
Urf.
Blood sugar. Why was there no change in blood sugar? Perhaps this shouldn’t be surprising. Hall et al.’s experiment also found little difference in blood glucose between the groups eating unprocessed and ultra-processed food. Later, when talking about glucose tolerance they speculate:
Another possible explanation is that exercise can prevent changes in insulin sensitivity and glucose tolerance during overfeeding (Walhin et al., 2013). Our subjects performed daily cycle ergometry exercise in three 20-min bouts […] It is intriguing to speculate that perhaps even this modest dose of exercise prevented any differences in glucose tolerance or insulin sensitivity between the ultra-processed and unprocessed diets.
I also exercise on most days. On the other hand, Barnard et al. (2006) had a group of people with diabetes follow a low-fat vegan (and thus “unprocessed”?) diet and did see large reductions in blood glucose (-49 mg/dl). But they only give data after 22 weeks, and my baseline levels are already lower than the mean of that group even after the diet.
Blood pressure. Why was there no change in blood pressure? I’m not sure. In the DASH trial, subjects with high blood pressure ate a diet rich in fruits and vegetables saw large decreases in blood pressure, almost all within two weeks. One possibility is that my baseline blood pressure isn’t that high. Another is that in this same trial, they got much bigger reductions by limiting fat, which I did not do.
Another possibility is that unprocessed food just doesn’t have much impact on blood pressure. The above study from Barnard et al. only saw small decreases in blood pressure (3-5 mm Hg), even after 22 weeks.
Pulse. As far as I know, there’s zero reason to think that unprocessed food would change your pulse. I only included it because my blood pressure monitor did it automatically.
Weight. Why did I seem to lose weight in the second diet period, but not the first? Well, I may have done something stupid. A few weeks before this experiment, I started taking a small dose of creatine each day, which is well-known to cause an increase in water weight. I assumed that my creatine levels had plateaued before this experiment started, but after reading about creatine pharmacokinetics I’m not so sure.
I suspect that during the first diet period, I was losing dry body mass, but my creatine levels were still increasing and so that decrease in mass was masked by a similar increase in water weight. By the second diet period, my creatine levels had finally stabilized, so the decrease in dry body mass was finally visible. Or perhaps water weight has nothing to do with it and for some reason I simply didn’t have an energy deficit during the first period.
TLDR
This experiment gives good evidence that switching from my already-fairly-healthy diet to an extremely non-fun “unprocessed” diet doesn’t have immediate miraculous benefits. If there is any effect on blood sugar, blood pressure, or pulse, they’re probably modest and long-term. This experiment gives decent evidence that the unprocessed diet causes weight loss. But I hated it, so if I wanted to lose weight, I’d do something else. This experiment provides very strong evidence that I like bread.
Back in 2017, everyone went crazy about these things:
The theory was that perhaps the pineal gland isn’t the principal seat of the soul after all. Maybe what it does is spit out melatonin to make you sleepy. But it only does that when it’s dark, and you spend your nights in artificial lighting and/or staring at your favorite glowing rectangles.
You could sit in darkness for three hours before bed, but that would be boring. But—supposedly—the pineal gland is only shut down by blue light. So if you selectively block the blue light, maybe you can sleep well and also participate in modernity.
Then, by around 2019, blue-blocking glasses seemed to disappear. And during that brief moment in the sun, I never got a clear picture of if they actually work.
So, do they? To find out, I read all the papers.
Light and eyes
Before getting to the papers, please humor me while I give three excessively-detailed reminders about how light works. First, it comes in different wavelengths.
Color
Wavelength (nm)
violet
380–450
blue
450–485
cyan
485–500
green
500–565
yellow
565–590
orange
590–625
red
625–750
Outside the visible spectrum, infrared light and microwaves and radio waves have even longer wavelengths, while ultraviolet light and x-rays and gamma rays have even shorter wavelengths. Shorter wavelengths have more energy. Do not play around with gamma rays.
Other colors are hallucinations made up by your brain. When you get a mixture of all wavelengths, you see “white”. When you get a lot of yellow-red wavelengths, some green, and a little violet-blue, you see “brown”. Similar things are true for pink/purple/beige/olive/etc. (Technically, the original spectral colors and everything else you experience are also hallucinations made up by your brain, but never mind.)
Second, the ruleset of our universe says that all matter gives off light, with a mixture of wavelengths that depends on the temperature. Hotter stuff has atoms that are jostling around faster, so it gives off more total light, and shifts towards shorter (higher-energy) wavelengths. Colder stuff gives off less total light and shifts towards longer wavelengths. The “color temperature” of a lightbulb is the temperature some chunk of rock would have to be to produce the same visible spectrum. Here’s a figure, with the x-axis in kelvins.
The sun is around 5800 K. That’s both the physical temperature on the surface and the color temperature of its light. Annoyingly, the orange light that comes from cooler matter is often called “warm”, while the blueish light that comes from hotter matter is called “cool”. Don’t blame me.
You can’t sense most of those differences because you only have three types of cone cells. Rated color temperatures just reflect how much those cells are stimulated.
Your eyes probably see the frequencies they do because that’s where the sun’s spectrum is concentrated. In dim light, cones are inactive, so you rely on rod cells instead. You’ve only got one kind of rod, which is why you can’t see color in dim light. (Though you might not have noticed.)
In summary, you get widely varying amounts of different wavelengths of light in different situations, and the sun is very powerful. It’s reasonable to imagine your body might regulate its sleep schedule based that input.
Experiments
OK, but do blue-blocking glasses actually work? Let’s read some papers.
Kayumov
Kayumov et al. (2005) had 19 young healthy adults stay awake overnight for three nights, first with dim light (<5 lux) and then with bright light (800 lux), both with and without blue-blocking goggles. They measured melatonin in saliva each hour.
The goggles seemed to help a lot. With bright light, subjects only had around 25% as much melatonin as with dim light. Blue-blocking goggles restored that to around 85%.
I rate this as good evidence for a strong increase in melatonin. Sometimes good science is pretty simple.
Burkhart
Burkhart and Phelps (2009) first had 20 adults rate their sleep quality at home for a week as a baseline. Then, they were randomly given either blue-blocking glasses or yellow-tinted “placebo” glasses and told to wear them for 3 hours before sleep for two weeks.
Oddly, the group with blue-blocking glasses had much lower sleep quality during the baseline week, but this improved a lot over time.
I rate this as decent evidence for a strong improvement in sleep quality. I’d also like to thank the authors for writing this paper in something resembling normal human English.
Van der Lely
Van der Lely et al. (2014) had 13 teenage boys wear either blue-blocking glasses or clear glasses from 6pm to bedtime for one week, followed by the other glasses for a second week. Then they went to a lab, spent 2 hours in dim light, 30 minutes in darkness, and then 3 hours in front of an LED computer, all while wearing the glasses from the second week. Then they were asked to sleep, and their sleep quality was measured in various ways.
The boys had more melatonin and reported feeling sleepier with the blue-blocking glasses.
However, the sleep quality measurements show no real effect. They are all pretty close in the two groups, sometimes slightly better with the blue-blocking glasses and sometimes slightly worse.
I rate this as decent evidence for a moderate increase in melatonin, and weak evidence for near-zero effect on sleep quality.
Gabel
Gabel et al. (2017) took 38 adults and first put them through 40 hours of sleep deprivation under white light, then allowed them to sleep for 8 hours. Then they were subjected to 40 more hours of sleep deprivation under either white light (250 lux at 2800K), blue light (250 lux at 9000K), or very dim light (8 lux, color temperature unknown).
Their results are weird. In younger people, dim light led to more melatonin that white light, which led to more melatonin that blue light. That carried over to a tiny difference in sleepiness. But in older people, both those effects disappeared, and blue light even seemed to cause more sleepiness than white light. The cortisol and wrist activity measurements basically make no sense at all.
I rate this as decent evidence for a moderate effect on melatonin, and very weak evidence for a near-zero effect on sleep quality. (I think its decent evidence for a near-zero effect on sleepiness, but they didn’t actually measure sleep quality.)
Esaki
Esaki et al. (2017) gathered 20 depressed patients with insomnia. They first recorded their sleep quality for a week as a baseline, then were given either blue-blocking glasses or placebo glasses and told to wear them for another week starting at 8pm.
The changes in the blue-blocking group were a bit better for some measures, but a bit worse for others. Nothing was close to significant. Apparently 40% of patients complained that the glasses were painful, so I wonder if they all wore them as instructed.
I rate this was weak evidence for near-zero effect on sleep quality.
Shechter
Shechter et al. (2018) gave 14 adults with insomnia either blue-blocking or clear glasses and had them wear them for 2 hours before bedtime for one week. Then they waited four weeks and had them wear the other glasses for a second week. They measured sleep quality through diaries and wrist monitors.
The blue-blocking glasses seemed to help with everything. People fell asleep 5 to 12 minutes faster, and slept 30 to 50 minutes longer, depending on how you measure. (SOL is sleep onset latency, TST is total sleep time).
I rate this as good evidence for a strong improvement in sleep quality.
Knufinke
Knufinke et al. (2019) had 15 young adult athletes either wear blue-blocking glasses or transparent glasses for four nights.
The blue-blocking group did a little better on most measures (longer sleep time, higher sleep quality) but nothing was statistically significant.
I rate this as weak evidence for a small improvement in sleep quality.
Janků
Janků et al. (2019) took 30 patients with insomnia and had them all go to therapy. They randomly gave them either blue-blocking glasses or placebo glasses and asked the patients to wear them for 90 minutes before bed.
The results are pretty tangled. According to sleep diaries, total sleep time went up by 37 minutes in the blue-blocking group, but slightly decreased in the placebo group. The wrist monitors show total sleep time decreasing in both groups, but it did decrease less with the blue-blocking glasses. There’s no obvious improvement in sleep onset latency or the various questionnaires they used to measure insomnia.
I rate this as weak evidence for a moderate improvement in sleep quality.
Esaki (again)
Esaki et al. (2020) followed up on their 2017 experiment from above. This time, they gathered 43 depressed patients with insomnia. Again, they first recorded their sleep quality for a week as a baseline, then were given either blue-blocking glasses or placebo glasses and told to wear them for another week starting at 8pm.
The results were that subjective sleep quality seemed to improve more in the blue-blocking group. Total sleep time went down by 12.6 minutes in the placebo group, but increased by 1.1 minutes in the blue-blocking group. None of this was statistically significant, and all the other measurements are confusing. Here are the main results. I’ve added little arrows to show the “good” direction, if there is one.
These confidence intervals don’t make any sense to me. Are they blue-blocking minus placebo or the reverse? When the blue-blocking number is higher than placebo, sometimes the confidence interval is centered above zero (VAS), and sometimes it’s centered below zero (TST). What the hell?
Anyway, they also had a doctor estimate the clinical global impression for each patient, and this looked a bit better for the blue-blocking group. The doctor seemingly was blinded to the type of glasses the patients were wearing.
This is a tough one to rate. I guess I’ll call it weak evidence for a small improvement in sleep quality.
Guarana
Guarana et al. (2020) sent either blue-blocking glasses or sham glasses to 240 people, and asked them to wear them for at least two hours before bed. They then had them fill out some surveys about how much and how well they slept.
Wearing the blue-blocking glasses was positively correlated with both sleep quality and quantity with a correlation coefficient of around 0.20.
This paper makes me nervous. They never show the raw data, there seem to be huge dropout rates, and lots of details are murky. I can’t tell if the correlations they talk about weight all people equally, all surveys equally, or something else. That would make a huge difference if people dropped out more when they weren’t seeing improvements.
I rate this as weak evidence for a moderate effect on sleep. There’s a large sample, but I discount the results because of the above issues and/or my general paranoid nature.
Domagalik
Domagalik et al. (2020) had 48 young people wear either blue-blocking contact lenses or regular contact lenses for 4 weeks. They found no effect on sleepiness.
They also found that blocking blue light was harmful for attention and working memory.
I rate this as very weak evidence for near-zero effect on sleep. The experiment seems well-done, but it’s testing the effects of blocking blue light all the time, not just at night. Given the effects on attention and working memory, don’t do that.
Bigalke
Bigalke et al. (2021) had 20 healthy adults wear either blue-blocking glasses or clear glasses for a week from 6pm until bedtime, then switch to the other glasses for a second week. They measured sleep quality both through diaries (“Subjective”) and wrist monitors (“Objective”).
The differences were all small and basically don’t make any sense.
I rate this weak evidence for near-zero effect on sleep quality. Also, see how in the bottom pair of bar-charts, the y-axis on the left goes from 0 to 5, while on the right it goes from 30 to 50? Don’t do that, either.
See also
I also found a couple papers that are related, but don’t directly test what we’re interested in:
Appleman et al. (2013) either exposed people to different amounts of blue light at different times of day. Their results suggest that early-morning exposure to blue light might shift your circadian rhythm earlier.
Sasseville et al. (2015) had people stay awake from 11pm to 4am on two consecutive nights, while either wearing blue-blocking glasses or not. With the blue-blocking glasses there was more overall light to equalizing the total incoming energy. I can’t access this paper, but apparently they found no difference.
Drumroll
For a synthesis, I scored each of the measured effects according to this rubric:
Rating
Meaning
↑↑↑
large increase
↑↑
moderate increase
↑
small increase
↔
no effect
↓
small decrease
↓↓
moderate decrease
↓↓↓
large decrease
And I scored the quality of evidence according to this one:
Rating
Meaning
★☆☆☆☆
very weak evidence
★★☆☆☆
weak evidence
★★★☆☆
decent evidence
★★★★☆
good evidence
★★★★★
great evidence
Here are the results for the three papers that measured melatonin:
Study
Effect on melatonin
Quality of evidence
Kayumov
↑↑↑
★★★★☆
Van der Lely
↑↑
★★★☆☆
Gabel
↑
★★★☆☆
And here are the results for the papers that measured sleep quality:
Study
Effect on sleep
Quality of evidence
Burkhart
↑↑↑
★★★☆☆
Van der Lely
↔
★★☆☆☆
Gabel
↔
★☆☆☆☆
Esaki
↔
★★☆☆☆
Shechter
↑↑↑
★★★☆☆
Knufinke
↑
★★☆☆☆
Janků
↑↑
★★☆☆☆
Esaki (again)
↑
★★☆☆☆
Guarana
↑↑
★★☆☆☆
Domagalik
↔
★☆☆☆☆
Bigalke
↔
★★☆☆☆
We should adjust all that a bit because of publication bias and so on. But still, here are my final conclusions after staring at those tables:
There is good evidence that blue-blocking glasses cause a moderate increase in melatonin. It could be large, or it could be small, but I’d say there’s an ~85% chance it’s not zero.
There is decent evidence that blue-blocking glasses cause a small improvement in sleep quality. This could be moderate (or even large) or it could be zero. It might be inconsistent and hard to measure. But I’d say there’s an ~75% chance there is some positive effect.
I’ll be honest—I’m surprised.
So….
If those effects are real, do they warrant wearing stupid-looking glasses at night for the rest of your life? I guess that’s personal.
But surely the sane thing is not to block blue light with headgear, but to not create blue light in the first place. You can tell your glowing rectangles to block blue light at night, but lights are harder. Modern LED lightbulbs typically range in color temperature from 2700K for “warm” lighting to 5000 K for “daylight” bulbs. Judging from this animation that should reduce blue frequencies to around 1/3 as much.
Old-school incandescent bulbs are 2400 K. But to really kill blue, you probably want 2000K or even less. There are obscure LED bulbs out there as low as 1800K. They look extremely orange, but candles are apparently 1850K, so probably you’d get used to it?
So what do we do then? Get two sets of lamps with different bulbs? Get fancy bulbs that change color temperature automatically? Whatever it is, I don’t feel very optimistic that we’re going to see a lot of RCTs where researchers have subjects install an entire new lighting setup in their homes.
AI 2027 forecasts that AGI could plausibly arrive as early as 2027. I recently spent some time looking at both the timelines forecast and some critiques [1, 2, 3].
Initially, I was interested in technical issues. What’s the best super-exponential curve? How much probability should it have? But I found myself drawn to a more basic question. Namely, how much value is the math really contributing?
This provides an excuse for a general rant. Say you want to forecast something. It could be when your hair will go gray or if Taiwan will be self-governing in 2050. Whatever. Here’s one way to do it:
Think hard.
Make up some numbers.
Don’t laugh—that’s the classic method. Alternatively, you could use math:
Think hard.
Make up a formal model / math / simulation.
Make up some numbers.
Plug those numbers into the formal model.
People are often skeptical of intuition-based forecasts because, “Those are just some numbers you made up.” Math-based forecasts are hard to argue with. But that’s not because they lack made-up numbers. It’s because the meaning of those numbers is mediated by a bunch of math.
So which is better, intuition or math? In what situations?
Here, I’ll look at that question and how it applies to AI 2027. Then I’ll build a new AI forecast using my personal favorite method of “plot the data and scribble a bunch of curves on top of it”. Then I’ll show you a little tool to make your own artisanal scribble-based AI forecast.
Two kinds of forecasts
To get a sense of the big picture, let’s look at two different forecasting problems.
First, here’s a forecast (based on the IPCC 2023 report) for Earth’s temperature. There are two curves, corresponding to different assumptions about future greenhouse gas emissions.
Those curves look unassuming. But there are a lot of moving parts behind them. These kinds of forecasts model atmospheric pressure, humidity, clouds, sea currents, sea surface temperature, soil moisture, vegetation, snow and ice cover, surface albedo, population growth, economic growth, energy, and land use. They also model the interactions between all those things.
That’s hard. But we basically understand how all of it works, and we’ve spent a ludicrous amount of effort carefully building the models. If you want to forecast global surface temperature change, this is how I’d suggest you do it. Your brain can’t compete, because it can’t grind through all those interactions like a computer can.
OK, but here’s something else I’d really like to forecast: Where is this blue line going to go?
You could forecast this using a “mechanistic model” like with climate above. To do that, you’d want to model the probability Iran develops a nuclear weapon and what Saudi Arabia / Turkey / Egypt might do in response. And you’d want to do the same thing for Poland / South Korea / Japan and their neighbors. You’d also want to model future changes in demographics, technology, politics, technology, economics, military conflicts, etc.
In principle, that would be the best method. As with climate, there are too many plausible futures for your tiny brain to work through. But building that model would be very hard, because it basically requires you to model the whole world. And if there’s an error anywhere, it could have serious consequences.
In practice, I’d put more trust in intuition. A talented human (or AI?) forecaster would probably take an outside view like, “Over the last 80 years, the number of countries has gone up by 9, so in 2105, it might be around 18.” Then, they’d consider adjusting for things like, “Will other countries learn from the example of North Korea?” or “Will chemical enrichment methods become practical?”
Intuition can’t churn through possible futures the way a simulation can. But if you don’t have a reliable simulator, maybe that’s OK.
Broadly speaking, math/simulation-based forecasts shine when the phenomena you’re interested in has two properties.
It evolves according to some well-understood rule-set.
The behavior of the ruleset is relatively complex.
The first is important because if you don’t have a good model for the ruleset (or at least your uncertainty about the ruleset), how will you build a reliable simulator? The second is important because if the behavior is simple, why do you even need a simulator?
The ideal thing to forecast with math is something like Conway’s game of life. Simple known rules, huge emergent complexity. The worst thing to forecast with math is something like the probability that Jesus Christ returns next year. You could make up some math for that, but what would be the point?
The AI 2027 forecast
This post is (ostensibly) about AI 2027. So how does their forecast work? They actually have several forecasts, but here I’ll focus on the Time horizon extension model.
That forecast builds on a recent METR report. They took a set of AIs released over the past 6 years, and had them attempt a set of tasks of varying difficulty. They had humans perform those same tasks. Each AI was rated according to the human task length that it could successfully finish 50% of the time.
The AI 2027 team figured that if an AI could successfully complete long-enough tasks of this type, then the AI would be capable of itself carrying AI research, and AGI would not be far away. Quantitatively, they suggest that the necessary task length is probably somewhere between 1 month and 10 years. They also suggest you’d need a success rate of 80% (rather than 50% in the above figure).
So, very roughly speaking, the forecast is based on predicting how long it will take these dots to get up to one of the horizontal lines:
(It's a bit more complicated than that, but that's the core idea.)
Technical notes:
The AI 2027 team raises the success rate to 80%, rather than 50% in the original figure from the METR report. That’s why the dots in the above figure are a bit lower.
I made the above graph using the data that titotal extracted from the AI 2027 figures.
The AI 2027 forecast creates a distribution over the threshold that needs to be reached rather than considering fixed thresholds.
The AI 2027 forecast also adds an adjustment based on the theory that companies have internal models that are better than they release externally. They also add another adjustment on the theory that public-facing models are using limited compute to save money. In effect, these add a bit of vertical lift to all the points.
I think this framing is great. Instead of an abstract discussion about the arrival of AGI, suddenly we’re talking about how quickly a particular set of real measurements will increase. You can argue if “80% success at a 1-year task horizon” really means AGI is imminent. But that’s kind of the point—no matter what you think about broader issues, surely we’d all like to know how fast those dots are going to go up.
So how fast will they go up? You could imagine building a mechanistic model or simulation. To do that, you’d probably want to model things like:
How quickly is the data + compute + money being put into AI increasing?
How quickly is compute getting cheaper?
How quickly is algorithmic progress happening?
How does data + compute + algorithmic progress translate into improvements on the METR metrics?
How long will those trends hold? How do all those things interact with each other? How do they interact with AI progress itself.
In principle, that makes a lot of sense. Some people predict a future where compute keeps getting cheaper pretty slowly and we run out of data and new algorithmic ideas and loss functions stop translating to real-world performance and investment drops off and everything slows down. Other people predict a future where GPUs accelerate and we keep finding better algorithms and AI grows the economy so quickly that AI investment increases forever and we spiral into a singularity. In between those extremes are many other scenarios. A formal model could churn through all of them much better than a human brain.
But the AI 2027 forecast is not like that. It doesn’t have separate variables for compute / money / algorithmic progress. It (basically) just models the best METR score per year.
That’s not bad, exactly. But I must admit that I don’t quite see the point of a formal mathematical model in this case. It’s (basically) just forecasting how quickly a single variable goes up on a graph. The model doesn’t reflect any firm knowledge about subtle behavior other than that the curve will probably go up.
In a way, I think this makes the AI 2027 forecast seem weaker than it actually is. Math is hard. There are lots of technicalities to argue with. But their broader point doesn’t need math. Say you accept their premise that 80% success on tasks that take humans 1 year means that AGI is imminent. Then you should believe AGI is around the corner unless those dots slow down. An argument that their math is flawed doesn’t imply that the dots are going to stop going up.
Scribble-based forecasting
So, what’s going to happen with those dots? The ultimate outside view is probably to not think at all and just draw a straight line. When I do that, I get something like this:
I guess that’s not terrible. But personally, I feel like it’s plausible that the recent acceleration continues. I also think it’s plausible that in a couple of years we stop spending ever-larger sums on training AI models and things slow down. And for a forecast, I want probabilities.
So I took the above dots and I scribbled 50 different curves on top, corresponding to what I felt were 50 plausible futures:
Then I treated those lines as a probability distribution over possible futures. For each of three task-horizon thresholds, I calculated what percentage of the lines had reached them in a given year.
Here’s a summary as a table:
Threshold
10th Percentile
50th Percentile
90th Percentile
% Reached by 2050
1 month
2028.7
2032.3
2039.3
94%
1 year
2029.5
2034.8
2041.4
88%
10 year
2029.2
2037.7
2045.0
54%
My scribbles may or may not be good. But I think the exercise of drawing the scribbles is great, because it forces you to be completely explicit, and your assumptions are completely legible.
I recommend it. In fact, I recommend it so strongly that I’ve created a little tool that you can use to do your own scribbling. It will automatically generate a plot and table like you see above. You can import or export your scribbles in CSV format. (Mine are here if you want to use them as a starting point.)
Here’s a video demo:
While scribbling, you may reflect on the fact that the tool you’re using is 100% AI-generated.
I haven’t followed AI safety too closely. I tell myself that’s because tons of smart people are working on it and I wouldn’t move the needle. But I sometimes wonder, is that logic really unrelated to the fact that every time I hear about a new AI breakthrough, my chest tightens with a strange sense of dread?
AI is one of the most important things happening in the world, and possibly the most important. If I’m hunkering in a bunker years from now listening to hypersonic kill-bots laser-cutting through the wall, will I really think, boy am I glad I stuck to my comparative advantage?
So I thought I’d take a look.
I stress that I am not an expert. But I thought I’d take some notes as I try to understand all this. Ostensibly, that’s because my outsider status frees me from the curse of knowledge and might be helpful for other outsiders. But mostly, I like writing blog posts.
So let’s start at the beginning. AI safety is the long-term problem of making AI be nice to us. The obvious first question is, what’s the hard part? Do we know? Can we say anything?
To my surprise, I think we can: The hard part is making AI want to be nice to us. You can’t solve the problem without doing that. But if you can do that, then the rest is easier.
This is not a new idea. Among experts, I think it’s somewhere between “the majority view” and “near-consensus”. But I haven’t found many explicit arguments or debates, meaning I’m not 100% sure why people believe it, or if it’s even correct. But instead of cursing the darkness, I thought I’d construct a legible argument. This may or may not reflect what other people think. But what is a blog, if not an exploit on Cunningham’s Law?
My argument, at a high level
Here’s my argument that the hard part of AI safety is making AI want to do what we want:
To make an AI be nice to you, you can either impose restrictions, so the AI is unable to do bad things, or you can align the AI, so it doesn’t choose to do bad things.
Restrictions will never work.
You can break down alignment into making the AI know what we want, making it want to do what we want, and making it succeed at what it tries to do.
Making an AI want to do what we want seems hard. But you can’t skip it, because then AI would have no reason to be nice.
Human values are a mess of heuristics, but a capable AI won’t have much trouble understanding them.
True, a super-intelligent AI would likely face weird “out of distribution” situations, where it’s hard to be confident it would correctly predict our values or the effects of its actions.
But that’s OK. If an AI wants to do what we want, it will try to draw a conservative boundary around its actions and never do anything outside the boundary.
Drawing that boundary is not that hard.
Thus, if an AI system wants to do what we want, the rest of alignment is not that hard.
Thus, making AI systems want to do what we want is necessary and sufficient-ish for AI safety.
I am not confident in this argument. I give it a ~35% chance of being correct, with step 8 the most likely failure point. And I’d give another ~25% chance that my argument is wrong but the final conclusion is right.
(Y’all agree that a low-confidence prediction for a surprising conclusion still contains lots of information, right? If we learned there was a 10% chance Earth would be swallowed by an alien squid tomorrow, that would be important, etc.? OK, sorry.)
My argument, in more detail
I’ll go quickly through the parts that seem less controversial.
1. There are two conceivable paths to AI safety.
Roughly speaking, to make AI safe you could either impose restrictions on AI so it’s not able to do bad things, or align AI so it doesn’t choose to do bad things. You can think of these as not giving AI access to nuclear weapons (restrictions) or making the AI choose not to launch nuclear weapons (alignment).
2. Restrictions will never work.
I advise against giving AI access to nuclear weapons. Still, if an AI is vastly smarter than us and wants to hurt us, we have to assume it will be able to jailbreak any restrictions we place on it. Given any way to interact with the world, it will eventually find some way to bootstrap towards larger and larger amounts of power. Restrictions are hopeless. So that leaves alignment.
3. You can break down alignment into three parts.
Here’s a simple-minded decomposition:
The Knowing problem: Making AI know what we want.
The Wanting problem: Making AI want to do what we want.
The Success problem: Making AI succeed at what it tries to do.
I sometimes wonder if that’s a useful decomposition. But let’s go with it.
4. Wanting is necessary.
The Wanting problem seems hard, but there’s no way around it. Say an AI knows what we want and succeeds at everything it tries to do, but doesn’t care about what we want. Then, obviously, it has no reason to be nice. So we can’t skip Wanting.
Also, notice that even if you solve the Knowing and Success problems really well, that doesn’t seem to make the Wanting problem any easier. (See also: Orthogonality)
5. Human values are a shallow mess.
My take on human values is that they’re a big ball of heuristics. When we say that some action is right (wrong) that sort of means that genetic and/or cultural evolution thinks that the reproductive fitness of our genes and/or cultural memes is advanced by rewarding (punishing) that behavior.
Of course, evolution is far from perfect. Clearly our values aren’t remotely close to reproductively optimal right now, what with fertility rates crashing around the world. But still, values are the result of evolution trying to maximize reproductive fitness.
Why do we get confused by trolley problems and population ethics? I think because… our values are a messy ball of heuristics. We never faced evolutionary pressure to resolve trolley problems, so we never really formed coherent moral intuitions about them.
So while our values have lots of quirks and puzzles, I don’t think there’s anything deep at the center of them, anything that would make learning them harder than learning to solve Math Olympiad problems or translating text between any pair of human languages. Current AI already seems to understand our values fairly well.
Arguably, it would be hard to prevent AI from understanding human values. If you train an AI to do any sufficiently difficult task, it needs a good world model. That’s why “predicting the next token” is so powerful—to do it well, you have to model the world. Human values are an important and not that complex part of that world.
6. Distribution shift may make it harder for AI to Know or Succeed.
The idea of “distribution shift” is that after super-intelligent AI arrives, the world may change quite a lot. Even if we train AI to be nice to us now, in that new world it will face novel situations where we haven’t provided any training data.
This could conceivably create problems both for AI knowing what we want, or for AI succeeding at what it tries to do.
For example, maybe we teach an AI that it’s bad to kill people using lasers, and that it’s bad to kill people using viruses, and that it’s bad to kill people using radiation. But we forget to teach it that it’s bad to write culture-shifting novels that inspire people to live their best lives but also gradually increase political polarization and lead after a few decades to civilizational collapse and human extinction. So the AI intentionally writes that book and causes human extinction because it thinks that’s what we want, oops.
Alternatively, maybe a super-powerful AI knows that we don’t like dying and it wants to help us not die, so it creates a retrovirus that spreads across the globe and inserts a new anti-cancer gene in our DNA. But it didn’t notice that this gene also makes us blind and deaf, and we all starve and die. In this case, the AI accidentally does something terrible, because it has so much power that it can’t correctly predict all the effects of its actions.
7. But all AI needs to do is draw a conservative boundary.
What are your values? Personally, very high on my list would be:
If an AI is considering doing anything and it’s not very sure that it aligns with human values, then it should not do it without checking very carefully with lots of humans and getting informed consent from world governments. Never ever do anything like that.
And also:
AIs should never release retroviruses without being very sure it’s safe and checking very carefully with lots of humans and getting informed consent from world governments. Never ever, thanks.
That is, AI safety doesn’t require AIs to figure out how to generalize human values to all weird and crazy situations. And it doesn’t need to correctly predict the effects of all possible weird and crazy actions. All that’s required is that AIs can recognize that something is weird/crazy and then be conservative.
Clearly, just detecting that something is weird/crazy is easier than making correct predictions in all possible weird/crazy situations. But how much easier?
8. Drawing that boundary isn’t that hard.
(I think this is the weakest part of this argument. But here goes.)
Would I trust an AI to correctly decide if human flourishing is more compatible with a universe where up quarks make up 3.1% of mass-energy and down quarks 1.9% versus one where up quarks make up 3.2% and down quarks 1.8%? Probably not. But I wouldn’t trust any particular human to decide that either. What I would trust a human to do is say, “Uhhh?” And I think we can also trust AI to know that’s what a human would say.
Arguably, “human values” are a thing that only exist for some limited range of situations. As you get further from our evolutionary environment, our values sort of stop being meaningful. Do we prefer an Earth with 100 billion moderately happy people, or one with 30 billion very happy people? I think the correct answer is, “No”.
When we have coherent answers, AI will know what they are. And otherwise, it will know that we don’t have coherent answers. So perhaps this is a better picture:
And this seems… fine? AI doesn’t need to Solve Ethics, it just needs to understand the limited range of human values, such as they are.
That argument (if correct) resolves the issue of distribution shift for values. But we still need to think about how distribution shift might make it harder for AI to succeed at what it tries to do.
If AI attains godlike power, maybe it will be able to change planetary orbits or remake our cellular machinery. With this gigantic action space, it’s plausible that there would be many actions with bad but hard-to-predict effects. Even if AI only chooses actions that are 99.999% safe, if it makes 100 such actions per day, calamity is inevitable.
Sure, but surely we want AI to take false discovery rates (“calamitous discovery rates”?) into account. It should choose a set of actions such that, taken together, they are 99.999% safe.
Something that might work in our favor here is that verification is usually much easier than generation. Perhaps we could ask the AI to create a “proof” that all proposed actions are safe and run that proof by a panel of skeptical “red-team” AIs. If any of them find anything confusing at all, reject.
I find the idea that “drawing a safe boundary is not that hard” fairly convincing for human values, but not only semi-convincing for predicting the effects of actions. So I’d like to see more debate on this point. (Did I mention that this is the weakest part of my argument?)
9. Thus, if an AI system wants to do what we want, the rest of alignment is not that hard.
It AI truly wants to do what we want, then the only thing it really needs to know about our values is “be conservative”. This makes the Knowing and Success problems much easier. Instead of needing to know how good all possible situations are for humans, it just needs to notice that it’s confused. Instead of needing to succeed at everything it tries, it just needs to notice that it’s unsure.
10. Thus, making AI systems want to do what we want is necessary and sufficient for AI safety.
Since restrictions won’t work, you need to do alignment. Wanting is hard, but if you can solve Wanting, then you only need to solve easier version of Knowing and Success. So Wanting is the hard part.
Consistency with other views
Again, I think the idea that “wanting is the hard part” is the majority view. Paul Christiano, for example, proposes to call an AI “intent aligned” if it is trying to do what some operator wants it to do and states:
[The broader alignment problem] includes many subproblems that I think will involve totally different techniques than [intent alignment] (and which I personally expect to be less important over the long term).
Rather, my main concern is that AGIs will understand what we want, but just not care, because the motivations they acquired during training weren’t those we intended them to have.
Many people have also told me this is the view of MIRI, the most famous AI-safety organization. As far as I can see, this is compatible with the MIRI worldview. But I don’t feel comfortable stating as a fact that MIRI agrees, because I’ve never seen any explicit endorsement, and I don’t fully understand how it fits together with other MIRI concepts like corrigibility or coherent extrapolated volition.
Counterarguments
Why might this argument be wrong?
Maybe restrictions would work.
(I don’t think so, but it’s good to be comprehensive.)
Maybe Wanting is easy for some reason
Wanting seems hard, to me. And most experts seem to agree. But who knows, maybe it’s easy.
Here’s one esoteric possibility. Above, I’ve implicitly assumed that an AI could in principle want anything. But it’s conceivable that only certain kinds of wants are stable. That might make Wanting harder or even quasi-impossible. But it could also conceivably make it easy. Maybe once you cross some threshold of intelligence, you become one with the universal mind and start treating all other beings as a part of yourself? I wouldn’t bet on it.
Maybe drawing the boundary is hard
A crucial part of my argument is the idea that it would be easy for AI to draw a conservative boundary when trying to predict human values or effects of actions. I find that reasonably convincing for values, but less so for actions. It’s certainly easier than correctly generalizing to all situations, but it might still be very hard.
It’s also conceivable that AI creates such a large action space that even if humans were allowed to make every single decision, we would destroy ourselves. For example, there could be an undiscovered law of physics that says that if you build a skyscraper taller than 900m, suddenly a black hole forms. But physics provides no “hints”. The only way to discover that is to build the skyscraper and create the black hole.
More plausibly, maybe we do in fact live in a vulnerable world, where it’s possible to create a planet-destroying weapon with stuff you can buy at the hardware store for $500, we just haven’t noticed yet. If some such horrible fact is lurking out there, AI might find it much sooner than we would.
Maybe these are the wrong abstractions
Finally, maybe the whole idea of an AI “wanting” things is bad. It seems like a useful abstraction when we think about people. But if you try to reduce the human concept of “wanting” to neuroscience, it’s extremely difficult. If an AI is a bunch of electrons/bits/numbers/arrays flying around, is it obvious that the same concept will emerge?
Who is “we”?
I’ve been sloppy in this post in talking about AIs respecting “our” values or “human values”. That’s probably not going to happen. Absent some enormous cultural development, AIs will be trained to advance the interests of particular human organizations. So even if AI alignment is solved, it seems likely that different groups of humans will seek to create AIs that help them, even at some expense to other groups.
That’s not technically a flaw in the argument, since it just means Wanting is even harder. But it could be a serious problem, because…
Arms races might destroy conservatism
Suppose you live in Country A. Say you’ve successfully created a super-intelligent AI that’s very conservative and nice. But people in Country B don’t like you, so they create their own super-intelligent AI and ask it to hack into your critical systems, e.g. to disable your weapons or to prevent you from making an even-more-powerful AI.
What happens now? Well, their AI is too smart to be stopped by the humans in Country A. So your only defense will be to ask your own AI to defend against the hacks. But then, Country B will probably notice that if they give their AI more leeway, it’s better at hacking. This forces you to give your AI more leeway so it can defend you. The equilibrium might be that both AIs are told that, actually, they don’t need to be very conservative at all.
Things I read
Finally, here’s some stuff I found useful, from people who may or may not agree with the above argument:
A couple of months ago (April 2025), a group of prominent folks released AI 2027, a project that predicted that AGI could plausibly be reached in 2027 and have important consequences. This included a set of forecasts and a story for how things might play out. This got a lot of attention. Some was positive, some was negative, but it was almost all very high level.
More recently (June 2025) titotal released a detailed critique, suggesting various flaws in the modeling methodology.
I don’t have much to say about AI 2027 or the critique on a technical level. It would take me at least a couple of weeks to produce an opinion worth caring about, and I haven’t spent the time. But I would like to comment on the discourse. (Because “What we need is more commentary on the discourse”, said no one.)
Very roughly speaking, here’s what I remember: First, AI 2027 came out. Everyone cheers. “Yay! Amazing!” Then the critique came out. Everyone boos. “Terrible! AI 2027 is not serious! This is why we need peer-review!”
This makes me feel simultaneously optimistic and depressed.
Should AI 2027 have been peer-reviewed? Well, let me tell you a common story:
Someone decides to write a paper.
In the hope of getting it accepted to a journal, they write it in arcane academic language, fawningly cite unrelated papers from everyone who could conceivably be a reviewer, and make every possible effort to hide all flaws.
This takes 10× longer than it should, results in a paper that’s very boring and dense, and makes all limitations illegible.
They submit it to a journal.
After a long time, some unpaid and distracted peers give the paper a quick once-over and write down some thoughts.
There’s a cycle where the paper is revised to hopefully make those peers happy. Possibly the paper is terrible, the peers see that, and the paper is rejected. No problem! The authors resubmit it to a different journal.
Twelve years later, the paper is published. Oh happy day!
You decide to read the paper.
After fighting your way through the writing, you find something that seems fishy. But you’re not sure, because the paper doesn’t fully explain what they did.
The paper cites a bunch of other papers in a way that implies they might resolve your question. So you read those papers, too. It doesn’t help.
You look at the supplementary material. It consists of insanely pixelated graphics and tables with labels like Qetzl_xmpf12 that are never explained.
In desperation, you email the authors.
They never respond.
The end.
And remember, peer review is done by peers from the same community who think in similar ways. Different communities settle on somewhat random standards for what’s considered important or what’s considered an error. In much of the social sciences, for example, quick-and-dirty regressions with strongly implied causality are A+ supergood. Outsiders can complain, but they aren’t the ones doing the reviewing.
I wouldn’t say that peer review is worthless. It’s something! Still, call me cynical—you’re not wrong—but I think the number of mistakes in peer-reviewed papers is one to two orders of magnitude higher than generally understood.
Why are there so many mistakes to start with? Well I don’t know if you’ve heard, but humans are fallible creatures. When we build complex things, they tend to be flawed. They particularly tend to be flawed when—for example—people have strong incentives to produce a large volume of “surprising” results, and the process to find flaws isn’t very rigorous.
Aren’t authors motivated by Truth? Otherwise, why choose that life over making lots more money elsewhere? I personally think this is an important factor, and probably the main reason the current system works at all. But still, it’s amazing how indifferent many people are to whether their claims are actually correct. They’ve been in the game so long that all they remember is their h-index.
And what happens if someone spots an error after a paper is published? This happens all the time, but papers are almost never retracted. Nobody wants to make a big deal because, again, peers. Why make enemies? Even when publishing a contradictory result later, people tend to word their criticisms so gently and indirectly that they’re almost invisible.
As far as I can tell, the main way errors spread is: Gossip. This works sorta-OK-ish for academics, because they love gossip and will eagerly spread the flaws of famous papers. But it doesn’t happen for obscure papers, and it’s invisible to outsiders. And, of course, if seeing the flaws requires new ideas, it won’t happen at all.
If peer review is so imperfect, then here’s a little dream. Just imagine:
Alice develops some ideas and posts them online, quickly and with minimal gatekeeping.
Because Alice is a normal human person, there are some mistakes.
Bob sees it and thinks something is fishy.
Bob asks Alice some questions. Because Alice cares about being right, she’s happy to answer those questions.
Bob still thinks something is fishy, so he develops a critique and posts it online, quickly and with minimal gatekeeping.
Bob’s critique is friendly and focuses entirely on technical issues, with no implications of bad faith. But at the same time, he pulls no punches.
Because Bob is a normal human person, he makes some mistakes, too.
Alice accepts some parts of the critique. She rejects other parts and explains why.
Carol and Eve and Frank and Grace see all this and jump in with their own thoughts.
Slowly, the collective power of many human brains combine to produce better ideas than any single human could.
Wouldn’t that be amazing? And wouldn’t it be amazing if some community developed social norms that encouraged people to behave that way? Because as far as I can tell, that’s approximately what’s happening with AI 2027.
I guess there’s a tradeoff in how much you “punish” mistakes. Severe punishment makes people defensive and reduces open discussion. But if you’re too casual, then people might get sloppy.
My guess is that different situations call for different tradeoffs. Pure math, for example, might do well to set the “punishment slider” fairly high, since verifying proofs is easier than creating the proofs.
The best choice also depends on technology. If it’s 1925 and communication is bottlenecked by putting ink on paper, maybe you want to push most of the verification burden onto the original authors. But it’s not 1925 anymore, and surely it’s time to experiment with new models.
Tea is a little-known beverage, consumed for flavor or sometimes for conjectured effects as a stimulant. It’s made by submerging the leaves of C. Sinensis in hot water. But how hot should the water be?
To resolve this, I brewed the same tea at four different temperatures, brought them all to a uniform serving temperature, and then had four subjects rate them along four dimensions.
Subjects
Subject A is an experienced tea drinker, exclusively of black tea w/ lots of milk and sugar.
Subject B is also an experienced tea drinker, mostly of black tea w/ lots of milk and sugar. In recent years, Subject B has been pressured by Subject D to try other teas. Subject B likes fancy black tea and claims to like fancy oolong, but will not drink green tea.
Subject C is similar to Subject A.
Subject D likes all kinds of tea, derives a large fraction of their joy in life from tea, and is world’s preeminent existential angst + science blogger.
Tea and brewing
For a tea that was as “normal” as possible, I used pyramidal bags of PG Tips tea (Lipton Teas and Infusions, Trafford Park Rd., Trafford Park, Stretford, Manchester M17 1NH, UK).
I brewed it according to the instructions on the box, by submerging one bag in 250ml of water for 2.5 minutes. I did four brews with water at temperatures ranging from 79°C to 100°C (174.2°F to 212°F). To keep the temperature roughly constant while brewing, I did it in a Pyrex measuring cup (Corning Inc., 1 Riverfront Plaza, Corning, New York, 14831, USA) sitting in a pan of hot water on the stove.
After brewing, I poured the tea into four identical mugs with the brew temperature written on the bottom with a Sharpie Pro marker (Newell Brands, 5 Concourse Pkwy Atlanta, GA 30328, USA). Readers interested in replicating this experiment may note that those written temperatures still persist on the mugs today, three months later. The cups were dark red, making it impossible to see any difference in the teas.
After brewing, I put all the mugs in a pan of hot water until they converged to 80°C, so they were served at the same temperature.
Serving
I shuffled the mugs and placed them on a table in a random order. I then asked the subjects to taste from each mug and rate the teas for:
“Aroma”
“Flavor”
“Strength”
“Goodness”
Each rating was to be on a 1-5 scale, with 1=bad and 5=good.
Subjects A, B, and C had no knowledge of how the different teas were brewed. Subject D was aware, but was blinded as to which tea was in which mug.
During taste evaluation, Subjects A and C remorselessly pestered Subject D with questions about how a tea strength can be “good” or “bad”. Subject D rejected these questions on the grounds that “good” cannot be meaningfully reduced to other words and urged Subjects A and C to review Wittgenstein’s concept of meaning as use, etc. Subject B questioned the value of these discussions.
After ratings were complete, I poured tea out of all the cups until 100 ml remained in each, added around 1 gram (1/4 tsp) of sugar, and heated them back up to 80°C. I then re-shuffled the cups and presented them for a second round of ratings.
Results
For a single summary, I somewhat arbitrarily combined the four ratings into a “quality” score, defined as
Here is the data for Subject A, along with a linear fit for quality as a function of brewing temperature. Broadly speaking, A liked everything, but showed weak evidence of any trend.
And here is the same for Subject B, who apparently hated everything.
Here is the same for Subject C, who liked everything, but showed very weak evidence of any trend.
And here is the same for Subject D. This shows extremely strong evidence of a negative trend. But, again, while blinded to the order, this subject was aware of the brewing protocol.
Finally, here are the results combining data from all subjects. This shows a mild trend, driven mostly by Subject D.
Thoughts
This experiment provides very weak evidence that you might be brewing your tea too hot. Mostly, it just proves that Subject D thinks lower-middle tier black tea tastes better when brewed cooler. I already knew that.
There are a lot of other dimensions to explore, such as the type of tea, the brew time, the amount of tea, and the serving temperature. I think that ideally, I’d randomize all those dimensions, gather a large sample, and then fit some kind of regression.
Creating dozens of different brews and then serving them all blinded at different serving temperatures sounds like way too much work. Maybe there’s an easier way to go about this? Can someone build me a robot?
If you thirst to see Subject C’s raw aroma scores or whatever, you can download the data or click on one of the entries in this table:
This is an article that just appeared in Asimov Press, who kindly agreed that I could publish it here and also humored my deep emotional need to use words like “Sparklepuff”.
Do you like information theory? Do you like molecular biology? Do you like the idea of smashing them together and seeing what happens? If so, then here’s a question: How much information is in your DNA?
When I first looked into this question, I thought it was simple:
Human DNA has about 3.1 billion base pairs.
Each base pair can take one of four values (A, T, C, or G)
It takes 2 bits to encode one of four possible values (00, 01, 10, or 11)
Thus, human DNA contains 6.2 billion bits.
Easy, right? Sure, except:
You have two versions of each base pair, one from each of your parents. Should you count both?
All humans have almost identical DNA. Does that matter?
DNA can be compressed. Should you look at the compressed representation?
It’s not clear how much of our DNA actually does something useful. The insides of your cells are a convulsing pandemonium of interacting “hacks”, designed to keep working even as mutations constantly screw around with the DNA itself. Should we only count the “useful” parts?
Such questions quickly run into the limits of knowledge for both biology and computer science. To answer them, we need to figure out what exactly we mean by “information” and how that’s related to what’s happening inside cells. In attempting that, I will lead you through a frantic tour of information theory and molecular biology. We’ll meet some strange characters, including genomic compression algorithms based on deep learning, retrotransposons, and Kolmogorov complexity.
Ultimately, I’ll argue that the intuitive idea of information in a genome is best captured by a new definition of a “bit”—one that’s unknowable with our current level of scientific knowledge.
On counting
What is “information”? This isn’t just a pedantic question, as there are actually several different mathematical definitions of a “bit”. Often, the differences don’t matter, but for DNA, they turn out to matter a lot, so let’s start with the simplest.
In the storage space definition, a bit is a “slot” in which you can store one of two possible values. If some object can represent 2ⁿ possible patterns, then it contains n bits, regardless of which pattern actually happens to be stored.
So here’s a question we can answer precisely: How much information could your DNA store?
A few reminders: DNA is a polymer. It’s a long chain of chunks of ~40 atoms called “nucleotides”. There are four different chunks, commonly labeled A, T, C, and G. In humans, DNA comes in 23 pieces of different lengths, called “chromosomes”. Humans are “diploid”, meaning we have two versions of each chromosome. We get one from each of our parents, made by randomly weaving together sections from the two chromosomes they got from their parents.
At least, that's true for the first 22 chromosomes. For the last, females have two "X" chromosomes, while males have one "X" and one "Y" chromosome. There's no mixing between these, so men pass on one to their children pretty much unchanged.
Technically there’s also a tiny amount of DNA in the mitochondria. This is neat because you get it from your mother basically unchanged and so scientists can trace tiny mutations back to see how our great-great-…-great grandmothers were all related. If you go far enough back, our maternal lines all lead to a single woman, Mitochondrial Eve, who probably lived in East Africa 120,000 to 156,000 years ago. But mitochondrial DNA is tiny so I won’t mention it again.
Chromosomes 1-22 have a total of 2.875 billion nucleotides; the X chromosome has 156 million, and the Y chromosome has 62 million. From here, we can calculate the total storage space in your DNA. Remember, each nucleotide has 4 options, corresponding to 2 bits. So if you’re female, your total storage space is:
For comparison, a standard single-layer DVD can store 37.6 billion bits or 4.7 GB. The code for your body, magnificent as it is, takes up as much space as around 40 minutes of standard definition video.
So in principle, your DNA could represent around 212,000,000,000 different patterns. But hold on. Given human common ancestry, the chromosome pair you got from your mother is almost identical to the one you got from your father. And even ignoring that, there are long sequences of nucleotides that are repeated over and over in your DNA, enough to make up a significant fraction of the total. It seems weird to count all this repeated stuff. So perhaps we want a more nuanced definition of “information.”
On compression
A string of 12 billion zeros is much longer than this article. But most people would (I hope) agree that this article contains more information than a string of 12 billion zeros. Why?
One of the fundamental ideas from information theory is to define information in terms of compression. Roughly speaking, the “information” in some string is the length of the shortest possible compressed representation of that string.
So how much can you compress DNA? Answers to this question are all over the place. Some people claim it can be compressed by more than 99 percent, while others claim the state of the art is only around 25 percent. This discrepancy is explained by different definitions of “compression”, which turn out to correspond to different notions of “information”.
If you pick any two random people on Earth, almost all of their DNA will be exactly the same. It's often said that people are 99.9 percent genetically identical, but this is wrong—it only measures substitutions and neglects things like insertions, deletions, and transpositions. If you account for all these things, the best estimate is that we are ~99.6 percent identical.
Fun facts: Because of these deletions and insertions, different people have slightly different amounts of DNA. In fact, each of your chromosome pairs have DNA of slightly different lengths. When your body creates sperm/ova it uses a crazy machine to align the chromosomes in a sensible way so different sections can be woven together without creating nonsense. Also, those same measures of similarity would say that we’re around 96 percent identical with our closest living cousins, the bonobos and chimpanzees.
The fact that we share so much DNA is key to how some algorithms can compress DNA by more than 99 percent. They do this by first storing a reference genome, which includes all the DNA that’s shared by all people and perhaps the most common variants for regions of DNA where people differ. Then, for each individual person, these algorithms only store the differences from the reference genome. Because that reference only has to be stored once, it isn’t counted in the compressed representation.
That’s great if you want to cram as many of your friends’ genomes on a hard drive as possible. But it’s a strange definition to use if you want to measure the “information content of DNA”. It implies that any genomic content that doesn’t change between individuals isn’t important enough to count as “information”. However, we know from evolutionary biology that it’s often the most crucial DNA that changes the least precisely because it’s so important. Heritability tends to be lower for genes more closely related to reproduction.
The best compression without a reference seems to be around 25 percent. (I expect this number to rise a bit over time, as the newest methods use deep learning and research is ongoing.) That’s not a lot of compression. However, these algorithms are benchmarked in terms of how well they compress a genome that includes only one copy of each chromosome. Since your two chromosomes are almost identical (at least, ignoring the Y chromosome), I’d guess that you could represent the other half almost for free, meaning a compression rate of around 50 percent + ½ × 25 percent ≈ 62 percent.
On information
So if you compress DNA using an algorithm with a reference genome, it can be compressed by more than 99 percent, down to less than 120 million bits. But if you compress it without a reference genome, the best you can do is 62 percent, meaning 4.6 billion bits.
Which of these is right? The answer is that either could be right. There are two different definitions of a “bit” in information theory that correspond to different types of compression.
In the Kolmogorov complexity definition, named after the remarkable Soviet mathematician Andrey Kolmogorov, a bit is a property of a particular string of 1s and 0s. The number of bits of information in the string is the length of the shortest computer program that would output that string.
In the Shannon information definition, named after the also-remarkable American polymath Claude Shannon, a bit is again a property of a particular sequence of 1s and 0s, but it’s only defined relative to some large pool of possible sequences. In this definition, if a given sequence has a probability p of occurring, then it contains n bits for whatever value of n satisfies 2ⁿ=1/p. Or, equivalently, n=-log₂ p.
The Kolmogorov complexity definition is clearly related to compression. But what about Shannon’s?
Well, say you have three beloved pet rabbits, Fluffles, Marmalade, and Sparklepuff. And say you have one picture of each of them, each 1 MB large when compressed. To keep me updated on how you’re feeling, you like to send me these same pictures over and over again, with different pets for different moods. You send a picture of Fluffles ½ the time, Marmalade ¼ of the time, and Sparklepuff ¼ of the time. (You only communicate in rabbit pictures, never with text or images.)
But then you decide to take off in a spacecraft, and your data rates go way up. Continuing the flow of pictures is crucial, so what’s the cheapest way to do that? The best thing would be that we agree that if you send me a 0, I should pull up the picture of Fluffles, while if you send 10 I should pull up Marmalade, and if you send 11, I should pull up Sparklepuff. This is unambiguous: If you send 0011100, that means Fluffles, then Fluffles again, then Sparklepuff, then Marmalade, then Fluffles one more time.
It all works out. The “code length” for Fluffles is the number n so that 2ⁿ=1/p:
pet
probability p
code
code length n
2ⁿ
1/p
Fluffles
½
0
1
2
2
Marmelade
¼
10
2
4
4
Sparklepuff
¼
11
2
4
4
Intuitively, the idea is that if you want to send as few bits as possible over time, then you should give short codes to high-probability patterns and long codes to low-probability patterns. If you do this optimally (in the sense that you’ll send the fewest bits over time), it turns out that the best thing is to code a pattern with probability p with about n bits, where 2ⁿ=p. (In general, things don’t work out quite this nicely, but you get the idea.)
In the Fluffles scenario, the Kolmogorov complexity definition would say that each of the images contains 1 MB of information since that’s the smallest each image can be compressed. But under the Shannon information definition, the Fluffles image contains 1 bit of information, and the Marmalade and Sparklepuff images contain 2 bits. This is quite a difference!
Now, let’s return to DNA. There, the Kolmogorov complexity definition basically corresponds to the best possible compression algorithm without a reference. As we saw above, the best-known current algorithm can compress by 62 percent. So, under the Kolmogorov complexity definition, DNA contains at most 12 billion × (1-0.62) ≈ 4.6 billion bits of information.
Meanwhile, under the Shannon information definition, you can assume that the distribution of all human genomes is known. The information in your DNA only includes the bits needed to reconstruct your genome. That’s essentially the same as compressing with a reference. So, under the Shannon information definition, your DNA contains less than 12 billion × (1-0.01) ≈ 120 million bits of information.
While neither of these is “wrong” for DNA, I prefer the Kolmogorov complexity definition for its ability to best capture DNA that codes for features and functions shared by all humans. After all, if you’re trying to measure how much “information” our DNA carries from our evolutionary history, surely you want to include that which has been universally preserved.
On biology
At some point, your high-school biology teacher probably told you (or will tell you) this story about how life works:
First, your DNA gets transcribed into matching RNA.
Next, that RNA gets translated into protein.
Then the protein does Protein Stuff.
If things were that simple, we could easily calculate the information density of DNA just by looking at what fraction of your DNA ever becomes a protein (only around 1 percent). But it’s not that simple. The rest of your DNA does other important things, like regulating what proteins get made. Some of it seems to exist only for the purpose of copying itself. Some of it might do nothing, or it might do important things we don’t even know about yet.
So let me tell you that story again with slightly more detail:
In the beginning, your DNA is relaxing in the nucleus.
Some parts of your DNA, called promoters, are designed so that if certain proteins are nearby, they’ll stick to the DNA.
If that happens, then a hefty little enzyme called “RNA polymerase” will show up, crack open the two strands of DNA, and start transcribing the nucleotides on one side into “pre-messenger RNA” (pre-mRNA).
Eventually, for one of several reasons—none of which make any sense to me—the enzyme will decide it’s time to stop transcribing, and the pre-mRNA will detach and float off into the nucleus. At this point, it’s a few thousand or a few tens of thousands of nucleotides long.
Then, my personal favorite macromolecular complex, the “spliceosome”, grabs the pre-mRNA, cuts away most of it, and throws those parts away. The sections of DNA that code for the parts that are kept are called exons, while the sections that code for parts that are thrown away are called introns.
Next, another enzyme called “RNA guanylyltransferase” (we can’t all be beautiful) adds a “cap” to one end, and an enzyme called “poly(A) polymerase” adds a “tail” to the other end.
The pre-mRNA is now all grown up and has graduated to being regular mRNA. At this point, it is a few hundred or a few thousand nucleotides long.
Then, some proteins notice that the mRNA has a tail, grab it, and throw it out of the nucleus into the cytoplasm, where the noble ribosome lurks.
The ribosome grabs the mRNA and turns it into a protein. It does this by starting at one end and looking at chunks of three nucleotides at a time, called "codons". When it sees a certain "start" pattern, it starts translating each chunk into one of 20 amino acids and continues until it sees a chunk with a "stop" pattern.Since there are 4 kinds of nucleotides, there are 4³=64 possible chunks, while your body only uses 20 amino acids. So the ribosome, logically, gives some amino acids (like leucine) six different codons, and others (like tryptophan) only one codon. Also there are three different stop codons, but only one start codon, and that start codon is also the codon for methionine. So all proteins have methionine at one end unless something else comes and removes it later. Biology is layer after layer of this kind of exasperating complexity, totally indifferent to your desire to understand it.
The resulting protein lives happily ever after.
It’s thought that ~1 percent of your DNA is exons and ~24 percent is introns. What’s the rest of it doing?
Well, while the above dance is happening, other sections of DNA are “regulating” it. Enhancers are regions of DNA where a certain protein can bind and cause the DNA to physically bend so that some promoter somewhere else (typically within a million nucleotides) is more likely to get activated. Silencers do the opposite. Insulators block enhancers and silencers from influencing regions they shouldn’t influence.
While that might sound complicated, we’re just warming up. The same region of DNA can be both an intron and an enhancer and/or a silencer. That’s right, in the middle of the DNA that codes for some protein, evolution likes to put DNA that regulates some other, distant protein. When it’s not regulating, it gets transcribed into (probably useless) pre-RNA and then cut away and recycled by the spliceosome.
There's also structural DNA that's needed to physically manipulate the chromosomes. Centromeres are "attachment points" used when copying DNA during cell division. Telomeres are "extra" DNA at the ends of the chromosomes.
Telomeres shrink as we age. The body has mechanisms to re-lengthen them, but it mostly only uses these in stem cells and reproductive cells. Longevity folks are interested in activating these mechanisms in other tissues to fight aging, but this is risky since the body seems to intentionally limit telomere repair as a strategy to prevent cancer cells from growing out of control.
Further complicating this picture are many regions of DNA that code for RNA that’s never translated into a protein but still has some function. Some regions make tRNA, whose job is to bring amino acids to the ribosome. Other regions make rRNA, which bundle together with some proteins to become the ribosome. There’s siRNA, microRNA, and piRNA that screw around with mRNA produced. And there’s scaRNA, snoRNA, rRNA, lncRNA, and mrRNA. Many more types are sure to be defined in the future, both because it’s hard to know for sure if DNA gets transcribed, it’s hard to know what functions RNA might have, and because academics have strong incentives to invent ever-finer subcategories.
There are also pseudogenes. These are regions of DNA that almost make proteins, but not quite. Sometimes, this happens because they lack a promoter, so they never get transcribed into mRNA. Other times, they might lack a start codon, so after their mRNA makes it to the ribosome, it never actually starts making a protein. Then, there are instances when the DNA has an early stop codon or a "frameshift" mutation meaning the alignment of the RNA into chunks of three gets screwed up. In these cases, the ribosome will often detect that something is wrong and call for help to destroy the protein. In other cases, a short protein is made that doesn't do anything.
In more serious cases, these mutations might make the organism non-viable, or lead to problems like Tay-Sachs disease or Cystic fibrosis. But this wouldn’t be considered a pseudogene.
On messiness
Why? Why is this all such a mess? Why is it so hard to say if a given section of DNA does anything useful?
Biologists hate “why” questions. We can’t re-run evolution, so how can we say “why” evolution did things the way it did? Better to focus on how biological systems actually work. This is probably wise. But since I’m not a biologist (or wise), I’ll give my theory: Cells work like this because DNA is under constant attack from mutations.
Mutations most commonly arise during cell replication. Your DNA is composed of around 250 billion atoms. Making a perfect copy of all those atoms is hard. Your body has amazing nanomachines with many redundant mechanisms to try to correct errors, and it’s estimated that the error rate is less than one per billion nucleotides. But with several billion nucleotides, mutations happen.
There are also environmental sources of mutations. Ultraviolet light has more energy than visible light. If it hits your skin, that energy can sort of knock atoms out of place. The same thing happens if you’re exposed to radiation. Certain chemicals, like formaldehyde, benzene, or asbestos, can also do this or can interfere with your body’s error correction tricks.
Finally, we return to the huge fraction of your DNA (~50-60 percent) that is repeats of the same sequences. Some of this is caused by the machinery "slipping" while making a copy, leading to a loss or repetition of some DNA. There are also little sections of DNA called "transposons" that sort of trick your machinery into making another copy of those sections and then inserting them somewhere else in the genome.
“DNA transposons” get cut out and stuck back in somewhere else, while “retrotransposons” create RNA that’s designed to get reverse-transcribed back into the DNA in another location. There are also “retroviruses” like HIV that contain RNA that they insert into the genome. Some people theorize that retrotransposons can evolve into retroviruses and vice-versa.
It’s rare for retrotransposons to actually succeed in making a copy of themselves. They seem to have only a 1 in 100,000 or in 1,000,000 chance of copying themselves during cell division. But this is perhaps 10 times as high in the germ line, so the sperm from older men is more likely to contain such mutations.
Mutations in your regular cells will just affect you, but mutations in your sperm/eggs could affect all future generations. Evolution helps manage this through selection. Say you have 10 bad mutations, and I have 10 bad mutations, but those mutations are in different spots. If we have some babies together, some of them might get 13 bad mutations, but some might only get 7, and the latter babies are more likely to pass on their genes.
But as well as selection, cells seem designed to be extremely robust to these kinds of errors. Instead of just relying on selection, there are many redundant mechanisms to tolerate them without much issue.
And remember, evolution is a madman. If it decides to tolerate some mutation, everything else will be optimized against it. So even if a mutation is harmful at first, evolution may later find a way to make use of it.
On information again
So, in theory, how should we define the “information content” of DNA? I propose a definition I call the “phenotypic Kolmogorov complexity”. (This has surely been proposed by someone before, but I can’t find a reference, try as I might.) Roughly speaking, this is how short you could make DNA and still get a “human”.
The “phenotype” of an animal is just a fancy way of referring to its “observable physical characteristics and behaviors”. So this definition says, like Kolmogorov complexity, to try and find the shortest compressed representation of the DNA. But instead of needing to lead to the same DNA you have, it just needs to lead to an embryo that would look and behave like you do.
The idea is this: Take a single-cell human embryo with your DNA, and imagine all the different ways you can modify the DNA. This would include not only removing useless sections but also moving things around. Limit yourself to changes that still lead to a "person" that would still look like you and have all the same capabilities you do. Now, compress each of those representations. The smallest compressed representation is the "information" in your DNA.
This definition isn’t totally precise, because I’m not saying how precisely the phenotype needs to match. Even if there’s some completely useless section of DNA and we remove it, that would make all your cells a tiny bit lighter. We need to tolerate some level of approximation. The idea is that it should be very close, but it’s hard to make this precise.
So what would this number be? My guess is that you could reduce the amount of DNA by at least 75 percent, but not by more than 98 percent, meaning the information content is:
12 billion bits
× 2 bits / nucleotide
× (2 to 25 percent)
= 480 million to 6 billion bits
= 60 MB to 750 MB
But in reality, nobody knows. We still have no idea what (if anything) lots of DNA is doing, and we’re a long way from fully understanding how much it can be reduced. Probably, no one will know for a long time.