The response prior
What our track record can tell us about existential risk
We can think of the chance that a particular threat ends in an existential disaster as determined by two factors:
The difficulty of the problem (what it would take to prevent disaster)
The effectiveness of our response (whether we’ll do enough)
In my view, most attempts to estimate existential risk concentrate on the difficulty of the problem, and pay much less attention to our likely response. There are probably several reasons for this, but I think one has to do with the nature of typical existential risk research. While part of the job is to estimate the risk of an existential disaster – where both factors come into play – a more central part is to work out how to prevent it. And here researchers naturally focus on the first factor, since the response is something to bring about rather than to predict. I think this focus easily carries over when they turn to estimating the risk.
An overlooked asymmetry
This lopsided emphasis is unfortunate, since it’s arguably easier to assess the likely effectiveness of the response than the difficulty of the problem. There’s little precedent we can use to gauge the difficulty of preventing threats like AI misalignment. Instead, most analyses are based on causal models of how the threat might unfold and what it would take to stop it. It’s difficult to build such models, and easy to become overconfident in conclusions that rest on long chains of conjectures.
On the other hand, we’re in a better position to judge the response than many seem to think. It’s true that we can’t say exactly how we will respond to specific threats, but there is another way of approaching this question. Instead of asking what particular measures we’ll take, we can ask whether we’ll take whatever measures turn out to be necessary. We can consult history: to what extent did we take the right measures against past threats? While the difficulty of past threats tells us little about future threats, past responses can teach us much about future responses. Together with general knowledge of human psychology and societal institutions, this track record gives us a useful response prior.
Implications
Of course, a prior is only a starting point. Predictions of our likely response to, say, AI misalignment must also draw on the specifics of the case. But the response prior can still have major implications for estimates of existential risk. So what are those implications? I think there’s room for reasonable disagreement about that. I’m more confident in the importance of the response prior than in what exactly it implies.
But my own response prior is shaped by my view of human agency. I think we often fall for the sleepwalking illusion, seeing ourselves as more passive than we really are. In fact, threats of disaster alarm us so much that we probably overshoot more than we undershoot. Moreover, our response tends to scale with the perceived threat. We are doing far more to stop climate change than we did to close the ozone hole – and if AI misalignment proves harder still, I’d expect an even more forceful response. Consequently, I think existential risk is lower than many researchers believe.
Granted, there are ways our response could fall short. We might not recognise the threat in time. Or we might fail to coordinate, even though we know what we need to do. These concerns are real, but I think they are sometimes given too much weight. There are many researchers scanning for new threats, and governments listen to them more than it might seem. And nuclear arms control suggests that when the alternative is death, even sworn enemies can coordinate. Human agency reaches further than we intuitively think.


