The Trolley Problem Changed Jobs Without Changing Its Name
A thought experiment built to test a moral principle was repurposed to locate a brain process, and answering the second question was mistaken for answering the first
Philippa Foot’s 1967 essay on abortion runs to eleven pages, and the trolley occupies barely one paragraph of it. A runaway tram is heading for five men on the track; the driver can turn it onto a side track, where it will kill one. Foot’s interest was not the tram but the asymmetry: why does it seem permissible to divert the trolley, killing one to save five, when it seems impermissible for a surgeon to kill one healthy patient to harvest organs for five dying ones? The trolley was a control case, built to isolate whether the doctrine of double effect, and not simple arithmetic, was doing the moral work. Judith Jarvis Thomson multiplied the cases through the 1970s and into 1985, adding a bystander at a switch, a fat man pushed from a footbridge, a loop track that returns the trolley toward the five unless the one it strikes stops it. Her question stayed the same kind of question Foot’s had been: which principle, if any, correctly explains a moral difference competent judges seem to feel. Neither philosopher treated the scenario as a stimulus to be administered. It was an argument, compressed into a hypothetical, whose premises a reader could accept or reject.
That changed in 2001, when Joshua Greene and his colleagues put variants of the switch and footbridge cases in front of subjects in an MRI scanner. Their finding, reported in Science, was that footbridge-type dilemmas, which require bringing about harm through personal, hands-on force, activated brain regions associated with emotional processing more strongly than switch-type dilemmas did, and that this differential engagement tracked which judgment subjects made. Greene went on to build this into a dual-process account of moral cognition: an automatic, emotionally driven system that recoils from up-close harm, and a slower, controlled system capable of the cost-benefit reasoning that favours the numbers. The trolley and the footbridge stopped being a test of which normative principle explains a moral difference. They became a stimulus pair for locating which neural and cognitive systems produce a judgment at all.
This is a genuine change of subject, not a change of method within the same one. Foot and Thomson were asking a normative question: what makes an act of killing permissible or impermissible, given the details of how the harm comes about. Greene’s programme asks a causal question: what psychological process generates the verdict a person gives when presented with a description of harm. A theory that correctly answers the second question is not thereby a theory that answers the first, any more than a correct account of why people believe a mathematical proof settles whether the proof is valid. The two questions can certainly bear on each other. But they are not the same question wearing different instruments, and the trolley cases migrated from philosophy to cognitive science carrying an ambiguity about which one they were now being used to settle.
The ambiguity has not stayed harmless. Greene’s own further step was to argue that the emotional origin of footbridge-type refusals is evidence against trusting them: if a judgment is produced by an evolved aversion to close-range violence rather than by tracking any morally relevant feature of the act, its emotional pedigree undermines its authority as data for ethical theory, and the deliberative, arithmetic-tracking judgment deserves more weight. This is a debunking argument, and debunking arguments are a legitimate form of philosophical reasoning; discovering that a belief was produced by a process insensitive to the truth of the matter is a standard way of discrediting it. The trouble is that debunking arguments of this kind are only as good as the psychology they rest on, and the psychology has not held up as cleanly as the 2001 result suggested it would.
Guy Kahane’s 2015 review is the sharpest statement of the problem. What the sacrificial dilemmas actually track, he argues, is willingness to endorse instrumental harm to a person as a means, not endorsement of utilitarian aggregation as such; a judgment can favour pushing the man off the footbridge for reasons that have nothing to do with maximising welfare, and a judgment can refuse to push him for reasons that have nothing to do with a Kantian prohibition on treating persons as means. Sorting responses to a two-case contrast into a “utilitarian” bin and a “deontological” bin imports a normative taxonomy the empirical results do not themselves establish. Bauman and colleagues raised a related worry the same decade: sacrificial dilemmas are often reported by participants as more amusing than sobering, are unrepresentative of the moral situations people actually navigate, and may not engage whatever psychological processes ordinary moral judgment relies on outside the laboratory. If the switch and footbridge cases do not cleanly isolate a utilitarian-versus-deontological contrast, and do not generalise past their own artificial staging, then the debunking argument built on their neural correlates is debunking a target that may not be the target it claims.
The strongest reply available to the psychological programme is that this criticism proves too little. Even a flawed operationalisation of “utilitarian” and “deontological” judgment can still show that some judgments are more emotionally driven than others, and that finding alone, whatever we call the two categories, is philosophically interesting: it tells us something about the causal etiology of a class of moral intuitions that reflective equilibrium treats as evidence. Kahane does not deny that the dilemmas produce a real and replicable behavioural contrast; he denies only the label conventionally pinned to it. Rename the categories “harm-averse” and “outcome-focused” and the neuroscience survives, even if the debunking argument built on top of it needs a narrower conclusion than the one Greene originally drew.
That reply is fair, and it is also the concession that limits the original migration rather than vindicates it. What survives is a psychological finding about which situations recruit which cognitive processes; what does not automatically survive is the further philosophical claim that this recruitment pattern tells us which of the two resulting judgments we ought to trust. Thomson’s arguments about the footbridge case were conceptual: they turned on what it is to use a person as a means, on whether the physical route from act to death matters morally, on cases constructed precisely to strip away confounds like intention and foreseeability so that one variable could be examined at a time. None of that argument was ever hostage to what a scanner would show, and none of it is rescued or threatened by discovering that footbridge judgments recruit the amygdala more than switch judgments do. The normative question Foot and Thomson posed is answerable, if it is answerable at all, by the same philosophical means it always was. The causal question Greene’s programme poses is a real question too, and worth answering on its own psychological merits. Fifty years of using the same runaway trolley to ask both has made it easy to mistake progress on one for progress on the other.
References
Foot, P. (1967). The problem of abortion and the doctrine of the double effect. Oxford Review, 5, 5–15.
Thomson, J. J. (1976). Killing, letting die, and the trolley problem. The Monist, 59(2), 204–217.
Thomson, J. J. (1985). The trolley problem. Yale Law Journal, 94(6), 1395–1415.
Greene, J. D., Sommerville, R. B., Nystrom, L. E., Darley, J. M., & Cohen, J. D. (2001). An fMRI investigation of emotional engagement in moral judgment. Science, 293(5537), 2105–2108.
Bauman, C. W., McGraw, A. P., Bartels, D. M., & Warren, C. (2014). Revisiting external validity: Concerns about trolley problems and other sacrificial dilemmas in moral psychology. Social and Personality Psychology Compass, 8(9), 536–554.
Kahane, G. (2015). Sidetracked by trolleys: Why sacrificial moral dilemmas tell us little (or nothing) about utilitarian judgment. Social Neuroscience, 10(5), 551–560.