The Trial Surgery Never Has to Pass
Drugs cannot reach patients without beating a placebo; the few operations tested the same way have not always beaten sham, and most are never tested at all.
A new antihypertensive cannot be prescribed on the strength of a single-arm study showing that patients who took it got better. Regulators require a trial against placebo, blinded on both sides, because getting better after treatment is compatible with regression to the mean, natural remission, and the therapeutic charge of being cared for at all. A new way of operating on a knee or a spine faces no equivalent requirement. It can enter routine practice on case series, surgeon testimony, and the plausibility of its underlying mechanism, and once adopted it accumulates decades of use before anyone asks whether patients who received a sham version of the same procedure would have done just as well.
The few times the sham-controlled trial has actually been run on an established operation, the answer has not been reassuring. Moseley and colleagues randomised patients with osteoarthritic knees to arthroscopic debridement, arthroscopic lavage, or a sham procedure consisting of skin incisions and simulated instrument sounds with no joint instrumentation; at two years, the sham group reported pain and function outcomes statistically indistinguishable from either active arm (Moseley et al. 2002). Sihvonen and colleagues ran the same design on arthroscopic partial meniscectomy for degenerative meniscal tears, a procedure performed on hundreds of thousands of patients a year, and again found no significant advantage for real surgery over sham at twelve months on the trial’s primary outcome measures (Sihvonen et al. 2013). Freed and colleagues went further, drilling burr holes in the skulls of sham-control patients without penetrating the dura, to test fetal dopaminergic neuron transplantation for Parkinson’s disease; the transplant produced genuine improvement on some measures in younger patients but also caused disabling dyskinesias the sham arm did not experience, a result only interpretable because a sham arm existed to separate the treatment’s own harms from the disease’s natural course (Freed et al. 2001).
The pattern is not a single trial’s fluke. In the same week in August 2009, the New England Journal of Medicine published two independent, differently led trials of vertebroplasty, the injection of bone cement into a fractured vertebra to relieve pain. Kallmes and colleagues randomised 131 patients to vertebroplasty or a simulated procedure that reproduced the positioning, the needle pressure, and even the smell of the cement without injecting any, and found no significant difference in disability or pain scores at one month (Kallmes et al. 2009). Buchbinder and colleagues, working independently with a different patient population, ran a near-identical design and found no advantage for vertebroplasty over sham at any time point measured (Buchbinder et al. 2009). A procedure performed on tens of thousands of patients a year, reimbursed and taught as standard practice for the better part of a decade, failed to beat sham twice, in the same month, from two unconnected research groups. A systematic review by Wartolowska and colleagues of fifty-three published surgical trials with a placebo arm found that the sham group improved in the large majority of them, and that in roughly half the real procedure showed no advantage over sham at all (Wartolowska et al. 2014).
None of this means that most operations are worthless; the review’s other half of trials found real procedures beating sham ones, and unlike a sugar pill, a sham operation under anaesthesia is not an inert comparator. It carries its own powerful ritual and expectation effects, which is exactly the confound a trial exists to isolate. The claim is narrower and more structural: surgery has never been made to clear the bar that would tell anyone, in advance, which of its procedures belong in that worthwhile half. Pharmaceutical regulation treats the placebo-controlled trial as the default evidentiary unit, with departures from it requiring justification. Surgical innovation treats it as an occasional luxury, undertaken years or decades into a procedure’s life, usually at the initiative of an unusually sceptical investigator rather than as a condition of entry.
The institutional reasons for the asymmetry are real enough to name. A drug is a fungible molecule that a regulator can hold constant across a trial population; an operation is a skill that varies by surgeon, by hospital, and by the thousand small adaptations a procedure undergoes as it is taught from one surgeon to the next, which makes it harder to specify what exactly a trial would be testing. The case for forcing every new procedure through a dedicated sham-controlled trial before any adoption at all would also multiply the time and cost of surgical innovation by an order most health systems could not absorb, and demanding it of every arthroscopic refinement or suturing variant would freeze exactly the kind of incremental, surgeon-driven improvement that has made modern operations safer than their nineteenth-century ancestors. Sham surgery is also not free of cost in the way an inert pill is: an anaesthetic, an incision, and a recovery carry real risk for patients who by definition cannot benefit from the sham arm’s intervention, which is why Horng and Miller’s ethical analysis of the Moseley-era debate concluded that sham surgery is justifiable only under tightly bounded conditions, not as a routine hurdle for any operation whatsoever (Horng and Miller 2002).
That constraint narrows the claim rather than defeating it. The case for a mandatory sham-controlled trial is strongest exactly where the institutional objection is weakest: a procedure that is standardised enough to be taught and billed as a single code, performed on a large and growing population, invasive enough to carry meaningful surgical risk, and resting on a mechanistic rationale that has not itself been tested against the possibility that benefit comes from the incision and the expectation rather than the specific intervention inside it. Arthroscopic partial meniscectomy and vertebroplasty both met every one of those conditions for the better part of a decade before anyone tested them against sham, each performed at a scale that dwarfs almost any single drug’s annual prescription volume, each exposure carrying anaesthesia risk, infection risk, and cost, each resting on a mechanical rationale, shaving a torn meniscus, stabilising a fractured vertebra with cement, that turned out not to predict the trial’s result. Procedures meeting that description are not edge cases calling for occasional scrutiny; they are the paradigm case the sham-controlled trial was built to answer, and the fact that someone had to wait for an outside investigator’s initiative, years into each procedure’s dominance, rather than a professional or regulatory requirement built into its adoption, is the asymmetry this argument is against.
Surgery does not need the full apparatus of drug regulation, and most of what surgeons do will never need a sham arm to justify it. But procedures that reach that scale and that risk, resting on a mechanism nobody has tested against its own absence, should not get to wait for a sceptic with research funding and an unfashionable hypothesis. The bar drugs already clear is not an arbitrary bureaucratic inheritance; it exists because the alternative is operating, for years or decades, on the strength of a story that turns out to belong to the incision rather than to the intervention performed through it.
References
Buchbinder, R., Osborne, R. H., Ebeling, P. R., Wark, J. D., Mitchell, P., Wriedt, C., Graves, S., Staples, M. P., & Murphy, B. (2009). A randomized trial of vertebroplasty for painful osteoporotic vertebral fractures. New England Journal of Medicine, 361(6), 557–568.
Freed, C. R., Greene, P. E., Breeze, R. E., Tsai, W. Y., DuMouchel, W., Kao, R., Dillon, S., Winfield, H., Culver, S., Trojanowski, J. Q., Eidelberg, D., & Fahn, S. (2001). Transplantation of embryonic dopamine neurons for severe Parkinson’s disease. New England Journal of Medicine, 344(10), 710–719.
Horng, S., & Miller, F. G. (2002). Is placebo surgery unethical? New England Journal of Medicine, 347(2), 137–139.
Kallmes, D. F., Comstock, B. A., Heagerty, P. J., et al. (2009). A randomized trial of vertebroplasty for osteoporotic spinal fractures. New England Journal of Medicine, 361(6), 569–579.
Moseley, J. B., O’Malley, K., Petersen, N. J., Menke, T. J., Brody, B. A., Kuykendall, D. H., Hollingsworth, J. C., Ashton, C. M., & Wray, N. P. (2002). A controlled trial of arthroscopic surgery for osteoarthritis of the knee. New England Journal of Medicine, 347(2), 81–88.
Sihvonen, R., Paavola, M., Malmivaara, A., et al. (2013). Arthroscopic partial meniscectomy versus sham surgery for a degenerative meniscal tear. New England Journal of Medicine, 369(26), 2515–2524.
Wartolowska, K., Judge, A., Hopewell, S., Collins, G. S., Dean, B. J. F., Rombach, I., Brindley, D., Savulescu, J., Beard, D. J., & Carr, A. J. (2014). Use of placebo controls in the evaluation of surgery: systematic review. BMJ, 348, g3253.