Nudge Theory: The Architecture of Choice, and What the Evidence Audit Left Standing

Every arrangement of options moves behaviour, so somebody is always choosing the arrangement. That is the doctrine's strongest argument, and it is a philosophical one that no experiment settles. The empirical half went through an unusually public audit in 2022: a contested meta-analysis of more than 200 studies reporting an effect its own authors called small to medium, a re-analysis of the very same dataset reporting almost nothing left once publication bias was corrected, two further published objections, an authors' reply, and a separate study that opened two of the largest United States nudge units' complete files and found real effects roughly a sixth the published size. This article carries the founding science that replicates across 19 countries, the famous organ donation default that moves registered consent enormously and, in one 35-country comparison, not the transplant rates, two research-integrity cases in the adjacent honesty literature, and it refuses the overclaim in both directions.
A form arrives with one box already ticked. A canteen puts the fruit where a hand reaches first and the pastries further down the counter. An app opens on a screen somebody chose, offering a plan somebody set as standard. None of these takes an option away from you, none of them changes what anything costs, and all of them change what people do. That is the subject of this article. The argument that turned it into a government programme is the observation that there was never a version of the form, the counter or the screen that did not do this.
Two claims sit at the centre of the field and separating them is most of the work. The first is descriptive: people do not weigh options the way the older economic models assumed, and the evidence for that is strong and has been directly replicated at scale. The second is a policy claim built on the first: that cheap rearrangements of a choice reliably change behaviour where it matters. In 2022 the second claim went through an unusually public audit. Five items appeared in one journal in a single year, and a separate study went and opened the filing cabinets of two of the largest United States nudge units, which had kept every trial, including the ones nobody published. Some of the programme survived that. Much of what had been in print turned out to be the visible part of a literature with a hole in it.
Two boundaries and one rule before we start. The bias programme underneath all of this, System 1 and System 2, anchoring, the Linda problem, the mechanics of loss aversion and the bias catalogue, belongs to this wing's article on the mind's blind spots, and enters here in two short passages and a pointer, because this article's subject begins where that one ends. Belief engineered from outside, by states and agencies and campaigns, belongs to propaganda. And the rule: every effectiveness claim below arrives with its own audit attached, in the same paragraph wherever the audit will fit there, because on this literature a confident sentence is nearly always a compressed one. Where our own research file states something its own sources do not support, this article says so at the point the claim would have been made. That happens six times below, and two further identifier defects are recorded in the source notes at the foot of the page.
01The Mind The Models Left Out

Behavioural economics begins as a repair to an assumption. The classical models gave their agent stable preferences and unlimited computation. Herbert Simon, in the Quarterly Journal of Economics in 1955, argued that a realistic model of choice has to be built instead for an organism with limited knowledge and limited computing ability. His term for the result is bounded rationality, and its working form is satisficing, his own coinage: under cognitive limits, time pressure and incomplete information, people take an option that is good enough against an internal aspiration level rather than optimise across the whole set. Note the date. This is 1955, roughly two decades before the heuristics and biases programme and from a different tradition, and bounded rationality is not Kahneman's idea.
The empirical foundation under everything that follows is prospect theory. Daniel Kahneman and Amos Tversky, in Econometrica in 1979, reported that people evaluate outcomes as gains and losses measured from a reference point rather than as absolute states, that losses weigh more heavily than equivalent gains, and that small probabilities are overweighted while large ones are underweighted. It is among the most cited papers in economics. That is the whole treatment it gets here, deliberately: the value function, the mechanics of loss aversion and the catalogue of biases built on top of them are the subject of the mind's blind spots, live in this wing, and re-teaching them here would be a tax on a reader who can simply follow the link.
Put the foundation's own replication status first, because most of this article is about results that did not hold, and a reader who meets those first will draw the wrong conclusion. Kai Ruggeri and thirty-one co-authors, in Nature Human Behaviour in 2020, re-ran Kahneman and Tversky's original 1979 items across 4,098 participants in 19 countries and 13 languages, adjusting only for current and local currencies and requiring every participant to answer every item. Their result: the results replicated for 94 percent of items, with some attenuation, and twelve of thirteen theoretical contrasts replicated. Both hedges inside that good news are theirs and both are kept here: some attenuation, and one contrast in thirteen that did not hold. The descriptive science of how people weigh gains against losses is among the better-replicated findings in psychology. What is contested in this article is the leap from that science to the claim that a cheap intervention reliably changes behaviour at scale.
The size of loss aversion has been measured many times, and the field has pooled the measurements. Alexander Brown, Taisuke Imai, Ferdinand Vieider and Colin Camerer, in the Journal of Economic Literature in 2024, gathered 607 empirical estimates of the loss aversion coefficient from 150 articles across economics, psychology and neuroscience, covering 1992 to 2017, and put the mean at 1.955, with a 95 percent probability that the true value lies between 1.820 and 2.102. Our own research file T_4_08 states the coefficient as approximately 2.25. That figure originates in a single 1992 paper and lies outside the meta-analytic interval above, so this page does not carry it as the value. Losses weigh roughly twice as much as equivalent gains: that is what the pooled evidence supports, and the interesting part is that the folk version was directionally right and numerically overprecise, which is the pattern this whole article traces.
Loss aversion is also not settled, and the disagreement is technical rather than rhetorical. David Gal and Derek Rucker, in the Journal of Consumer Psychology in 2018, argued that the evidence for a universal ratio near two to one is weaker than commonly assumed, particularly outside laboratory gambling paradigms. Eldad Yechiam and Dana Zeif went further in the Journal of Economic Psychology in 2025: restricting the pool to studies with symmetric gains and losses and no ordering of items, they report a loss aversion parameter of about 1.07, not significantly above 1.0, which would mean gains and losses weighted about equally. That is one recent paper against a pooled estimate across all designs, and the two are not answering quite the same question, because the disagreement is partly about which designs count. No winner is crowned here, and the single-paper result is not this page's verdict on loss aversion.
There is a neuroimaging result underneath this and it is more interesting than the version in circulation. Sabrina Tom, Craig Fox, Christopher Trepel and Russell Poldrack, in Science in 2007, scanned people deciding whether to accept fifty-fifty gambles. Potential losses were not represented by a separate alarm system firing harder. They were represented by the reward system going quieter: the same midbrain dopaminergic regions and their targets that ramped up as potential gains increased ramped down as potential losses increased, and the steepness of that asymmetry in the ventral striatum and prefrontal cortex predicted how loss-averse a person was behaviourally. Our own file T_4_08 reports this paper as showing greater amygdala and striatal activation for losses. The direction is inverted and the amygdala does not appear in the paper's abstract at all, which names the ventral striatum and the prefrontal cortex. Loss aversion in the brain looks like the absence of pleasure rather than the presence of pain.
02The Doctrine, And The Test With Teeth
The policy doctrine has a name and a date. Richard Thaler and Cass Sunstein set out libertarian paternalism in the American Economic Review in 2003, five years before the book that made it famous: choice architects may steer people toward what is good for them, so long as no option is removed and opting out stays cheap. The phrase is a deliberate contradiction in terms and its authors know it. Whether the libertarian half is doing real work or standing in as a fig leaf is the ethical argument of this entire subject, and it is argued out in section 08 rather than assumed here.

Nudge, published by Yale University Press in 2008, is where the doctrine reached policymakers, and its two working definitions are the ones the rest of this article uses. A choice architecture is any arrangement of options that a person must encounter in order to choose at all. A nudge is any feature of that arrangement that predictably changes behaviour without forbidding an option and without significantly changing its economic cost. Both definitions are narrower than the way the word is used in public debate, and holding them tightly is what makes the rest of this page possible.
The freedom-preserving test is doing all the ethical work, and it has teeth. A nudge must leave every option available and must not significantly change the economic cost of any of them. So a tax is not a nudge. A ban is not a nudge. A subsidy is not a nudge. Moving the fruit to eye level is. A great deal of what gets called nudging in public argument fails this test on the doctrine's own terms, and this article polices the definition rather than repeating it.
The argument that makes nudging hard to refuse is that there is no neutral option. Someone has to decide which item the canteen puts first, which box on the form is already ticked, which screen the app opens on. Every one of those arrangements moves behaviour, so the choice is never between influencing and not influencing; it is between influencing thoughtlessly and influencing deliberately. The standard reply is Daniel Hausman and Brynn Welch, in the Journal of Political Philosophy in 2010, and it is a good one: from the fact that some arrangement is unavoidable it does not follow that any particular deliberate arrangement is legitimate, and shaping behaviour by exploiting a known flaw in someone's reasoning is different in kind from shaping it by giving them a reason. This is a philosophical dispute rather than an empirical one. No study settles it, and this page does not pretend one does, which is why it is filed at the credible tier and not the verified one.
One mechanism does need naming here, because it is the tool that makes a nudge possible without changing a single fact. Amos Tversky and Daniel Kahneman, in Science in 1981, showed that logically equivalent descriptions of the same outcome produce different choices: a treatment described by its 90 percent survival rate is chosen more often than the same treatment described by its 10 percent mortality rate. The phenomenon itself belongs to the mind's blind spots. What matters here is that framing is the exact point at which nudging and misleading become hard to tell apart, because both consist of choosing a true description.
03From A Book To A Government Unit
The idea became machinery. The Organisation for Economic Co-operation and Development published a survey of behavioural-insights units inside governments in 2017, Behavioural Insights and Public Policy: Lessons from Around the World, by which point dedicated teams applying this literature were operating inside national administrations and international bodies across several continents. No count of such units appears on this page. The research for this article did not read the report's own tally, and a number nobody here checked is exactly the kind of statistic that would open a policy section well and be wrong.
The best known of these teams is the British one, universally called the nudge unit. Its entry on the UK government's own website now records that the Behavioural Insights Team is independent of government and that Behavioural Insights Limited is fully owned by Nesta, a charity. That transition is the citable fact and it is the interesting one: a behavioural science team inside a government became a company outside it. The page carries no founding date, no original mandate and no statement of which department the team once sat in, so this article states none of those.
Two of the people in this article hold the Sveriges Riksbank Prize in Economic Sciences in Memory of Alfred Nobel: Daniel Kahneman in 2002 and Richard Thaler in 2017. The prize years are not in serious doubt and they are carried here in prose. The official prize records were unreachable when the research for this article was done and are reachable now, so both are linked at the foot of this page. They are cited for the prize and the year, and for nothing else.
Thaler also named the inverse of his own idea. Writing in Science in 2018 he called it sludge: choice architecture deployed to make a good outcome harder rather than easier. The friction in a cancellation flow. The rebate that requires a posted form. The unsubscribe link three screens down. That one-page essay is not a study and is not treated here as one, but the concept is the most immediately useful thing in this article for a reader's own life, and it is the honest bridge to the ethics: the same knowledge that lets an architect smooth a path lets another obstruct one, and commercial actors had the knowledge first.
04The Most Famous Default In The World
Every popular account of nudging opens with the organ donation register, and this one does too, because it is also the clearest case of the thing this article is really about. Eric Johnson and Daniel Goldstein, in Science in 2003, set European countries side by side and found effective consent rates clustering at two extremes. Where a citizen must actively opt in, the rates ran from about 4 percent in Denmark to about 28 percent in the Netherlands, with Germany near 12 percent and the United Kingdom near 17 percent. Where consent is presumed unless a citizen opts out, they ran from about 86 percent in Sweden to about 99.98 percent in Austria, with Belgium, France, Hungary, Poland and Portugal all above 98 percent. Austria and Germany are the argument in one line: neighbouring countries, one difference in the law, near-opposite consent rates. The measured variable throughout is registered consent. It is not donation, and it is not a transplant.
| Country | Consent Regime | Effective Consent Rate |
|---|---|---|
| Denmark | Opt-in, consent must be given actively | About 4 percent |
| Germany | Opt-in, consent must be given actively | About 12 percent |
| United Kingdom | Opt-in, consent must be given actively | About 17 percent |
| Netherlands | Opt-in, consent must be given actively | About 28 percent |
| Sweden | Opt-out, consent presumed | About 86 percent |
| Belgium, France, Hungary, Poland and Portugal | Opt-out, consent presumed | All above 98 percent |
| Austria | Opt-out, consent presumed | About 99.98 percent |
The consent figures are not donation figures, and that gap is the whole argument. Adam Arshad, Benjamin Anderson and Adnan Sharif, in Kidney International in 2019, compared 35 OECD countries for 2016, 17 of them opt-out and 18 opt-in, using donation and transplantation rates from the Global Observatory on Donation and Transplantation in a multivariate regression model. Opt-out countries had fewer living donors per million population, 4.8 against 15.7. There was no significant difference in deceased donors, 20.3 against 15.4. And there was no significant difference in kidney transplantation, 35.2 against 42.3, in non-renal transplantation, 28.7 against 20.9, or in total solid organ transplantation, 63.6 against 61.7. In their model an opt-out system independently predicted fewer living donors and was not associated with deceased donors or with transplantation rates at all. Their own conclusion is that other barriers to donation must be addressed even where consent is presumed. Read this carefully in the other direction too: it is one cross-sectional comparison at one time point, and it does not show that opt-out systems do harm. It shows no advantage on the outcomes measured.
The editorial published alongside that comparison put it more sharply. Rafael Matesanz and Beatriz Dominguez-Gil, two figures central to the Spanish transplant system, wrote that there are "no clear examples of countries with a real sustained increase in organ donation after modifying the law". Their argument is that what distinguishes a high-donation country is organisational: the coordinators, the intensive care pathways, the family conversations, and that beside all that the consent statute is close to a rounding error. This is an editorial rather than a study, and the tier reflects that. It is also the most interesting possible source for the claim, because Spain is the country everyone cites as the presumed-consent success story.
Refused: that switching a country from opt-in to opt-out organ donation increases the number of transplants. The consent statistics above are real and they are enormous. The outcome evidence does not license the inference, and the contemporary 35-country comparison finds no significant difference in deceased donors or in total solid organ transplantation. A related refusal is procedural rather than substantive. Our own research file, and the build plan that commissioned this article, both carry a vivid sentence putting the effect of an opt-out switch at a specific pair of percentages. Neither figure appears in the paper cited for them, and the paper's measure is registered consent rather than donation, so this page does not print that pair even in the act of correcting it. A registry is not a surgical team, a coordinator, an intensive care bed, a family conversation or a matching organ.
05The Other Defaults, And What Ownership Does
The other default with real money behind it is retirement saving. Brigitte Madrian and Dennis Shea, in the Quarterly Journal of Economics in 2001, studied a single large American employer that switched new hires from having to enrol in a 401(k) plan to being enrolled unless they declined. Participation was significantly higher under automatic enrolment. Their second finding is the one that should worry a choice architect: a substantial fraction of those automatically enrolled kept both the default contribution rate and the default fund allocation, even though almost nobody hired before the change had chosen that particular combination for themselves. Our own file states a before-and-after pair of participation percentages for this study. The research for this article could not verify them, because the paper's own abstract states the direction and not the figures and the published article sat behind a paywall, so no percentage for this study appears anywhere on this page. The persistence finding is the better story in any case. The nudge did not merely get people to save. It got them to save at a rate a company picked, which almost nobody would have picked for themselves, and that is the entire ethical problem arriving as an empirical result.
There is a theoretical reason a default might do more than save effort. Richard Thaler named the endowment effect in the Journal of Economic Behavior and Organization in 1980: people demand more to give up something they already own than they would have paid to acquire it. Daniel Kahneman, Jack Knetsch and Richard Thaler made it a laboratory staple in the Journal of Political Economy in 1990, showing the gap between willingness to accept and willingness to pay across repeated market experiments. If ownership itself changes valuation, then being defaulted into something does not just spare a person the paperwork; it changes what the option is worth to them. That is the strongest bridge from the bench to the policy, which is exactly why the next paragraph matters.
The endowment effect has been argued over for twenty years, and the argument runs both ways. Charles Plott and Kathryn Zeiler, in the American Economic Review in 2005, contended that the gap between what people demand to sell and what they will pay to buy is produced by subjects misunderstanding the elicitation procedure, and that under a training protocol designed to remove those misconceptions the gap vanishes. Dietmar Fehr, Rustamdjan Hakimov and Dorothea Kuebler, in the European Economic Review in 2015, then ran Plott and Zeiler's own protocol and did not reproduce their result. So the finding is contested and the critique of the finding is contested, and this page does not adjudicate an experimental-economics methods dispute it cannot resolve. Keep the shape of it, though, because it inoculates against a reflex the rest of this article could otherwise encourage: a replication failure is not automatically a verdict against the original claim. Here it is a verdict against the critique.
06The Social Norm, Measured At Scale
The social-norms nudge has the largest clean field test in the literature and it is the best calibration device this article can offer. Hunt Allcott, in the Journal of Public Economics in 2011, evaluated letters that told American households how their electricity use compared with their neighbours', across randomised natural field experiments at 600,000 treatment and control households. The average programme reduced energy consumption by 2.0 percent. That is a small number and a real one: the paper puts it as equivalent to a short-run electricity price increase of 11 to 20 percent, at a cost-effectiveness that compares favourably with traditional conservation programmes. Note the claim being made. Not that the effect is large, which it is not, but that it is cheap, which it is. The effect was also very unevenly spread: households in the highest decile of pre-treatment consumption cut usage by 6.3 percent, while the lowest decile cut 0.3 percent.
The obvious follow-up question is whether an effect that small also fades, and the best-powered answer is more interesting than either decay or permanence. Hunt Allcott and Todd Rogers, in the American Economic Review in 2014, followed the same programme over years. Each new report produced a burst of saving followed by backsliding, and those short cycles flattened out over time. The underlying effect did not evaporate: households whose reports were discontinued after two years kept most of the saving, losing it at 10 to 20 percent a year, and households that kept receiving reports kept responding, with no sign of habituating away. Our own file T_4_08 asserts, with no citation attached, that nudge effects attenuate or disappear over time as people adapt. On the one intervention the file itself names, the best longitudinal evidence says the opposite about the level. What attenuates is the sawtooth, not the effect, and this article will not assert general attenuation. This is a place where the literature comes out better than our own summary of it, and saying so is part of the same discipline as saying the reverse.
07The Reckoning
The attempt that set the terms of the dispute came from Stephanie Mertens, Mario Herberz, Ulf Hahnel and Tobias Brosch, in the Proceedings of the National Academy of Sciences in 2022: more than 200 studies reporting over 450 effect sizes, a combined sample of 2,149,683 people. They reported an overall effect of Cohen's d = 0.45, with a 95 percent confidence interval from 0.39 to 0.52, which they described as small to medium. Interventions that restructure the choice itself outperformed those that merely describe the options or reinforce an intention, and food choices responded up to two and a half times more strongly than other domains. Their own abstract also carries a sentence that our file drops and that the rest of that year's literature seized on, and it is quoted directly below. Our file T_4_08 additionally reports this work as 212 randomised controlled trials with a median effect of 0.43; the unit is studies and effect sizes rather than trials, and the statistic is a mean with an interval rather than a median.
Six months later, in the same journal, six authors re-analysed that same dataset with a method built to correct for publication bias, and the effect very nearly disappeared. Maximilian Maier, Frantisek Bartos, T. D. Stanley, David Shanks, Adam Harris and Eric-Jan Wagenmakers, using robust Bayesian model-averaging, reported a model-averaged estimate of 0.04, with a credible interval from 0.00 to 0.14. Taking only the single most precise estimate from each study raised it to 0.11, interval 0.00 to 0.24. Split by category, the evidence ran against an effect for information-based nudges and for assistance-based nudges, and came out simply inconclusive for the category that restructures the choice itself, which is where defaults live: 0.12, with an interval from 0.00 to 0.43 and a Bayes factor close to 1. That last point is the one most often misread, so it gets its own sentence. A Bayes factor near 1 means the data cannot tell, not that the effect is zero. Their title says no evidence for nudging after adjusting for publication bias, and no evidence for is a different claim from evidence against.
Two further objections appeared in the same issue, and this page names them without summarising arguments it did not read. Nine authors including the statistician Andrew Gelman and Daniel Goldstein, co-author of the organ donation study in section 04, published "No reason to expect large and consistent effects of nudge interventions". Goldstein's presence on it is worth a clause: the co-author of the most famous default study in the world signed a paper saying not to expect large and consistent effects. Jonathan Bakdash and Laura Marusich published "Left-truncated effects and overestimated meta-analytic means", arguing that the distribution of effects in the meta-analysis was truncated at the low end in a way that inflates a pooled mean. Both are stated here at the level of their titles, which is all the research for this article read, and neither is elaborated into a statistical argument this page cannot vouch for.
The original authors replied in the same issue, under the title "The present and future of choice architecture research". This article did not read that reply either, and it names it for a reason that is not politeness: an article that gives a re-analysis the last word on a dispute where the original authors published an answer has quietly taken a side in a live scientific argument. Five items in one journal in one year, and the honest closing position is narrower than any of the headlines. What is established is that the published literature carries substantial publication bias. Where the true average effect sits, somewhere between very small and moderate, nobody yet knows.
The cleanest explanation of the gap came from a different direction entirely, and it is the strongest evidence on this page. Stefano DellaVigna and Elizabeth Linos, in Econometrica in 2022, obtained every trial run by two of the largest nudge units in the United States: 126 randomised controlled trials covering 23 million individuals, including the ones nobody wrote up because they did not work. In the academic journal papers, the average nudge raised take-up by 8.7 percentage points, a 33.4 percent increase over the control group. In the units' complete record, the average effect was 1.4 percentage points, an 8.0 percent increase. Still statistically significant. Still real. Roughly a sixth the published size. Their conclusion is that publication bias in the academic journals, made worse by low statistical power, can account for the full difference. This is not a modelling assumption or a re-weighting. It is the filing cabinet, opened, and it delivers the fairest verdict available: the effects are real, and they were oversold by something like a factor of six.
| The Item | What It Pooled Or Measured | What It Reported |
|---|---|---|
| Mertens, Herberz, Hahnel and Brosch (2022), PNAS | More than 200 studies reporting over 450 effect sizes, combined sample 2,149,683 | Cohen's d = 0.45, 95 percent confidence interval 0.39 to 0.52, described by the authors as small to medium, alongside their own report of a moderate publication bias toward positive results |
| Maier, Bartos, Stanley, Shanks, Harris and Wagenmakers (2022), PNAS | A robust Bayesian model-averaging re-analysis of that same dataset, correcting for publication bias | 0.04, credible interval 0.00 to 0.14. Most precise estimate per study: 0.11, interval 0.00 to 0.24. The category containing defaults: 0.12, interval 0.00 to 0.43, inconclusive rather than null |
| Szaszi and eight co-authors (2022), PNAS | A separate objection in the same issue, read here as a record and a title only | Title carried, argument not summarised: No reason to expect large and consistent effects of nudge interventions |
| Bakdash and Marusich (2022), PNAS | A third objection in the same issue, read here as a record and a title only | That the distribution of effects was truncated at the low end in a way that inflates a pooled mean. Stated at the level of the title and no further |
| Mertens, Herberz, Hahnel and Brosch (2022), reply, PNAS | The original authors' response to all three, read here as a record and a title only | The present and future of choice architecture research. Named so that the re-analysis does not get the last word by default |
| DellaVigna and Linos (2022), Econometrica | 126 randomised controlled trials covering 23 million individuals: every trial run by two of the largest United States nudge units, published or not | Academic journal papers, 8.7 percentage points, a 33.4 percent increase. The units' full record, 1.4 percentage points, an 8.0 percent increase, still statistically significant |
None of this arrived as a surprise to the doctrine's own authors. Cass Sunstein published "Nudges that fail" in the first issue of Behavioural Public Policy in 2017, five years before the meta-analytic argument broke. This page names that paper and does not enumerate its reasons, because the research behind this article read the record and not the text. The bare fact is worth carrying by itself: the criticism did not come only from outside, and the caricature in which nudge advocates conceded nothing is false.
One reply to the whole dispute is that averaging was the mistake. Christopher Bryan, Elizabeth Tipton and David Yeager argued in Nature Human Behaviour in 2021 that behavioural science is unlikely to change the world without a heterogeneity revolution: that the useful question is not how big the average effect is, but which people, in which situations, respond at all. This article already has the instance to hand. Allcott's heaviest-using tenth of households cut 6.3 percent while the lightest-using tenth cut 0.3 percent, a twenty-fold difference inside one intervention, and it is precisely what an average of 2.0 percent conceals. That position is carried here as its authors' title states it and no further, and it lets this section end without either triumph or nihilism.
Refused: that nudges do not work. Three things forbid it, and the third is the most easily forgotten. The publication-bias correction returned a Bayes factor close to 1 for the category that contains defaults, which means the data cannot tell rather than that the effect is absent. The complete unpublished record of two nudge units still shows an average effect of 1.4 percentage points that is statistically significant, on interventions that cost almost nothing. And the original authors published a reply this page did not read, which means the dispute is open on the record and not closed by the re-analysis. An article landing on it was all nonsense would be repeating the original error at the opposite sign: taking a striking result from one side of a live literature and stating it as settled.
08The Objections That Do Not Turn On Effect Sizes
The ethical objection is old and has never been answered to everyone's satisfaction. Even libertarian paternalism requires that someone decide what is good for other people, and the deciding is done by whoever controls the form, the app or the shelf. Hausman and Welch sharpen it: a nudge that works by exploiting a known defect in someone's reasoning is not the same kind of act as an argument, even when the destination is identical and the exit stays open. The second half is mandatory, though, and it is usually missing. If choice architecture is genuinely unavoidable, then declining to design deliberately protects nobody's autonomy; it hands the design to whoever is already doing it commercially, which is sludge. So the strongest form of the argument is not that nudging is wrong. It is that nudging is an exercise of power and needs the accountability of one: who chose, on what authority, with what published evidence, reversible by whom and how.
The deepest objection is not about ethics or effect sizes but about the premise. Gerd Gigerenzer's programme holds that fast and frugal heuristics, among them the recognition heuristic set out with Daniel Goldstein in Psychological Review in 2002 and the rule called take-the-best, are not defective approximations of optimisation but adaptations to specific environments, and that under uncertainty and limited information a simple rule can outperform a complex model. On that reading, much of what the heuristics and biases tradition catalogued as error is the right answer to a harder question than the laboratory was asking. Gigerenzer put the argument directly against this doctrine in the Review of Philosophy and Psychology in 2015, in a paper titled "On the Supposed Evidence for Libertarian Paternalism". This page states the position and cites the paper; it reports no specific argument, example or comparison from inside either, because the research behind this article did not read them.
There is an alternative with a name and a clean contrast behind it. Ralph Hertwig and Till Gruene-Yanoff, in Perspectives on Psychological Science in 2017, set boosting against nudging: instead of arranging the environment so that a person with unchanged competences behaves better, teach the competence, so that the improvement travels with the person and survives the removal of the scaffolding. That distinction is one a reader can actually use, and it sharpens the whole effectiveness question. A nudge that works only while the designer maintains it is a different kind of policy from a skill that persists. No specific boosting intervention or effect size is attributed here; the distinction is what this page carries.
The broadest objection came from within behavioural economics itself. Nick Chater and George Loewenstein, in Behavioral and Brain Sciences in 2023, argued that behavioural public policy went astray by framing problems at the level of the individual rather than the system, and that an emphasis on individual-level solutions has diverted attention and political energy from structural reform. Loewenstein is one of the founders of the field, which makes the source of this critique matter about as much as its content. The position is stated here as the paper's title states it and no further: the research behind this article read neither the target article nor the peer commentaries that accompany every paper in that journal.
09The Adjacent Collapse, And What It Does Not Touch
The most widely deployed choice-architecture intervention of the 2010s was a signature box moved to the top of a form. A 2012 paper in the Proceedings of the National Academy of Sciences, by Lisa Shu, Nina Mazar, Francesca Gino, Dan Ariely and Max Bazerman, reported that signing an honesty declaration before filling in a form rather than after reduced dishonest self-reporting, including in a field study with an insurance company. Governments and firms adopted it. The paper was retracted in September 2021, and its own publisher's metadata now carries the retraction marker in the title.
The order of events matters more than the headline. Eighteen months before the retraction, the original authors themselves published a paper in the same journal titled "Signing at the beginning versus at the end does not decrease dishonesty". The effect had already failed to reproduce in their own hands before any question about the data arose. Carrying the retraction without this paper tells a story about fraud. Carrying both tells a truer and more useful one: an intervention can be adopted at scale, fail to replicate under its own authors' testing, and only later be found to rest on a compromised dataset. Those are two separate events with two separate lessons, and the first one arrived first. Publishing a null against your own famous result is the behaviour a reader should want from a field, and every author of the 2012 paper is on it, with two more.
Two further papers in Psychological Science, both first-authored by Francesca Gino, carry retraction notices in their own publisher records: a 2014 paper with Scott Wiltermuth on dishonesty and creativity, and a 2015 paper with Maryam Kouchaki and Adam Galinsky on inauthenticity and feelings of immorality. The retraction status of both is established by the publisher's deposited metadata rather than by reporting. Two verified retractions in her own first-authored work is what this page states. No count of how many papers are affected appears here; counts circulate, and the research behind this article verified two.
The institutional aftermath is journalism rather than literature, and it is carried as such. The Harvard Crimson reported in May 2025 that Harvard revoked the professor's tenure following a Harvard Business School investigation, and other outlets reported the same. The investigation report became public through litigation rather than through the university. That is the extent of what this page says about it: it concerns a living person, the university has not published its report, and this article has read the reporting and not the record.
About the dataset itself, the record supports four things and this page states those four and no more. The insurance-company data behind the retracted 2012 study were found to contain fabricated responses. All five original authors agreed the data were not genuine. Dan Ariely, who obtained the dataset from the company, has consistently said that he did not fabricate it. Duke University investigated and has not made its findings public. This article names nobody as the fabricator, does not clear anybody either, and does not characterise a conclusion that has never been published. Everything past that boundary is an accusation about a living person that nothing available here supports.
And now the proportion, because two true stories are easy to fuse into one false one. Not one of Herbert Simon, Daniel Kahneman, Amos Tversky, Richard Thaler, Cass Sunstein, Brigitte Madrian, Dennis Shea, Eric Johnson, Daniel Goldstein, Hunt Allcott, Todd Rogers, Stephanie Mertens or Maximilian Maier is implicated in any research-integrity finding anywhere in the material behind this article. The cases above belong to the honesty and dishonesty research programme, which is genuinely adjacent, because the signature intervention is itself a choice-architecture nudge and two of the authors are among the most cited names in behavioural science. It is not the foundation of nudge theory, and its collapse does not touch prospect theory, defaults, social norms or the effectiveness dispute in section 07. The two stories are about different failures. The effectiveness story is about publication bias and low statistical power, which are structural and involve no allegation of misconduct against anyone. The integrity story is about specific compromised datasets in specific named papers. A sentence of the form the field has been rocked by fraud fuses them, and fusing them would be exactly the error this article spends its length accusing popular accounts of making.
10What Replicated, And What Did Not
The replication crisis in this neighbourhood was selective rather than general, and one project shows both halves at once. Anchoring, from Tversky and Kahneman's 1974 paper in Science, is the finding that an arbitrary number presented before a numerical judgment drags the judgment toward it. It is one of the sturdiest results in the area: the Many Labs project, reported by Richard Klein and fifty co-authors in Social Psychology in 2014, tested thirteen classic and contemporary effects across 36 independent samples totalling 6,344 participants, ten of the thirteen replicated consistently, and all four anchoring items were among the largest effects measured. The phenomenon itself belongs to the mind's blind spots; it earns its paragraph here only because of what happened to its neighbours in the same study.
The clearest casualty in the same neighbourhood is social priming. John Bargh, Mark Chen and Lara Burrows reported in the Journal of Personality and Social Psychology in 1996 that participants exposed to words associated with old age subsequently walked more slowly down a corridor. Stephane Doyen, Olivier Klein, Cora-Lise Pichon and Axel Cleeremans failed to reproduce it in PLoS ONE in 2012, and found the effect appeared only when the experimenters knew which condition a participant was in. In the Many Labs project, two behavioural priming effects, flag priming and currency priming, did not replicate at all, in the very study in which all four anchoring items did. Two corrections to our own file travel with this paragraph. These are social psychology studies rather than behavioural economics ones, and the association is loose enough to mislead about which literature the failure damages. And our file attributes the elderly-priming failure jointly to Doyen and colleagues and to Many Labs; the research for this article could not verify that any Many Labs project attempted elderly priming, so the two attributions are kept separate here. The sorting was not random, which is the part worth taking away: the fragile effects were mostly the ones claiming that a subtle cue produces a large behavioural change, and that is uncomfortably close to the nudge claim itself.
11When The Architect Is An Algorithm
The open question is what happens when the choice architect is a system that watches you. A personalised default, adjusted in real time against an individual profile, may not be the same object as a rearranged canteen: it is dynamic, it is invisible, and it is not the same for any two people, so the opt-out that makes a nudge legitimate on the doctrine's own test may be present in principle and unfindable in practice. Karen Yeung named this hypernudge in Information, Communication and Society in 2017. Whether it would dramatically improve health, financial and environmental outcomes, or constitute a form of manipulation without precedent, is an active area of ethical and empirical argument and is not settled. Every hedge in this paragraph is load-bearing, and the tier is our own file's, not a downgrade applied here.
Fast Facts
- The Definition, Held Tightly
- A choice architecture is any arrangement of options a person must encounter in order to choose at all. A nudge changes behaviour predictably without forbidding an option and without significantly changing its economic cost. A tax is not a nudge, a ban is not a nudge, and much of what public argument calls nudging fails the doctrine's own test
- The Foundation Replicates
- Kahneman and Tversky's 1979 items were re-run across 4,098 participants in 19 countries and 13 languages (Ruggeri and 31 co-authors, 2020): 94 percent of items replicated with some attenuation, and twelve of thirteen theoretical contrasts held
- The Size Of Loss Aversion
- 607 estimates from 150 articles give a mean coefficient of 1.955, 95 percent interval 1.820 to 2.102 (Brown, Imai, Vieider and Camerer, 2024). A 2025 re-meta-analysis restricted to symmetric designs reports about 1.07, not significantly above 1.0. One pooled estimate, one restricted one, and a live technical dispute about which designs count
- The Organ Donation Default
- Johnson and Goldstein (2003) measured registered CONSENT: about 4 percent in Denmark to about 28 percent in the Netherlands under opt-in, about 86 percent in Sweden to about 99.98 percent in Austria under opt-out. Arshad, Anderson and Sharif (2019), across 35 OECD countries, found no significant difference in deceased donors or in total solid organ transplantation, and fewer living donors in opt-out countries
- The Best-Measured Field Nudge
- Home energy reports across 600,000 treatment and control households reduced consumption by 2.0 percent, equivalent to a short-run price rise of 11 to 20 percent, at favourable cost (Allcott, 2011). The heaviest-using tenth cut 6.3 percent; the lightest-using tenth cut 0.3 percent
- The Two Numbers From One Dataset
- Mertens and colleagues (2022) pooled more than 200 studies and over 450 effect sizes to a contested Cohen's d = 0.45, 95 percent interval 0.39 to 0.52, reporting a moderate publication bias in the same abstract. Maier and colleagues (2022) re-analysed that same dataset correcting for publication bias and estimated 0.04, credible interval 0.00 to 0.14. Two further objections and an authors' reply appeared in one journal in one year, and the question is open
- The Filing Cabinet Opened
- 126 randomised controlled trials covering 23 million individuals, every trial run by two large United States nudge units (DellaVigna and Linos, 2022): 8.7 percentage points in the academic papers, 1.4 percentage points in the complete record, still statistically significant. Roughly a sixth the published size
- Sludge
- Thaler's own name, in Science in 2018, for choice architecture used to make a good outcome harder: the cancellation flow, the posted rebate form, the buried unsubscribe link
- The Integrity Cases, And Their Limits
- A 2012 signature-box study was retracted in September 2021, eighteen months after its own authors published a failed self-replication. Two further papers first-authored by Francesca Gino carry retraction notices in their publishers' records. This page names nobody as the fabricator of any dataset, and no figure in the effectiveness story above is implicated in any integrity finding
- Refused In Both Directions
- That switching to opt-out organ donation increases transplants, which the outcome evidence does not support. And that nudges do not work, which the Bayes factor near 1 for the default category, the significant 1.4 percentage points in the unpublished record, and the unread authors' reply all forbid
- What Our Own File Got Wrong
- A loss aversion coefficient of 2.25, outside the field's own meta-analytic interval; a neuroimaging result stated backwards, with a brain region that is not in the paper's abstract; a pair of organ donation percentages that appear in no cited source, attached to donation when the measure is consent; general attenuation over time, contradicted by the best study of the one programme the file names; a contested meta-analysis reported as 212 trials with a median of 0.43, where the paper itself gives a disputed mean of 0.45 with an interval of 0.39 to 0.52, and where its own publication-bias sentence was dropped; and social priming misfiled as behavioural economics. Each is corrected in the open above
- Numbers This Page Will Not Print
- The organ donation percentage pair our file and the build plan both carry, refused in section 04 without reprinting it; the 401(k) participation percentages our file states, which could not be verified against the paper; a count of nudge units anywhere in the world; and any result at all from the three 2022 items and the seven further works this page names for their existence and does not read
What We Can Actually Stand Behind
The descriptive foundation holds. People evaluate outcomes as gains and losses from a reference point, losses weigh more than equivalent gains, and small probabilities are overweighted, and a 2020 multinational study re-ran the original 1979 items across 4,098 participants in 19 countries and 13 languages with 94 percent of items replicating, with some attenuation, and twelve of thirteen theoretical contrasts holding. Bounded rationality is older still, from 1955. Whatever else is contested below, behavioural economics is not a house of cards.
Defaults move registration enormously. Effective consent for organ donation runs from about 4 percent in Denmark to about 28 percent in the Netherlands where a citizen must opt in, and from about 86 percent in Sweden to about 99.98 percent in Austria where consent is presumed. Automatic enrolment raised participation in a 401(k) plan at a single large employer, and a substantial fraction of those enrolled kept both the contribution rate and the fund allocation the employer had picked. No percentage for that second study appears on this page, because none could be verified against the paper.
Real nudges have real and small effects. Home energy reports across 600,000 households cut consumption by 2.0 percent, which the paper puts as equivalent to a short-run electricity price rise of 11 to 20 percent at favourable cost, and the effect persisted after the reports stopped, decaying at 10 to 20 percent a year. The complete trial record of two nudge units, 126 randomised controlled trials covering 23 million individuals, gives an average of 1.4 percentage points, an 8.0 percent increase, still statistically significant and roughly a sixth of the 8.7 percentage points the published academic papers showed. Cheap and small is the honest shape, and cheap and small is worth having.
The average effect of nudging is unresolved in print. The field's own meta-analysis of more than 200 studies gave a contested Cohen's d = 0.45, 95 percent interval 0.39 to 0.52, alongside its own report of a moderate publication bias in the same abstract; a re-analysis of that same dataset correcting for publication bias estimated 0.04, credible interval 0.00 to 0.14, and, for the category containing defaults, an unresolved 0.12 with an interval of 0.00 to 0.43, where the Bayes factor sits close to 1 and the data therefore cannot tell. Two more objections and the original authors' reply were published in the same journal in the same year. What is established is that the published literature carries substantial publication bias. Where the true average sits is not established by anyone yet.
The magnitude of loss aversion is contested too, and the two disputes are independent of each other. A meta-analysis of 607 estimates from 150 articles puts the mean coefficient at 1.955, 95 percent interval 1.820 to 2.102. A 2025 re-meta-analysis restricted to symmetric designs with no item ordering reports about 1.07, not significantly above 1.0. That is one recent paper against a pooling of all designs, the disagreement is partly about which designs belong in the pool, and neither is treated here as the answer.
The ethics are unresolved because they are not an empirical question. That some arrangement of options is unavoidable is a strong argument and it does not license any particular arrangement, and shaping behaviour by exploiting a known defect in reasoning is different in kind from giving somebody a reason. The reply is equally strong: refusing to design deliberately does not protect anyone, it hands the design to whoever is already doing it commercially. Both halves stand, unresolved, and the useful demand that survives both is accountability: who chose, on what authority, on what published evidence, reversible how.
Whether a personalised, dynamic, individually-computed default is a nudge at all is genuinely open. It may leave the opt-out present in principle and unfindable in practice, which would break the freedom-preserving test from the inside, and whether the result is a large improvement in outcomes or a form of manipulation without precedent is argued and not settled. The tier is our own file's and this page does not upgrade it.
No, switching a country from opt-in to opt-out organ donation is not shown to increase transplants. The consent statistics are real and enormous, and the outcome evidence does not license the inference: a 35-country comparison for 2016 found fewer living donors in opt-out countries, no significant difference in deceased donors, and no significant difference in kidney, non-renal or total solid organ transplantation. This is one cross-sectional comparison at one time point and it does not show harm. It shows no advantage on the outcomes measured.
No, and equally, nudges are not shown to do nothing. The publication-bias correction returned a Bayes factor near 1 for the category containing defaults, which means the data cannot tell rather than that the effect is absent; the complete unpublished record of two nudge units still carries a statistically significant 1.4 percentage points; and the original authors published a reply this page did not read. The defensible position is that the effects are real, much smaller than advertised, very unevenly distributed, and worth having where they are cheap.
No to three specific claims in our own research file, each refused above at the point it would have been made. That neuroimaging shows greater amygdala activation for losses: the paper reports decreasing activity in gain-sensitive regions and its abstract does not mention the amygdala. That losses loom about 2.25 times as large as gains: the meta-analytic mean is 1.955 with a 95 percent interval of 1.820 to 2.102, which excludes it. And that a contested meta-analysis of 212 randomised controlled trials found a median effect of 0.43: the disputed original pooled more than 200 studies and over 450 effect sizes, its statistic is an estimated mean of 0.45 with an interval of 0.39 to 0.52, and the paper's own publication-bias sentence was dropped from the summary.
So what is left is smaller than the promise and more durable than the backlash. The doctrine's central argument was never an empirical one, and nothing in the 2022 exchange touched it: if no arrangement of options is neutral, then somebody is always designing the choice, and the question of who designed it and on what authority does not go away whichever way the effect sizes land. What unavoidability does not settle is whether any particular design is legitimate, and the objection that a nudge working by exploiting a defect in reasoning differs in kind from an argument has never been answered to everyone's satisfaction. The interventions themselves are cheap, and the complete trial record of two nudge units says they do something, at roughly a sixth the size the published papers showed, and very unevenly across people, which the averages conceal. The most valuable thing this literature produced may turn out not to be any single nudge at all, but the audit: a field that opened its own filing cabinets, published the trials that failed, re-analysed its own flagship result against itself, and let the argument stay unfinished in print. And it leaves a reader with the question the doctrine cannot answer from inside: when the arrangement in front of you was chosen by somebody, how would you tell whether it was chosen for you?
Sources & further reading
This article was written from one file in our own research library, T_4_08, Behavioral Economics and Nudge Theory. That file is where the work started; it is not where the work can be checked. The 50 external entries below are where it can be checked: 44 works cited by digital object identifier, one book cited by ISBN and verified against Open Library, one official United Kingdom government page, two items of news reporting, and two official prize records, each named with the section it supports and the claim it carries there. Several things are better said plainly here than left for a reader to discover. First, this page corrects our own file in the open, above, at six places: a loss aversion coefficient of 2.25 that lies outside the field's own meta-analytic interval (§01); a neuroimaging result stated with its direction inverted and a brain region that is not in the paper's abstract (§01); an organ donation percentage pair that appears in no cited source and is attached to donation when the paper measures registered consent (§04); a general claim of attenuation over time, asserted with no citation and contradicted by the best longitudinal study of the one programme the file itself names (§06); a meta-analysis reported as 212 randomised controlled trials with a median of 0.43, when it pooled more than 200 studies and over 450 effect sizes to a mean of 0.45 with an interval, and when its own publication-bias sentence was dropped (§07); and social priming misfiled as behavioural economics, with a replication attribution that could not be verified (§10). Second, two identifier defects in that same file are recorded here rather than in the body, because they are bibliographic rather than substantive: the file cites the book Nudge with a digital object identifier that resolves to Thomas Leonard's five-page review of the book in Constitutional Political Economy, a different work by a different author in a different venue, so the entry below ships the 2008 Yale University Press ISBN instead; and it cites the endowment effect experiments with an identifier that resolves to the 2011 reprint chapter rather than the 1990 Journal of Political Economy original, which the file's own body cites correctly, so the entry below ships the original. Third, what was not read. Most of the literature here was read at the level of a resolved bibliographic record and a published abstract rather than in full. Ten works are named on this page for their existence, their title and their authorship, with no result and no magnitude taken from any of them, and with nothing attributed to any of them beyond what its own title asserts: the two 2022 objections by Szaszi and colleagues and by Bakdash and Marusich, the Mertens reply, Sunstein's Nudges that fail, Bryan, Tipton and Yeager on heterogeneity, the recognition heuristic paper of Goldstein and Gigerenzer, Gigerenzer's 2015 paper on libertarian paternalism, Hertwig and Gruene-Yanoff on boosting, Chater and Loewenstein on the i-frame and the s-frame, and Yeung on hypernudging. Johnson and Goldstein's 2003 paper was not opened either: the publisher returned an access refusal, and its country figures at §04 are corroborated against independent descriptions of the paper's own figure rather than read first-hand, which is why they are given in rounded form. The OECD report's own count of behavioural units was not read, so no such count appears. Fourth, two notes on the list itself. The two official Nobel prize records were refused by the server during the research for this page and were re-checked before publication, when both served normally, so they are listed below and are cited for the prize and its year only. What is deliberately absent, by contrast: Science news articles about the research-integrity cases were identified and are not listed, because the server refused access and their content was never read; the two news items that are listed, from the Harvard Crimson and from Retraction Watch, are the better-supported of them and are used only for what §09 attributes to reporting. Finally, three of the works below were retracted by their publishers and one is a retraction notice. Their presence is deliberate: §09 is about those retractions, and an article about a withdrawn paper has to cite the withdrawn paper.
Image credits
- A German organ donation card lying at a shallow angle on a pale wood surface with a ballpoint pen resting above it Photograph by Mediatrotter, 29 August 2014, via Wikimedia Commons; licensed CC BY-SA 4.0, a ShareAlike licence. CC BY-SA 4.0 Source.
- The bowl of a urinal seen straight on, with a small dark fly figure on the porcelain above the drain cover Photograph by Gustav Broennimann, 20 February 2010, via Wikimedia Commons; licensed CC BY 3.0 CH, the Switzerland port of the Attribution licence. CC BY 3.0 CH Source.
- A painted half length portrait of Herbert Simon, arms folded, worked almost entirely in reds Painting by Richard Rappaport, 1987: a commissioned portrait of Herbert A. Simon in the collection of Carnegie Mellon University, via Wikimedia Commons; dual licensed CC BY 3.0 Unported and GFDL 1.2 or later. CC BY 3.0 Unported / GFDL 1.2 or later Source.
- Card crop of A German organ donation card lying at a shallow angle on a pale wood surface with a ballpoint pen resting above it Photograph by Mediatrotter, 29 August 2014, via Wikimedia Commons; licensed CC BY-SA 4.0, a ShareAlike licence. CC BY-SA 4.0 Source.