A famous experiment was read as proof that too much choice paralyzes us. A meta-analysis put the average effect near zero. Both findings hold, and the reason is that the number of options was never the whole story.
On two consecutive Saturdays, a tasting booth inside Draeger’s, an upscale grocery in Menlo Park, California, offered shoppers either 6 or 24 flavors of jam, switching the display every hour. The larger display drew more people: 60 percent of passing shoppers stopped, against 40 percent for the small one. Then the pattern flipped. Of those who stopped at the six-jam table, nearly 30 percent later used their coupon to buy a jar. Of those who stopped at the 24-jam table, 3 percent did (Iyengar & Lepper, 2000).
Sheena Iyengar and Mark Lepper ran two more studies in the same paper. Students offered 6 essay topics for extra credit were more likely to write the essay than students offered 30 (74 versus 60 percent). Participants choosing one Godiva chocolate from 6 rather than 30 were more satisfied with what they tasted and four times as likely to take chocolates instead of cash as payment (48 versus 12 percent). In all three, more choice was more attractive at first and less motivating in the end.
That finding became a rule: too much choice is bad, so cut options. The rule is not what the paper says. Iyengar and Lepper noted that in “preference-matching contexts,” where people arrive knowing what they want, larger assortments should help, and they deliberately screened out participants with strong prior preferences to give the effect a chance to appear. Their chocolate participants also enjoyed choosing from 30 more than from 6, while finding it more difficult and frustrating. The original result was already conditional. The culture kept the headline and dropped the conditions.
The average that came out near zero
Ten years later, Benjamin Scheibehenne, Rainer Greifeneder and Peter Todd pooled 63 conditions from 50 published and unpublished experiments, 5,036 participants in all, and asked how large the choice-overload effect was on average. The answer was D = 0.02, with a 95 percent confidence interval from −0.09 to 0.12: a mean effect “of virtually zero” (Scheibehenne, Greifeneder & Todd, 2010).
It is tempting to stop there and declare the jam study a fluke. That reading misses the second half of the result. The variance between studies was well above what chance alone would produce; some experiments found strong overload, others found that more options helped. The meta-regression pointed toward conditions rather than absence: people with well-defined prior preferences or expertise did better with more options, and published studies reported somewhat larger overload effects than unpublished ones, with more recent studies less likely to find the effect at all. What the authors could not find was any sufficient condition, a variable that reliably produced overload whenever present. So the honest summary of 2010 is not “choice overload is false.” It is “there is no single, dependable average effect, and the interesting action is in the spread.” An average near zero across heterogeneous studies is consistent with a real effect that appears under some conditions and reverses under others, and the difference between that and “the effect does not exist” is the whole article.
Conditions, and a fight about how to count
Alexander Chernev, Ulf Böckenholt and Joseph Goodman took the next step in 2015 with 99 observations from 53 studies and 7,202 participants. Rather than asking whether overload happens on average, they asked what predicted it, and identified four moderators: the complexity of the choice set (whether one option dominates, how attractive the options are, how they relate to one another), the difficulty of the decision task (time constraints, accountability, how many attributes describe each option), how uncertain the chooser’s preferences are, and the chooser’s goal, such as buying rather than browsing. Each pushed the effect in the expected direction, and once the moderators were in the model, the overall effect of assortment size was significant (Chernev, Böckenholt & Goodman, 2015; see the evidence note on the corrigendum).
That looked like a resolution: the 2010 average was zero because the studies mixed conditions that produce overload with conditions that produce its opposite. Then in 2018 Blakeley McShane and Ulf Böckenholt reanalyzed the same 57 studies from 21 papers with a model built to respect how the data were structured: 172 observations, six dependent measures, four moderators, nested within conditions, groups of participants, studies and papers. Their conclusion was less tidy than either predecessor. Choice overload “varies substantially as a function of the dependent measure and moderator,” with interactions between them: reliable for some combinations, reversed for others, mixed for the rest. Chernev and colleagues had reported that satisfaction, regret, deferral and switching were “equally powerful” measures that “can be used interchangeably.” McShane and Böckenholt found that the measures carried different amounts of variation, and that two-fifths of the possible measure-by-moderator combinations had never been examined at all (McShane & Böckenholt, 2018).
Notice what this does and does not license. It does not show that regret and deferral are separate phenomena; the data are too sparse for that. It shows something methodological and, for a practitioner, more useful: what you measure changes the choice-overload story you get. Whether people buy and how they feel afterward are not necessarily the same thing, and the literature has too often assumed they are.
What is the person actually being asked to do?
If option count alone predicts little, what does? The most consistent thread across the conditional evidence is not a number but a kind of work.
Start with preference. Chernev (2003) found across four experiments that people with an “available ideal point,” a clear sense of the attribute combination they wanted, formed stronger preferences from larger assortments, while people without one formed stronger preferences from smaller ones. The same large set was a resource for one group and a burden for the other. This is the boundary Iyengar and Lepper drew in their discussion, now tested directly. It does not mean that asking users what they want dissolves overload; articulated preferences are something people bring, not something a form reliably extracts.
Then comparison structure. John Gourville and Dilip Soman (2005) distinguished “alignable” assortments, where variants differ along one dimension and choosing means trading more of something against less of it, from “nonalignable” ones, where variants differ across features that cannot be lined up. In their first study, a brand of microwave ovens gained share as it added variants that differed only in capacity and lost share as it added variants that differed in bundles of special features. The count was the same; the comparison the shopper had to perform was not. They tied the loss to effort and the potential for regret, and found that simplifying the presentation, making the choice reversible or reducing the nonalignability weakened it. The idea has an older root: Amos Tversky and Eldar Shafir (1992) showed that adding or improving options can increase deferral and default-taking when the additions create conflict between attractive alternatives. Their studies were not tests of choice overload in the modern sense, but they make the underlying point that difficulty depends on the relations among options, not only their number.
Regret, the mechanism popular accounts lean on hardest, is more interesting than “more options, more counterfactuals.” Yoel Inbar, Simona Botti and Karlene Hanko (2011) proposed that people apply a lay theory, “a quick choice is a bad choice,” when judging their own decisions; large sets under ordinary time pressure feel rushed, and the feeling of having rushed produces regret. In one of their four studies, choosing from 30 chocolates rather than 6 produced more regret only when the experimenter waited in the room. Told to take their time, participants choosing from 30 regretted no more than those choosing from 6; in a separate study, changing the lay theory itself eliminated the effect. This is one regret mechanism among several plausible ones, not the settled account. It matters because it locates part of the cost in how the chooser interprets the process, a different lever from the assortment.
Categorization fits the pattern. Cassie Mogilner, Tamar Rudnick and Sheena Iyengar (2008) found that dividing an assortment into more categories raised satisfaction among choosers unfamiliar with the domain, by increasing how much variety they perceived, and did nothing for familiar choosers who could see the variety without help. Even uninformative labels worked for the unfamiliar. That is a finding about perceived variety and satisfaction in particular populations, not a finding that categories cure overload.
Put these together and the question changes. Instead of “how many options is too many,” ask what the chooser has to do with them: search, compare along dimensions they can hold in mind, form a preference they did not arrive with, commit to a choice they may later judge. A large set can be easy when preferences are clear, differences are alignable and the person is exploring rather than committing. A small set can be hard when attributes conflict, preferences are unformed and the stakes are high. These are inferences from the conditional evidence, not laws.
Friction before comparison
Most of this literature studies people who have already engaged with the options. A field experiment on Alibaba’s Tmall and Taobao marketplaces looked earlier. When a shopper asked a merchant about a product and did not buy within five minutes, an automated message recommended products from that merchant. Xiaoyang Long and colleagues randomized 1.6 million such shoppers to receive one, two, three or four recommendations, with the product they had asked about always first (Long et al., 2025).
Purchase probability rose and then fell: 3.08 percent with one product, 5.13 with two, 4.53 with three, 4.29 with four. What the researchers could observe was whether the shopper clicked on anything at all. Going from two recommendations to three, 64 percent of the drop in purchases was accounted for by fewer shoppers starting a search; from three to four, the figure was 30 percent. The authors interpret this through a model in which anticipated regret about not searching everything, combined with search cost, deters people from beginning. That is their interpretation; the experiment does not adjudicate among mechanisms.
Two boundaries matter. The choice sets were tiny, one to four items, in a retargeting message to people who had already shown interest in a specific product. Nothing here speaks to 24 jams or 500 search results, and “64 percent” describes one transition in one setting, not the share of choice overload anywhere that is caused by search costs. What the study does establish is narrower and still valuable: in this context, friction appeared before comparison, because a share of people never opened the door. A designer who assumes overload happens at the moment of final selection would be looking in the wrong place.
Same change, different verdicts
The last complication is that “does more choice help or hurt” has no single answer even within one business, because it depends on whom you measure and when. Olivia Natan (2025) studied an online restaurant-delivery platform in the Los Angeles area between 2015 and 2018, using the staggered entry of restaurants into different delivery zones to compare households whose assortment grew with neighbors whose did not. This is a difference-in-differences design, not a randomized experiment, and it rests on assumptions Natan spells out. Within those limits, assortment growth increased adoption by new users and reduced how often existing users ordered. The reduction ran through engagement: the elasticity of weekly search with respect to assortment size was about −0.5, while conversion from search to purchase was unaffected. Users who had explored many restaurants were hurt most; users who ordered from one or two were not. The effects were small, and Natan says so.
Read alongside Long, this is a different friction in a similar location: in both cases the cost showed up as not starting, not as choosing badly. Read alongside the meta-analyses, it is the measurement point made concrete. The same expansion looks like growth on an acquisition dashboard and like decay on a frequency one, and neither reading is wrong.
The broader retail evidence agrees. A review of 177 studies from 95 papers published between 1970 and 2021 found that assortment size carries both benefits, including favorable evaluation, confidence, freedom of choice and purchase, and costs, including effort, uncertainty, difficulty and deferral, with an aggregate relationship that was positive but concave (Sethuraman, Gázquez-Abad & Martínez-López, 2022). A 2024 systematic review counted 92 articles on choice overload across 22 years (Jacob, Thomas & Joseph, 2024): the scale of a serious literature, not a debunked one. Large assortments are neither inherently harmful nor inherently helpful. They are trade-offs whose sign depends on what is being measured.
What a practitioner can actually infer
Do not optimize blindly for fewer options. The evidence does not say fewer converts better; it says the cost of choice, when it appears, comes from specific kinds of decision work, and the intervention should match the work.
If people are not starting, as in the Alibaba and delivery-platform studies, the problem is the entry cost of engaging, and the levers are the size and framing of the first step, not the depth of the catalog. If people start but stall, ask whether the options are comparable along dimensions they can hold; Gourville and Soman’s results suggest that simplifying presentation and making choices reversible reduce the cost of nonalignable sets without removing options. If choosers are new to the domain, structure that helps them perceive variety may raise satisfaction, as it did for Mogilner’s unfamiliar choosers, while doing little for experts. If the chooser knows what they want, a larger set is more likely to help than hurt, and trimming it may remove the option they came for. If regret is the concern, the metacognitive evidence points toward pacing rather than pruning.
Every one of these is contingent. None is a rule. Categories, filters, recommendations and smaller sets each help under some conditions and are inert or harmful under others, and the literature does not support presenting any of them as a fix. The defensible move is diagnostic: identify which kind of work this assortment is asking this chooser to do, and whether that work is what your outcome measure is sensitive to.
No magic number
A quarter-century after the jam study, the question “how many options is too many” has no answer, and the research is clear about why. Six was not optimal; 24 was not the threshold; the effect that made the study famous was conditional in the original paper and averaged out across the literature that followed. What survived is more demanding. Choice becomes costly when it asks people to search without a foothold, compare across dimensions they cannot align, form preferences they do not yet have or commit to a decision they expect to second-guess. Whether that cost shows up depends on the chooser, the structure and the outcome you decided to count.
The useful question is not how many options you are offering. It is what you are asking someone to do with them.
References
Chernev, A. (2003). When more is less and less is more: The role of ideal point availability and assortment in consumer choice. Journal of Consumer Research, 30(2), 170–183. https://doi.org/10.1086/376808
Chernev, A., Böckenholt, U., & Goodman, J. (2015). Choice overload: A conceptual review and meta-analysis. Journal of Consumer Psychology, 25(2), 333–358. https://doi.org/10.1016/j.jcps.2014.08.002. Corrigendum (2016): Journal of Consumer Psychology, 26(2), 312. https://doi.org/10.1016/j.jcps.2015.07.001
Gourville, J. T., & Soman, D. (2005). Overchoice and assortment type: When and why variety backfires. Marketing Science, 24(3), 382–395. https://doi.org/10.1287/mksc.1040.0109
Inbar, Y., Botti, S., & Hanko, K. (2011). Decision speed and choice regret: When haste feels like waste. Journal of Experimental Social Psychology, 47(3), 533–540. https://doi.org/10.1016/j.jesp.2011.01.011
Iyengar, S. S., & Lepper, M. R. (2000). When choice is demotivating: Can one desire too much of a good thing? Journal of Personality and Social Psychology, 79(6), 995–1006. https://doi.org/10.1037/0022-3514.79.6.995
Jacob, B. M., Thomas, S., & Joseph, J. (2024). Over two decades of research on choice overload: An overview and research agenda. International Journal of Consumer Studies, 48(2), e13029. https://doi.org/10.1111/ijcs.13029
Long, X., Sun, J., Dai, H., Zhang, D., Zhang, J., Chen, Y., Hu, H., & Zhao, B. (2025). The choice overload effect in online recommender systems. Manufacturing & Service Operations Management, 27(1), 249–268. https://doi.org/10.1287/msom.2022.0659
McShane, B. B., & Böckenholt, U. (2018). Multilevel multivariate meta-analysis with application to choice overload. Psychometrika, 83(1), 255–271. https://doi.org/10.1007/s11336-017-9571-z
Mogilner, C., Rudnick, T., & Iyengar, S. S. (2008). The mere categorization effect: How the presence of categories increases choosers’ perceptions of assortment variety and outcome satisfaction. Journal of Consumer Research, 35(2), 202–215. https://doi.org/10.1086/588698
Natan, O. R. (2025). Choice frictions in large assortments. Marketing Science, 44(3), 593–625. https://doi.org/10.1287/mksc.2023.0415
Scheibehenne, B., Greifeneder, R., & Todd, P. M. (2010). Can there ever be too many options? A meta-analytic review of choice overload. Journal of Consumer Research, 37(3), 409–425. https://doi.org/10.1086/651235
Sethuraman, R., Gázquez-Abad, J. C., & Martínez-López, F. J. (2022). The effect of retail assortment size on perceptions, choice, and sales: Review and research directions. Journal of Retailing, 98(1), 24–45. https://doi.org/10.1016/j.jretai.2022.01.001
Tversky, A., & Shafir, E. (1992). Choice under conflict: The dynamics of deferred decision. Psychological Science, 3(6), 358–361. https://doi.org/10.1111/j.1467-9280.1992.tb00047.x
Evidence note
The jam-study purchase figures (30 versus 3 percent) are shares of the 249 shoppers who stopped at the booth, not of everyone who passed it. The chocolate figures (48 versus 12 percent) are shares of participants in each choice condition who took chocolates rather than cash.
Chernev, Böckenholt and Goodman’s 2015 paper carries a corrigendum. Three statistics on page 348 were mislabeled: values reported as t statistics were regression coefficients. The corrected record reads b = .17, t(39) = 4.5, p < .001 for the overall effect across studies of choice from a given assortment; b = .04, t(39) = .9, p > .4 for studies with moderators; and b = .41, t(39) = 5.3, p < .001 for studies without. The correction changes the reporting, not the direction of any conclusion drawn here.
McShane and Böckenholt (2018) reanalyzed the study set assembled by Chernev and colleagues rather than collecting new studies; their 172 observations come from the same 57 studies and 21 papers.
In Long et al. (2025), the 64 percent figure refers specifically to the drop in purchase probability between two and three recommendations. For the drop between three and four, the corresponding share was 30 percent. Both are decompositions within one experiment, not general estimates.
Natan (2025) identifies effects from staggered restaurant entry using a difference-in-differences design with household and time fixed effects. The author notes that if restaurants draw delivery zones strategically, the estimates would be upper bounds, and describes the estimated sensitivity of purchase frequency to assortment size as very small for the average consumer.
Independent writing sample. Self-initiated and not commissioned by or affiliated with any client.