A familiar bias is not automatically a robust scientific finding, since some prominent effects weaken or fail under replication. A cognitive bias evidence map distinguishes well-supported results from contested ones, making the strength of the underlying evidence part of the claim.
Cognitive biases are popular - they fill listicles, business books, and slide decks. But not all of them are equally well-supported, and some famous effects have weakened or failed when researchers tried to replicate them.
This is an evidence map: it rates the best-known biases and decision effects by how strong and replicable the evidence is, with each entry linked to a primary source.
The goal is to let you tell the difference between a robust, well-replicated finding you can build on and a famous-but-shaky effect you should cite with caution. Where the research is strong, we say so.
Where it is contested or has failed to replicate, we say that plainly - because a bias map that ignores the replication crisis is itself misleading.
Why replication status matters here
Replication status matters because psychology went through a "replication crisis" in the 2010s in which large coordinated efforts found that a substantial share of published findings, including well-known cognitive biases, did not reproduce at their original effect sizes. A bias map that ignores this and lists every effect as equally solid is itself misleading, so this map flags which effects survived scrutiny and which did not.
Psychology went through a "replication crisis" in the 2010s: large coordinated efforts found that a substantial share of published findings did not reproduce at their original effect sizes.
The Open Science Collaboration's 2015 attempt to replicate 100 psychology studies successfully reproduced well under half at full strength. Decision and social psychology were among the hardest hit.
So a responsible cognitive-bias reference can't just list effects - it has to flag which ones survived scrutiny. That is what this map does.
"Robust" means the effect replicates widely and is broadly accepted; "supported with caveats" means real but bounded or sensitive to conditions; "contested / weak replication" means the original claim has been seriously challenged.
The evidence map
In one line: framing, sunk cost, and classic-number anchoring have survived large replication tests; confirmation bias, overconfidence, loss aversion, availability, and the planning fallacy are real but narrower than the popular version; ego depletion, incidental anchoring, and dramatic social priming have failed retesting and should not be treated as established.
| Bias / effect | What it claims | Evidence status | Honest caveat | Key source |
|---|---|---|---|---|
| Loss aversion | Losses loom larger than equivalent gains in decision-making. | Robust (but debated in scope) | Well-replicated as a phenomenon; researchers debate whether it is as universal/large as once claimed, and it can attenuate in some contexts. | Kahneman & Tversky (1979) |
| Framing effects | The same choice, framed as a gain vs. a loss, changes preferences. | Robust | Widely replicated across domains; magnitude varies with how the framing is constructed. | Tversky & Kahneman (1981) |
| Anchoring (classic estimation) | An initial number disproportionately influences a subsequent numerical estimate on the same topic. | Robust | One of the more reliably replicated effects; size depends on relevance/plausibility of the anchor. | Tversky & Kahneman (1974) |
| Anchoring (incidental / irrelevant anchor) | A number with no logical connection to the estimate (a dice roll, the last digits of an ID) still shifts the estimate. | Contested / weak replication | The classic and incidental versions are not the same claim; see the replication ledger below for why they are split here. | Critcher & Gilovich (2008); tested in Many Labs 2 |
| Availability heuristic | We judge probability by how easily examples come to mind. | Supported with caveats | The original letter-frequency demonstration has a documented replication challenge; the broader "ease of recall shapes judgment" mechanism is better supported than the specific classic demo. | Tversky & Kahneman (1973) |
| Confirmation bias | We seek, interpret, and recall information that confirms prior beliefs. | Supported | Best quantified by meta-analysis rather than a single replication project; effect is moderate, not the sweeping force pop-psychology implies. It is also an umbrella term covering several distinct paradigms. | Nickerson (1998); Hart et al. (2009) meta-analysis |
| Overconfidence / miscalibration | People are more confident in their judgments than accuracy warrants. | Supported with caveats | Reliable in calibration studies, but "overconfidence" is several distinct phenomena (overestimation, overplacement, overprecision), and part of the classic hard-easy pattern has been argued to be a scoring artifact. | Moore & Healy (2008) |
| Planning fallacy | We underestimate the time/cost of our own projects despite past evidence. | Supported, evidence base thin | The core finding is well-known and intuitive, but rests on the original study rather than a large-scale replication we could find. Untested-at-scale is not the same as contested. | Buehler, Griffin & Ross (1994) |
| Sunk cost fallacy | Past, unrecoverable investment irrationally drives future decisions. | Supported | Real, and one of the ten Many Labs 1 effects that replicated; size and moderators still vary by decision type. | Arkes & Blumer (1985) |
| Ego depletion | Self-control is a finite resource that "depletes" with use. | Contested / failed replication | A 23-lab pre-registered replication found an effect indistinguishable from zero. Treat the strong "willpower as fuel" claim as unsupported, despite an older meta-analysis of roughly 200 studies once suggesting otherwise. | Hagger et al. (2016), Registered Replication Report |
| Social/behavioral priming (e.g. "elderly words slow walking") | Subtle cues unconsciously and strongly shape behavior. | Contested / failed replication | The flagship elderly-priming and several other behavioral-priming demonstrations failed with larger samples and stricter controls. Treat dramatic priming claims with strong skepticism. | Doyen et al. (2012); failures in Many Labs 1 and 2 |
The replication ledger: evidence base x outcome
A single "robust / contested" label hides an important distinction: an effect that was tested at large scale and failed is not the same situation as one that was never tested at large scale at all. The table above already reflects this, but the ledger below makes the underlying replication record explicit and sourced, so the rating isn't just our word for it. This is original synthesis: the ratings are ours, computed from the cited studies, not copied from any single source.
In one line: the biggest surprise in the underlying record is ego depletion, tested at large scale and clearly failed, versus the planning fallacy, never tested at large scale at all. Both can sound "shaky" from a one-word label, but they are not the same kind of shaky.
| Effect | Evidence base | Documented outcome |
|---|---|---|
| Framing effects | Multi-lab: Many Labs 1 (2014) and Many Labs 2 (2018), 36+ samples | Replicated in both; Many Labs 2 found the effect about half as strong as the original |
| Anchoring, classic estimation | Multi-lab: Many Labs 1 (2014) | Replicated strongly; among the largest effect sizes (d > 1.0) in the whole project |
| Anchoring, incidental/irrelevant | Multi-lab: Many Labs 2 (2018) plus several independent studies | Failed in Many Labs 2 (d = 0.04, indistinguishable from zero); mixed across independent attempts, some succeeded, several failed |
| Sunk cost fallacy | Multi-lab: Many Labs 1 (2014); meta-analysis of 98 effect sizes (2014) | Replicated in Many Labs 1; meta-analysis confirms the effect, contingent on decision type |
| Ego depletion | Multi-lab: Hagger et al. (2016), 23 labs, n = 2,141, pre-registered | Failed: d = 0.04, 95% CI [-0.07, 0.15], crossing zero |
| Social/behavioral priming (elderly words) | Independent replications: Doyen et al. (2012) and others, larger samples with objective timing | Failed to replicate the original effect; original study also linked to experimenter-expectancy effects |
| Confirmation bias | Meta-analysis: Hart et al. (2009), Psychological Bulletin | Moderate, real effect: d = 0.36 preference for congenial over uncongenial information |
| Loss aversion | Independent replications: Zeif & Yechiam (2023), ~2,001 participants; critical review by Gal & Rucker (2018) | Did not appear for small stakes, weak at moderate stakes, present but smaller than historical estimates at larger stakes |
| Availability heuristic | Single independent challenge: Sedlmeier, Hertwig & Gigerenzer (1998) on the original letter-frequency demo | Reported failure to replicate the classic demonstration specifically; the broader mechanism was not the target |
| Overconfidence / miscalibration | Methodological critique: Juslin, Winman & Olsson (2000) and others on the hard-easy effect | Argue part of the classic pattern is a scale-end/artifact issue rather than a pure cognitive bias; not a replication failure of overconfidence broadly |
| Planning fallacy | Single foundational study only | No large-scale, multi-lab, or pre-registered replication attempt found; absence of testing, not evidence of failure |
Two rows were deliberately left off the confidence claims above: we could not verify Google Scholar citation counts for the founding papers well enough to publish a number, and none of these ten effects were part of the original, randomly sampled Open Science Collaboration (2015) 100-study project, so that project is cited here for context on the replication crisis generally, not as direct evidence for any single row.
The search transparency gap: an original measurement
The replication ledger above tells you what the scientific record actually says. It does not tell you what a typical reader actually encounters when they search for these terms. So we measured that directly: for each effect, we pulled the top organic web search results and checked, page by page, whether the result disclosed any replication concern, contested status, or evidence-strength caveat at all, anywhere on the page, even a single sentence.
This is original data. It is not copied from any source; it is our own one-time measurement, dated below, and the methodology is disclosed so it can be checked or repeated.
In one line: nearly every top result for ego depletion admits it failed replication, while almost none of the top results for overconfidence, the planning fallacy, or the sunk cost fallacy admit anything about the strength of the underlying evidence. Most biases sit somewhere in between, closer to silent than transparent.
| Effect | Results checked | Disclose replication concerns |
|---|---|---|
| Overconfidence / miscalibration | 6 | 0% (0 of 6) |
| Planning fallacy | 12 | 0% (0 of 12) |
| Sunk cost fallacy | 10 | 0% (0 of 10) |
| Anchoring, incidental/irrelevant | 9 | 11.1% (1 of 9) |
| Availability heuristic | 12 | 16.7% (2 of 12) |
| Anchoring, classic estimation | 11 | 18.2% (2 of 11) |
| Framing effect | 11 | 18.2% (2 of 11) |
| Confirmation bias | 9 | 33.3% (3 of 9) |
| Loss aversion | 13 | 46.2% (6 of 13) |
| Social/behavioral priming | 7 | 57.1% (4 of 7) |
| Ego depletion | 13 | 100% (13 of 13) |
Citable finding: "Every usable top search result for ego depletion (13 of 13) discloses its 23-lab failed replication, while none of the top search results for the planning fallacy, overconfidence, or the sunk cost fallacy (0 of 28 combined) disclose the actual strength of their evidence base." Source: When Notes Fly, Search Transparency Gap measurement, August 2026.
Three results stand out. Ego depletion is the one effect where disclosure is near-universal: every usable top result we checked, including Wikipedia and every general-audience explainer site, mentioned the 23-lab failed replication. This is likely because the failure is so large (d = 0.04, a null result) and well-covered that it has become part of the effect's own story. Overconfidence, the planning fallacy, and the sunk cost fallacy sit at the opposite extreme, disclosed in none of the results we checked for any of the three. Per the ledger above, these three are not the same kind of gap: the planning fallacy has genuinely never been tested at large scale, overconfidence has a live methodological critique of its classic demonstration, and sunk cost actually replicated in Many Labs 1, so its 0% is a case of real, positive evidence going uncited, not a hidden weakness.
A reader searching any of these terms today would have no way to tell those three situations apart.
The broader pattern: outside of ego depletion, disclosure of replication concerns in the top search results is the exception, not the rule: under 20% for six of the eleven effects checked, and never above 60% for any effect except ego depletion.
Methodology for this measurement
- What we searched: the plain common name of each effect (e.g. "ego depletion willpower psychology," "anchoring effect cognitive bias"), with a second rephrased query per effect where needed to build a larger result pool, general web search.
- What counted as a result: the top 15 organic (non-ad, non-video) results per query that were real general-audience articles, encyclopedia entries, or reference pages. Pages that could not be fetched or parsed were excluded from the denominator rather than counted as non-disclosing (marked "unreadable" and dropped), so the percentages reflect usable pages only, not all 15 slots. Sample sizes actually checked ranged from 6 to 13 usable pages per effect.
- What counted as disclosure: any sentence anywhere on the page noting a replication failure, contested status, or a caveat that the effect is weaker, narrower, or less settled than the popular framing implies. A single disclosing sentence on an otherwise uncritical page still counted as disclosed, which is a low bar, so the low percentages above likely understate, not overstate, the real transparency gap.
- What this does not claim: this is one search snapshot, not a tracked index; rankings and page content change, and a different query phrasing would return a different set of pages. We are not claiming these are the only or the definitive top results, only what we observed for these specific queries on the date below. Sample sizes of 6 to 13 per effect are still small enough that a single-digit swing in disclosed count meaningfully moves the percentage; treat these as directional, not precise to the decimal.
- This measurement was revised once already: an initial pass at 4 to 9 results per effect found sunk cost fallacy at 28.6% disclosure; rerunning at a larger sample (10 usable results) found 0%. We are showing the larger, more recent sample rather than averaging the two, and flag this specifically because it is exactly the kind of small-sample instability the rest of this page argues readers should watch for.
- When this was measured: August 2026. We plan to re-run this measurement periodically and will note the date of the most recent run here.
How to use this responsibly
- Robust effects (framing, classic estimation anchoring, sunk cost) replicated in large multi-lab projects and are safe to teach and design around, while remembering effect sizes are conditional.
- Supported / supported-with-caveats effects (confirmation bias, overconfidence, availability, loss aversion, planning fallacy) are real but should be described with their boundaries, not as iron laws - and in the planning fallacy's case, with the caveat that nobody appears to have run the large-scale replication yet.
- Contested effects (ego depletion, dramatic social priming, incidental/irrelevant anchoring) should be cited only with the replication failure noted - or not used as load-bearing evidence at all.
- The meta-lesson: the existence of a bias does not mean any specific intervention reliably "debiases" it. Debiasing evidence is its own, often weaker, literature.
Methodology and scope notes
- Evidence basis: foundational papers for each effect, cross-checked against large replication efforts (Many Labs 1 and 2, Registered Replication Reports) and meta-analyses where one exists. The replication ledger records, for each effect, what kind of evidence exists (single study, meta-analysis, or multi-lab replication) and what it found - both axes are reported so a thin evidence base is never disguised as a failed one, or vice versa.
- What we did not do: we did not run new experiments and we did not estimate any number we could not trace to a specific cited study. Where a claim about a study (a citation count, a meta-analysis' study count) could not be independently confirmed, it was cut rather than published as a guess.
- The Open Science Collaboration (2015) 100-study project is cited for context on the replication crisis, not as direct evidence for any single row - none of these ten effects were part of that randomly sampled project.
- "Robust" is not "unlimited." Even well-replicated biases have moderators, cultural variation, and context limits. Treat ratings as evidence strength, not universal magnitude.
- Not individual advice. This summarises general findings about typical decision-makers; it is not personalised, clinical, financial, or legal guidance.
- Maintenance: updated when major new replication evidence changes a rating, not on a fixed schedule.
Sources checked: the raw audit
For transparency, every page we checked for the search transparency gap measurement above is listed below, one collapsible section per row, with our verdict and the exact disclosing sentence where one was found. This is the full working, not a summary of it.
Overconfidence / miscalibration: 0 of 6 disclosed
- helpfulprofessor.com: excluded, page unreadable
- masterclass.com: excluded, page unreadable
- profrjstarr.com: not disclosed
- newristics.com: excluded, page unreadable
- en.wikipedia.org: not disclosed
- scribbr.com: excluded, page unreadable
- ethicsunwrapped.utexas.edu: not disclosed
- vaia.com: excluded, page unreadable
- examples.yourdictionary.com: not disclosed
- psychologytoday.com: not disclosed
- sciencedirect.com: excluded, page unreadable
- fiveable.me: excluded, page unreadable
- developdiverse.com: not disclosed
Planning fallacy: 0 of 12 disclosed
- en.wikipedia.org: not disclosed
- thedecisionlab.com: not disclosed
- spsp.org: not disclosed
- monday.com: not disclosed
- ohai.ai: not disclosed
- memtime.com: not disclosed
- calendar.com: not disclosed
- nesslabs.com: not disclosed
- qz.com: excluded, page unreadable
- psychologytoday.com: not disclosed
- profrjstarr.com: not disclosed
- everydaypsych.com: not disclosed
- michelleporterfit.com: not disclosed
- scribbr.com: excluded, page unreadable
- pwc.pl: excluded, page unreadable
Sunk cost fallacy: 0 of 10 disclosed
- positivepsychology.com: not disclosed
- asana.com: not disclosed
- training.nih.gov: excluded, page unreadable
- scribbr.com: excluded, page unreadable
- behavioraleconomics.com: excluded, page unreadable
- thedecisionlab.com: not disclosed
- wallstreetprep.com: not disclosed
- schwab.com: excluded, page unreadable
- cognitivebiaslab.com: not disclosed
- simple.wikipedia.org: not disclosed
- bachelorprint.com: not disclosed
- vaia.com: not disclosed
- wiobyrne.com: excluded, page unreadable
- wallstreetmojo.com: not disclosed
- facet.com: not disclosed
Anchoring, incidental/irrelevant: 1 of 9 disclosed
- scribbr.com: excluded, page unreadable
- masterclass.com: excluded, page unreadable
- medium.com: not disclosed
- cogn-iq.org: excluded, page unreadable
- nudgingfinancialbehaviour.com: not disclosed
- decodethefuture.org: excluded, page unreadable
- simplypsychology.org: disclosed. "The well-powered replication did not reproduce the large original effect. Instead of a 31% shift, the anchor produced only about a 3.4% increase."
- psychologytoday.com: not disclosed
- nelson.edu: not disclosed
- betterup.com: not disclosed
- explorepsychology.com: not disclosed
- abtasty.com: not disclosed
- yukaichou.com: not disclosed
- courses.eller.arizona.edu: excluded, page unreadable
Availability heuristic: 2 of 12 disclosed
- en.wikipedia.org: disclosed. "Some researchers have claimed that the classic studies on the availability heuristic are too vague in that they fail to account for people's underlying mental processes."
- simplypsychology.org: not disclosed
- scribbr.com: excluded, page unreadable
- thedecisionlab.com: not disclosed
- masterclass.com: excluded, page unreadable
- researcher.life: not disclosed
- forbes.com: not disclosed
- betterup.com: not disclosed
- quillbot.com: not disclosed
- cognitivebiaslab.com: not disclosed
- yukaichou.com: not disclosed
- sciencedirect.com: excluded, page unreadable
- dovetail.com: not disclosed
- vaia.com: not disclosed
- openjdm.github.io: disclosed. "Schwarz et al demonstrate a context where experienced ease of recall, not availability, drives intuitive estimates of frequency, critiquing and refining the Kahneman and Tversky findings."
Anchoring, classic estimation: 2 of 11 disclosed
- en.wikipedia.org: disclosed. "A small number of studies used procedures that were clearly random, such as Excel random generator button and die roll, and failed to replicate anchoring effects."
- simplypsychology.org: disclosed. "Well-powered replications suggest anchoring's direction is robust, but its size is often smaller than early, underpowered studies reported."
- scribbr.com: excluded, page unreadable
- psychcentral.com: not disclosed
- cogn-iq.org: excluded, page unreadable
- ebsco.com: not disclosed
- sciencedirect.com: excluded, page unreadable
- sciencedaily.com: not disclosed
- berkeleywellbeing.com: not disclosed
- mindmax.me: not disclosed
- stlouisfed.org: excluded, page unreadable
- pon.harvard.edu: not disclosed
- pon.harvard.edu: not disclosed
- yukaichou.com: not disclosed
- michaelgearon.medium.com: not disclosed
Framing effect: 2 of 11 disclosed
- simplypsychology.org: disclosed. "Goal framing is the smallest and least reliable effect."
- forbes.com: not disclosed
- en.wikipedia.org: not disclosed
- ebsco.com: not disclosed
- study.com: excluded, page unreadable
- hanshow.com: not disclosed
- helpfulprofessor.com: excluded, page unreadable
- scribbr.com: excluded, page unreadable
- thedecisionlab.com: not disclosed
- nesslabs.com: not disclosed
- vaia.com: not disclosed
- sciencedirect.com: excluded, page unreadable
- boycewire.com: not disclosed
- simplypsychology.org: not disclosed
- en.wikipedia.org: disclosed. "Multiple studies have questioned the existence of loss aversion. In several studies examining the effect of losses in decision-making, no loss aversion was found."
Confirmation bias: 3 of 9 disclosed
- en.wikipedia.org: disclosed. "Subsequent research has since failed to replicate findings supporting the backfire effect."
- britannica.com: excluded, page unreadable
- dictionary.apa.org: excluded, page unreadable
- simplypsychology.org: disclosed. "More recent, larger studies suggest the backfire effect is rarer and less robust than early research implied."
- scribbr.com: excluded, page unreadable
- thedecisionlab.com: not disclosed
- sciencedaily.com: not disclosed
- ebsco.com: disclosed. "his original studies have been subject to various critiques from later researchers"
- study.com: excluded, page unreadable
- helpfulprofessor.com: excluded, page unreadable
- usemultiplier.com: excluded, page unreadable
- bayareacbtcenter.com: not disclosed
- subjectguides.lib.neu.edu: not disclosed
- theuncertaintyproject.org: not disclosed
- psychologynoteshq.com: not disclosed
Loss aversion: 6 of 13 disclosed
- en.wikipedia.org: disclosed. "Multiple studies have questioned the existence of loss aversion. In several studies examining the effect of losses in decision-making, no loss aversion was found under risk and uncertainty."
- psychologytoday.com: not disclosed
- thedecisionlab.com: not disclosed
- ebsco.com: disclosed. "Losses can have a greater psychological impact than equivalent gains, although the strength of this effect varies across people, situations, and study designs."
- masterclass.com: excluded, page unreadable
- behavioraleconomics.com: disclosed. "Some researchers have questioned the robustness or even existence of loss aversion (Gal & Rucker, 2018)."
- tutor2u.net: not disclosed
- economicshelp.org: not disclosed
- insidebe.com: disclosed. "Evidence seems to consistently show that loss aversion does not occur in all situations, and not in the same way in different situations."
- nngroup.com: not disclosed
- yukaichou.com: disclosed. "Loss aversion still exists; it is just smaller and more context-dependent than the 2.0 figure implies."
- fiveable.me: excluded, page unreadable
- wallstreetprep.com: not disclosed
- psychologyfanatic.com: disclosed. "their sensitivity drops significantly or even disappears once they've experienced those gains or losses"
- simplypsychology.org: not disclosed
Social/behavioral priming: 4 of 7 disclosed
- en.wikipedia.org: disclosed. "The result is that the efficacy of priming may have been greatly overstated in earlier literature, or have been entirely illusory."
- ebsco.com: not disclosed
- en.wikipedia.org: excluded, off-topic
- study.com: excluded, page unreadable
- en.wikipedia.org: not disclosed
- psypost.org: excluded, page unreadable
- sciencedirect.com: excluded, page unreadable
- betterhelp.com: disclosed. "Experts have raised concerns about replication failures, selective reporting, and other problematic practices in areas of priming research."
- psychologytoday.com: disclosed. "Follow-up tests of a number of these "social priming" effects have cast doubt on whether they are genuine."
- en.wikipedia.org: excluded, off-topic
- mytherapist.com: disclosed. "The results of some studies on priming were not able to be replicated or verified in follow-up or similar studies."
- zimbardo.com: excluded, page unreadable
- communicationtheory.org: not disclosed
- helpfulprofessor.com: excluded, page unreadable
- bravethinkinginstitute.com: excluded, page unreadable
Ego depletion: 13 of 13 disclosed
- en.wikipedia.org: disclosed. "In 2016, a major multi-lab replication study carried out at two dozen labs across the world using a single protocol failed to find any evidence for ego depletion."
- simplypsychology.org: disclosed. "There have both been studies to support and question the validity of ego depletion as a theory."
- thedecisionlab.com: disclosed. "It recently came under fire when an attempt to replicate Baumeister's landmark study did not yield results in support of ego depletion."
- psychologytoday.com: disclosed. "The first published failure to replicate the effect came out in 2004."
- sciencedirect.com: excluded, page unreadable
- study.com: excluded, page unreadable
- betterhelp.com: disclosed. "The studies that supported ego depletion may have been flawed."
- reachlink.com: disclosed. "Twenty-three laboratories across the world tested 2,141 participants using the same standardized protocol. The result was an effect size of just d=0.04."
- scienceinsights.org: disclosed. "A major coordinated effort involving 23 laboratories and over 2,100 participants failed to replicate the effect."
- nirandfar.com: disclosed. "Plenty of new research has found that willpower is actually not "used up" like gas in a gas tank or charge in a battery."
- psychologyfanatic.com: disclosed. "Ego-depletion and the strength model recently have come under fire. There is scientific backing to some of the opposition."
- teuxdeux.com: disclosed. "Studies by other orgs and researchers have either found no ego depletion effect or even gone as far as claiming publication bias"
- psychologs.com: disclosed. "This perspective is controversial in modern psychology, as motivation plays a significant role."
- stillmindflorida.com: disclosed. "A large-scale replication effort published in Psychological Science found limited support for the theory, sparking discussions about its reliability and broader applicability."
- speakandregret.michaelinzlicht.com: disclosed. "24 labs from around the world and tested over 2,000 participants, all to replicate a previously published ego depletion study...And then the results came in: no effect."
Sources
- Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124-1131. DOI: 10.1126/science.185.4157.1124
- Kahneman, D., & Tversky, A. (1979). Prospect theory: An analysis of decision under risk. Econometrica, 47(2), 263-291. DOI: 10.2307/1914185
- Tversky, A., & Kahneman, D. (1981). The framing of decisions and the psychology of choice. Science, 211(4481), 453-458. DOI: 10.1126/science.7455683
- Tversky, A., & Kahneman, D. (1973). Availability: A heuristic for judging frequency and probability. Cognitive Psychology, 5(2), 207-232. DOI: 10.1016/0010-0285(73)90033-9
- Nickerson, R. S. (1998). Confirmation bias: A ubiquitous phenomenon in many guises. Review of General Psychology, 2(2), 175-220. DOI: 10.1037/1089-2680.2.2.175
- Moore, D. A., & Healy, P. J. (2008). The trouble with overconfidence. Psychological Review, 115(2), 502-517. DOI: 10.1037/0033-295X.115.2.502
- Buehler, R., Griffin, D., & Ross, M. (1994). Exploring the "planning fallacy." Journal of Personality and Social Psychology, 67(3), 366-381. DOI: 10.1037/0022-3514.67.3.366
- Arkes, H. R., & Blumer, C. (1985). The psychology of sunk cost. Organizational Behavior and Human Decision Processes, 35(1), 124-140. DOI: 10.1016/0749-5978(85)90049-4
- Hagger, M. S., et al. (2016). A multilab preregistered replication of the ego-depletion effect. Perspectives on Psychological Science, 11(4), 546-573. DOI: 10.1177/1745691616652873
- Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. DOI: 10.1126/science.aac4716
- Klein, R. A., et al. (2014). Investigating variation in replicability: A "many labs" replication project. Social Psychology, 45(3), 142-152. DOI: 10.1027/1864-9335/a000178
- Klein, R. A., et al. (2018). Many Labs 2: Investigating variation in replicability across samples and settings. Advances in Methods and Practices in Psychological Science, 1(4), 443-490. DOI: 10.1177/2515245918810225
- Hart, W., Albarracín, D., Eagly, A. H., Brechan, I., Lindberg, M. J., & Merrill, L. (2009). Feeling validated versus being correct: A meta-analysis of selective exposure to information. Psychological Bulletin, 135(4), 555-588. DOI: 10.1037/a0015701
- Zeif, N., & Yechiam, E. (2023). Loss aversion simply does not materialize for smaller losses. Judgment and Decision Making, 18, e19. DOI: 10.1017/jdm.2023.19
- Gal, D., & Rucker, D. D. (2018). The loss of loss aversion: Will it loom larger than its gain? Journal of Consumer Psychology, 28(3), 497-516. DOI: 10.1002/jcpy.1047
- Doyen, S., Klein, O., Pichon, C.-L., & Cleeremans, A. (2012). Behavioral priming: It's all in the mind, but whose mind? PLOS ONE, 7(1), e29081. DOI: 10.1371/journal.pone.0029081
- Sedlmeier, P., Hertwig, R., & Gigerenzer, G. (1998). Are judgments of the positional frequencies of letters systematically biased due to availability? Journal of Experimental Psychology: Learning, Memory, and Cognition, 24(3), 754-770. DOI: 10.1037/0278-7393.24.3.754
- Juslin, P., Winman, A., & Olsson, H. (2000). Naive empiricism and dogmatism in confidence research: A critical examination of the hard-easy effect. Psychological Review, 107(2), 384-396. DOI: 10.1037/0033-295X.107.2.384
Related reading on When Notes Fly
- Better Decisions Under Uncertainty - the decision-making guide this map supports
- Probabilistic Thinking for Better Decisions
- Common Decision Traps
- Learning Science Evidence Map - the companion evidence map
Frequently Asked Questions
Which cognitive biases are best supported by evidence?
Framing effects and classic-estimation anchoring are the most robust, replicated in large multi-lab projects like Many Labs 1 and 2. Confirmation bias, overconfidence, the availability heuristic, and loss aversion are real and well-documented but supported with caveats: effect sizes are smaller or narrower than the popular version implies. Sunk cost is supported. Planning fallacy is real but rests on a thin evidence base, since no large-scale replication of it appears to exist. Incidental or irrelevant anchoring (the classic ‘random number example’) is contested and failed a large multi-lab replication.
Which famous biases failed to replicate?
Ego depletion (the idea that willpower is a finite fuel) showed little to no effect in a large preregistered multi-lab replication of 23 labs and over 2,100 participants. Incidental or irrelevant anchoring (an unrelated number, like a dice roll, shifting an estimate) also failed a large multi-lab replication. Several dramatic social/behavioral priming results (e.g. ‘elderly words slow walking’) have also failed to replicate. These should be cited only with the replication failure noted.
What is the replication crisis and why does it matter for bias claims?
In the 2010s, large efforts found many published psychology findings did not reproduce at their original strength - the Open Science Collaboration (2015) replicated well under half of 100 studies at full effect. Decision and social psychology were heavily affected, so any responsible cognitive-bias reference must flag which effects survived scrutiny.
Does knowing about a bias remove it?
Not reliably. The existence of a bias does not mean any particular ‘debiasing’ intervention works - debiasing is its own, often weaker, literature. Structural countermeasures (e.g. reference-class forecasting for the planning fallacy) tend to beat simply being aware of the bias.
How often do popular explanations of these biases mention replication problems?
We measured it directly by checking the top search results for each bias. Disclosure varies enormously: every usable top result for ego depletion mentioned its failed replication, while none of the top results for the planning fallacy mentioned that its evidence base is a single study with no large-scale replication. Most biases fall well under 50% disclosure, meaning most popular explanations of most biases say nothing about how strong the underlying evidence actually is.