When AI outperforms humans, who bears the burden of proof?
Abstract
If an AI system reliably outperforms human judgment by the standards we actually care about, when does insisting on human control become ethically questionable? We argue that many debates about AI governance begin from an unexamined default: treating human decision-making as the natural and morally superior baseline, independent of comparative performance. We call this the human gold standard fallacy and trace it to four recurring drivers (i.e., omission bias, agent-relative ethics, intuitive exceptionalism, and the control premium) which make status-quo human authority feel better even when it performs worse by our stated standards. We do not defend blind automation. Rights- and process-based constraints (e.g., contestability, nondiscrimination, recourse, and accountable oversight) can legitimately limit any decision system, human or machine. Nonetheless, our claim is about symmetry: when audited evidence shows AI outperforming humans on the criteria we endorse and can satisfy core procedural constraints, the justificatory burden shifts. As such, asking only whether AI is safe enough to deploy leaves the harder question unasked: whether human control is safe enough to keep.
Data availability
No datasets were generated or analysed during the current study.
References
Bansak K, Paulson E (2024) Public attitudes on performance for algorithmic and human decision-makers. PNAS Nexus. https://doi.org/10.1093/pnasnexus/pgae520
Baron J, Ritov I (2004) Omission bias, individual differences, and normality. Organ Behav Hum Decis Process 94:74–85. https://doi.org/10.1016/j.obhdp.2004.03.003
Baumard N, Hyafil A, Morris I, Boyer P (2015) Increased affluence explains the emergence of ascetic wisdoms and moralizing religions. Curr Biol 25:10–15. https://doi.org/10.1016/j.cub.2014.10.063
Baumeister RF, André N, Southwick DA, Tice DM (2024) Self-control and limited willpower: current status of ego depletion theory and research. Curr Opin Psychol 60:101882. https://doi.org/10.1016/j.copsyc.2024.101882
Bigman YE, Gray K (2018) People are averse to machines making moral decisions. Cognition 181:21–34. https://doi.org/10.1016/j.cognition.2018.08.003
Bobadilla-Suarez S, Sunstein CR, Sharot T (2017) The intrinsic value of choice: the propensity to under-delegate in the face of potential gains and losses. J Risk Uncertain 54:187–202. https://doi.org/10.1007/s11166-017-9259-x
Bortolotti L (2018) Stranger than fiction: costs and benefits of everyday confabulation. Rev Philos Psychol 9:227–249. https://doi.org/10.1007/s13164-017-0367-y
Boyer P (2001) Religion explained: the human instincts that fashion gods, spirits and ancestors. Heinemann, London
Brehm JW (1966) A theory of psychological reactance. Academic Press, New York
Brynjolfsson E, Li D, Raymond L (2025) Generative AI at work. Q J Econ 140:889–942. https://doi.org/10.1093/qje/qjae044
Castelo N, Bos MW, Lehmann DR (2019) Task-dependent algorithm aversion. J Mark Res 56:809–825. https://doi.org/10.1177/0022243719851788
Ceglarek A, Hubalewska-Mazgaj M, Lewandowska K et al (2021) Time-of-day effects on objective and subjective short-term memory task performance. Chronobiol Int 38:1330–1343. https://doi.org/10.1080/07420528.2021.1929279
Chen C, Liu H, Yang J et al (2025) Can domain experts rely on AI appropriately? a case study on AI-assisted prostate cancer MRI diagnosis. arXiv:2502.03482. https://doi.org/10.48550/arXiv.2502.03482
Citron DK, Pasquale F (2014) The scored society: due process for automated predictions. Wash Law Rev 89:1–34
Cobbe J, Lee MSA, Singh J (2021) Reviewable automated decision-making: a framework for accountable algorithmic systems. In: Proceedings of the 2021 ACM Conference on fairness, accountability, and transparency. ACM, Virtual Event Canada, pp 598–609
Cooper AF, Moss E, Laufer B, Nissenbaum H (2022) Accountability in an algorithmic society: relationality, responsibility, and robustness in machine learning. In: 2022 ACM Conference on fairness accountability and transparency. ACM, Seoul Republic of Korea, pp 864–876
Croskerry P, Singhal G, Mamede S (2013) Cognitive debiasing 2:impediments to and strategies for change. BMJ Qual Saf 22:ii65–ii72. https://doi.org/10.1136/bmjqs-2012-001713
Davidovic J (2023) On the purpose of meaningful human control of AI. Front Big Data 5:1017677. https://doi.org/10.3389/fdata.2022.1017677
Dietvorst BJ, Simmons JP, Massey C (2015) Algorithm aversion: people erroneously avoid algorithms after seeing them err. J Exp Psychol Gen 144:114–126. https://doi.org/10.1037/xge0000033
Dietvorst BJ, Simmons JP, Massey C (2018) Overcoming algorithm aversion: people will use imperfect algorithms if they can (even slightly) modify them. Manag Sci 64:1155–1170. https://doi.org/10.1287/mnsc.2016.2643
Douglas M (1966) Purity and danger: an analysis of concept of pollution and taboo, Repr. Routledge, London, UK
Elish MC (2019) Moral crumple zones: cautionary tales in human-robot interaction. SSRN Electron J. https://doi.org/10.2139/ssrn.2757236
Enarsson T, Enqvist L, Naarttijärvi M (2022) Approaching the human in the loop—legal perspectives on hybrid human/algorithmic decision-making in three contexts. Inf Commun Technol Law 31:123–153. https://doi.org/10.1080/13600834.2021.1958860
Endsley MR, Kiris EO (1995) The out-of-the-loop performance problem and level of control in automation. Hum Factors 37:381–394. https://doi.org/10.1518/001872095779064555
Enqvist L (2023) ‘Human oversight’ in the EU artificial intelligence act: what, when and by whom? Law Innov Technol 15:508–535. https://doi.org/10.1080/17579961.2023.2245683
Fiske AP, Tetlock PE (1997) Taboo trade-offs: reactions to transactions that transgress the spheres of justice. Polit Psychol 18:255–297. https://doi.org/10.1111/0162-895X.00058
Foot P (1967) The problem of abortion and the doctrine of the double effect. Oxf Rev Econ Policy 5:5–15
Gogoll J, Uhl M (2018) Rage against the machine: automation in the moral domain. J Behav Exp Econ 74:97–103. https://doi.org/10.1016/j.socec.2018.04.003
Green B (2022) The flaws of policies requiring human oversight of government algorithms. Comput Law Secur Rev 45:105681. https://doi.org/10.1016/j.clsr.2022.105681
Grove WM, Zald DH, Lebow BS et al (2000) Clinical versus mechanical prediction: a meta-analysis. Psychol Assess 12:19–30. https://doi.org/10.1037/1040-3590.12.1.19
Grundke A (2024) If machines outperform humans: status threat evoked by and willingness to interact with sophisticated machines in a work-related context*. Behav Inf Technol 43:1348–1364. https://doi.org/10.1080/0144929X.2023.2210688
Johansson P, Hall L, Sikström S, Olsson A (2005) Failure to detect mismatches between intention and outcome in a simple decision task. Science 310:116–119. https://doi.org/10.1126/science.1111709
Johnson DG (2021) Algorithmic accountability in the making. Soc Philos Policy 38:111–127. https://doi.org/10.1017/S0265052522000073
Johnson EJ, Goldstein D (2003) Do defaults save lives? Science 302:1338–1339. https://doi.org/10.1126/science.1091721
Justen L (2025) LLMs Outperform experts on challenging biology benchmarks. arXiv:2505.06108. https://doi.org/10.48550/arXiv.2505.06108
Kahneman D, Sibony O, Sunstein CR (2021) Noise: a flaw in human judgment. Little, Brown Spark, New York
Kern C, Gerdon F, Bach RL et al (2022) Humans versus machines: who is perceived to decide fairer? Experimental evidence on attitudes toward automated decision-making. Patterns (NY) 3:100591. https://doi.org/10.1016/j.patter.2022.100591
Kestin G, Miller K, Klales A et al (2025) AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Sci Rep 15:17458. https://doi.org/10.1038/s41598-025-97652-6
Langer EJ (1975) The illusion of control. J Pers Soc Psychol 32:311–328. https://doi.org/10.1037/0022-3514.32.2.311
Laux J (2024) Institutionalised distrust and human oversight of artificial intelligence: towards a democratic design of AI governance under the European Union AI Act. AI Soc 39:2853–2866. https://doi.org/10.1007/s00146-023-01777-z
Levinson W, Kao A, Kuby A, Thisted RA (2005) Not all patients want to participate in decision making: a national study of public preferences. J Gen Intern Med 20:531–535. https://doi.org/10.1111/j.1525-1497.2005.04101.x
Lima G, Grgić-Hlača N, Jeong JK, Cha M (2022) The conflict between explainable and accountable decision-making algorithms. 2022 ACM Conference on Fairness Accountability and Transparency. ACM, Seoul Republic of Korea, pp 2103–2113
Lind EA, Tyler TR (1988) The social psychology of procedural justice. Plenum Press, New York
Liu J, Hsu F-C, Yu J et al (2025) When humble AI meets narcissistic customers: a terror management perspective. Int J Inf Manag 83:102904. https://doi.org/10.1016/j.ijinfomgt.2025.102904
Logg JM, Minson JA, Moore DA (2019) Algorithm appreciation: people prefer algorithmic to human judgment. Organ Behav Hum Decis Process 151:90–103. https://doi.org/10.1016/j.obhdp.2018.12.005
Longoni C, Bonezzi A, Morewedge CK (2019) Resistance to medical artificial intelligence. J Consum Res 46:629–650. https://doi.org/10.1093/jcr/ucz013
Madrian BC, Shea DF (2001) The power of suggestion: inertia in 401(k) participation and savings behavior. Q J Econ 116:1149–1187. https://doi.org/10.1162/003355301753265543
Mandel DR, Vartanian O (2008) Taboo or tragic: effect of tradeoff type on moral choice, conflict, and confidence. Mind Soc 7:215–226. https://doi.org/10.1007/s11299-007-0037-3
Matthias A (2004) The responsibility gap: ascribing responsibility for the actions of learning automata. Ethics Inf Technol 6:175–183. https://doi.org/10.1007/s10676-004-3422-1
McNaughton D, Rawling P (1991) Agent-relativity and the doing-happening distinction. Philos Stud 63:167–185. https://doi.org/10.1007/BF00381686
Meehl PE (1954) Clinical versus statistical prediction: a theoretical analysis and a review of the evidence. University of Minnesota Press, Minneapolis
Mei P, Cannon R, Everett J et al (2025) Public trust and blame attribution in human-AI interactions: a comparison between air traffic control and vehicle driving. Transp Res Interdiscip Perspect 32:101545. https://doi.org/10.1016/j.trip.2025.101545
Nagel T (1986) The view from nowhere. Oxford Univ. Press, Oxford
Nisbett RE, Wilson TD (1977) Telling more than we can know: verbal reports on mental processes. Psychol Rev 84:231–259. https://doi.org/10.1037/0033-295X.84.3.231
Novelli C, Taddeo M, Floridi L (2024) Accountability in artificial intelligence: what it is and how it works. AI Soc 39:1871–1882. https://doi.org/10.1007/s00146-023-01635-y
Noy S, Zhang W (2023) Experimental evidence on the productivity effects of generative artificial intelligence. Science 381:187–192. https://doi.org/10.1126/science.adh2586
Oeberst A, Imhoff R (2023) Toward parsimony in bias research: a proposed common framework of belief-consistent information processing for a set of biases. Perspect Psychol Sci 18:1464–1487. https://doi.org/10.1177/17456916221148147
Othman K (2023) Public attitude towards autonomous vehicles before and after crashes: a detailed analysis based on the demographic characteristics. Cogent Eng. https://doi.org/10.1080/23311916.2022.2156063
Parasuraman R, Manzey DH (2010) Complacency and bias in human use of automation: an attentional integration. Hum Factors 52:381–410. https://doi.org/10.1177/0018720810376055
Parfit D (1987) Reasons and persons, 1. issued in paperback (with corr.), reprinted with further corr. Clarendon Press, Oxford
Reed C (2018) How should we regulate artificial intelligence? Philos Trans R Soc A Math Phys Eng Sci 376:20170360. https://doi.org/10.1098/rsta.2017.0360
Reich T, Kaju A, Maglio SJ (2023) How to overcome algorithm aversion: learning from mistakes. J Consum Psychol 33:285–302. https://doi.org/10.1002/jcpy.1313
Reis M, Reis F, Kunde W (2024) Influence of believed AI involvement on the perception of digital medical advice. Nat Med 30:3098–3100. https://doi.org/10.1038/s41591-024-03180-7
Rhodes M, Gelman SA, Karuza JC (2014) Preschool ontology: the role of beliefs about category boundaries in early categorization. J Cogn Dev 15:78–93. https://doi.org/10.1080/15248372.2012.713875
Rozenblit L, Keil F (2002) The misunderstood limits of folk science: an illusion of explanatory depth. Cogn Sci 26:521–562. https://doi.org/10.1207/s15516709cog2605_1
Rozin P (2005) The meaning of “natural”: process more important than content. Psychol Sci 16:652–658. https://doi.org/10.1111/j.1467-9280.2005.01589.x
Santoni De Sio F, Mecacci G (2021) Four responsibility gaps with artificial intelligence: why they matter and how to address them. Philos Technol 34:1057–1084. https://doi.org/10.1007/s13347-021-00450-x
Santoni De Sio F, Van Den Hoven J (2018) Meaningful human control over autonomous systems: a philosophical account. Front Robot AI 5:15. https://doi.org/10.3389/frobt.2018.00015
Scheffler S (1982) The rejection of consequentialism: a philosophical investigation of the considerations underlying rival moral conceptions, Rev. ed., reprint. in paperback. Clarendon Press, Oxford
Shay LA, Lafata JE (2015) Where is the evidence? A systematic review of shared decision making and patient outcomes. Med Decis Making 35:114–131. https://doi.org/10.1177/0272989X14551638
Shenhav A, Musslick S, Lieder F et al (2017) Toward a rational and mechanistic account of mental effort. Annu Rev Neurosci 40:99–124. https://doi.org/10.1146/annurev-neuro-072116-031526
Simon HA (1955) A behavioral model of rational choice. Q J Econ 69:99. https://doi.org/10.2307/1884852
Simon HA (2000) Bounded rationality in social science: today and tomorrow. Mind Soc 1:25–39. https://doi.org/10.1007/BF02512227
Spatola N, Normand A (2021) Human vs. machine: the psychological and behavioral consequences of being compared to an outperforming artificial agent. Psychol Res 85:915–925. https://doi.org/10.1007/s00426-020-01317-0
Spranca M, Minsk E, Baron J (1991) Omission and commission in judgment and choice. J Exp Soc Psychol 27:76–105. https://doi.org/10.1016/0022-1031(91)90011-T
Stahl BC, Antoniou J, Ryan M et al (2022) Organisational responses to the ethical issues of artificial intelligence. AI Soc 37:23–37. https://doi.org/10.1007/s00146-021-01148-6
Stojilović D, Franklin M, Malle BF, Fernandez-Basso C, Awad E, Lagnado D (2024) Are autonomous vehicles blamed differently? In: Proceedings of the 46th Annual Meeting of the Cognitive Science Society, pp 466–472. Cognitive Science Society
Sunstein C, Gaffe J (2025) An anatomy of algorithm aversion. Sci Technol Law Rev 26:. https://doi.org/10.52214/stlr.v26i1.13339
Tetlock PE, Kristel OV, Elson SB et al (2000) The psychology of the unthinkable: taboo trade-offs, forbidden base rates, and heretical counterfactuals. J Pers Soc Psychol 78:853–870. https://doi.org/10.1037/0022-3514.78.5.853
Tversky A, Kahneman D (1974) Judgment under uncertainty: heuristics and biases: biases in judgments reveal some heuristics of thinking under uncertainty. Science 185:1124–1131. https://doi.org/10.1126/science.185.4157.1124
Tversky A, Kahneman D (1981) The framing of decisions and the psychology of choice. Science 211:453–458. https://doi.org/10.1126/science.7455683
Vallor S (2015) Moral deskilling and upskilling in a new machine age: reflections on the ambiguous future of character. Philos Technol 28:107–124. https://doi.org/10.1007/s13347-014-0156-9
Van Wynsberghe A, Robbins S (2019) Critiquing the reasons for making artificial moral agents. Sci Eng Ethics 25:719–735. https://doi.org/10.1007/s11948-018-0030-8
Wagner B (2019) Liable, but not in control? Ensuring meaningful human agency in automated decision-making systems. Policy Internet 11:104–122. https://doi.org/10.1002/poi3.198
Willemsen P, Reuter K (2016) Is there really an omission effect? Philos Psychol 29:1142–1159. https://doi.org/10.1080/09515089.2016.1225194
Yam KC, Eng A, Gray K (2025) Machine replacement: a mind-role fit perspective. Annu Rev Organ Psychol Organ Behav 12:239–267. https://doi.org/10.1146/annurev-orgpsych-030223-044504
Yeung SK, Yay T, Feldman G (2022) Action and inaction in moral judgments and decisions: meta-analysis of omission bias omission-commission asymmetries. Pers Soc Psychol Bull 48:1499–1515. https://doi.org/10.1177/01461672211042315
Zerilli J, Bhatt U, Weller A (2022) How transparency modulates trust in artificial intelligence. Patterns (n Y) 3:100455. https://doi.org/10.1016/j.patter.2022.100455
Zhang Q, Wallbridge CD, Jones DM, Morgan P (2021) The blame game: double standards apply to autonomous vehicle accidents. In: Stanton N (ed) Advances in human aspects of transportation. Springer International Publishing, Cham, pp 308–314
Author information
Authors and Affiliations
Contributions
B.A. and D.G. jointly developed the central concept and argument of the manuscript. B.A. wrote the initial draft and led the overall manuscript development. D.G. contributed to ideation, theoretical framing, and substantive revisions, including expanding key sections and refining the structure.
Corresponding author
Ethics declarations
Conflict of interest
The authors declare no competing interests.
Additional information
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
About this article
Cite this article
Aeon, B., Gruda, D. When AI outperforms humans, who bears the burden of proof?. AI & Soc (2026). https://doi.org/10.1007/s00146-026-03375-1
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s00146-026-03375-1
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.