Why We Need Epistemically, Not Morally, Trustworthy AI: Decision Support and the Need for Collective Accountability
Abstract
Calls for âtrustworthy AIâ have become ubiquitous in policy and industry, yet the term remains conceptually underspecified. In earlier work (Dorsch & Deroy, 2024), we argued that moral trustworthiness is neither possible nor necessary for AI decision-support systems (AI-DSS), and indeed unethical to pursue, since it risks the category mistake of attributing benevolence to systems that lack it. This paper revisits that claim in light of recent critiques and developments, arguing that once moral trustworthiness is set aside, the central normative question shifts from whether AI systems can be trusted to how their epistemic influence on human decision-making ought to be justified and governed. We begin by situating our earlier thesis alongside complementary arguments that converge on the dangers of anthropomorphising AI. Second, we respond to critics who seek residual roles for moral trustworthiness and who clarify the epistemic conditions of trust. Third, we advance an account of epistemic trustworthiness as a property of AI outputs, rather than a property of the systems themselves. We contend that guaranteeing epistemic trustworthiness depends on human institutions, and is necessary to mitigate vulnerabilities introduced by integrating AI-DSS into hybrid decision-making. We connect this property to legal frameworks, showing how the EU AI Actâs âRight to Explanationâ codifies it as a necessary condition for deploying AI in high-risk domains.
1 Introduction
âTrustworthy AIâ is everywhere: in policy documents, corporate white papers, and marketing campaigns (European Commission, 2019; Biden, 2023; ISACA, 2025). But for all its rhetorical appeal, the phrase remains underspecified. What does it really mean to call a machine trustworthy, and why should this be our goal? Without such clarity, the discourse risks becoming both conceptually confused and socio-politically expedient, potentially undergirding arbitrary or even unethical design choices. Building on previous discussions of trustworthiness in AI (e.g., Ryan, 2020; Bryson, 2018), we here focus on defending a positive account of epistemic trustworthiness that specifies how an AI decision-support system (AI-DSS) should be designed to support, rather than override, human reasoning.
Our starting point is with philosophical accounts of trust, which have long emphasized that mere reliability, while necessary, is not sufficient for trustworthiness. Reliability means consistency in performance: an agent counts as reliable when its actions tend to succeed not only in actual situations but also in relevant counterfactual ones (see Dorsch & Deory, 2024 for more). Reliability in this sense provides a baseline of competence, but it does not yet justify trust. Something moreâsome âT-factorâ (i.e., trust factor)âis required to ground it. Yet decades of debate have yielded no consensus about what this T consists in: goodwill (Baier, 1986), emotional sensitivity to moral values (Lahno, 2001), or having the trustorâs interests as oneâs own (Hardin, 2002). What unites these views is the claim that trustworthiness involves more than mere dependability; it entails an evaluative stance toward reasons or shared values.
When this discourse turns toward emerging technology, âtrustworthy AIâ is often cast as a solution to the alignment problem (Christian, 2020), which, put simply, is the challenge of ensuring that AI systems reliably act in ways that reflect our values (such as fairness, safety, or non-discrimination) rather than in ways that are opaque or misaligned with human interests.Footnote 1 But critics argue that such talk risks a category mistake: importing moralized notions of trust into contexts where they cannot apply, thereby encouraging anthropomorphism and distorting design and governance priorities (Bryson 2018; Ryan 2020; Deroy 2023; Dorsch & Deroy, 2024). The irony is that this very rhetoric risks eroding public trust in the institutions that promote it. Our concern, however, is not merely terminological. The deeper issue is that framing AI-DSS in terms of trustworthiness without further qualification risks encouraging both design goals and governance expectations that implicitly treat such systems as possessing human-like moral capacities they cannot instantiate.
In earlier work (Dorsch & Deroy, 2024), we sought to draw sharper conceptual boundaries by distinguishing between different kinds of trust, and by arguing that one prominent formâmoral trustworthinessâis neither feasible nor required in the case of AI-DSS. AI-DSS are systems such as predictive algorithms or recommender models that generate scores, rankings, or recommendations to guide human judgment, while leaving final decisions with the human user.Footnote 2 We defined moral trustworthiness as the capacity to do the right thing for the right reasons, which we argued requires conceptual or emotional understandingâcapacities that artificial systems do not possess.Footnote 3 Our conclusion was that calls for trustworthy AI, in an explicitly moral sense, rest on a category mistake: they ascribe a distinctively human capacity to systems that, even if highly reliable and useful, cannot instantiate it. The downstream implication is a communicative and policy-oriented one: AI governance discourse should avoid framing AI-DSS in ways that imply or encourage anthropomorphic interpretations. Other forms of trustâmost notably epistemic trustâwe set aside at the time and have since pursued in later work (Dorsch & Moll, 2024; Moll & Dorsch, 2025).
What matters for AI-DSS is not that they act for reasons in the normative, socially accountable sense distinctive of human agents, but that they are maximally reliable and, crucially, that this reliability is made intelligible in ways that human users can integrate into their own decision-making (Dorsch & Deroy, 2024). Because such systems often function as quasi-partners in collective deliberation (Dorsch & Moll, 2024), reliable output alone is insufficient: they must also disclose the grounds of recommendations. Only then can they begin to mitigate the vulnerabilities of those who rely on them, and only then does the question of epistemic trustworthiness arise, that is, how AI-DSS might be designed to communicate their reliability in ways that justify uptake in human reasoning. Crucially, unlike moral trustworthiness, epistemic trustworthiness in this sense does not require attributing human-like moral agency to AI systems, but instead concerns the relational and institutional conditions under which system outputs may justifiably acquire epistemic weight within human deliberation.
Thus, this paper takes up that question in light of both subsequent critiques and new developments. We begin by reinforcing our original thesis that moral trustworthiness is unnecessary and inappropriate by situating it alongside complementary contributions that converge on the same caution against anthropomorphising AI (Sect. 2). We then turn to criticisms, both by correcting misreadings of our earlier work and by engaging with substantive challenges that advance the debate (Sect. 3). Finally, we advance a positive account of epistemic trustworthiness intended precisely to avoid the anthropomorphic and category-mistake concerns raised earlier in the paper, by exploring whether and under what conditions AI-system outputs may be justifiably integrated into human deliberation (Sect. 4). Our account is grounded in the philosophical ideal of reason-giving and reason-receiving practices, modeled on collaborative norms that govern collective deliberation and decision-making, and externally reinforced by the EU AI Actâs Right to Explanation (European Parliament, 2024), which lends democratic legitimacy to this standard.
2 Complementary Contributions and Converging Arguments
Before turning to the complementary contributions that our earlier paper has attracted, it is useful to recall its central theses. We argued, first, that in the context of AI-DSS, morally trustworthy AI is neither feasible nor required, and second, that unreflective use of the language of âtrustworthinessâ in relation to AI risks importing human-specific capacities into technological contexts.This second claim follows from the first: our concern is not simply with linguistic usage itself, but with the way particular forms of policy and governance language can encourage anthropomorphic assumptions about the nature of AI systems.
More specifically, the concern is that such discourse encourages four related pathologies: anthropomorphic projections of moral agency onto AI systems, unreflective uses of trust discourse detached from clear normative standards, managerial âbox-checkingâ approaches to ethics, and techno-solutionist assumptions that trust can be engineered independently of human oversight and institutional responsibility. As we argue later in Sect. 4, our positive account of epistemic trustworthiness is intended precisely as a response to these concerns. In this section, we highlight how a growing body of literature has echoed and reinforced these claims, often from different disciplinary standpoints.
Upstream from our caution about describing machines as âtrustworthyââa term that too easily slides into implying moral trustworthinessâthere is a growing body of literature that criticizes certain linguistic descriptors in science and policy communication, words that are argued to be nothing more than mere metaphor when applied to artificial entities. Consider the ever-expanding use of the term âintelligentâ to describe various algorithmic functions and capabilities: what was once labelled simply âimage processingâ is now routinely sold as AI. In this light, Lanier and Weyl (2020) argue that AI is not a kind of technology at all, but a powerful marketing tool.
In this literature, anthropomorphizing artifacts has a pernicious flipside: ascribing human-like qualities to machines reduces human qualities to machine-like ones, fostering dehumanization (Bender, 2024). Baria and Cross (2021) show how the metaphor âthe brain is a computerâ both understates human complexity and inflates machine capacities, privileging cold rationality over emotionality and casting AI as the more âtrustworthyâ intelligence. This concern about anthropomorphic projection forms one of the central motivations for our later attempt to develop a notion of epistemic trustworthiness that does not depend upon attributing human-like moral agency to AI systems.
Recent work allows us to examine more finely the idea that anthropomorphism and trust in AI go hand in hand. In a recent study by Colombatto, Birch, and Fleming (2025), participants rated ChatGPTâs capacity for various mental states and then completed an advice-taking task. Attributions of intelligence-related traits (reasoning, planning, memory) were strongly associated with greater acceptance of the systemâs advice, whereas attributions of experience-related traits (emotions, feelings) predicted lower advice-taking. Closer to our point, attributions of consciousness increased self-reported trust, though the increase was weak. These were also unrelated to behavioural reliance. This suggests that anthropomorphism does not uniformly foster trust, but that specific dimensions of perceived mentality play divergent roles (see also here our own work in Geiselmann et al., 2023).
Whereas these analyses highlight the cultural dangers of anthropomorphizing machines, other scholars have turned to the political uses of this language, showing how the framing functions ideologically. Espinosa-Leal and Stocchetti (2025) offer a political critique of the âtrustworthy AIâ discourse, exposing how the term functions as a rhetorical device within broader ideological projects. In their account, trustworthy AI is a cultivated narrative that aligns with policymaking that treats markets, competition, and managerial oversight as the primary mechanisms of social order and with technological solutionism, the assumption that complex social and political problems can be resolved through technical fixes.
Importantly, this critique is not merely terminological. Its force lies in showing how the appeal to trustworthiness becomes a troubling proxy for sound moral reasoning. This proxy is effective because âtrustworthy AIâ serves to humanize the technology, presenting it as though it were a genuine partner in social life. At the same time, it promotes managerial ethics, that is, an ethics focused narrowly on compliance, risk management, and efficiency, as the only relevant mode of ethical reflection, often restricting instrumental reasoning as the only relevant normative regulatory force. The result is a discourse that sidelines alternative ethical traditions, such as virtue ethics, deontology, or care ethics, which otherwise highlight dimensions that might neither be accounted for by managerial oversight, nor reduced to mere means-ends calculus.
Our concern, then, is not simply that the language of âtrustworthy AIâ may be rhetorically misleading, but that its underspecified use can materially shape governance practices, design priorities, and public expectations in normatively problematic ways. These concerns also motivate our later attempt to articulate a more explicit and constrained account of epistemic trustworthiness, rather than leaving âtrustâ to function as an intuitive placeholder.
While our earlier paper sought to demonstrate that morally trustworthy AI was an unnecessary design goal, Carroll (2025) approaches this issue from a prudential standpoint: is trustworthy AI even something we should value? In Carrolâs reasoning, moral trustworthiness implies that an agent is positioned to serve as a substitute for our own moral judgment under certain conditions. In human relationships, this can be appropriate: we sometimes entrust moral decisions to others precisely because we recognise their moral character. But in the case of AI, Carroll warns, such deference risks becoming hazardous. The more a system appears to warrant moral trust, the stronger the pull to allow it to settle questions that should remain under human deliberation.
Carrollâs caution aligns closely with one of the practical concerns that motivated our original argument: namely, that anthropomorphising AI in moral terms is not only a conceptual mistake, but also a dangerous design goal. Carrollâs analysis reinforces this by showing that the danger does not vanish even if the attribution were somehow accurate. Instead, the very fact of having a morally trustworthy AI would reconfigure the decision-making environment in ways that could diminish human moral responsibility and undermine the institutional safeguards that are currently premised on human oversight. Additionally, pursuing systems with human-level, âdomain-generalâ intelligenceâor even consciousnessâcreates ethical spillovers, from environmental costs to safety risks, without a proportional gain in instrumental value. Specialized, distributed, and more machine-like systems already achieve equal or superior performance without incurring these risks (Deroy et al., 2024).
The opacity of the term âtrustworthiness,â together with its industry-wide promotion as a design goal, has yielded serious problems for empirical research, where it is frequently operationalised in inconsistent ways. Baltes et al. (2024) document this trend by directing attention to how âtrustâ is actually measured in studies across humanâAI interaction. Their analysis surveys a range of studies, finding that âtrustâ is often used as a catch-all dependent variable, frequently conflated with measures of satisfaction or perceived usefulness.
This matters for two reasons. First, it produces a feedback loop in which findings appear to confirm that âtrustâ is a meaningful construct in the AI context, while in fact the construct being measured varies widely and often fails to meet even the minimal conceptual criteria for what should be construed as trust. Second, by treating âtrustâ as an unproblematic target for optimisation, such research risks embedding anthropomorphic assumptions into system design, with the consequence that systems are optimised to score highly on trust metrics without any corresponding improvement in reliability, transparency, or accountability. Our later discussion of the epistemic trustworthiness is intended in part to address this problem by making explicit the substantive epistemic conditions under which deliberative uptake of AI outputs may be justified.
Hence, authoritative use of the term trustworthy requires careful restraint, a restraint captured in our earlier articulation of the Principle of the Non-equivalence of Trustworthiness and Reliability (PNTR). PNTR states, in short, that one ought to speak of âtrustâ or âtrustworthinessââin the context of science or policy communicationâonly where these notions imply something more than mere reliability; otherwise, the language of reliability should be used. In this respect, PNTR is not at odds with the stronger critiques surveyed above, which warn against âtrustworthy AIâ functioning as a mere rhetorical placeholder; rather, it specifies a baseline condition under which such terminology would cease to function in that way. Failure to adhere to this principle risks more than conceptual confusion. It may signal that vulnerability toward a system is justified when it is not, and this mis-signalling ought to be construed as a harm in its own right, independent of any further harms that might result from vulnerable reliance on the system.
Admittedly, much of the debate over whether AI can be trustworthy risks sounding like a dispute about labels. Of course, it is more than that, since it concerns the underlying conditions of trustworthiness and whether a machine could ever satisfy them. Still, the question arises whether labels themselves matter. Our earlier argument suggested they do, since we argued that describing AI as âtrustworthyâ risks a category mistake that might lead users to falsely regard the system as a moral agent in a socially meaningful sense. In subsequent work, we investigated this claim empirically: does calling emerging technology âtrustworthyâ correlate with stronger anthropomorphic associations than calling it merely âreliableâ? To answer this question, we conducted a study that examined the anthropomorphic pull of the âtrustworthyâ label when applied to automotive AI systems ( Dorsch & Deory, 2025).
Participants were presented with three vignettes describing distinct automotive AI systems; each vignette was mirrored across the two groups, differing only in whether the system was labeled âtrustworthyâ (experimental group) or âreliableâ (control group). Participants then completed an adapted validated technology acceptance questionnaire (Choung et al., 2022). This enabled us to measure differences in how the two groups evaluated the system along various dimensions that distinguish between reliability and trustworthiness. We found that the âtrustworthy AIâ label had a significant positive effect on the perception that automotive AI cares about our well-being (p = 0.00323, 95% CI [0.1626, 0.8097]). Within the distribution of responses, significant shifts were observed between categories: participants in the âtrustworthyâ condition were far less likely to select Strongly Disagree rather than Disagree (p = 0.0013) and more likely to move from Neutral to Agree (p = 0.011). This suggests that the label systematically shifted responses toward greater perceived benevolence, defined in Mayer et al. (1995) classic model as the extent to which a trustee is believed to wish to do good to the trustor (see Fig. 1).
These findings are important for two reasons. First, they demonstrate that anthropomorphic ascriptions can be triggered in ordinary users simply by the addition of a âtrustworthyâ label as a minimal linguistic cue. Second, they expose a potentially dangerous form of category mistake: treating an AI system as if it possesses the capacity to care for human well-being. If this mistake is left unchecked, it opens the door to a subtle but serious risk: moral dependence on a system that is fundamentally incapable of reciprocating that relationship, the likes of which we have already begun to see with AI romantic partners (Apple, 2025) and therapists (Richter, 2025). For this reason, the positive account developed in this paper does not attempt to rehabilitate morally trustworthy AI, but instead seeks a more narrowly epistemic and institutionally constrained framework for understanding when reliance on AI-DSS may be justified without inviting anthropomorphic distortions.
One might object that introducing qualifiers such as âepistemicâ does little to mitigate the risks associated with âtrustworthy AIâ discourse, and that the safer strategy would therefore be to abandon the terminology altogether. We agree that merely relabelling a concept without altering its content would be insufficient. However, qualifiers can fundamentally alter what is being attributed to a system, particularly when they introduce a technical and restricted domain of evaluation. âEpistemic trustworthiness,â as we use the term, does not attribute moral character or autonomous agency to AI systems, but instead refers narrowly to the conditions under which outputs may justifiably acquire epistemic weight within human deliberation. Moreover, because âepistemicâ is itself a technical term rather than an intuitive consumer-facing descriptor, it invites further specification rather than functioning as an mere rhetorical placeholder. The positive account developed in Sect. 4 is intended precisely to show how this constrained notion of epistemic trustworthiness can avoid the pathologies surveyed above by operationalising the normative surplus beyond reliability through explicit epistemic conditions.
3 Residual Roles and Epistemic Misreadings: Responses to Franke and Mitchell
Two articles have recently challenged our claim that we do not need morally trustworthy AI. Although they share much of our conceptual framework, they press on different points of disagreement. Franke (2024) seeks to preserve a residual role for moral trustworthiness in cases where calibration through reliability metrics can be limited, while Mitchell (2025) argues reliability requires trust, clarified through transparency distinctions.
Frankeâs strategy is to identify residual roles for moral trust, where calibration through reliability metrics may fall short. Reliability metrics, such as precision, recall, and the F1-score, are standard tools for evaluating the performance of machine learning models. Precision measures how many of the cases the system identified as positive were in fact correct, while recall measures how many of the actual positive cases were successfully identified. These correspond, respectively, to avoiding false positives (Type I errors) and avoiding false negatives (Type II errors). The F1-score combines both into a single measure that balances precision and recall.
To understand what is at stake, it is useful to clarify what calibration means in this context. By calibration we refer to the alignment between when users ought to rely on an AI-DSS and when they in fact do so (also referred to as âappropriate relianceâ or âtrust calibrationâ; see Lee & See, 2004). Reliance on the outputs of an AI-DSS is well-calibrated when users accept the systemâs outputs when they are correct and reject them when they are not. The vulnerability at issue arises precisely when this alignment fails. Reliability metrics are therefore indispensable: they estimate the likelihood of miscalibrated reliance and thereby help operators avoid it.
To illustrate, consider an AI-DSS used for medical screening. Deploying a high precision model means that when the system flags a patient as at risk, it is rarely wrong; while a model with high recall means that it rarely misses patients actually at risk. An ideal system, in this sense, would approach a high F1-score, since that indicates reliable success in avoiding both Type I and Type II errors. But as Franke points out, in practice improving one often comes at the expense of the other (for a survey, see Terven et al., 2023).
Franke agrees with our claim that moral trustworthiness means doing the right thing for the right reasons, and that trust should be reserved for agents capable of moral rationality. He also concedes that in high-stakes systems trained on datasets with ground truth, reliance can typically be well-calibrated through standard metrics. The disagreement, then, is whether such calibration can fully eliminate the need for morally trustworthy AI.
Frankeâs first objection is that well-calibrated reliance, while necessary, is insufficient for mitigating vulnerabilities incurred by the deployment of AI-DSS. Reliability scores can mislead if inappropriate metrics are chosen, if training data is outdated or imbalanced. He describes this as a problem of âsecond-level biasâ: bias not in the modelâs outputs themselves, but in the selection of the reliability criteria by which those outputs are evaluated. Even when ground truth is available and calibration is efficiently performed, an epistemic asymmetry persists, as there remains significant uncertainty about whether the appropriate features have been measured. For present purposes, we take it that this criticism can apply to further reliability metrics tested and proposed in the literature (e.g., Steyvers et al., 2025).
In framing this as an epistemic asymmetry, Franke is explicitly drawing on the terminology of our original paper. There we used the term to describe situations in which there is a gap between what an agent actually knows and what they ought to know about the reliability of another agent or system. In such cases, we argued, moral trustworthiness can serve as a scaffold: it can allow one to justify vulnerability to the decision of another in cases where their reliability cannot be fully assessed. Franke adapts this notion to argue that even when reliance is well-calibrated, a residual asymmetry remains: what one knows is how the model performs relative to a chosen metric, but what one ought to know is whether that chosen metric is the right one for the context. On this basis, he assigns to morally trustworthy AI a minimal residual role: assisting in the selection of appropriate reliability metrics.
To make this concrete, consider again the distinction between Type I and Type II errors. A Type I error occurs when the system predicts a condition that is not actually present, as in diagnosing a healthy patient with a disease. A Type II error occurs when the system fails to detect a condition that is in fact present, as in missing a genuine case of illness. In machine learning terms, prioritizing the avoidance of Type I corresponds to maximizing precision, whereas prioritizing the avoidance of Type II corresponds to maximizing recall. Reliability scores tell us how well the model performs on whichever metric has been selected. What it cannot tell us is whether the chosen metric is the right one to begin with.
Whether to prioritize minimizing Type I or Type II errors depends on context. In blood supply testing, avoiding Type I prevents wasting safe donations; in cancer screening, avoiding Type II prevents missing deadly cases. During COVID-19, early focus on Type II errors (undetected cases) later shifted as Type I errors (false positives) became equally disruptive, cancelling procedures and distorting data (Surkova et al., 2020). The key question is when, and on what grounds, to shift between these priorities.
At this point, the question is no longer one of technical expertise alone but of normative judgment: who decides which errors matter more, and according to what principles? Here our disagreement with Franke becomes clear. While he sees this residual gap as a role for morally trustworthy AI, our view is that what is needed is not a morally trustworthy machine but cohorts of morally trustworthy humans. And crucially, this is not a novel prescription but a recognition of what has long been established. These decisions are already, and quite deliberately, embedded in collective structures that embody the understanding that no single âcorrectâ answer lies hidden, awaiting discovery by statistical methods alone.
Medical practice offers a clear illustration of this. Frameworks such as the CDCâs (2025) Diagnostic Excellence Framework mandates multidisciplinary input across laboratory experts, clinicians, quality officers, and hospital leadership. What counts as doing the right thing, finding the right balance between errors, is not determined solely by procedures aimed at accurate correspondence with ground truth, but crucially also by collective deliberation that brings diverse perspectives, professional norms, and social values to bear.
That point is crucial for understanding why we do not need morally trustworthy AI to select such metrics. To build a system with that role would be to collapse distributed, socially accountable decision-making into a single decisive source. This is a problem because the legitimacy of these choices derives largely from their collective character. A mathematical proof is correct if it follows from its axioms; an empirical claim is correct if it represents reality. One crucial component of what makes normative judgments of this kind correct is the process by which they are reached: a process that remains answerable to oversight and contest. Hence, the soundness of a decision about whether to minimize false negatives or false positives does not lie in its mirroring of some ground truth, but in its having been forged through transparent, socially accountable deliberation among responsible and invested agents. This is not to deny the importance of empirical factsâepidemiological data, prevalence rates, and risk profiles all inform how trade-offs are framed. But when those facts pull in different directions, it is what we collectively value that determines the right course of action.
This line of thought is hardly novel. Philosophers across traditions have long emphasized that the validity of normative judgments lies in procedures of collective accountability. In ethics and political theory, Habermasâs (1990) discourse ethics and Cohenâs (1997) deliberative democracy argue that legitimacy arises from inclusive processes of reason-giving among free and equal participants. In the philosophy of science, Longino (1990) stresses that objectivity depends on diverse communities engaged in critical interaction. What unites these accounts is the claim that what makes such normative judgments right is not their correspondence to some ground truth, but their having been arrived at for the right reasonsâreasons that have survived community scrutiny, having been justified through accountable collective processes.
Imagine, for the sake of argument, a panel of morally trustworthy AIs debating false positives versus false negatives, overseen by another panel of equally trustworthy machines. The obvious problem, however, is whether such cohorts of machines could ever be socially accountable, that is, capable of being held morally and legally responsible in cases of negligence or corruption. But even if such a system existed, it would mimic the very structure we already have, and hence an âartificial panel of expertsâ is not what we need.Footnote 4 What we need is not AI to determine our values for us, but to help us sift through the data, so that human institutions can enact our values responsibly.
Consider, by contrast, the unsettled debate over AI welfare, the claim that we may soon have moral obligations to safeguard the well-being of artificial systems (Long et al., 2024). Advocates warn that we might soon develop conscious AI models capable of suffering, and that failing to recognize this possibility risks committing a grave moral wrong. This is the fear of a Type II error: failing to detect moral status where it does in fact exist, thereby allowing preventable harm. But not everyone is convinced that this is the right way of unpacking this novel (and speculative) risk. Two authors of this paper, along with others (Dorsch et al., 2025), counter that a Type I error threatens the more substantial risk: ascribing moral status where none exists, and thus diverting scarce resources away from entities whose vulnerability already places them in urgent need of care.
This is where Frankeâs second counterargument comes into view. He argues that in novel domains without data about the ground truth, reliability metrics cannot be established to properly resolve the relevant epistemic asymmetries, and so a morally trustworthy AI might play a residual role by guiding us to find the right balance of errors. But this inference, again, misconstrues what makes such determinations justified.
Even in novel domains, the problem is not the absence of facts but the contested ways of interpreting them. Returning to the question of AI welfare, whether a specific configuration of computational architecture is sufficient for valenced consciousness may well be an empirical matter, one we have yet to determine. But even if every observable indicator has been exhausted, theory and evidence can still point in different directions. At that point, the problem is no longer one of determining what there is, but of determining what we value.
This point matters because problems of value are not easily resolved by technical means. Extrapolating regularities from data or simulating outcomes may reveal behavioral patterns, but they do not reveal the reasons that make a decision right or wrong. One might object that if situations recur and values are applied consistently, they should emerge in the data. Yet this assumes a stability that normative decision-making rarely affords. Values are historically contingent and evolve; practices once prohibited, such as divorce or interfaith marriage, are now widely accepted. Because value-based decision-making is open-ended and shifting, technical strategies alone cannot determine what is at stake. What is needed instead is a socially responsible, collective process capable of providing genuine justification.
Frankeâs paper is valuable for highlighting two structural limits of calibrated reliance: second-level bias and novel domains. But these are not technical deficits that we need a âmoral machineâ to remedy; they are governance challenges that require socially responsible and accountable collectives of morally trustworthy humans. By proposing residual roles for morally trustworthy AI, Franke suggests systems that, lacking social accountability, could never supply what is in fact required to address these challenges.
While Frankeâs strategy was to identify possible residual roles for morally trustworthy AI, Mitchellâs (2025) approach is of a different kind. He sets out four objections to our claim that morally trustworthy AI is unnecessary. Yet upon closer analysis, only one of these rises to the level of a substantive critique. One is a repetition of Frankeâs concern with novel domains. Another rests on a misconstrual of our view as a wholesale rejection of trust. The remaining two objections are, in fact, variations of the same worry: that our proposal of coupling outputs with reliability scores âpushes the problem back a step,â since one must then trust the scores.
Mitchell develops this objection in two registers: likening reliability scores to human confidence reports and, elsewhere, to reputations. If these analogies held, the scores would indeed demand trust, but both misfire. Human confidence reports are shaped by opaque metacognitive processes and are often biased (Fleming & Lau, 2014), as in the DunningâKruger Effect (Dunning & Kruger, 1999), or strategically inflated to gain influence (Kurvers et al., 2021; Moll et al., 2022). Hence, justifying vulnerability to another personâs confidence requires assurance not only that their calibration is accurate, but also that they are neither deceiving us nor themselves. Reputations, too, are socially constructed, opaque, and manipulable. Of course, reliability scores are, in reality, nothing like human confidence reports, nor social reputations: they are generated by transparent, well-defined computational procedures, auditable in their derivation, and devoid of possible ulterior motives. To the extent that moral trust is required, it should not be directed at the scores themselves, but at the human institutions who are socially responsible for ensuring scores track reliable performance.Footnote 5
4 On the Possibility and Limits of Epistemically Trustworthy AI
The aim of this section is to clarify the normative conditions under which the outputs of an AI-DSS may justifiably enter human deliberation, once moral trustworthiness has been ruled out. As established in our earlier argument, moral trustworthiness is neither possible nor necessary for AI-DSS. In that context, we introduced PNTR, which holds that whenever âtrustworthyâ can be replaced without semantic loss by âreliable,â the latter is the appropriate term for policy and scientific communication when describing the capacity of AI. PNTR is not merely linguistic housekeeping. It is a safeguard against two predictable distortions discussed already in detail above: first, the ideological drift that comes from indiscriminate use of the language of âtrustworthy AI,â which risks legitimizing technocratic narratives; second, the conceptual error of anthropomorphizing systems by attributing to them capacities for norm-following they cannot possess. More broadly, the concerns raised in Sect. 2 also included the tendency for trust discourse to become an underspecified managerial placeholder and for techno-solutionist framings to treat trust as something that can be engineered into systems independently of human oversight and institutional responsibility. The account developed in this section is intended precisely as a constrained and non-anthropomorphic response to these pathologies.
4.1 Conceptual Foundations: Epistemic Trustworthiness, Transparency, and PNTR
Having ruled out moral trustworthiness as a requirement for AI-DSS, we now ask whether a more disciplined notion of epistemic trustworthiness can be articulated without reproducing the same conceptual pitfalls. The criticism we pressed earlier was directed mainly at trustworthy as a stand-alone descriptor, because in that form it strongly invites anthropomorphic projection. By contrast, epistemic trustworthiness names a relational property governing the reasonable integration of information from a source into an agentâs deliberation. It refers to the conditions under which it is reasonable for agents to integrate a systemâs outputs into their own reasoning, that is, to treat those outputs as eligible for deliberative uptake at all. The degree or weight of influence those outputs should have is to be determined in a graded, context-sensitive way by the human decision-maker, given appropriate norms of reliability and transparency.Footnote 6 Crucially, this is not a capacity of the AI as an autonomous system, but a normative status conferred within collective and institutional frameworks that monitor and regulate the systemâs performance. In this way, epistemic trustworthiness avoids the anthropomorphic pathology identified in Sect. 2: the normative status at issue does not reside in the machine itself, but in the broader deliberative and institutional relationship governing how its outputs are evaluated and used.
Epistemic trustworthiness also marks a genuine normative surplus beyond reliability. Transparency illustrates this surplus. By transparency in the epistemically relevant sense, we mean the extent to which a system enables users to form an accurate understanding of how its outputs depend on inputs and relevant counterfactual conditions, such that they can reliably anticipate when, why, and with what degree of confidence those outputs should change. A car engine can be highly reliable yet not transparent: it starts every morning, but the average driver has no idea how exactly. Conversely, a person can be fully transparent yet utterly unreliable: they may declare openly that they will fail you or that they will change the reasons for their actions every time they act, such that their behavior resists any stable dependence.
But what does transparency amount to in an epistemically relevant sense for AI-DSS, especially in domains of high vulnerability? To answer this, it is useful to step back. Epistemic trustworthiness is itself a multi-faceted concept, and whether AI might support it depends crucially on which dimension is in view. One sense concerns an agentâs stance toward truth, whether it takes truth seriously as a norm, and consciously regulates its beliefs toward that ideal (e.g., Moran, 2005; Faulkner, 2011). In human cases, this involves reflexive metacognitive monitoring and a willingness to revise beliefs when faced with disconfirming evidence. In AI, however, there is at present no clear way to ensure that a system explicitly regulates itself with respect to the epistemic value of truth, as opposed to optimising for a surrogate objective that may only contingently align with it (though for a discussion of this, see Dorsch, 2025). This is an area where extreme caution is warranted, because alignment between output optimisation and truth-tracking may be domain-specific at best.
That said, one way to understand machine learning might be framed in terms of aligning outputs with truth: labels preserve the ground truth and serve as truth-conditions for model outputs, and the objective (or loss) function quantifies the error between the modelâs predictions and those conditions, serving as a guide for the model to reduce that error during training. Yet truth alignment in the robustly normative sense is more than this. It is an open-ended process of subjecting oneâs beliefs to both internal and external scrutiny, with the possibility of their continual revision. In this robust sense, truth is approached asymptotically, in an ongoing, often holistic process of acquiring true beliefs, discarding false ones, and strengthening justification.
That knowledge unfolds this way has long been recognized in epistemology. The idea is already implicit in the Socratic method, which treats inquiry as an unending process of examination (Plato, ca. 350 B.C.E./1997). It finds modern formulations in Peirceâs (1878) pragmatist conception of truth as the ideal limit of inquiry and Popperâs (1962) fallibilist account of scientific progress. By contrast, an algorithmâs optimization against labelled data or surrogate objectives is a closed procedure constrained by predefined metrics, whereas epistemic inquiry is an open-ended enterprise responsive to culturally evolving standards of justification.
Another sense of epistemic trustworthiness is realized in virtue of the capacities of the trustee to serve as a provider of information or as a conveyer of knowledge (Wilholt, 2013), wherein notions of reliability, transparency, and explainability become deeply relevant (Alvarado 2023a, b; Dorsch & Moll, 2024). This is the sense most relevant to AI-DSS, where the question is whether human operators can justifiably treat the systemâs outputs as carrying epistemic weight within their own deliberation. This requires more than reliability; it requires transparency about the limits and counterfactual robustness of the outputs. By making these epistemic conditions explicit, the account also avoids the problem identified in Sect. 2 of treating âtrustâ as an underspecified placeholder that can be operationalised through arbitrary proxy measures. It is, de facto, possible to suggest and test such solutions (see, for an example, Steyvers et al., 2025).
This is where a thought experiment presented in our earlier work becomes illuminating (Dorsch & Moll, 2024). Imagine two partners navigating mountainous terrain. One is a domain expert, familiar with the land and hiking techniques; the other is a technical expert, reading GPS data and the like. When their judgments diverge about which path to take, what would justify the domain expert in changing their mind? What information would the technical partner need to provide in order for their stance to be justifiably adopted? We argued that three things matter: reasons for the recommendation, the partnerâs confidence in those reasons, and counterfactuals showing how the judgment, and ideally their confidence levels as well, would shift under relevantly different conditions. Together, these constitute the RCC framework (Dorsch & Moll, 2024; Moll & Dorsch, 2025).
In our view, these elements capture the core of epistemic trustworthiness. Epistemic trust, as we use the term, is a willingness to let the epistemic stance of an external source (namely, its outputs as situated within norms of reliability and transparency) inform, guide, and sometimes shape oneâs own beliefs. Epistemic trustworthiness, by contrast, refers to the properties that justify placing such trust in an agent or informational source. In AI-DSS contexts, this means that when a doctor disagrees with an AIâs cancer diagnosis, or when a child welfare screener hesitates to follow an AIâs risk score, whether one is justified in revising or maintaining oneâs view will depend on the availability of the reasons, confidence, and its response to counterfactuals. Because these conditions are substantive, contestable, and embedded within broader practices of collective human deliberation, they are intended to resist the managerial and techno-solutionist tendencies discussed earlier, where trust becomes reduced either to a compliance label or to a supposedly self-sufficient technical property of the system itself. Together, these normsâreasons, confidence, and counterfactual sensitivityâoperationalise transparency in the epistemic sense, and thereby specify the conditions under which the outputs of an AI-DSS may be treated as epistemically trustworthy.
4.2 Epistemic Trustworthiness and the Right to Explanation
Crucially, this is not only an ethical requirement but also a legal one. Article 86 of the EU AI Act stipulates the âRight to Explanationâ that:
âAny affected person subject to a decision which is taken by the deployer on the basis of the output from a high-risk AI system⊠shall have the right to obtain from the deployer clear and meaningful explanations of the role of the AI system in the decision-making procedure and the main elements of the decision takenâ (EU Parliament, Article 86).
To be clear, our aim is not to enter into debates about the correct legal interpretation of Article 86, but to explore what it would mean for the outputs of an AI-DSS to be epistemically trustworthy to the extent that they can satisfy this legally grounded conception of explanation. For the purposes of this paper, we take Kaminski and Malgieriâs (2025) interpretation as a promising working understanding of the Right to Explanation. Their analysis situates Article 86 within broader European data protection jurisprudence (most notably the GDPR) and argues that the provision should be read as imposing four distinct requirements: that explanations be clear, meaningful, concerned with the main elements of the decision, and specify the role of the AI system.Footnote 7
First, an explanation is clear if it is comprehensible to the affected person. The standard is intelligibility rather than technical precision: the recipient must be able to grasp what the system did without needing expert training in computer science or law. Second, an explanation is meaningful if it enables action. In particular, the individual must be placed in a position to contest the decision or to exercise related rights, such as those protecting against discrimination. Third, the requirement to explain the main elements of the decision goes beyond listing correlations in the data. It encompasses both counterfactuals (how the outcome would have changed under different inputs) and reasons (the substantive grounds for the decision, explaining why the outcome followed from the systemâs operation rather than how it might have differed). Finally, Article 86 requires that the role of the AI system be explained. This entails clarifying how the automated component interacted with human oversight: whether the AI was decisive, advisory, or merely one input among others. A key point is to render visible the contribution of âhumans in the loopâ (De-Arteaga et al., 2020) and to prevent the anthropomorphic misperception that the system itself bears responsibility for the decision.
It might be objected, as Taylor (2024) has, that these requirements set the bar unrealistically high: that demanding explanations of this kindâwhat Taylor calls ânormative explanationsâ in that they provide justification for the decision takenâis either technically infeasible or institutionally overburdensome. Yet these challenges, while serious, should not be treated as decisive reasons to abandon the project.
On the technical side, it is premature to dismiss normative explanations as infeasible. Normative explanations consist of elements, such as those represented by reliability metrics and counterfactual information, that are technically well within reach. Methods for producing these are an active area of research in explainable AI, with concrete proposals demonstrating how small changes in inputs would alter outputs, thereby grounding judgments about what features of a case made a difference (see Mathew et al., 2025 for a state-of-the-art review, as well as Moll & Dorsch, 2025 for explainable reinforcement learning).
Whether such explanations succeed, however, does not depend solely on their technical sophistication, but on how they shape operatorsâ judgment. This matters because epistemic trustworthiness is not about making systems more accurate in isolation, but about supporting operators in judging when reliance is warranted. Steyvers et al. (2025) illustrate this point clearly. They show that, under current conditions, operators tend to overestimate the accuracy of LLM outputs when relying on default explanations, those which LLMs ordinarily provide when asked to explain why a given output was produced. In practice, this manifests in two related failures: operatorsâ overall sense of the systemâs accuracy drifts away from its actual performance, and, more locally, they struggle to discern, case by case, when an answer is likely to be correct. Crucially, these failures persist even when users are given default explanations. What improves matters is not making the model more accurate, but changing how epistemically relevant cues are communicated. When explanations are adapted to reflect the systemâs own confidenceâfor example, by explicitly signaling low, medium, or high certaintyâusers become better at judging when outputs are likely to be reliable, even though the answers themselves do not improve.
A second and deeper mistake is to misplace the normative burden. The claim that normative explanations are technically infeasible often assumes that the AI system alone must fulfil the function of determining what counts as the right reasons, as though it were a morally trustworthy agent. This misframes the issue. As we have argued, the proper site of normative responsibility lies not in the machine, but in the humans-in-the-loop and, crucially, in the institutional structures of accountability, oversight, and regulation that govern the use of AI-DSS. It is these structures that determine whether the technically feasible building blocks of normative explanation are integrated into a justificatory framework aligned with domain-specific norms and principles. Put differently, the AI can provide raw epistemic materials, and when it does this transparently in a manner conducive for our epistemic practices around justification, it does it in an epistemically trustworthy manner. Yet this status is never self-standing: epistemic trustworthiness depends on being embedded in a framework of collective accountability, where human agents and institutions determine whether these materials add up to genuine normative explanations that justify reliance on the systemâs outputs.
Consider a medical DSS that recommends a biopsy. The system can supply a reliability score (e.g., a 92% F1-score) and a counterfactual explanation showing that without the patientâs family history of cancer, the case would not have been flagged as high risk. Technically, both elements are feasible and already in development. However, they remain, and ought to remain, elements of decision support rather than decision replacement, taken up by the physician in consultation with colleagues, and then translated into a normative explanation.
Ultimately, the decision to perform the biopsy is not justified by the machine learning techniques. It is justified because acting on this information aligns with established medical principles and norms, which treat family history as a salient risk factor and require interventions to be grounded in reliable evidence. Crucially, further justification is secured because the decision is made within structures of oversight and accountability. In this way, technically feasible outputs are only fragments within a much larger decision architecture, where genuinely normative justification emerges through human judgment and review.
On the institutional side, the objection that the right to normative explanation would demand an unprecedented degree of transparency is compelling only insofar as it would oblige companies to reveal corporate secrets. In high-risk domains, for instance, the precise formula for classifying an applicant as high- or low-risk for investment may legitimately remain proprietary. But this does not entail that all explanatory content must be withheld, and related provisions are already included in the EU AI Act (see Article 78, âConfidentialityâ). What matters is distinguishing between those elements that are genuinely protected trade secrets and those that must be disclosed if individuals are to contest decisions that affect their rights. Moreover, scholars such as Bryson (2022) emphasize that transparency requirements can be structured as due diligence, akin to routine logging, without undermining innovation.
Taken together, the requirements set out by the Right to Explanation resonate with the RCC framework. Reasons, confidence, and counterfactuals are precisely the kinds of resources that make explanations clear, meaningful, and concerned with the main elements of the decision. They enable intelligibility, they allow decisions to be contested and acted upon, showing how outcomes would differ under relevantly different conditions. In this way, the RCC framework captures the general thrust of what is legally demanded, ensuring that epistemic trustworthiness is not just a philosophical ideal but grounded in European law.
4.3 Design Implications for AI-DSS
What follows, then, are the design implications of this framework. If the right to explanation is to be guaranteed in practice, human operators and deployers must be equipped with interfaces that make the systemâs reasons, confidence, and counterfactuals accessible at the right level of abstraction. The question now becomes one of humanâmachine interaction: how should information be presented so that operators can genuinely facilitate the affected personâs right to explanation, rather than being overwhelmed, misled, or deskilled?
In developing the RCC framework, we examined a wide range of empirical studies on how the presentation of reliability metrics shapes the humanâAI decision-making relationship Dorsch & Moll, 2024). This review revealed a consistent pattern: the value of reliability metrics depends as much on their raw informational content as on how and when they are presented. One of the clearest findings comes from work comparing propositional statements of confidence to more technical visualisations such as feature-importance graphs. Expressing a modelâs reliability in a clear, declarative sentenceâfor example, âThe model is 70% confident in this predictionââconsistently produced better reliance calibration than showing visual breakdowns of feature weights, even when participants had been trained to read such graphs. The advantage seems to stem from epistemic clarity: propositional statements map directly onto how people typically articulate and assess reasons in everyday epistemic practices.
That said, model confidence, by itself, does not automatically improve decision accuracy. For that reason, reliability metrics are most effective when paired with counterfactuals or concrete examples (e.g., Le et al., 2022). Several studies have shown that when confidence scores are presented in isolation, they can produce an anchoring effect, leading usersâparticularly those without strong domain knowledgeâto overweight the systemâs recommendation, even when it is wrong (Ma et al., 2023). Pairing reliability with a counterfactual (âIf this factor were different, the decision would changeâ) re-engages the userâs critical reasoning and reduces the risk of unreflective deference. This positive effect is amplified when reliability scores are revealed only after the human user has made a preliminary judgment (Chen et al., 2023). Thus, sequence matters: showing confidence scores too early can short-circuit independent reasoning, while withholding them until after the userâs own assessment can preserve deliberative engagement.
Hence, epistemically trustworthy AI involves the system being situated within a communicative and evaluative relationship with its operator, such that the factors determining its outputs are accessible and intelligible at the right level of abstraction. Mitchellâs (2025) type/token distinction can be instructive here.Footnote 8 Type transparency provides general, model-level information: reliability metrics, operational boundaries, etc. Token transparency provides case-specific explanations: why this output was produced for this input. Mitchellâs argument is that while type transparency can meaningfully support epistemic trust, token transparency can in some contexts undermine it. The worry is that if every output is exhaustively inspected, the relationship becomes one of constant auditing, devoid of trust. This is a real concern, especially in decision-support settings where human oversight capacity is limited.
Our own view is that both forms of transparency can be relevant, but their roles differ. Type transparency is essential for justifying a general stance of epistemic openness toward a system. It is what justifies users in treating the systemâs outputs as having prima facie validity. Token transparency, by contrast, is a tool for calibration and contestation: it enables users to override the system when specific outputs seem suspect or when domain knowledge suggests a mismatch. The challenge is thus a technical one, how best to design systems that provide both forms without overwhelming the userâs cognitive capacities.Footnote 9
Even with reliability and transparency in place, epistemic trust in AI-DSS remains fragile. This fragility stems from the fact that trustworthinessâeven epistemic trustworthinessâis not a property of a system or an agent alone but of the relationship embedded within a wider collective ecology. Without governance structures to ensure accountability, the same technical properties that might justify epistemic trust in principle can easily become vehicles for misplaced trust in practice. This fragility also helps explain the concerns raised earlier about unqualified appeals to âtrustworthy AIâ: when trust is left underspecified, it readily invites anthropomorphic interpretations, substitutes labels for normative analysis, and risks displacing responsibility onto technical systems. By contrast, restricting trust to the epistemic domain and operationalising it through the RCC framework is meant to avoid these pathologies by making explicit the conditions under which AI outputs may justifiably inform human reasoning.
To see what is at stake, consider an AI-DSS that appears to comply with the RCC framework and yet lacks the required oversight to ensure that the epistemic fragments it provides genuinely correspond to the mechanisms that produced its outputs. A related issue arises with LLMs, flagged above: when asked to explain why they produce a particular claim, the explanation is generated by the same probabilistic process that produced the claim itself, conditioned on that claim, rather than by reference to the actual evidential basis that would justify it. Hence, human oversight remains essential if epistemically trustworthy AI is to be achieved. Recognizing both the possibility and the limits of epistemically trustworthy AI sets the stage for our concluding claim: that while moral trustworthiness should be taken off the table, a disciplined account of epistemic trustworthiness should still guide the responsible design and governance of AI-DSS.
5 Conclusion
The debate over trustworthy AI often falters by importing moralized notions of trust into contexts where they cannot apply. Our central claim is that moral trustworthiness is neither feasible nor required for AI-DSS. Challenges from Franke and Mitchell do not overturn this claim but instead clarify its scope and implications. What emerges is a sharper distinction: where moral trust is a category mistake, epistemic trustworthinessâanchored in reliability, transparency, and collective governanceâremains a meaningful and necessary standard. The RCC framework and the EU AI Actâs right to explanation show how this standard can be operationalized in practice. Yet epistemic trustworthiness is fragile, depending not on AI systems alone but on the human and institutional contexts in which they are embedded. Recognizing both its promise and its limits allows us to replace a misleading discourse of âtrustworthy AIâ with a clearer focus on reliability, transparency, and collective responsibility in decision-making.
Data Availability
Not applicable.
Notes
As we argued in the original paper, agency in this context can be understood in a minimal sense: an entity need only be capable of resolving âmany-many problemsâ that arise when acting (Wu, 2013). On this minimal account, a typical action for an AI decision-support system is to make a recommendation (Dorsch & Deroy, 2024).
Here and throughout, we use âAIâ in a deliberately restricted sense to refer to machine learning algorithms, primarily for supervised and reinforcement learning, potentially including neural networks.
As laid down in our earlier paper, we understand moral trustworthiness as a form of sensitivity to moral norms, such that an agentâs behavior is guided by an appreciation of what is morally required and by reasons that are themselves morally appropriate, as opposed to mere rule-following. This sensitivity may take different forms. In some cases, it involves characteristic emotional responses and motivational dispositions through which violations or upholding of moral norms are experienced as salient and action-guiding. In other cases, it involves a more conceptual grasp of moral reasons, their justificatory force, and the normative significance of conforming to them, which presupposes a reflective, metacognitive capacity to evaluate oneâs own reasons as reasons. In both forms, moral trustworthiness consists not merely in conforming to moral norms, but in doing so for reasons that can be endorsed as genuinely moral, rather than merely tracking outcomes.
It is worth acknowledging that such a system could be desirable, since it might increase efficiency and accelerate processes. The relevant question, however, is whether morally trustworthy is necessary.
Readers may wonder why our response to Mitchell is comparatively brief. This mostly reflects the structure of his contribution: whereas Franke offers a sustained response to our argument, Mitchellâs paper is a standalone account of typeâtoken trust in which our position figures only peripherally. We discuss this type-token distinction later on.
Although our discussion focuses on AI systems, the account of epistemic trustworthiness we develop is not AI-specific. It is intended as a general account of the conditions under which it is reasonable for agents to integrate the outputs of a source or system into their own reasoning. Understood in this broad sense, such systems may include human informants, institutional sources (such as scientific bodies or media organisations), or technological artefacts more generally. AI systems constitute a particularly salient case, but not a theoretically exceptional one.
While Article 86 explicitly vests the right in the affected person (e.g., the applicant denied a loan), its realization presupposes enabling conditions on the side of the human operator. Operators must be in a position to request, obtain, and communicate intelligible accounts of the AI systemâs role and the main elements of the decision; otherwise, the affected personâs entitlement cannot be guaranteed.
A similar distinction can be found in our (Dorsch & Moll, 2024) differentiating between interpretability and explainability.
Note that systems need not fix one mode of expressing certainty; adaptive presentation can accommodate user preferences and cultural differences (Bang et al., 2014).
References
Alvarado, R. (2023a). What kind of trust does AI deserve, if any? AI Ethics, 3, 1169â1183. https://doi.org/10.1007/s43681-022-00224-x
Alvarado, R. (2023b). AI as an Epistemic Technology. Science and Engineering Ethics, 29, 32. https://doi.org/10.1007/s11948-023-00451-3
Apple, S. (2025). My Couples Retreat With 3 AI Chatbots and the Humans Who Love Them. Wired. https://www.wired.com/story/couples-retreat-with-3-ai-chatbots-and-humans-who-love-them-replika-nomi-chatgpt
Baier, A. (1986). Trust and antitrust. Ethics, 96(2), 231â260. https://doi.org/10.1086/292745
Bang, D., Fusaroli, R., TylĂ©n, K., Olsen, K., Latham, P. E., Lau, J. Y. F., Roepstorff, A., Rees, G., Frith, C. D., & Bahrami, B. (2014). Does interaction matter? Testing whether a confidence heuristic can replace interaction in collective decision-making. Consciousness and Cognition, 26, 13â23 https://doi.org/10.1016/j.concog.2014.02.002
Baria, A. T., & Cross, K. (2021). The brain is a computer is a brain: neuroscienceâs internal debate and the social significance of the Computational Metaphor. arXiv preprint arXiv:2107 14042. https://doi.org/10.48550/arXiv.2107.14042
Bender, E. M. (2024). Resisting Dehumanization in the Age of AI. Current Directions in Psychological Science, 33(2), 114â120. https://doi.org/10.1177/09637214231217286
Biden, J. R. (2023, October 30). Executive Order 14110: Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence. Federal Register. https://www.federalregister.gov/documents/2023/11/01/2023-24283/safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence
Bryson, J. (2018). AI & global governance: no one should trust AI. United Nations Centre for Policy Research. https://unu.edu/cpr/blog-post/ai-global-governance-no-one-should-trust-ai
Bryson, J. (2022). Belgian and Flemish policy makersâ guide to AI regulation. KCDSâCiTiP Fellow Lecture Series. https://data-en-maatschappij.ai/en/publications/paper-belgian-and-flemish-policy-makers-guide-to-ai-regulation
Carroll, N. G. (2025). Artificial Goodwill and Human Vulnerability: The Case for Building Merely Reliable, Rather than Trustworthy, Artificial Intelligence Technologies. Philosophy & Technology, 38, 65. https://doi.org/10.1007/s13347-025-00881-w
Centers for Disease Control and Prevention. (2025). Core elements of hospital diagnostic excellence (DxEx). U.S. Department of Health & Human Services. https://www.cdc.gov/patient-safety/hcp/hospital-dx-excellence/index.html
Chen, V., Liao, Q. V., Vaughan, J. W., & Bansal, G. (2023). Understanding the Role of Human Intuition on Reliance inHuman-AI Decision-Making with Explanations. Proceedings of the ACM on Human-Computer Interaction, 7(CSCW2), Article370. https://doi.org/10.1145/3610219
Choung, H., David, P., & Ross, A. (2022). Trust in AI and Its Role in the Acceptance of AI Technologies. International Journal of HumanâComputer Interaction, 39(9), 1727â1739. https://doi.org/10.1080/10447318.2022.2050543
Christian, B. (2020). The Alignment Problem: Machine learning and human values. WW Norton & Company.
Cohen, J. (1997). Deliberation and democratic legitimacy. In J. Bohman, & W. Rehg (Eds.), Deliberative democracy: Essays on reason and politics ,67â91. MIT Press.
Council of the European Union. (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 on artificial intelligence and amending certain legislative acts (Artificial Intelligence Act). Official Journal of the European Union, L(2024), 1â178. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A32024R1689
De-Arteaga, M., Fogliato, R., & Chouldechova, A. (2020). A case for humans-in-the-loop: Decisions in the presence of erroneous algorithmic scores. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems. https://doi.org/10.1145/3313831.3376638
Deroy, O. (2023). The Ethics of Terminology: Can we use human terms to describe AI? Topoi, 42(3), 881â889. https://doi.org/10.1007/s11245-023-09934-1
Deroy, O., Bacciu, D., Bahrami, B., Della Santina, C., & Hauert, S. (2024). Shared Awareness Across Domain-Specific Artificial Intelligence: An Alternative to DomainâGeneral Intelligence and Artificial Consciousness. Advanced Intelligent Systems, 6(10), 2300740. https://doi.org/10.1002/aisy.202300740
Dorsch, J. (2025). Mindshaping and AI: Will mindshaping a robot create an artificial person? In T. Zawidzki, & R. Tison (Eds.), The Routledge Handbook of Mindshaping ,406â417. Routledge.
Dorsch, J. & Deroy, O. (2024). Quasi-Metacognitive Machines: Why We Donât Need Morally Trustworthy AI and Communicating Reliability is Enough. Philosophy & Technology, 37, 62. https://doi.org/10.1007/s13347-024-00752-w
Dorsch J. & Deroy, O. (2025). The impact of labeling automotive AI as trustworthy or reliable on user evaluation and technology acceptance. Scientific Reports, 15, 1481. https://doi.org/10.1038/s41598-025-85558-2
Dorsch, J., Goddu, M. K., Nave, K., Vierkant, T. Coeckelbergh, M., GĂŒrtler, P., Urban, P., Spang, F., & Moll, M. (2025). Against AI Welfare: Care Practices Should Prioritize Living Beings Over AI. AI Magazine https://doi.org/10.1002/aaai.70016
Dorsch, J.& Moll, M. (2024). Explainable and human-grounded ai for decision support systems: The theory of epistemic quasi-partnerships. arXiv preprint. . arXiv:2409.14839. https://doi.org/10.48550/arXiv.2409.14839
Dunning, D., & Kruger, J. (1999). Unskilled and unaware of it: How difficulties in recognizing oneâs own incompetence lead to inflated self-assessments. Journal of Personality and Social Psychology, 77(6), 1121â1134. https://doi.org/10.1037/0022-3514.77.6.1121
Espinosa-Leal, L., & Stocchetti, M. (2025). On the Meaning of Trust, Reasons of Fear and the Metaphors of AI: Ideology, Ethics, and Fear. In Social Robots with AI: Prospects, Risks, and Responsible Methods ,227â237. IOS Press. https://doi.org/10.3233/FAIA241508
European Commission, High-Level Expert Group on Artificial Intelligence (2019, April 8). Ethics guidelines for trustworthy AI. Publications Office of the European Union. https://digital-strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai
Faulkner, P. (2011). Knowledge on trust. Oxford University Press. https://doi.org/10.1093/acprof:oso/9780199589784.001.0001
Fleming, S. M., & Lau, H. C. (2014). How to measure metacognition. Frontiers in Human Neuroscience. https://doi.org/10.3389/fnhum.2014.00443
Franke, U. (2024). The Limits of Calibration and the Possibility of Roles for Trustworthy AI. Philosophy & Technology, 37, 82. https://doi.org/10.1007/s13347-024-00771-7
Geiselmann, R., Tsourgianni, A., Deroy, O., & Harris, L. T. (2023). Interacting with agents without a mind: the case for artificial agents. Current Opinion in Behavioral Sciences, 51, 101282. https://doi.org/10.1016/j.cobeha.2023.101282
Habermas, J. (1990). Moral consciousness and communicative action (C. Lenhardt & S. W. Nicholsen, Trans.). MIT Press. (Original work published 1983)
Hardin, R. (2002). Trust and Trustworthiness. Russell Sage Foundation.
ISACA (2025). Using the Digital Trust Ecosystem Framework to Achieve Trustworthy AI [White paper]. ISACA. https://www.isaca.org/resources/white-papers/2024/using-dtef-to-achieve-trustworthy-ai
Kaminski, M. E., & Malgieri, G. (2025). The Right to Explanation in the AI Act. U of Colorado Law Legal Studies Research Paper No. 25 â 9. https://ssrn.com/abstract=5194301
Kurvers, R. H., Hertz, U., Karpus, J., Balode, M. P., Jayles, B., Binmore, K., & Bahrami, B. (2021). Strategic disinformation outperforms honesty in competition for social influence. IScience, 24(12). https://doi.org/10.1016/j.isci.2021.103505
Lahno, B. (2001). On the emotional character of Trust. Ethical Theory and Moral Practice, 4(2), 171â189. https://doi.org/10.1023/A:1011425102875
Le, T., Miller, T., Singh, R., & Sonenberg, L. (2022). Improving model understanding and trust with counterfactualexplanations of Model confidence. arXiv Preprint arXiv:220602790. https://doi.org/10.48550/arXiv.2206.02790
Lee, J. D., & See, K. A. (2004). Trust in Automation: Designing for Appropriate Reliance. Human Factors, 46(1), 50â80. https://doi.org/10.1518/hfes.46.1.50_30392
Long, R., Sebo, J., Butlin, P., Finlinson, K., Fish, K., Harding, J., Pfau, J., Sims, T., Birch, J., & Chalmers, D. (2024). Taking AI welfare seriously. arXiv preprint. https://doi.org/10.48550/arXiv.2411.00986
Longino, H. (1990). Science as social knowledge: Values and objectivity in scientific inquiry. Princeton University Press.
Ma, S., Lei, Y., Wang, X., Zheng, C., Shi, C., Yin, M., & Ma, X. (2023). Who Should I Trust: AI or Myself? LeveragingHuman and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-Making. Proceedings ofthe 2023 CHI Conference on Human Factors in Computing Systems Hamburg Germany. https://doi.org/10.1145/3544548.3581058
Mathew, D. E., Ebem, D. U., Ikegwu, A. C., Ukeoma, P. E., & Dibiaezue, N. F. (2025). Recent Emerging Techniques in Explainable Artificial Intelligence to Enhance the Interpretable and Understanding of AI Models for Human. Neural Processing Letters, 57(1), 16. https://doi.org/10.1007/s11063-025-11732-2
Mayer, R. C., Davis, J. H., & Schoorman, F. D. (1995). An integrative model of organizational trust. The Academy of Management Review, 20(3), 709â734. https://doi.org/10.2307/258792
Mitchell, T. (2025). Trust and Transparency in Artificial Intelligence. Philosophy & Technology, 38, 87. https://doi.org/10.1007/s13347-025-00916-2
Moll, M. & Dorsch, J. (2025). A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness. Human-Intelligent Systems Integration. https://doi.org/10.1007/s42454-025-00084-w
Moll, M., Karpus, J., & Bahrami, B. (2022). Do Artificial Agents Reproduce Human Strategies in the Advisersâ Game? In International Conference on Operations Research ,603â609. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-031-24907-5_72
Moran, R. (2005). Getting told and being believed. Philosophersâ Imprint, 5, 1â29. http://hdl.handle.net/2027/spo.3521354.0005.005
Nguyen, C. T. (2022). Trust as an unquestioning attitude. Oxford Studies in Epistemology, 7, 214â244.
Peirce, C. S. (1878). How to make our ideas clear. Popular Science Monthly, 12, 286â302.
Plato. (ca. 380 B.C.E./1997). Meno (G. M. A. Grube, Trans.). In J. M. Cooper & D. S. Hutchinson (Eds.), Plato: Complete works (pp.870â897). Hackett Publishing Company.
Popper, K. (1962). Conjectures and refutations: The growth of scientific knowledge. Routledge.
Richter, H. (2025). âIt saved my life.â The people turning to AI for therapy. Reuters. https://www.reuters.com/lifestyle/it-saved-my-life-people-turning-ai-therapy-2025-08-23/
Ryan, M. (2020). In AI we trust: Ethics, Artificial Intelligence, and reliability. Science and Engineering Ethics, 26(5), 2749â2767. https://doi.org/10.1007/s11948-020-00228-y
Steyvers, M., Tejeda, H., Kumar, A., Belem, C., Karny, S., Hu, X., & Smyth, P. (2025). What large language models know and what people think they know. Nature Machine Intelligence, 7(2), 221â231. https://doi.org/10.1038/s42256-024-00976-7
Surkova, E., Nikolayevskyy, V., & Drobniewski, F. (2020). False-positive COVID-19 results: hidden problems and costs. The Lancet Respiratory Medicine, 8(12), 1167â1168. https://doi.org/10.1016/S2213-2600(20)30453-7
Taylor, E. (2024). Explanation and the Right to Explanation. Journal of the American Philosophical Association, 10(3), 467â482. https://doi.org/10.1017/apa.2023.7
Terven, J., Cordova-Esparza, D. M., Ramirez-Pedraza, A., Chavez-Urbiola, E. A., & Romero-Gonzalez, J. A. (2023). Loss functions and metrics in deep learning. arXiv preprint arXiv, 230702694. https://doi.org/10.48550/arXiv.2307.02694
Wilholt, T. (2013). Epistemic trust in science. The British Journal for the Philosophy of Science, 64(2), 233â253. https://doi.org/10.1093/bjps/axs007
Wu, W. (2013). Mental Action and the threat of Automaticity. In A. Clark, J. Kiverstein, & T. Vierkant (Eds.), Decomposing the Will, 244â261. Oxford University Press.
Funding
Open Access funding enabled and organized by Projekt DEAL. This publication is an outcome of John Dorsch's participation in the project âEstablishing the Center for Environmental and Technology EthicsâPrague (CETE-P),â which has received funding from the European Unionâs HORIZON EUROPE Framework Pro-gramme under grant agreement no. 101086898/HORIZON-WIDERA-2022-TALENTS-01.
Author information
Authors and Affiliations
Contributions
John Dorsch prepared the original draft of the manuscript. Maximilian Moll and Ophelia Deroy provided critical revisions and substantive feedback. All authors reviewed and approved the final version of the manuscript.
Corresponding author
Ethics declarations
Ethical Approval
Not applicable.
Competing Interests
The authors declare no competing interests.
Consent for Publication
The authors consent to the publication of this manuscript.
Consent to Participate
Not applicable.
Additional information
Publisherâs Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the articleâs Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the articleâs Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
About this article
Cite this article
Dorsch, J., Moll, M. & Deroy, O. Why We Need Epistemically, Not Morally, Trustworthy AI: Decision Support and the Need for Collective Accountability. Philos. Technol. 39, 189 (2026). https://doi.org/10.1007/s13347-026-01166-6
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s13347-026-01166-6
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content â general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached â you'll always get the same 5 for this article.