Should we trust AI? The wrong question may be holding us back

Imagine you apply for a loan and are rejected. You ask why. The bank explains that their AI system – let’s call it SmartCredit – has assigned you a low credit score, but beyond pointing to the input data, nobody can tell you more. Not the bank. Not the software engineers. The system makes reliable predictions, yet its inner workings are simply too complex for any human to comprehend. How do you feel? Understandably frustrated, and probably deeply reluctant to trust the verdict.

Professor Mona Simion, Professor of Philosophy at the University of Oxford and Young Academy of Scotland member

Cases like this have fuelled a major movement in AI research: the push for explainable AI, or XAI. The core idea is straightforward – if people can’t understand how an AI reaches its conclusions, they won’t trust it, and if they won’t trust it, they won’t use it. Given how much AI has to offer (cancer-screening algorithms, for instance, have proven remarkably effective). The prescription from XAI advocates follows naturally: make AI systems transparent and trust will follow.

My team1 and I think this diagnosis, though understandable, is wrong – and that getting it wrong has practical consequences for how we think about building trustworthy AI.

Consider the following. You stop a stranger on the street and ask for directions. They tell you the nearest tube station is two streets away. You thank them and set off. At no point do you ask them to explain how they know this, let alone demand a justification of their answer. You simply trust them and, in ordinary circumstances, that’s entirely rational.

This points to something fundamental that the XAI movement tends to overlook: rational trust almost never requires explanation. When we defer to experts in fields we’ll never fully master – quantum physicists, oncologists, structural engineers – we are typically in a situation much like you were with the SmartCredit example above: the inner workings are, for us, a black box. Yet trusting expert testimony remains perfectly rational.

The real question, then, isn’t whether AI can explain itself. It’s whether AI can be trustworthy.

Reliability is not enough

Here things get philosophically interesting. Trustworthiness, most people agree, is more than mere reliability. Trust proper requires that the thing being trusted is not just reliably doing something, but evidence that it is disposed to fulfil an obligation to do it. This is where most philosophical accounts of trustworthiness run into trouble when applied to AI. Accounts that require goodwill, virtuous character, or sensitivity to the fact that someone is depending on you all presuppose features of inner mental life that current AI systems simply don’t have.

Our proposal cuts through this difficulty. AI systems are subject to what we call functional oughts – generated by norms derived from what they are designed, or have come, to do. A cancer-screening AI ought to detect cancer; that is its function. A credit-scoring system ought to assess creditworthiness accurately and fairly. These functional obligations are real, and a system with a robust disposition to meet them is, in a perfectly meaningful sense, trustworthy.

This reframing has practical upshots. If trustworthiness – not explainability – is the core concept, then the best way to increase rational trust in AI mirrors how we learn to trust any artefact: through exposure. We learn to trust a sharp knife by using it, consistently and successfully, in appropriate circumstances. The same applies to AI: allowing people to witness and interact with AI systems in lower stakes situations, where those systems can demonstrate their disposition to fulfil their function reliably, is likely to be more effective than demanding explanations that may be technically unattainable.

None of this is an argument that people should trust AI uncritically. Whether any given system deserves trust depends on whether it actually is trustworthy – and that is a question about design, deployment, and accountability. But by asking the right question, we put ourselves in a far better position to answer it.

References

1 ERC KnowledgeLab www.knowledgelabresearch. com/ is a European Research Council-funded project based at the University of Oxford and the University of Glasgow. For an academic article defending this claim at length, see Simion. M and Willard-Kyle, C. (forthcoming). Trusting AI: Explainability vs. Trustworthiness. Communication with AI: Philosophical Perspectives, (eds. Herman Cappelen and Rachel Sterken), Oxford University Press.


Professor Mona Simion, Professor of Philosophy at the University of Oxford and Young Academy of Scotland member

This article originally appeared in ReSourcE Spring/Summer 2026.

The RSE’s blog series offers personal views on a variety of issues. These views are not those of the RSE and are intended to offer different perspectives on a range of current issues.