Trust Calibration in Human–AI Decision Support: Confidence Displays, Failure-Mode Explanation and Appropriate Reliance
Olivia Whitmore1 · Hannah Beaumont2 · Marcus Doyle1
- 1 University of Essex, United Kingdom
- 2 Toronto Metropolitan University, Canada
Identifiers: this article has no DOI. JIPCET does not yet deposit metadata with a DOI registration agency, so we publish a JIPCET article reference and a permanent article URL instead of an identifier that would not resolve. Please cite the URL below. Registration and indexing status is described on the peer review and publishing page.
Abstract
Background. Decision-support interfaces routinely display model confidence on the assumption that numerical uncertainty enables users to rely on advice selectively. Evidence that confidence alone improves the quality of reliance decisions is mixed. Objective. We test whether confidence displays support appropriate reliance — accepting correct advice and rejecting incorrect advice — and whether pairing confidence with an explanation of the model's characteristic failure modes changes that relationship. Methods. Two pre-registered within-subjects experiments (total n = 214; Study 1 n = 96 clinicians and nurses, Study 2 n = 118 insurance adjudicators) used a triage task with 48 cases per participant. We manipulated confidence display (absent, numeric, verbal) and failure-mode explanation (absent, present), holding model accuracy at 0.78 with a stratified error distribution. Primary outcomes were appropriate-reliance rate and a signal-detection measure of reliance discrimination; we also recorded decision time and self-reported trust. Results. Numeric confidence alone raised overall agreement with the model from 0.61 to 0.74 but did not improve reliance discrimination, because agreement rose on incorrect advice as well as correct advice — over-reliance rather than calibration. Adding failure-mode explanation improved discrimination substantially and reduced agreement on incorrect high-confidence advice by 19 percentage points. The interaction between confidence and explanation was significant in both studies and larger among more experienced participants. Explanation added a mean 4.1 seconds per decision. Conclusion. Confidence information is not self-interpreting. Interfaces should pair uncertainty with an account of when the model is characteristically wrong; displaying confidence without it measurably increases misplaced reliance.
Keywords human-AI interaction · trust calibration · decision support · appropriate reliance · explainability · uncertainty communication
Key research findings
- Numeric confidence raised agreement with the model by 13 points without improving the ability to tell correct from incorrect advice.
- Confidence paired with failure-mode explanation cut agreement on incorrect high-confidence advice by 19 points.
- The benefit of explanation was larger for more experienced practitioners, not smaller.
- Explanation cost an average of 4.1 seconds per decision — a real but modest price for better reliance.
Cite this research article
Olivia Whitmore, Hannah Beaumont, Marcus Doyle. Trust Calibration in Human–AI Decision Support: Confidence Displays, Failure-Mode Explanation and Appropriate Reliance. Journal of Innovation, Product, Computing & Emerging Technologies (JIPCET). 2026;3(1):5–34. https://jipcet.org/articles/jipcet-2026-0036
1.Introduction
Appropriate reliance, not maximal reliance, is the design target for decision aids. A support system that increases agreement with its own advice has improved nothing if the additional agreement is distributed evenly across correct and incorrect recommendations; in safety-relevant settings it has made matters worse, because the errors it introduces are now endorsed by a human decision-maker.
Confidence displays are the standard interface response to this problem. Their implicit theory is that users can convert a probability into a reliance decision. That conversion requires knowing not only how often the model is wrong but where — which cases lie in its weak region. A calibrated confidence number does not carry that information, and our central hypothesis is that without it, confidence functions as an endorsement cue rather than an uncertainty cue.
We test this in two experiments with domain practitioners rather than crowdworkers, because reliance behaviour depends on the participant having independent competence to exercise.
2.Related Work
The reliance literature distinguishes trust as an attitude from reliance as a behaviour, and has repeatedly found the two to dissociate. Measures of self-reported trust respond readily to interface changes that do not alter decision quality, which is why we treat discrimination, not agreement or attitude, as the primary outcome.
Studies of explanation in decision support report benefits that vary with explanation type. Feature-attribution explanations often increase agreement without improving discrimination, mirroring what we find for confidence alone. Explanations that describe model limitations rather than justify individual outputs have been proposed as a remedy but have been tested mainly in laboratory tasks with lay participants; our contribution is a pre-registered test with practitioners and a stratified error distribution designed to make over-reliance detectable.
3.Method
Participants. Study 1 recruited 96 clinicians and senior nurses from three hospital systems; Study 2 recruited 118 insurance adjudicators from two firms. Both samples had a median of over six years of relevant experience. Both studies were pre-registered and approved by the institutional ethics committees at both universities.
Task and materials. Participants triaged 48 cases into one of four priority bands, receiving a model recommendation on each. The model operated at 0.78 accuracy with errors deliberately stratified so that a third of them occurred at high stated confidence — the condition under which over-reliance is consequential and which naturally occurring error distributions make rare enough to be hard to measure.
Design. Confidence display (absent, numeric percentage, verbal band) and failure-mode explanation (absent, present) were varied within subjects in a counterbalanced Latin-square order. The failure-mode explanation was a persistent two-sentence statement of the case characteristics on which the model was known to underperform, written from held-out validation data and identical across participants.
Measures and analysis. Appropriate reliance was scored per decision as accepting correct advice or overriding incorrect advice. We computed a signal-detection discrimination index over reliance decisions, and fitted mixed-effects models with random intercepts for participant and case. Decision time and a standard trust scale were recorded as secondary outcomes. Analyses followed the pre-registered plan; two exploratory analyses are labelled as such.
4.Results
Confidence alone. Adding numeric confidence raised agreement with the model from 0.61 to 0.74. Discrimination did not improve in either study: agreement rose by comparable amounts on correct and incorrect recommendations. Verbal confidence bands produced a smaller version of the same pattern. This is the signature of over-reliance, and self-reported trust rose in step with agreement, which is why attitude measures would have reported this condition as a success.
Confidence with explanation. In the paired condition, discrimination improved significantly in both studies, and agreement on incorrect high-confidence recommendations fell by 19 percentage points. Agreement on correct recommendations was essentially unchanged, so the improvement came from selective rejection rather than general scepticism — the outcome the design is meant to produce.
Interaction and experience. The confidence-by-explanation interaction was significant in both studies. In an exploratory analysis, the benefit of explanation increased with participant experience: more experienced practitioners had the domain knowledge to act on a statement about the model's weak region, whereas less experienced participants more often deferred regardless of condition.
Costs. Explanation added a mean 4.1 seconds per decision (median 3.4 s). Participants reported the persistent format as unobtrusive, and no condition produced a detectable change in overall task completion.
5.Discussion
The results argue against treating confidence display as an uncertainty-communication solution. A probability tells a user how often to doubt; it does not tell them when, and in our data users resolved that ambiguity by treating high confidence as endorsement. Failure-mode explanation supplies the missing 'when', and does so cheaply because it is written once from validation data rather than generated per case.
For practice, three recommendations follow. Report discrimination rather than agreement when evaluating decision-support interfaces. Do not ship a confidence display without an accompanying statement of characteristic failure. And revisit that statement whenever the model is retrained, since it is the component of the interface that encodes model-specific truth.
The experience finding complicates a common assumption that explanation mainly assists novices. In a task where the human must exercise independent judgement, explanation was most useful to those best able to use it.
6.Limitations
Both studies used a stratified error distribution to make high-confidence errors measurable; effect sizes therefore should not be read as expected field magnitudes, though the direction and interaction should transfer. Participants knew they were in a study and bore no consequence for error, which likely raises override willingness relative to practice.
We tested one style of failure-mode explanation, held constant across participants. Case-specific or interactive variants may behave differently, and our data cannot separate the effect of content from the effect of its persistent placement. Longitudinal effects — whether the benefit survives months of routine exposure — remain untested and are our next study.
7.Conclusion
Across 214 practitioners and 10,272 decisions, confidence displays increased reliance on AI advice without improving the ability to distinguish good advice from bad. Pairing confidence with a short, persistent account of the model's characteristic failure modes improved that discrimination and cut endorsement of incorrect high-confidence recommendations by 19 percentage points, at a cost of roughly four seconds per decision. Confidence information is necessary but not sufficient; interfaces should state where the model is weak, not only how sure it is.
8.References
- [1] Whitmore, O. (2025). Reliance and calibration in clinical decision aids. Human Factors 67(2), 210–229.
- [2] Beaumont, H. & Doyle, M. (2025). Discrimination measures for reliance research. ACM Transactions on Computer–Human Interaction 32(4), 1–29.
- [3] Doyle, M. (2024). Trust attitudes and reliance behaviour: a dissociation. Ergonomics 67(9), 1188–1205.
- [4] Beaumont, H. & Whitmore, O. (2025). Surface completeness and reviewer judgement. Human–Computer Interaction 40(4), 388–412.
- [5] Paxton, D. & Okonkwo, R. (2026). Failure taxonomies as interface content. JIPCET 3(2), 1–40.
- [6] Hale, E. (2025). Uncertainty communication in operational settings. Safety Science 181, 106–124.
