← Back to Home

Explainable Dermatology AI Is More Trustworthy—and May Also Be More Misleading: Study Reveals Automation Bias

An experiment involving 623 members of the general public and 153 primary care physicians showed that diagnostic AI designed to account for fairness across skin tones can improve overall performance; however, when the model is wrong, fluent LLM explanations are the most likely to steer non-expert users toward the wrong answer.

By SURL BioNews

If medical AI not only provides an answer but also offers a seemingly reasonable explanation, it should theoretically make it easier for people to judge whether it is trustworthy. However, a dermatology study published in *Nature Medicine* indicates that explanations can also amplify persuasiveness: they help users make judgments when the model is correct, but make it harder for people without professional knowledge to resist incorrect recommendations when the model is wrong.

The research team conducted two online experiments. A total of 623 participants without medical backgrounds identified melanomas and benign moles, while 153 primary care physicians provided differential diagnoses for atopic dermatitis, pityriasis rosea, Lyme disease, and cutaneous T-cell lymphoma. Each person evaluated 12 clinical images, split evenly between darker and lighter skin tones, and was randomly assigned to one of four types of AI assistance: predictions with confidence scores only, heatmaps highlighting image regions, similar cases, or textual explanations generated by a multimodal large language model. The study also included 320 medical students to help compare differences by experience.

The underlying diagnostic models were trained with fairness constraints. For the model used to classify melanomas and moles, the gap in balanced accuracy between darker and lighter skin tones narrowed from 9.1 percentage points with the baseline model to 2.1 percentage points; AI assistance also increased the general public’s mean accuracy from 69.7% to 75.8%. After physicians received AI recommendations, their diagnostic performance across the four major skin diseases likewise improved, while the pre-existing skin-tone gap narrowed. The results show that reducing unfairness cannot rely solely on explanations in the interface; what matters more is whether the model itself maintains relatively balanced performance across different skin tones.

The risk emerged when the AI made mistakes. Among members of the general public, LLM-generated textual explanations produced a 13.4-percentage-point improvement when predictions were correct, exceeding the gains from the other three presentation methods. But when predictions were incorrect, accuracy fell by 21.1 percentage points, also the largest decline. The researchers suggested that even when fluent explanations citing skin features are based on ambiguous or incomplete visual cues, they may still lead users to mistake something that “sounds true” for a correct judgment.

Primary care physicians were more resistant to incorrect recommendations. Under the various explanation methods, their final performance was barely affected by incorrect AI predictions. LLM explanations also did not produce the largest gains in accuracy; their main benefit was instead to bring physicians’ confidence levels into closer alignment with whether their answers were actually correct. Medical students were more inclined than experienced physicians to follow the AI, suggesting that professional training may provide a line of defense, though not an absolute safeguard. When the system displayed the AI’s answer before users made their own judgments, both groups generally became more inclined to follow it, and physicians also exhibited a more pronounced anchoring effect when presented with LLM explanations.

These results cannot yet be directly equated with outcomes in real clinical settings. The tasks in the two experiments were different, so the figures for the general public and physicians cannot be compared directly. Each participant viewed only 12 images and received no clinical context such as age, medical history, or symptoms. Test cases were also deliberately sampled to control the proportions of correct and incorrect AI answers, while the LLM explanations were influenced by specific prompts and a single-generation approach. The study therefore does not support the idea that “adding another paragraph of explanation makes the system safer.” Instead, it supports designing medical AI interfaces that allow users to form their own judgments first, clearly indicate uncertainty, and introduce review by a second professional when the AI substantially changes the initial judgment.

References

  1. Nature Medicine
  2. arXiv