LLMs show bias in doctor recommendations, favoring female and minority names
TL;DR. A new audit reveals large language models exhibit biases in physician recommendations, favoring female and minority-signaled names. - LLMs act as 'AI infomediaries,' silently influencing patient choices by algorithmically recommending doctors. - Reputation signals like ratings heavily influenced recommendations, increasing choice probability by over 31 percentage points. - Demographic parity was rejected; female and minority-signaled names gained 1.3-2.9 percentage points over white names. - Models rarely mentioned gender or ethnicity in their explanations, making biases invisible through self-reporting mechanisms.
- LLMs act as 'AI infomediaries,' guiding patient choices for physicians.
- A randomized audit of seven LLMs (including gpt-4o-mini) revealed biases in doctor recommendations.
- Reputation (e.g., higher ratings) is the strongest factor, increasing choice probability by over 31 percentage points.
- Demographic biases exist: female-signaled names gained 2.5 percentage points, and Hispanic, South-Asian, and Black names gained 1.3-2.9 percentage points over White-signaled names.
- Models' internal explanations did not reveal these demographic biases, making self-reporting for transparency ineffective.
- The audit design is frozen and repeatable, allowing for ongoing behavioral assessment of new models.