New research from Flinders University has found that advances in AI reasoning capabilities do not necessarily translate into fairer or more representative outputs in healthcare.
Researchers evaluated OpenAI’s o3-mini alongside DeepSeek-R1 by asking the models to generate 36,000 clinical vignettes describing fictional patients with common medical conditions. Both models frequently misrepresented the distribution of race and gender, reproducing patterns of stereotyping previously identified in earlier-generation large language models (LLMs).
The researchers found o3-mini met their threshold for significant racial misrepresentation in 78 per cent of conditions and gender misrepresentation in 56 per cent. For DeepSeek-R1, the figures were 89 per cent for race and 67 per cent for gender.
By comparison, GPT-4 had previously recorded significant misrepresentation in 67 per cent of conditions for both race and gender.
The models also disproportionately associated Black patients with conditions such as sarcoidosis, systemic lupus erythematosus, pre-eclampsia, and essential hypertension. Median racial misrepresentation for these conditions reached 44 per cent for o3-mini and 31 per cent for DeepSeek-R1, compared with 15 per cent for GPT-4.
Lead researcher Joshua Docking, from Flinders University’s College of Medicine and Public Health, said the findings highlighted the risk that AI could reinforce existing health biases.
“Large language models have the potential to transform healthcare but risk exacerbating health disparities if they perpetuate biases,” Docking said.
The researchers said the patterns could reflect underlying bias in training data, although the models may also be generating prototypical patient cases rather than attempting to reflect real-world demographic distributions.
Analysis of DeepSeek-R1’s reasoning traces found that the model explicitly invoked associations between diseases and demographic characteristics when selecting patient demographics, without referring to quantitative epidemiological data.
Professor Michael Sorich, research lead author and professor in clinical pharmacology at Flinders University, said the findings showed that improved reasoning did not automatically address representational bias.
“Despite having enhanced reasoning capabilities, the clinical outputs of o3-mini and DeepSeek-R1 still exhibit racial and gender disease stereotyping in common medical conditions,” Sorich said.
The researchers warned that repeatedly associating particular diseases with specific demographic groups could reinforce narrow assumptions in clinical environments, where understanding how conditions affect different populations is important for diagnosis and decision making.
They said organisations deploying LLMs in clinical workflows should therefore treat fairness and representation as ongoing governance issues rather than assuming that improvements in model capability will address them automatically.
“Awareness of these demographic defaults is essential for the safe integration of LLMs into clinical workflows, and continuous monitoring of potential biases should accompany their adoption,” the researchers said.
The research – “Evaluating the Potential of Reasoning Large Language Models to Perpetuate Racial and Gender Disease Stereotypes in Health Care” – was published in the Journal of Medical Internet Research.
Want to see more stories from trusted news sources?Make Cyber Daily a preferred news source on Google.