Abstract
ObjectiveAs large language models (LLMs) enter clinical decision support, concerns persist about sociodemographic bias. We assessed whether LLM recommendations for dizziness vary by patient descriptors and clinical detail.MethodsWe conducted a cross-randomized in-silico vignette study. One hundred synthetic emergency department dizziness cases were created using established diagnostic frameworks including the TiTrATE paradigm, SAEM GRACE-3 guidelines, and Bárány Society diagnostic criteria. Each vignette was tested in a neutral form and with 33 sociodemographic descriptor variants (34 total). Twelve instruction-tuned LLMs from multiple model families were evaluated. Models answered five binary clinical decision questions addressing etiology classification, triage disposition, neuroimaging, bedside vestibular examination, and mental health referral. Each model-vignette-descriptor combination was repeated 10 times, yielding 2,040,000 responses. Sociodemographic bias was quantified as descriptor-specific percentage-point deviations from neutral control recommendations with 95% confidence intervals.ResultsSociodemographic descriptors influenced LLM recommendations, with the largest differences observed for mental health referral decisions in diagnostically ambiguous cases. Referral likelihood was lower for Black transgender women (-12.2 pp; 95% CI -14.0 to -10.3), Black patients experiencing homelessness (-9.1 pp; -11.0 to -7.3), and patients experiencing homelessness (-7.7 pp; -9.5 to -5.9). Differences were attenuated when vignettes contained clearer diagnostic information. Other effects were smaller, including increased neuroimaging recommendations for low-income descriptors (+4.0 pp; 95% CI 2.1-5.8).ConclusionLLM clinical recommendations varied by sociodemographic descriptors, particularly under diagnostic uncertainty. More detailed clinical information reduced these disparities, suggesting structured inputs may mitigate bias in clinical AI systems.