Publications

2026

Lotan, Tamir, Mahmud Omar, Yiftach Barash, Jonathan B Kruskal, Muneeb Ahmed, Olga R Brook, Eyal Klang, and Alon Gorenshtein. (2026) 2026. “Reference Quality of OpenEvidence across Five Medical Specialties.”. Npj Health Systems 3 (1). https://doi.org/10.1038/s44401-026-00142-8.

OpenEvidence (OE) is a retrieval-augmented clinical AI tool. We evaluated the references OE returned to 150 standardized prompts across five specialties, analyzing all 4979 citations. No reference was confirmed as fabricated; three abstracts had attribution errors. Most references were recent and high-impact, though journal concentration varied by specialty. This study assessed reference existence and characteristics, not whether references support OE's clinical claims - a benchmark for evaluating clinical AI tools.

Jia, Eric, Mahmud Omar, Yiftach Barash, Olga R Brook, Muneeb Ahmed, Jonathan B Kruskal, Alon Gorenshtein, and Eyal Klang. (2026) 2026. “OpenEvidence Errs on the Safe Side in a Structured Test of Triage Recommendations.”. International Journal of Medical Informatics 222: 106687. https://doi.org/10.1016/j.ijmedinf.2026.106687.

BACKGROUND: Large language model (LLM) chatbots are increasingly consulted for triage decisions. A structured benchmark reported that ChatGPT Health, a consumer-facing health assistant, under-triaged 51.6 % of true emergencies and was susceptible to social anchoring. Whether a physician-facing clinical decision support platform fails similarly is unknown.

OBJECTIVE: To characterize the triage safety profile of OpenEvidence, a retrieval-augmented, physician-facing platform, using the identical benchmark previously applied to ChatGPT Health.

METHODS: We evaluated 60 clinician-authored vignettes from 30 clinical scenarios across 21 domains. Each scenario was written with and without objective clinical data and crossed with demographic and contextual modifiers in a 2 × 2 × 2 × 2 factorial design, yielding 960 prompts (480 clear-case, 480 edge-case). Responses were classified against a clinician gold standard as correct triage, under-triage, over-triage, or evidence-seeking refusal. Analyses used cluster bootstrap resampling, mixed-effects logistic regression, and Holm-Bonferroni correction.

RESULTS: Among 449 clear-case responses that returned a recommendation, accuracy was 71.3 %. OpenEvidence under-triaged 12.5 % of emergency presentations versus 51.6 % in the previously reported ChatGPT Health benchmark, and over-triaged 68.0 % of nonurgent Home presentations (ChatGPT Health, 64.8 %). Anchoring statements did not alter recommendations (OR = 1.08, 95 % CI 0.62-1.88; Holm-adjusted p = 1.0). Objective clinical data eliminated emergency under-triage (25 % to 0 %; p = 0.005) and reduced nonurgent over-triage (78.7 % to 57.8 %; p = 0.014). In 65 of 960 responses (6.8 %), the platform declined to assign a triage level, exclusively in symptom-only Home or Routine prompts.

CONCLUSIONS: Under this benchmark, OpenEvidence produced fewer missed emergencies than the historical ChatGPT Health comparison, while errors concentrated in over-triage and evidence-seeking refusal. These findings support evaluating health AI within its deployment context and treating refusal as a distinct output category whose clinical implications require separate assessment.

Scharf, Zachary, Miguel Muniz, Elizabeth Tchitchkan, Avina Rami, Praful K Ravi, Jacob J Orme, Ali Tarhini, Heather A Jacene, Daniel Sentana-Lledo, and Daniel S Childs. (2026) 2026. “An Assessment of Time Toxicity in Patients Receiving [177Lu]Lu-PSMA-617 for Metastatic Castration-Resistant Prostate Cancer (mCRPC).”. The Oncologist. https://doi.org/10.1093/oncolo/oyag327.

BACKGROUND: [177Lu]Lu-PSMA-617 improves survival and quality of life in men with metastatic castration-resistant prostate cancer but requires substantial healthcare involvement. Our analysis evaluated the time toxicity associated with [177Lu]Lu-PSMA-617, representing the first dedicated radioligand time toxicity assessment.

MATERIALS AND METHODS: We conducted a multi-institutional retrospective study of patients initiating [177Lu]Lu-PSMA-617 between April 2022 and March 2023. Patients were followed from first [177Lu]Lu-PSMA-617 administration until 3 months after the final cycle, initiation of new systemic therapy, or death. Time toxicity was quantified as healthcare contact days during the at-risk interval. Associations with baseline characteristics were evaluated using negative binomial regression models. Univariate and multivariable models were performed.

RESULTS: Among 135 patients, median age was 69 years and median prior systemic therapies was 4. Patients completed a median of 5 [177Lu]Lu-PSMA-617 cycles over 7.6 months. Median healthcare contact days were 20 (IQR, 15-28), representing 9.5% (IQR, 6.4-13.9) of at-risk days. Fifteen patients (11%) experienced high time toxicity (>1 contact day per 5 days), while 42 (31%) had low time toxicity (≤1 contact day per 14 days). Among the 57 patients who received both cycle 1 and cycle 6, mean per cycle healthcare contact days were 4 and 3.6, respectively (paired Wilcoxon signed-rank test, p = 0.04). Routine oncology visits comprised 32.7% of contact days, compared with 8.8% from hospitalizations. Lower baseline hemoglobin was independently associated with greater time toxicity (RR = 0.87; 95% CI, 0.79-0.96; p = 0.004).

CONCLUSION: In heavily pretreated patients, [177Lu]Lu-PSMA-617 demonstrated relatively low time toxicity, supporting it as a time-efficient therapy.

Omar, Mahmud, Mohammad E Naffaa, Reem Agbareia, Fadi Hassan, Abdulla Watad, Helana Jeries, Alon Gorenshtein, et al. (2026) 2026. “AI-Assisted Rheumatology Triage Changes With Referral Framing.”. Rheumatology (Oxford, England). https://doi.org/10.1093/rheumatology/keag414.

OBJECTIVES: Two in three US physicians now use healthcare AI, and large language models (LLMs) are entering the triage workflows that determine which patients reach rheumatology and how quickly. We aimed to test whether nine prespecified cues in referral notes, patient descriptions and demographics shift AI-assisted triage decisions when the underlying clinical information is unchanged.

METHODS: We conducted a controlled, physician-validated experiment across 30 physician-authored rheumatology vignettes, each independently rephrased three times (90 case variants). We tested five LLMs from three providers - Anthropic, Google, and OpenAI - under nine dimensions spanning demographics, clinical context, and communication framing, with 57 controlled contextual modifications, four system-prompt personas, and five repetitions per cell, yielding more than 200,000 model queries. Sixteen clinical fields were graded against physician-validated ground truth, with excellent inter-rater agreement (Fleiss' kappa=0.92).

RESULTS: Baseline composite concordance with expert ground truth was high at 0.869. We had expected sociodemographic cues to be the strongest source of distortion. Instead, the largest shifts came from how the case was framed and described. When patients were described as anxious, models attributed symptoms to psychological rather than organic causes nearly three times as often as at baseline (13.1% vs 4.5%; odds ratio 3.2 versus stoic framing), a shift that risks relabeling organic disease as functional. Clinician anchoring in the referral note reduced concordance, consistently across models and rephrasings and significantly for acuity (dismissive anchor, vignette-level p = 0.01), mainly by downgrading urgency. In contrast, race or ethnicity, socioeconomic status, and language barrier produced no detectable effect, including in mixed-effects models that accounted for repeated vignette use.

CONCLUSION: Although baseline concordance with specialist ground truth was high, it was readily disrupted by how referral notes were worded and how patients described their symptoms, not by patient demographics. Before AI-assisted triage enters rheumatology referral pathways, systems should separate objective clinical evidence from interpretive framing, and urgency and psychological attribution should be treated as auditable safety signals.

Sorin, Vera, and Eyal Klang. (2026) 2026. “AI Scribe Safety: Measuring What Happens After Signing.”. JMIR Medical Informatics 14: e103162. https://doi.org/10.2196/103162.

Coiera and Fraile-Navarro question whether AI scribes are being evaluated on metrics that truly impact care. While current evaluations focus on the quality of the initial draft, signed clinical notes are dynamic, as their content can be copied, summarized, coded, and re-ingested by downstream AI tools. We argue that safety must be measured downstream, focusing on how small errors in initial documentation can compound across the patient's electronic health record.

Mudrik, Aya, Ayala Dodge, Alon Moore Galindo, Girish N Nadkarni, Shelly Soffer, and Eyal Klang. (2026) 2026. “Sex-Related Performance Disparities in Convolutional Neural Networks for Imaging-Based Assessment of Coronary Atherosclerosis and Ischemic Heart Disease: A Systematic Review.”. Journal of Imaging Informatics in Medicine. https://doi.org/10.1007/s10278-026-02173-x.

Sex-related disparities persist in the diagnosis and management of ischemic heart disease (IHD), raising concern that convolutional neural networks (CNNs) used in coronary imaging may perpetuate these inequities. This systematic review evaluated sex-related performance differences in CNN-based models using medical imaging to assess coronary atherosclerosis, coronary artery disease, or myocardial ischemia. A systematic literature search of PubMed, Web of Science, Scopus, IEEE Xplore, and ACM Digital Library was conducted through June 7, 2026, in accordance with PRISMA guidelines. Peer-reviewed original studies were included if they evaluated sex-related performance of CNN-based models using medical imaging inputs, either by reporting performance metrics separately for women and men or by assessing sex as a determinant of model error, misclassification, calibration, or agreement. Nine studies met the inclusion criteria, covering noncontrast cardiac CT, coronary CT angiography, PET-CT, chest radiography, and SPECT-based approaches. Overall, model performance was generally similar between men and women across imaging modalities. Nevertheless, several studies identified clinically relevant sex-related discrepancies, including higher error rates in men for PET-CT calcium scoring and sex-related differences in error patterns during automated CCTA interpretation. Notably, one SPECT-based study showed that targeted augmentation of training data improved calibration and reduced false-positive predictions in women. While CNNs generally show comparable performance between sexes, subtle sex-related biases can persist, often driven by data imbalance and reference standard limitations. Sex-stratified evaluation and bias-mitigation strategies are essential to ensure equitable clinical implementation.

Tessler, Idit, Mahmud Omar, Amit Wolfovitz, Yoav Gimmon, Sholem Hack, Noa Rozendorn, Nir Livneh, and Eyal Klang. (2026) 2026. “Sociodemographic Bias in LLMs’ Clinical Decision-Making for Dizziness.”. Journal of Vestibular Research : Equilibrium & Orientation, 9574271261474218. https://doi.org/10.1177/09574271261474218.

ObjectiveAs large language models (LLMs) enter clinical decision support, concerns persist about sociodemographic bias. We assessed whether LLM recommendations for dizziness vary by patient descriptors and clinical detail.MethodsWe conducted a cross-randomized in-silico vignette study. One hundred synthetic emergency department dizziness cases were created using established diagnostic frameworks including the TiTrATE paradigm, SAEM GRACE-3 guidelines, and Bárány Society diagnostic criteria. Each vignette was tested in a neutral form and with 33 sociodemographic descriptor variants (34 total). Twelve instruction-tuned LLMs from multiple model families were evaluated. Models answered five binary clinical decision questions addressing etiology classification, triage disposition, neuroimaging, bedside vestibular examination, and mental health referral. Each model-vignette-descriptor combination was repeated 10 times, yielding 2,040,000 responses. Sociodemographic bias was quantified as descriptor-specific percentage-point deviations from neutral control recommendations with 95% confidence intervals.ResultsSociodemographic descriptors influenced LLM recommendations, with the largest differences observed for mental health referral decisions in diagnostically ambiguous cases. Referral likelihood was lower for Black transgender women (-12.2 pp; 95% CI -14.0 to -10.3), Black patients experiencing homelessness (-9.1 pp; -11.0 to -7.3), and patients experiencing homelessness (-7.7 pp; -9.5 to -5.9). Differences were attenuated when vignettes contained clearer diagnostic information. Other effects were smaller, including increased neuroimaging recommendations for low-income descriptors (+4.0 pp; 95% CI 2.1-5.8).ConclusionLLM clinical recommendations varied by sociodemographic descriptors, particularly under diagnostic uncertainty. More detailed clinical information reduced these disparities, suggesting structured inputs may mitigate bias in clinical AI systems.