ObjectiveAs large language models (LLMs) enter clinical decision support, concerns persist about sociodemographic bias. We assessed whether LLM recommendations for dizziness vary by patient descriptors and clinical detail.MethodsWe conducted a cross-randomized in-silico vignette study. One hundred synthetic emergency department dizziness cases were created using established diagnostic frameworks including the TiTrATE paradigm, SAEM GRACE-3 guidelines, and Bárány Society diagnostic criteria. Each vignette was tested in a neutral form and with 33 sociodemographic descriptor variants (34 total). Twelve instruction-tuned LLMs from multiple model families were evaluated. Models answered five binary clinical decision questions addressing etiology classification, triage disposition, neuroimaging, bedside vestibular examination, and mental health referral. Each model-vignette-descriptor combination was repeated 10 times, yielding 2,040,000 responses. Sociodemographic bias was quantified as descriptor-specific percentage-point deviations from neutral control recommendations with 95% confidence intervals.ResultsSociodemographic descriptors influenced LLM recommendations, with the largest differences observed for mental health referral decisions in diagnostically ambiguous cases. Referral likelihood was lower for Black transgender women (-12.2 pp; 95% CI -14.0 to -10.3), Black patients experiencing homelessness (-9.1 pp; -11.0 to -7.3), and patients experiencing homelessness (-7.7 pp; -9.5 to -5.9). Differences were attenuated when vignettes contained clearer diagnostic information. Other effects were smaller, including increased neuroimaging recommendations for low-income descriptors (+4.0 pp; 95% CI 2.1-5.8).ConclusionLLM clinical recommendations varied by sociodemographic descriptors, particularly under diagnostic uncertainty. More detailed clinical information reduced these disparities, suggesting structured inputs may mitigate bias in clinical AI systems.
Publications
2026
INTRODUCTION: The increasing popularity of oral nicotine pouches (ONPs) has raised significant public health concerns regarding their impact on oral health. These products deliver nicotine through mucosal membranes without combustion, prompting the need for comprehensive research into their effects on the oral mucosa. The objective of this study was to investigate the association between duration of nicotine pouch use and oral mucosal lesion severity.
METHODS: This was a cross-sectional study of a convenience sample of 42 male students and interns at King Abdulaziz University Faculty of Dentistry (KAUFD), Saudi Arabia, who used ONPs. Clinical examinations were conducted to assess the severity of oral mucosal lesions, which were categorized ordinally. Duration of ONP use was measured using a combined index [duration (months) × frequency (pouches/day)]. Ordinal logistic regression was used to evaluate the association between duration of ONP use and severity of oral mucosal lesions, adjusting for age, smoking status, and pouch placement behavior.
RESULTS: In adjusted analyses, oral mucosal lesion severity was not significantly associated with duration of ONP use (AOR=1.001; 95% CI: 1.000-1.002, p=0.161), age (AOR=0.77; 95% CI 0.21-2.87, p=0.697), or smoking status (AOR=0.68; 95% CI: 0.20-2.25, p=0.523).
CONCLUSIONS: Oral mucosal lesion severity was not associated with the duration or frequency of ONP use. Future research should focus on examining the long-term effects of nicotine pouch use on oral health to better understand and mitigate health risks.
Left atrial (LA) function has emerged as a critical determinant of cardiovascular performance and disease progression across the spectrum of pediatric heart disease. Traditionally viewed as a passive conduit, the left atrium is now increasingly recognized as a dynamic biomechanical chamber that modulates ventricular filling, pulmonary venous pressures, cardiac output, and neurohormonal signaling. In children with congenital and acquired heart disease, alterations in atrial structure and mechanics frequently precede overt ventricular dysfunction, making LA dysfunction an early and potentially reversible marker of cardiovascular injury. Advances in cardiac imaging, particularly speckle-tracking echocardiography and cardiac magnetic resonance imaging, have enabled detailed characterization of atrial reservoir, conduit, and contractile function in pediatric populations. Emerging evidence suggests that impaired left atrial strain is associated with adverse hemodynamics, exercise intolerance, arrhythmogenesis, pulmonary hypertension, Fontan failure, cardiomyopathy progression, and heart failure severity. Mechanobiological processes, including atrial stretch, fibrosis, extracellular matrix remodeling, inflammation, and altered atrioventricular coupling, contribute to the development of pediatric left atrial myopathy. These changes are increasingly recognized in congenital heart lesions, single-ventricle physiology, hypertrophic and dilated cardiomyopathies, Kawasaki disease, obesity-related cardiovascular dysfunction, and cardio-oncology populations. Despite growing interest, important gaps remain regarding normative pediatric reference values, age-related physiological variation, imaging standardization, and long-term prognostic significance. This narrative review summarizes current understanding of left atrial mechanobiology, multimodality imaging assessment, and clinical implications of LA dysfunction in pediatric heart disease. We also discuss emerging technologies, including artificial intelligence-based imaging analysis and future directions for integrating left atrial phenotyping into pediatric cardiovascular risk stratification and therapeutic decision-making.
Food adulteration is a widespread global issue threatening consumer health and food safety. The Grocery Manufacturers Association estimates that 10-15% of food products worldwide are adulterated, posing significant health risks and substantial economic losses. Traditional detection methods, such as chemical analysis and sensory evaluation, are often slow sometimes. Recent advancements in sensor technology offer faster and more accurate detection. Techniques such as gas chromatography and mass spectrometry offer high sensitivity, while optical sensors, like near-infrared spectroscopy, enable rapid and non-destructive analysis. Biosensors offer selective and portable detection, and electronic noses and tongues mimic human senses to identify adulterants. Sensor-based systems can reduce detection times from days to minutes and detect adulterants at very low levels. Although challenges remain, these technologies have great potential to enhance food safety, reduce fraud, and protect public health globally. The present review critically evaluates scientific literature published between 2000 and 2025, focusing on the development and application of sensor-based technologies for detecting food adulteration across diverse matrices. It highlights the use of chemical, biological, optical, and electronic sensors to combat food adulteration. In contrast to earlier reviews, this work provides a comparative assessment of the detection mechanism, sensitivity, specificity, and commercial applications, while identifying key challenges and future directions.
Hepatocellular carcinoma (HCC) exhibits metabolic heterogeneity that is not fully characterized by glycolysis-focused spectroscopic profiling. This study investigated whether in vitro hyperpolarized (HP) [2-13C]pyruvate NMR spectroscopy can identify a mitochondria-active HCC phenotype and assess its association with sensitivity to mitochondrial metabolic inhibition. HP [2-13C]pyruvate NMR spectroscopy was used to evaluate mitochondrial metabolism in McA-RH7777 HCC cells, with N1S1 cells serving as a glycolysis-dominant reference. Cell viability following treatment with the glutaminase inhibitor BPTES and the mitochondrial metabolic inhibitor CPI-613 was assessed by MTT assay, and metabolic changes following CPI-613 treatment were further evaluated using HP [2-13C]pyruvate. HP [2-13C]pyruvate demonstrated enhanced pyruvate-to-glutamate conversion in McA-RH7777 cells, whereas N1S1 showed minimal glutamate labeling. CPI-613 treatment resulted in a dose-dependent reduction in cell viability, while BPTES produced limited effects. Although pyruvate-to-glutamate conversion did not significantly decrease following CPI-613 treatment, pyruvate-to-lactate conversion increased, indicating metabolic adaptation. These findings demonstrate that HP [2-13C]pyruvate enables functional identification of a mitochondria-active HCC phenotype characterized by enhanced pyruvate-to-glutamate conversion. This approach may facilitate metabolic subtype classification, help identify tumors susceptible to mitochondrial metabolic inhibition, and enable non-invasive monitoring of treatment-induced metabolic adaptation.
BACKGROUND: Rare and deleterious variants in the leptin-melanocortin pathway can underlie severe childhood obesity, but data from East Asian populations remain limited.
METHODS: We conducted a single-center prospective observational cohort study of 188 Taiwanese children (4-18 years) with nonsyndromic obesity recruited between Jan 2017 and Dec 2024 for targeted sequencing of 12 leptin-melanocortin genes and compared them with two cohorts from the Taiwan Biobank: 527 normal‑weight adults presumed to have normal weight in childhood and 213 obese adults. We compared carrier proportions and mean variant counts and applied the Optimal Sequence Kernel Association Test (SKAT-O) across groups at the pathway and gene levels for three variant sets: all exonic variants, rare variants, and potentially influential variants (PIVs, defined as rare protein-truncating variants or missense variants with high deleteriousness scores). We also analyzed questionnaire-derived eating behavior scores using SKAT-O within the childhood obesity cohort.
RESULTS: Children with obesity carried a higher PIV burden than both normal-weight adults (41.0% vs. 24.1%; p < 0.0001) and obese adults (41.0% vs. 25.8%; p = 0.001); mean PIV counts were also higher. Pathway-level SKAT-O supported enrichment driven by rare variants and PIVs. Gene-level enrichment for LEPR and MAGEL2, and variant-level enrichment for two missense variants, MAGEL2 p.Gly285Arg and LEPR p.Asn128Lys, were observed in the childhood obesity group. Analyses of eating‑behavior scores showed nominal signals for exonic variants in MAGEL2 and PIVs in SIM1, but neither met the Bonferroni‑corrected threshold.
CONCLUSIONS: Rare, putatively deleterious leptin-melanocortin variants, particularly in LEPR and MAGEL2, are enriched in Taiwanese children with nonsyndromic obesity, underscoring the value of pathway-focused genetic evaluation in pediatric obesity.
OBJECTIVE: Large language models (LLMs) are increasingly evaluated for rheumatology tasks, but their performance in inflammatory arthritis remains unclear. We systematically reviewed LLM performance across clinical tasks in inflammatory arthritis.
METHODS: We conducted a systematic review (PROSPERO: CRD420261359100), searching PubMed, Scopus, and PubMed Central (January 2022 to April 2026) for studies evaluating LLM performance on clinical tasks in inflammatory arthritis. Two reviewers (Y.A., A.G.) screened 113 records.
RESULTS: Eighteen studies covered rheumatoid arthritis (n=3), ankylosing spondylitis/axial spondyloarthritis (n=7), psoriatic arthritis (n=2), gout (n=1), juvenile idiopathic arthritis (n=1), and multiple diseases (n=4). Most diseases and tasks were represented by only one to a few studies, and the evidence base remains earlystage and uneven across conditions. Over 20 distinct LLMs were evaluated, including ChatGPT-3.5 to ChatGPT-4o, Gemini 2.0, DeepSeek-R1/V3, Claude, and Perplexity; ChatGPT/GPT variants were the most frequently tested models (16 of 18 studies), so the current evidence base is predominantly GPT/ChatGPT-based. Findings spanned patient education (n=11), guideline adherence (n=6), clinical reasoning (n=3), and other applications (n=1). All readability assessments exceeded recommended thresholds. Guideline concordance ranged from 48% to 96%. Accuracy was lower for case-based clinical scenarios (4.24/6) than FAQ and guideline-based questions (5.32-5.36/6; p=0.044). When compared with real clinical data, agreement was poor (Cohen and Fleiss κ ≈ 0).
CONCLUSION: LLMs may support patient education, factual medication queries, and structured guideline questions when used under clinician review, but should not be used for case-based reasoning, treatment selection, or autonomous clinical decisions. None of the 18 included studies evaluated retrieval-augmented or agent-based systems, and none prospectively validated LLMs in clinical workflows. Safe integration in rheumatology will require purpose-built, knowledge-grounded systems and prospective evaluation before routine clinical use.
BACKGROUND: Radiology impressions guide clinical care. Large Language Models (LLMs)-drafted impressions can drift into generic, off-style text. Retrieval-augmented generation (RAG) enables context-aware few-shot prompting during inference.
METHODS: This retrospective IRB-approved study included 11,998 CT pulmonary angiography (CTPA) reports. We built a retrieval bank from 11,399 reports and reserved 599 reports for testing. GPT-4o and LLaMA 3.1-70B generated impressions from the "findings" section using three setups: zero-shot, fixed random few-shot, and dynamic retrieval-selected few-shot (top-k semantic matches; k = 3/5/10). We ran temperatures 0, 0.7, 1. We scored outputs against the original impressions with ROUGE and BERTScore F1, report mean scores with 95% confidence intervals, and tested for statistical significance using Wilcoxon signed-rank test.
RESULTS: Dynamic retrieval-based few-shot prompting outperformed zero-shot and fixed few-shot prompting across all configurations (all p < 0.05). The highest scores were observed at temperature 0 and k = 10. ROUGE-1 F1 increased to 0.44-0.47 for GPT-4o and 0.37-0.50 for LLaMA, versus 0.35-0.37 and 0.25-0.37, respectively, in zero-shot prompting. Lower temperature and larger k were associated with higher similarity scores.
CONCLUSIONS: Dynamic, case-matched retrieval improved alignment of LLM-generated CTPA impressions with reference impressions on automated text-similarity metrics. Scores remained moderate, and radiologists' verification is still required before clinical deployment.
We tested state-of-the-art LLMs under clinical-scale workloads using two designs: a single agent handling all tasks and a multi-agent orchestrator assigning each task to a dedicated worker. Across retrieval, extraction, and dosing tasks, batch sizes ranged from 5-80. Multi-agent accuracy remained high (90.6% at 5 tasks; 65.3% at 80), while single-agent accuracy collapsed (73.1% to 16.6%; p < 0.01). Multi-agent runs used up to 65-fold fewer tokens and limited latency growth. These findings show that lightweight orchestration preserves accuracy and efficiency under mixed-task clinical loads.