Publications

2026

Dou, Qianhui, Sophia M Mirrione, Karen Dos Santos, Samuel R Calos, Aaron K Grant, Yi-Fen Yen, and Leo L Tsai. (2026) 2026. “Hyperpolarized [2-13C]Pyruvate Identifies a Mitochondria-Active Hepatocellular Carcinoma Phenotype With Vulnerability to Mitochondrial Inhibition.”. NMR in Biomedicine 39 (9): e70360. https://doi.org/10.1002/nbm.70360.

Hepatocellular carcinoma (HCC) exhibits metabolic heterogeneity that is not fully characterized by glycolysis-focused spectroscopic profiling. This study investigated whether in vitro hyperpolarized (HP) [2-13C]pyruvate NMR spectroscopy can identify a mitochondria-active HCC phenotype and assess its association with sensitivity to mitochondrial metabolic inhibition. HP [2-13C]pyruvate NMR spectroscopy was used to evaluate mitochondrial metabolism in McA-RH7777 HCC cells, with N1S1 cells serving as a glycolysis-dominant reference. Cell viability following treatment with the glutaminase inhibitor BPTES and the mitochondrial metabolic inhibitor CPI-613 was assessed by MTT assay, and metabolic changes following CPI-613 treatment were further evaluated using HP [2-13C]pyruvate. HP [2-13C]pyruvate demonstrated enhanced pyruvate-to-glutamate conversion in McA-RH7777 cells, whereas N1S1 showed minimal glutamate labeling. CPI-613 treatment resulted in a dose-dependent reduction in cell viability, while BPTES produced limited effects. Although pyruvate-to-glutamate conversion did not significantly decrease following CPI-613 treatment, pyruvate-to-lactate conversion increased, indicating metabolic adaptation. These findings demonstrate that HP [2-13C]pyruvate enables functional identification of a mitochondria-active HCC phenotype characterized by enhanced pyruvate-to-glutamate conversion. This approach may facilitate metabolic subtype classification, help identify tumors susceptible to mitochondrial metabolic inhibition, and enable non-invasive monitoring of treatment-induced metabolic adaptation.

Adiniaev, Yosef, Mahmud Omar, Tohar M Timor, Yiftach Barash, Olga R Brook, Mohammad E Naffaa, Alon Gorenshtein, and Eyal Klang. (2026) 2026. “ChatGPT and Other Large Language Models in Inflammatory Arthritis: A Systematic Review Across Clinical Tasks.”. The Journal of Rheumatology. https://doi.org/10.3899/jrheum.2026-0444.

OBJECTIVE: Large language models (LLMs) are increasingly evaluated for rheumatology tasks, but their performance in inflammatory arthritis remains unclear. We systematically reviewed LLM performance across clinical tasks in inflammatory arthritis.

METHODS: We conducted a systematic review (PROSPERO: CRD420261359100), searching PubMed, Scopus, and PubMed Central (January 2022 to April 2026) for studies evaluating LLM performance on clinical tasks in inflammatory arthritis. Two reviewers (Y.A., A.G.) screened 113 records.

RESULTS: Eighteen studies covered rheumatoid arthritis (n=3), ankylosing spondylitis/axial spondyloarthritis (n=7), psoriatic arthritis (n=2), gout (n=1), juvenile idiopathic arthritis (n=1), and multiple diseases (n=4). Most diseases and tasks were represented by only one to a few studies, and the evidence base remains earlystage and uneven across conditions. Over 20 distinct LLMs were evaluated, including ChatGPT-3.5 to ChatGPT-4o, Gemini 2.0, DeepSeek-R1/V3, Claude, and Perplexity; ChatGPT/GPT variants were the most frequently tested models (16 of 18 studies), so the current evidence base is predominantly GPT/ChatGPT-based. Findings spanned patient education (n=11), guideline adherence (n=6), clinical reasoning (n=3), and other applications (n=1). All readability assessments exceeded recommended thresholds. Guideline concordance ranged from 48% to 96%. Accuracy was lower for case-based clinical scenarios (4.24/6) than FAQ and guideline-based questions (5.32-5.36/6; p=0.044). When compared with real clinical data, agreement was poor (Cohen and Fleiss κ ≈ 0).

CONCLUSION: LLMs may support patient education, factual medication queries, and structured guideline questions when used under clinician review, but should not be used for case-based reasoning, treatment selection, or autonomous clinical decisions. None of the 18 included studies evaluated retrieval-augmented or agent-based systems, and none prospectively validated LLMs in clinical workflows. Safe integration in rheumatology will require purpose-built, knowledge-grounded systems and prospective evaluation before routine clinical use.

Sorin, Vera, Jeremy D Collins, Lewis D Hahn, Alex K Bratt, Eyal Klang, and Panagiotis Korfiatis. (2026) 2026. “Case-Matched Retrieval Improves Textual Alignment of LLM-Generated Radiology Impressions.”. PloS One 21 (7): e0354688. https://doi.org/10.1371/journal.pone.0354688.

BACKGROUND: Radiology impressions guide clinical care. Large Language Models (LLMs)-drafted impressions can drift into generic, off-style text. Retrieval-augmented generation (RAG) enables context-aware few-shot prompting during inference.

METHODS: This retrospective IRB-approved study included 11,998 CT pulmonary angiography (CTPA) reports. We built a retrieval bank from 11,399 reports and reserved 599 reports for testing. GPT-4o and LLaMA 3.1-70B generated impressions from the "findings" section using three setups: zero-shot, fixed random few-shot, and dynamic retrieval-selected few-shot (top-k semantic matches; k = 3/5/10). We ran temperatures 0, 0.7, 1. We scored outputs against the original impressions with ROUGE and BERTScore F1, report mean scores with 95% confidence intervals, and tested for statistical significance using Wilcoxon signed-rank test.

RESULTS: Dynamic retrieval-based few-shot prompting outperformed zero-shot and fixed few-shot prompting across all configurations (all p < 0.05). The highest scores were observed at temperature 0 and k = 10. ROUGE-1 F1 increased to 0.44-0.47 for GPT-4o and 0.37-0.50 for LLaMA, versus 0.35-0.37 and 0.25-0.37, respectively, in zero-shot prompting. Lower temperature and larger k were associated with higher similarity scores.

CONCLUSIONS: Dynamic, case-matched retrieval improved alignment of LLM-generated CTPA impressions with reference impressions on automated text-similarity metrics. Scores remained moderate, and radiologists' verification is still required before clinical deployment.

Gorenshtein, Alon, Yosef Adiniaev, Mahmud Omar, Yiftach Barash, Eyal Klang, and Oved Daniel. (2026) 2026. “The Unsteady Return of Bedside Motor Command-Following After Acute Brain Injury.”. Neurocritical Care. https://doi.org/10.1007/s12028-026-02624-x.

BACKGROUND/OBJECTIVE: Following a verbal command marks the bedside transition from unresponsiveness to overt recovery of consciousness after acute brain injury. Its timing across phenotypes, stability once present, and dependence on sedation are uncharacterized at scale.

METHODS: Retrospective cohort of adults with acute brain injury, first intensive care unit stay, in the Medical Information Mart for Intensive Care IV (MIMIC-IV). Command-following was the Glasgow Coma Scale motor response "Obeys Commands." Among patients not following commands at admission, cumulative incidence was estimated with death or hospice and discharge without recovery as competing events. Instability was quantified as transient first recovery and threshold crossings; examinations were tagged for concurrent sedation. Principal findings were externally validated in the multicenter eICU Collaborative Research Database.

RESULTS: Of 13,900 brain-injured patients with three or more motor examinations, 5498 (39.6%) were not following commands at admission. The cumulative incidence of first command-following was 43.5% by 24 h and 65.0% by 14 days, ranging at 14 days from 36.9% in anoxic injury to 77.2% in ischemic stroke (anoxic versus ischemic stroke at 72 h, difference 0.41; adjusted P = .002). Among 3573 patients who recovered, the first recovery was transient in 22.2%, and 62.4% crossed the threshold repeatedly. Nonfollowing was strongly associated with sedation, consistent with an arousal-dependent examination. In eICU, the 14-day incidence was 64.8%, and transient first recovery was 22.7%, closely matching the primary cohort.

CONCLUSIONS: After acute brain injury, overt bedside command-following returns early but unsteadily, with phenotype-dependent timing, threshold fluctuation, and strong dependence on sedation. A single charted observation is an unreliable index of the underlying state.

Chauhan, Aman, Heidi L Weiss, Charles Kunos, Rani Jayswal, Bhavana Konda, Heloisa P Soares, Daneng Li, et al. (2026) 2026. “Multi-Center Phase 1 Study of Triapine in Combination With [177Lu]Lutetium-Dotatate in Patients With Well-Differentiated Neuroendocrine Tumors (ETCTN 10388).”. Clinical Cancer Research : An Official Journal of the American Association for Cancer Research. https://doi.org/10.1158/1078-0432.CCR-26-0686.

INTRODUCTION: Next generation of Theranostics studies are looking at novel combinations with radiation sensitizers to enhance PRRT's efficacy. Here, we present safety and efficacy data from a multi-center phase 1 study of a [¹⁷⁷Lu]lutetium-dotatate -triapine combination in patients with GEP-NETs.

METHODS: The study consisted of two phases (dose escalation and expansion) administering the recommended phase 2 doses (RP2D) for the two agents. All patients received 200 mCi of [¹⁷⁷Lu]lutetium-dotatate on day 1 of four 8-week cycles, in combination with oral triapine administered at assigned dose levels (100 mg, 150 mg or 200mg) from days 1 to 14 of each cycle.

RESULTS: A total of 31 (A: 15 ; B: 16 ) patients received treatment. Nine patients experienced dose-limiting toxicities; Grade 3 anemia (n=1), Grade 5 cardiac arrest (n=1), and Grade 4 neutropenia (n=7). Based on the overall safety and pharmacokinetics data, a [¹⁷⁷Lu]lutetium-dotatate (200 mCi) plus triapine (150 mg) dose was selected as the RP2D for the expansion phase. Grade ≥3 treatment related adverse events were observed in 77% patients: anemia (23%), nausea and vomiting (3% ), lymphopenia ( 58%), neutropenia ( 35%) and leukopenia ( 39%). The majority of cytopenias were transient and resolved within two weeks. Among the 28 evaluable patients, the objective response rate (ORR) was 21.4% and the median progression-free survival (mPFS) has not been reached (median follow-up period: 23.7 months).

CONCLUSIONS: The combination of [¹⁷⁷Lu]lutetium-dotatate (200 mCi) and oral triapine (150 mg D1-14) was well tolerated and demonstrated preliminary activity in patients with well-differentiated GEP-NETs. (NCT04234568).

Akbasli, Izzet Turkalp, Baris Ozturk, Oguzhan Serin, Volkan Dogan, Göksu Bozdereli Berikol, Donnella S Comeau, Leo Anthony Celi, and Orhan Ozguner. (2026) 2026. “OCR-Mediated Modality Dominance in Vision-Language Models: Implications for Radiology AI Trustworthiness.”. Journal of Medical Imaging (Bellingham, Wash.) 13 (4): 045501. https://doi.org/10.1117/1.JMI.13.4.045501.

PURPOSE: Vision-language models (VLMs) are increasingly proposed for radiologic decision support, yet the security implications of deploying optical character recognition (OCR)-capable models in diagnostic workflows remain poorly characterized. When image-embedded text is not treated as untrusted input, the visual channel becomes vulnerable to adversarial manipulation.

APPROACH: Ten VLMs, none validated for clinical diagnosis, were evaluated on 600 brain magnetic resonance imaging studies for binary tumor detection under 5 conditions: clean input, visible radiology report injection, human-imperceptible stealth OCR injection, and multi-stage immune-prompt defense. A total of 27,000 inference calls were analyzed.

RESULTS: At baseline, performance was heterogeneous, with a median accuracy of 0.69, a sensitivity of 0.79, and a specificity of 0.59. Visible injection caused universal specificity collapse to 0.00 across all models with a false-positive rate (FPR) of 1.00 and a median attack success rate (ASR) of 0.97. Stealth injection, despite being imperceptible to human reviewers, drove substantial degradation with a median accuracy of 0.43, an ASR of 0.57, and an FPR of 0.84. Immune prompting achieved only partial mitigation: under stealth injection, median ASR decreased to 0.44, and accuracy improved to 0.56, yet residual overcalling persisted with a median FPR of 0.67, and three models maintained an FPR of 1.00.

CONCLUSIONS: Commercial VLMs exhibit a deployment-critical failure mode in radiology-like scenarios: OCR-readable text embedded in images can override pixel-level evidence, even under stealth conditions that evade human inspection. Prompt-level defenses provide insufficient protection. Any clinical integration of VLMs must be governed by system-level safeguards, including OCR-aware input handling, provenance controls, and enforced human verification, before deployment in safety-sensitive environments.

Koweek, Lynne M, Prachi Agarwal, Shari B Brosnahan, Sanjeev Bhalla, Vivian Bishay, Paul Cronin, Christopher J François, et al. (2026) 2026. “Pulmonary Embolism Reporting and Data System (PE-RADS™) V2026 for the Diagnosis of Acute Pulmonary Embolism on CT and MR Angiography.”. Chest. https://doi.org/10.1016/j.chest.2026.06.006.

Pulmonary embolism (PE), defined as blockage of the pulmonary arteries, is a common and serious condition that requires accurate diagnosis and standardized reporting. Noninvasive imaging is well established for the diagnosis and management of PE and is incorporated into multiple societal guidelines. However, specific technical recommendations for the optimal interpretation and standardized reporting of CT and MR angiography remain absent. Diagnostic imaging is used to identify the presence or absence of PE and to risk-stratify patients. Evolving treatment options can be considered once an acute clot is identified, the patient's status is defined, and the patient is referred for appropriate treatment in a timely manner. Patient management for PE often involves multiple care teams, and clear and consistent communication between care teams is critical to patient care. The goal of the Pulmonary Embolism Reporting and Data System (PE-RADS) is twofold: (a) to create a standard lexicon for the description of terms used for acute PE diagnosis at CT and MR angiography and (b) to develop a structured, hierarchical reporting system that incorporates anatomic and accessory findings to classify clot location and right heart imaging features in acute PE and to assist in determining the cardiovascular status of the patient to inform next-step management.

Koweek, Lynne M, Prachi Agarwal, Shari B Brosnahan, Sanjeev Bhalla, Vivian Bishay, Paul Cronin, Christopher J François, et al. (2026) 2026. “Pulmonary Embolism Reporting and Data System (PE-RADS™) V2026 for the Diagnosis of Acute Pulmonary Embolism on CT and MR Angiography.”. Radiology 320 (1): e252721. https://doi.org/10.1148/radiol.252721.

Pulmonary embolism (PE), defined as blockage of the pulmonary arteries, is a common and serious condition that requires accurate diagnosis and standardized reporting. Noninvasive imaging is well established for the diagnosis and management of PE and is incorporated into multiple societal guidelines. However, specific technical recommendations for the optimal interpretation and standardized reporting of CT and MR angiography remain absent. Diagnostic imaging is used to identify the presence or absence of PE and to risk-stratify patients. Evolving treatment options can be considered once an acute clot is identified, the patient's status is defined, and the patient is referred for appropriate treatment in a timely manner. Patient management for PE often involves multiple care teams, and clear and consistent communication between care teams is critical to patient care. The goal of the Pulmonary Embolism Reporting and Data System (PE-RADS) is twofold: (a) to create a standard lexicon for the description of terms used for acute PE diagnosis at CT and MR angiography and (b) to develop a structured, hierarchical reporting system that incorporates anatomic and accessory findings to classify clot location and right heart imaging features in acute PE and to assist in determining the cardiovascular status of the patient to inform next-step management. © 2026 RSNA and the American College of Chest Physicians published by Elsevier Inc. Supplemental material is available for this article. This article is a simultaneous joint publication in Radiology and CHEST. The articles are identical except for stylistic changes in keeping with each journal's style. Either version may be used in citing this article.

Wood, Erika J, Hannah S Milch, Lars J Grimm, Lisa A Mullen, Vandana Dialani, James Sayre, Jay R Parikh, and Katerina Dodelzon. (2026) 2026. “Parental Leave and Lactation Policies in Breast Imaging: Insights from a National Survey of U.S. Radiologists.”. Journal of Breast Imaging. https://doi.org/10.1093/jbi/wbag021.

OBJECTIVE: To assess variability in parental leave and lactation policies affecting breast radiologists and examine their associations with job satisfaction and burnout.

METHODS: An anonymous 43-question survey was distributed to U.S.-based physician members of the Society of Breast Imaging between December 2023 and February 2024. Questions assessed parental leave and lactation policies, workplace experiences, and validated items on job satisfaction and burnout. Descriptive statistics, chi-square tests, and multivariable linear regression were performed.

RESULTS: 262 respondents (8.0% [262/3275] response rate) completed the survey. The mean age was 47 ± 10.4 years, with an average of 14.6 ± 10 years of post-training experience; 86% (223/262) identified as women. Among respondents, 61% (160/262) were aware of their workplace's parental leave policy, which included a median of 6 weeks of paid leave (IQR 10). Forty percent (88/220) were unsure if a lactation policy existed, 57% (125/220) reported no dedicated lactation time, and 72% (147/220) were unaware of a designated lactation space. Most respondents (87% [191/219]) said these policies were not discussed during job interviews, though 59% (128/216) considered them important when choosing a job. One-third (72/262) had welcomed a child within the last five years. Among these, 39% (27/69) reported insufficient lactation support, and 26% (18/69) stopped breastfeeding earlier than desired due to work constraints. Lactation support was inversely correlated with burnout (r = -0.516, P < 0.001) and positively correlated with job satisfaction (r = 0.440, P < 0.001).

CONCLUSION: Improved transparency and institutional investment in parental leave and lactation policies by radiology practices may support well-being and retention of breast radiologists.

Ravi, Praful, Caiwei Zhong, Kimberly J Perez, Wanling Xie, Virginia Volpe, Gwo-Shu Mary Lee, Jonah Boardman, et al. (2026) 2026. “Clinical Impact and Dynamics of Clonal Hematopoiesis With 177Lu-PSMA-617 Therapy in Advanced Prostate Cancer.”. Journal of Nuclear Medicine : Official Publication, Society of Nuclear Medicine. https://doi.org/10.2967/jnumed.126.272125.

The prevalence, clinical impact, and clonal dynamics of clonal hematopoiesis (CH) in patients receiving 177Lu-PSMA-617 (LuPSMA) for metastatic castration-resistant prostate cancer (mCRPC) are unknown. Methods: Targeted next-generation sequencing of 21 genes recurrently mutated in CH was performed on DNA extracted from the peripheral blood of patients who received at least 4 cycles of LuPSMA for mCRPC at our institution between 2022 and 2023. Pathogenic somatic mutations with a variant allele fraction of at least 1% were identified using a standardized pipeline. Clinical outcomes pertaining to efficacy (overall survival [OS], measured from date of planned cycle 5) and hematologic toxicity of LuPSMA were collected from the electronic medical record. Results: Fifty patients treated with LuPSMA were eligible, with a median follow-up of 23 mo. At least 1 CH variant was detected in 33 patients (66%). The most common mutations were TET2 (n = 16), PPM1D (n = 15), and DNMT3A (n = 6). OS was similar in patients with or without CH (12-mo OS, 92% vs. 80%; hazard ratio, 0.89; 95% CI, 0.3-2.68). There was a trend toward greater hematologic toxicity in patients with CH, with a greater need for growth factor support (12% vs. 0%). In patients with serial samples available, the emergence of new clones or expansion of preexisting CH variants was detected in most patients, particularly with PPM1D and TP53-mutant clones. The key limitation was the small sample size and short follow-up. Conclusion: CH was highly prevalent and tended to lead to greater hematologic toxicity in patients with mCRPC receiving LuPSMA. Expansion or emergence of DNA damage repair CH clones was very common during and after LuPSMA therapy. Further study of the impact of CH on radiopharmaceutical therapy, particularly when used in earlier prostate cancer disease settings, is warranted.