Publications

2026

Pettersson, Samuel D, Jean Filo, Paulina Skrzypkowska, Thomas B Fodor, Kamil Siedlecki, Peter Liaw, Ciprian N Ionita, et al. (2026) 2026. “End-to-End Autonomous Quantification of Brain Aneurysm and Parent-Artery Morphology on CT Angiography.”. Radiology. Artificial Intelligence, e251093. https://doi.org/10.1148/ryai.251093.

Purpose To develop and validate an end-to-end autonomous platform for the quantification and visualization of brain aneurysm and parent-artery morphology on CT angiography (CTA). Materials and Methods A total of 2,980 CTA scans performed between 2004 to 2025 containing 2,585 aneurysms from 2980 patients obtained from five international high-volume stroke centers were included. The model was trained using expert hand-annotated vascular segmentations spanning the cervical internal carotid and vertebral arteries through the A4, M3, and P3 segments. Internal prospective and multicenter external testing assessed aneurysm detection performance and compared morphology measurements with expert-derived values obtained on digital subtraction angiography (DSA)-verified CTA. Statistical analysis used paired t tests or Wilcoxon signed-rank tests, with sensitivity, specificity, and 95% confidence intervals reported. A public web-based platform was developed to allow further validation. Results Internal and multicenter external testing yielded a patient-level sensitivity of 87.9% (290 of 330) (95% CI: 83.8, 91.3) and specificity of 86.6% (395 of 456) (95% CI: 83.3, 89.4). Physician-performed morphology extraction required 24.3 ± 6.6 minutes per scan, whereas the model completed the task autonomously in 90 ± 12 seconds (P < .001). At the aggregate level, no statistically significant differences were observed for aneurysm volume, neck diameter, parent artery diameter, flow angle, aspect ratio, size ratio, height-width ratio, undulating index, ellipticity index, or non-sphericity index. Dome height (mean difference (MD): 0.1 ± 1.1, P = .03) and surface area (MD: -5.3 mm2 ± 29.5 mm2, P = .04) differed, but with minimal effect sizes. Conclusion The developed autonomous system for rupture-related morphology metric acquisition substantially reduced physician workload. ©RSNA, 2026.

McCarthy, Colin J, Vijay Ramalingam, Yiftach Barash, Seetharam Chadalavada, Xiao Wu, Oleksandra Kutsenko, Daniel Raskin, Vera Sorin, and Ammar Sarwar. (2026) 2026. “Decreasing Administrative Effort Related to Non-Approval of Image-GuidED Procedures Using Large Language Models - The DENIED-AI Pilot Study.”. Academic Radiology. https://doi.org/10.1016/j.acra.2026.04.021.

RATIONALE AND OBJECTIVES: To evaluate whether large language models (LLMs) can generate accurate, clinically valid, and usable letters to appeal insurance denials for radiology procedures.

MATERIALS AND METHODS: This pilot study generated insurance appeal letters for a simulated clinical scenario. Four LLMs (Claude 3.5, Nova Pro, Llama-3.1-70B, ChatGPT-4o) were used with zero-shot, few-shot, and retrieval-augmented generation (RAG) techniques. Four board-certified interventional radiologists, blinded to model and technique, scored letters for content (accuracy, personalization, references), grammar and structure (readability, tone, persuasiveness), and usability (estimated editing time, usefulness as a template). References were verified for accuracy, and outputs were carefully examined for hallucinations. Statistical analyses included ANOVA, Chi-square, and Fleiss' Kappa for interrater reliability.

RESULTS: Mean content and grammar scores were 3.9 ± 0.95 and 4.3 ± 0.9 (out of 5), with no significant differences by model or technique (p >.05). Reviewer agreement was poor (Fleiss' Kappa -0.18 for content, -0.085 for grammar). Hallucinations were flagged by reviewers in 16/48 assessments, significantly more often with the online model (ChatGPT-4o: 58% vs offline 25%; p =.03). Of 44 references, 80% from the offline models were fabricated compared with 29% from ChatGPT-4o (p <.001). Estimated editing time was less than 10 min in 71% of responses, and the reviewers felt the letters would be useful as templates in 73% of cases.

CONCLUSION: LLM-generated appeal letters for insurance denials were generally well received, with high usability and adequate quality. However, fabricated references and hallucinations remain prevalent, necessitating careful human review before clinical use.

von Wedel, Dario, Simone Redaelli, Maxime Fosset, Joris Pensier, Denys Shay, Elena Ahrens, Luca J Wachtendorf, et al. (2026) 2026. “The Predicted Body Weight Equation Overestimates Lung Sizes of Female, Critically Ill Patients: An Analysis of Randomized, Controlled Trials and Real-World Clinical Data.”. Intensive Care Medicine. https://doi.org/10.1007/s00134-026-08442-1.

PURPOSE: Low tidal volume (Vt) ventilation is the standard of care among critically ill patients. Guidelines recommend scaling Vt to the predicted body weight (PBW) to avoid ventilator-induced lung injury (VILI). Concerns exist that the PBW overestimates lung volumes of critically ill females. We investigated whether this applies to clinically relevant measures of lung volume, whether PBW-guided mechanical ventilation yields comparable risk of lung stress among male and female patients, and whether this affects mortality.

METHODS: Mechanically ventilated, critically ill patients from ten randomized trials and two real-world retrospective clinical datasets were analyzed. Risk of high driving pressures (≥ 15 cmH2O) at comparable Vt/kg PBW as well as measures of anatomical and functional lung sizes, including computed tomography-measured lung volumes at the same PBW were compared between female and male patients.

RESULTS: Among 30,516 patients (39.4% female), ventilation with comparable tidal volumes standardized to PBW (ml/kg PBW) was associated with 4.2% (95% CI 3.2-5.3; aOR 1.26, 95% CI 1.19-1.33; p < 0.001) higher absolute risk of high driving pressures among females, mediating 8.4% of excess 28-day mortality (p < 0.001). At the same PBW, female patients had lower anatomical and aerated lung volumes (- 343 ml, 95% CI - 449 to - 237, p < 0.001; and - 188, 95% CI - 282 to - 94, p < 0.001, respectively) than males.

CONCLUSIONS: The widely used PBW equation overestimates lung volumes in female critically ill patients, resulting in excess risk of injurious driving pressures among females, mediating higher mortality. Personalized mechanical ventilation by using driving pressure-guided strategies might mitigate these disparities.

Metrouh, Oussama, Julie Bulman, Sarah Schroeppel DeBacker, Muneeb Ahmed, and Jeffrey Weinstein. (2026) 2026. “Does Residency Rank List Placement Predict Clinical Performance in Interventional Radiology Training?”. Academic Radiology. https://doi.org/10.1016/j.acra.2026.03.051.

RATIONALE AND OBJECTIVES: This study aimed to evaluate the correlation between Interventional Radiology (IR) trainees' clinical performance during residency and their final National Residency Matching Program (NRMP) rank order list (ROL) placement during the match application, and to identify application metrics predictive of strong clinical performance.

MATERIALS AND METHODS: A retrospective review of application data for IR residents and fellows graduating from a single academic center between 2020-2025 was conducted. Metrics included United States Medical Licensing Examination (USMLE) scores, number of clinical and research experiences, peer-reviewed publications, abstracts, awards and the final placement on the ROL. A structured survey aligned with Accreditation Council for Graduate Medical Education (ACGME) milestones was designed and distributed to faculty who directly worked with each trainee but were not involved in the program's NRMP rank order list formation to evaluate their clinical performance during IR training. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs), and a composite clinical score was calculated. Associations between application metrics, NRMP rank, and clinical performance were evaluated using Spearman correlation and univariate linear regression.

RESULTS: Moderate inter-rater reliability was observed for medical knowledge (ICC=0.60, p < 0.001), procedural competence (ICC=0.50, p < 0.001), and patient care (ICC = 0.50, p = 0.004). No significant correlation was found between NRMP rank list placement and clinical performance (Spearman ρ = 0.07, p = 0.82). USMLE Step 2 score was the only significant predictor of clinical performance (β = 0.17, p = 0.01), with the greatest separation observed at a cutoff score of 239.

CONCLUSION: These findings suggest that interview-driven rank placement may not reliably identify high-performing residents, whereas Step 2 scores may provide better predictive value for clinical performance during IR training.

Paolucci, Iwan, Christiaan G Overduin, Edward W Johnston, Gregor Laimer, Muneeb Ahmed, Ronald S Arellano, Marie Beerman, et al. (2026) 2026. “International Multisociety Delphi Consensus for Liver Tumour Thermal Ablation: Margin Assessment.”. The Lancet. Oncology 27 (5): e248-e258. https://doi.org/10.1016/S1470-2045(26)00143-9.

This multisociety, multidisciplinary consensus-formally endorsed by the European Society of Surgical Oncology, the Cardiovascular and Interventional Radiological Society of Europe, and the Society of Interventional Oncology-was developed to standardise the assessment of ablation margins in liver tumour thermal ablation. A modified Delphi process, consisting of two online surveys and a hybrid (online and in-person meeting in Innsbruk) consensus meeting of 72 experts from North America, South America, Europe, and Asia. Formal consensus was reached for 150 (75%) of 199 statements. Strong agreement was observed between interventional and surgical oncologists, with only 12 (6%) of 199 statements showing significantly different ratings. Participants agreed that ablation margins should be assessed and documented for every treated tumour. Margins should be assessed quantitatively in three dimensions, with contrast-enhanced CT or MRI, preferably intraprocedurally with ablation confirmation software. Ablation margins should be categorised as A0 (tumour completely covered with sufficient margin), A1 (tumour completely covered but insufficient margin), or A2 (portion of tumour remains unablated). This effort is, to our knowledge, the first international consensus initiative to define best-practice recommendations for margin assessment in liver tumour thermal ablation to standardise practices, aiming to improve and promote uniform outcomes.

Laimer, Gregor, Edward W Johnston, Christiaan G Overduin, Iwan Paolucci, Muneeb Ahmed, Ronald S Arellano, Marie Beermann, et al. (2026) 2026. “International Multisociety Delphi Consensus for Liver Tumour Thermal Ablation: Procedural and Practice Standards.”. The Lancet. Oncology 27 (5): e259-e270. https://doi.org/10.1016/S1470-2045(26)00114-2.

Thermal ablation offers a safer, less invasive, and more cost-effective curative-intent treatment for selected patients with primary and metastatic liver tumours than surgery; when done with appropriate technique, ablation can deliver similar oncological outcomes. However, effectiveness in routine practice varies because structured training, planning, and procedural governance remain scarce. These international multidisciplinary, multi-society guidelines-formally endorsed by the European Society of Surgical Oncology, the Cardiovascular and Interventional Radiological Society of Europe, and the Society of Interventional Oncology-define key domains contributing to procedural difficulty and practice variation in liver tumour thermal ablation. A Delphi consensus initiative held in Innsbruck, Austria, engaged 72 experts across three iterative rounds of scoring across 135 statements grouped into five domains: credentialing, indications, approach, procedural factors, and safety measures. Consensus was achieved for 94 (70%) of 135 statements. The least invasive route-typically percutaneous-should be prioritised, and margin adequacy was reaffirmed as the principal technical goal. Procedural difficulty was considered context-dependent, shaped by tumour factors, institutional infrastructure, and operator experience. Organ displacement techniques were endorsed to maintain safety and expand treatable indications. Complex ablations should be done by experienced operators (more than 100 previous cases), with programmes underpinned by structured training, multidisciplinary team participation, and routine audit. Future efforts should develop and validate practical tools such as difficulty scoring systems, standardised procedural reporting templates, and comprehensive training curricula to improve consistency, standardisation, and clinical outcomes globally.