Publications

2026

Balduzzi, Alberto, Luca Ghirotto, Massimo Tomasi, Giovanni Marchegiani, Marc Besselink, Marco J Bruno, Paolo Arcidiacono, et al. (2026) 2026. “Diagnosis and Follow-up of IPMNS of the Pancreas - Bringing the Ethical Issues into Focus.”. Pancreatology : Official Journal of the International Association of Pancreatology (IAP) . [et Al.]. https://doi.org/10.1016/j.pan.2026.08.002.

BACKGROUND: Intraductal papillary mucinous neoplasms (IPMNs) of the pancreas present a broad spectrum of biological behavior, ranging from benign to malignant. Their management poses significant ethical dilemmas, particularly concerning diagnostic uncertainty, the risk of overdiagnosis and overtreatment, and resource allocation.

METHODS: This qualitative study employed a constructivist approach to explore the ethical challenges faced by surgeons in managing IPMNs. Data were collected through focus group meetings (FGMs) with members of an expert working group at the Verona Evidence-Based Meeting on IPMNs (2020). Discussions were analyzed to identify key ethical concerns.

RESULTS: The analysis highlighted several major ethical concerns: (1) decision-making under diagnostic uncertainty, (2) ethical challenges in patient communication, (3) overdiagnosis and over-surveillance due to defensive medicine and patient anxiety, (4) overtreatment through unnecessary surgery, and (5) issues of distributive justice in access to care and healthcare resource utilization. Participants emphasized the difficulty of balancing transparency with the need to minimize psychological distress in patients, as well as the challenge of applying international guidelines in diverse healthcare settings.

CONCLUSIONS: Ethical decision-making in IPMN management requires balancing the risks of malignancy with the potential harms of overtreatment, while also considering patient autonomy and resource limitations. Enhancing decision-support tools, improving surgeon training in communication, and refining clinical guidelines to incorporate ethical considerations may help address these challenges. Further research is needed to develop strategies for more individualized and patient-centered care.

Scharf, Zachary, Miguel Muniz, Elizabeth Tchitchkan, Avina Rami, Praful K Ravi, Jacob J Orme, Ali Tarhini, Heather A Jacene, Daniel Sentana-Lledo, and Daniel S Childs. (2026) 2026. “An Assessment of Time Toxicity in Patients Receiving [177Lu]Lu-PSMA-617 for Metastatic Castration-Resistant Prostate Cancer (mCRPC).”. The Oncologist. https://doi.org/10.1093/oncolo/oyag327.

BACKGROUND: [177Lu]Lu-PSMA-617 improves survival and quality of life in men with metastatic castration-resistant prostate cancer but requires substantial healthcare involvement. Our analysis evaluated the time toxicity associated with [177Lu]Lu-PSMA-617, representing the first dedicated radioligand time toxicity assessment.

MATERIALS AND METHODS: We conducted a multi-institutional retrospective study of patients initiating [177Lu]Lu-PSMA-617 between April 2022 and March 2023. Patients were followed from first [177Lu]Lu-PSMA-617 administration until 3 months after the final cycle, initiation of new systemic therapy, or death. Time toxicity was quantified as healthcare contact days during the at-risk interval. Associations with baseline characteristics were evaluated using negative binomial regression models. Univariate and multivariable models were performed.

RESULTS: Among 135 patients, median age was 69 years and median prior systemic therapies was 4. Patients completed a median of 5 [177Lu]Lu-PSMA-617 cycles over 7.6 months. Median healthcare contact days were 20 (IQR, 15-28), representing 9.5% (IQR, 6.4-13.9) of at-risk days. Fifteen patients (11%) experienced high time toxicity (>1 contact day per 5 days), while 42 (31%) had low time toxicity (≤1 contact day per 14 days). Among the 57 patients who received both cycle 1 and cycle 6, mean per cycle healthcare contact days were 4 and 3.6, respectively (paired Wilcoxon signed-rank test, p = 0.04). Routine oncology visits comprised 32.7% of contact days, compared with 8.8% from hospitalizations. Lower baseline hemoglobin was independently associated with greater time toxicity (RR = 0.87; 95% CI, 0.79-0.96; p = 0.004).

CONCLUSION: In heavily pretreated patients, [177Lu]Lu-PSMA-617 demonstrated relatively low time toxicity, supporting it as a time-efficient therapy.

Smith, Martin R, and Bo Yang. (2026) 2026. “Directional Biases in Morphological Evolution: Implications for Phylogenetic Models.”. Systematic Biology. https://doi.org/10.1093/sysbio/syag063.

Even in the era of large molecular datasets, morphological evidence makes an integral contribution to reconstructing evolutionary history. Whereas traditional phylogenetic methods treat all morphological characters as if they evolve in the same way, we introduce a new analytical framework that explicitly distinguishes transformational characters, which document variation in form, from neomorphic characters, which record the presence or absence of discrete traits. Using 69 datasets that span major groups of animals and other organisms, we show that transformational characters typically change twice as fast as neomorphic characters; and that losses of neomorphic features occur more than twice as frequently as gains. Evolutionary models assuming long-term equilibrium are violated in the majority of datasets, suggesting ongoing directional trends. Accounting for these asymmetries improves model fit and can substantially alter inferred evolutionary relationships, sometimes producing posterior samples that occupy distinct regions of tree space. These results provide broad empirical evidence for directional biases in morphological evolution, supporting the view that developmental and ecological constraints favour trait loss over gain. By explicitly distinguishing different types of morphological change, this framework enhances model realism, recovers different phylogenies, and opens new avenues for interpreting macroevolutionary dynamics.

Chang, Chi-Chieh, Ying-Chang Lu, Yu-Ming Chang, Yen-Hui Chan, Hong-Yu Yan, Yi-Shan Lee, Pei-Hsuan Lin, et al. (2026) 2026. “A Novel Drug-Loaded Porous Parylene Electrode for Targeted Anti-Inflammatory Therapy and Hearing Preservation in Cochlear Implantation.”. Advanced Healthcare Materials, e71593. https://doi.org/10.1002/adhm.71593.

This study investigates the feasibility and therapeutic effectiveness of a novel drug‑loaded parylene electrode prototype (PEP) designed to reduce inflammation and hearing loss associated with cochlear implantation (CI). The porous PEP was fabricated using vapor‑phase sublimation and deposition, producing an interconnected structure with 20-30 µm pores and mechanical flexibility suitable for intracochlear placement. Dexamethasone was incorporated into the device (Dex/PEP), and in vitro characterization using ELISA and spectrophotometry demonstrated controlled first‑order release kinetics. Guinea pigs underwent microsurgical implantation of PEPs, with accurate positioning confirmed by micro‑CT imaging. Auditory outcomes were assessed through auditory brainstem response (ABR) measurements, and cochlear tissues were examined histologically to evaluate inflammatory cell infiltration and cytokine expression. The Dex/PEP device exhibited biosafety comparable to traditional silicone electrodes and provided sustained drug release. Animals receiving Dex/PEP showed significantly improved ABR thresholds at 1 and 4 weeks compared with blank PEP controls (p < 0.05), indicating reduced inflammation‑induced hearing loss. Histological analysis further revealed diminished inflammatory cell infiltration and markedly decreased TNF‑α expression at the round window membrane (p = 0.001). These findings support the drug‑loaded PEP as a promising targeted delivery platform for reducing cochlear inflammation and preserving hearing following CI.

Smith, Martin. (2026) 2026. “Privileged Worlds, Additivity and Maximal Risk.”. Synthese 208 (2): 103. https://doi.org/10.1007/s11229-026-05756-x.

In 'The unified theory of risk' (2024) Jaakko Hirvelä and Niall Paterson consider the recent debate between the probabilistic, modal and normic theories of risk, and propose an intriguing new account - the titular unified theory. In this paper I focus on Hirvelä and Paterson's criticisms of the normic theory, which are based on what they call the 'privileged world problem' and on the idea that risks 'add up'. I will argue that these criticisms can, to an extent, be avoided by adopting a certain reinterpretation of the normic theory, particularly as it pertains to the notion of maximal risk. I then turn to the unified theory of risk and argue that, on close inspection, the theory is very similar to one of the rival theories that it purports to unify - namely, the probabilistic theory - and may share its shortcomings.

Omar, Mahmud, Mohammad E Naffaa, Reem Agbareia, Fadi Hassan, Abdulla Watad, Helana Jeries, Alon Gorenshtein, et al. (2026) 2026. “AI-Assisted Rheumatology Triage Changes With Referral Framing.”. Rheumatology (Oxford, England). https://doi.org/10.1093/rheumatology/keag414.

OBJECTIVES: Two in three US physicians now use healthcare AI, and large language models (LLMs) are entering the triage workflows that determine which patients reach rheumatology and how quickly. We aimed to test whether nine prespecified cues in referral notes, patient descriptions and demographics shift AI-assisted triage decisions when the underlying clinical information is unchanged.

METHODS: We conducted a controlled, physician-validated experiment across 30 physician-authored rheumatology vignettes, each independently rephrased three times (90 case variants). We tested five LLMs from three providers - Anthropic, Google, and OpenAI - under nine dimensions spanning demographics, clinical context, and communication framing, with 57 controlled contextual modifications, four system-prompt personas, and five repetitions per cell, yielding more than 200,000 model queries. Sixteen clinical fields were graded against physician-validated ground truth, with excellent inter-rater agreement (Fleiss' kappa=0.92).

RESULTS: Baseline composite concordance with expert ground truth was high at 0.869. We had expected sociodemographic cues to be the strongest source of distortion. Instead, the largest shifts came from how the case was framed and described. When patients were described as anxious, models attributed symptoms to psychological rather than organic causes nearly three times as often as at baseline (13.1% vs 4.5%; odds ratio 3.2 versus stoic framing), a shift that risks relabeling organic disease as functional. Clinician anchoring in the referral note reduced concordance, consistently across models and rephrasings and significantly for acuity (dismissive anchor, vignette-level p = 0.01), mainly by downgrading urgency. In contrast, race or ethnicity, socioeconomic status, and language barrier produced no detectable effect, including in mixed-effects models that accounted for repeated vignette use.

CONCLUSION: Although baseline concordance with specialist ground truth was high, it was readily disrupted by how referral notes were worded and how patients described their symptoms, not by patient demographics. Before AI-assisted triage enters rheumatology referral pathways, systems should separate objective clinical evidence from interpretive framing, and urgency and psychological attribution should be treated as auditable safety signals.

Sorin, Vera, and Eyal Klang. (2026) 2026. “AI Scribe Safety: Measuring What Happens After Signing.”. JMIR Medical Informatics 14: e103162. https://doi.org/10.2196/103162.

Coiera and Fraile-Navarro question whether AI scribes are being evaluated on metrics that truly impact care. While current evaluations focus on the quality of the initial draft, signed clinical notes are dynamic, as their content can be copied, summarized, coded, and re-ingested by downstream AI tools. We argue that safety must be measured downstream, focusing on how small errors in initial documentation can compound across the patient's electronic health record.

Mudrik, Aya, Ayala Dodge, Alon Moore Galindo, Girish N Nadkarni, Shelly Soffer, and Eyal Klang. (2026) 2026. “Sex-Related Performance Disparities in Convolutional Neural Networks for Imaging-Based Assessment of Coronary Atherosclerosis and Ischemic Heart Disease: A Systematic Review.”. Journal of Imaging Informatics in Medicine. https://doi.org/10.1007/s10278-026-02173-x.

Sex-related disparities persist in the diagnosis and management of ischemic heart disease (IHD), raising concern that convolutional neural networks (CNNs) used in coronary imaging may perpetuate these inequities. This systematic review evaluated sex-related performance differences in CNN-based models using medical imaging to assess coronary atherosclerosis, coronary artery disease, or myocardial ischemia. A systematic literature search of PubMed, Web of Science, Scopus, IEEE Xplore, and ACM Digital Library was conducted through June 7, 2026, in accordance with PRISMA guidelines. Peer-reviewed original studies were included if they evaluated sex-related performance of CNN-based models using medical imaging inputs, either by reporting performance metrics separately for women and men or by assessing sex as a determinant of model error, misclassification, calibration, or agreement. Nine studies met the inclusion criteria, covering noncontrast cardiac CT, coronary CT angiography, PET-CT, chest radiography, and SPECT-based approaches. Overall, model performance was generally similar between men and women across imaging modalities. Nevertheless, several studies identified clinically relevant sex-related discrepancies, including higher error rates in men for PET-CT calcium scoring and sex-related differences in error patterns during automated CCTA interpretation. Notably, one SPECT-based study showed that targeted augmentation of training data improved calibration and reduced false-positive predictions in women. While CNNs generally show comparable performance between sexes, subtle sex-related biases can persist, often driven by data imbalance and reference standard limitations. Sex-stratified evaluation and bias-mitigation strategies are essential to ensure equitable clinical implementation.

Tessler, Idit, Mahmud Omar, Amit Wolfovitz, Yoav Gimmon, Sholem Hack, Noa Rozendorn, Nir Livneh, and Eyal Klang. (2026) 2026. “Sociodemographic Bias in LLMs’ Clinical Decision-Making for Dizziness.”. Journal of Vestibular Research : Equilibrium & Orientation, 9574271261474218. https://doi.org/10.1177/09574271261474218.

ObjectiveAs large language models (LLMs) enter clinical decision support, concerns persist about sociodemographic bias. We assessed whether LLM recommendations for dizziness vary by patient descriptors and clinical detail.MethodsWe conducted a cross-randomized in-silico vignette study. One hundred synthetic emergency department dizziness cases were created using established diagnostic frameworks including the TiTrATE paradigm, SAEM GRACE-3 guidelines, and Bárány Society diagnostic criteria. Each vignette was tested in a neutral form and with 33 sociodemographic descriptor variants (34 total). Twelve instruction-tuned LLMs from multiple model families were evaluated. Models answered five binary clinical decision questions addressing etiology classification, triage disposition, neuroimaging, bedside vestibular examination, and mental health referral. Each model-vignette-descriptor combination was repeated 10 times, yielding 2,040,000 responses. Sociodemographic bias was quantified as descriptor-specific percentage-point deviations from neutral control recommendations with 95% confidence intervals.ResultsSociodemographic descriptors influenced LLM recommendations, with the largest differences observed for mental health referral decisions in diagnostically ambiguous cases. Referral likelihood was lower for Black transgender women (-12.2 pp; 95% CI -14.0 to -10.3), Black patients experiencing homelessness (-9.1 pp; -11.0 to -7.3), and patients experiencing homelessness (-7.7 pp; -9.5 to -5.9). Differences were attenuated when vignettes contained clearer diagnostic information. Other effects were smaller, including increased neuroimaging recommendations for low-income descriptors (+4.0 pp; 95% CI 2.1-5.8).ConclusionLLM clinical recommendations varied by sociodemographic descriptors, particularly under diagnostic uncertainty. More detailed clinical information reduced these disparities, suggesting structured inputs may mitigate bias in clinical AI systems.