Publications

  • Karoly, Philippa J, Rachel E Stirling, Mark J Cook, Daniel M Goldenholz, Maxime O Baud, Vikram R Rao, Solveig Vieluf, and Benjamin H Brinkmann. (2026) 2026. “Seizure Forecasting: The Long and Winding Road to Clinical Translation.”. Epilepsia. https://doi.org/10.1002/epi.70394.

    Seizure forecasting has progressed from theoretical aspiration to a rapidly advancing research domain, yet clinical translation remains limited. Over the past decades, advances in algorithm development, chronic electroencephalography (EEG), wearable sensors, and the characterization of seizure cycles have demonstrated that seizure risk is not random but fluctuates according to identifiable biological rhythms and patient-specific patterns. Forecasting algorithms have shown promising performance across diverse retrospective datasets, including intracranial EEG, subscalp recordings, wearable physiological signals, and even self-reported diaries. However, the clinical value of forecasting is contentious. Prospective real-world validation and regulatory approval of patient-facing forecasting systems remain rare. This review incorporates perspectives presented at the 5th International Congress on Mobile Health and Digital Technology in Epilepsy (2025). We examine barriers impeding clinical translation and current attempts to address them. Crucially, forecasting performance cannot be evaluated in isolation from intended use. Applications range from low-risk uses, like scheduling diagnostic monitoring or visualization of historical trends, to higher risk interventions, including medication titration and adaptive neuromodulation. Each application entails distinct performance thresholds, ethical considerations, and regulatory requirements. Translational challenges include reliable seizure annotation, nonstationarity dynamics of biological cycles, and practical constraints for real-time deployment. Ethical concerns center on miscalibrated reliance on low-risk states, potential anxiety associated with high-risk advisories, and the heterogeneity of patient preferences and risk tolerance. Regulatory pathways are likely to depend on clearly defined use cases and clinically meaningful endpoints, which may extend beyond seizure counts to include quality of life, anxiety, locus of control, and other patient-reported outcomes. Ultimately, translation will require rigorous prospective evaluation against transparent benchmarks, sustainable scientific-commercial partnerships, and integration of probabilistic risk information into clinical workflows. With careful implementation, seizure forecasting may evolve from proof-of-concept research into a clinically meaningful component of epilepsy management, and we remain cautiously optimistic.

  • Dymm, Braydon, and Daniel M Goldenholz. (2026) 2026. “Prompting Is All You Need: How to Make LLMs More Helpful for Clinical Decision Support.”. MedRxiv : The Preprint Server for Health Sciences. https://doi.org/10.64898/2026.02.12.26346005.

    IMPORTANCE: Large language models (LLMs) offer potential decision support, but their accuracy varies. Prompt engineering can generally enhance LLM behavior in a clinical context, yet best practices have yet to be formally explored in realistic clinical contexts for neurology.

    OBJECTIVE: To evaluate the impact of structured prompting versus naive prompting on the performance of four LLMs (two closed-source: OpenAI GPT-4o, OpenAI o3; three open-source: Meta Llama-4-Scout-17B-16E-Instruct, Llama-3.3-70B-Instruct-Turbo, and the reasoning model r1-1776) for thrombolytic clinical decision support (CDS) in acute stroke.

    DESIGN: Models responded to three novel ischemic stroke vignettes using either a naive question ("Should this patient be offered thrombolytics?") or a five-step structured prompt (CARDS) guiding information extraction, timing analysis, contraindication checking, decision process explanation, and risk-benefit discussion. Outputs were assessed across seven domains: guideline adherence, unsafe recommendations, risk recognition, guideline grading accuracy, inclusion of conversational explanation, clarity, and overall helpfulness.

    RESULTS: Structured prompts significantly enhanced performance across most domains, with varying effects between model families. For closed-source models (GPT-4o, o3), prompts structured in the CARDS style improved guideline adherence from 83.3% to 100%, eliminated unsafe recommendations (16.7% to 0%), and increased specific guideline grading accuracy from 0% to 100%. Similarly, the open-source reasoning model r1-1776 achieved these top-tier outcomes (100% adherence, 0% unsafe, 100% grading, 100% conversation) when structured prompts were applied, with grading and conversation improving from 0%. In contrast, other open-source models (Llama-4-Scout, Llama-3.3-70B) showed more modest gains: risk recognition improved (83.3% to 100%) and guideline grading accuracy increased (0% to 66.7%), while guideline adherence (66.7%) and unsafe recommendations (33.3%) persisted. Overall, structured prompting yielded the largest improvements in guideline grading accuracy and conversational reasoning across multiple models.

    CONCLUSION AND RELEVANCE: Structured prompting substantially enhances LLM performance for acute stroke thrombolysis CDS. Notably, some models, including the proprietary GPT-4o and o3, and the open-source reasoning model r1-1776, achieved excellent safety and adherence with structured prompts. For clinical deployment of any LLM, structured prompts are crucial, and vigilant human oversight remains essential.