Please use this identifier to cite or link to this item: https://hdl.handle.net/1/3123
Title: 'Man vs. Machine: can ML algorithms diagnose headaches as accurately as clinicians? A systematic review.'
Authors: Manickam, Appukutty ;Valliappan, Abirami
Affliation: Central Coast Local Health District
Wyong Hospital
Issue Date: 26-May-2026
Source: 273(6):341
Journal title: Journal of Neurology
Department: Neurology
Abstract: Headache disorders are frequently misdiagnosed. We aimed to systematically evaluate the diagnostic accuracy, methodological quality, and clinical applicability of artificial intelligence (AI) and machine learning (ML) models for classifying adult headache disorders against clinician diagnoses using the International Classification of Headache Disorders (ICHD) criteria. In this systematic review, we searched PubMed, Embase, and the Cochrane Library (January 2015-December 2025) for AI/ML diagnostic headache studies. Two reviewers extracted data and assessed risk of bias using the QUADAS-2 tool with the QUADAS-AI extension. Main outcomes were sensitivity, specificity, area under the receiver operating characteristic curve (AUC-ROC), and risk of bias. We included 74 studies encompassing 154,856 participants. Models utilized traditional ML (n = 47), deep learning (n = 18), and hybrid or rule-based approaches (n = 9). Data inputs included neuroimaging (n = 27), multimodal datasets (n = 20), neurophysiological signals (n = 17), and clinical questionnaires (n = 10). Only 4 studies performed independent external validation. Overall sensitivity ranged from 47.5% to 100.0%, specificity from 50.4% to 100.0%, and AUC-ROC from 0.658 to 1.000. Models using structured questionnaires reported realistic accuracies (74-86%), whereas neuroimaging models frequently produced near-perfect, likely overfit estimates. Most studies (65 of 74) exhibited a high risk of bias driven by artificial case-control designs, data leakage, and absent external validation. AI-based headache diagnostic systems demonstrate promising but highly variable accuracy. Pervasive methodological flaws-specifically severe spectrum bias, data leakage, and a lack of independent external validation-currently preclude clinical implementation. Future studies require prospective, real-world validation to safely integrate these tools into practice.
URI: https://hdl.handle.net/1/3123
DOI: 10.1007/s00415-026-13853-7
Pubmed: https://pubmed.ncbi.nlm.nih.gov/42189256
Publicaton type: Journal Article
Keywords: Neurology
Appears in Collections:Health Service Research

Show full item record

Page view(s)

24
checked on Jul 11, 2026

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.