The Illusion of Detection: Why Machine Learning Can Overestimate Mental Disorder Recognition from Audio
DOI:
https://doi.org/10.31224/8100Keywords:
Applied Machine LearningAbstract
Machine learning (ML) research has attempted to detect mental disorders from people’s audio recordings. Despite seemingly impressive reported results, it remains unclear whether those results reflect genuine clinical signal or artefacts related to how ML models are trained and evaluated. We re-examine 25 studies included in a recent systematic review and identify three major recurring methodological problems that are likely to inflate the ability of ML models to detect mental disorders: (1) data leakage, (2) lack of multiple seeds, and (3) lack of large, naturalistic datasets. We show via an experiment how ML practices can produce apparently strong results even in the absence of real, predictive signals. We also conduct robust analyses using a large, novel dataset, where we find no reliable evidence that mental disorders can be detected from audio across multiple modelling approaches. Our findings suggest that current evidence may substantially overestimate the real-world capability of audio-based detection of mental health disorders, highlighting the need for evaluation standards aligned with best practices for clinical and policy use.
Downloads
Downloads
Posted
Versions
- 2026-09-04 (2)
- 2026-08-29 (1)
License
Copyright (c) 2026 Sia Shah, Pamela Qualter, Yongchao Huang, Nina Shahrizad, Dario Krpan, Matteo Galizzi

This work is licensed under a Creative Commons Attribution 4.0 International License.