Study shows mental health AI still risky, lacks real-world clinical trials
An international research review reveals LLMs have potential for mental health screening, but ethical and safety challenges remain.
A research team led by Yisong Chen published a systematic review titled "Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges" in the Journal of Industrial Integration and Management, and on the arXiv preprint repository (cs.AI).
Using PRISMA standards, the team screened 304 related research papers and selected 92 high-quality studies to analyze the application of Large Language Models (LLMs) in mental health.
The study found that current LLM applications range from detecting depression signals and early suicide risk assessment via social media text and electronic health records, to developing conversational therapy agents and decision-support systems for medical professionals. Innovations like Prompt Engineering for specialized medicine and Multimodal Integration—combining text, audio, and sensor data—are also used to improve accuracy.
However, the research highlights significant ethical challenges and limitations, noting that most studies remain at an experimental stage and lack real-world clinical validation. There are also accuracy and safety risks, such as hallucination and sycophancy, which could lead to diagnostic errors, alongside issues regarding patient data privacy, model transparency (black-box nature), and a lack of clear legal regulatory frameworks.
This highlights that while AI and LLM technologies are playing a bigger role in mental health screening, users and medical professionals must remain cautious as the models are not yet ready for independent clinical use due to safety limitations.