Benchmarking state-of-the-art artificial intelligence models against human psychiatrists on the national subspecialty examination in Peru: A cross-sectional study
- ,
- Jeff Huarcaya-Victoria,
- Cesar Copaja-Corzo
- ,
- Hospital Víctor Larco Herrera,
- Universidad Privada San Juan Bautista,
- Yachay - Centro de Excelencia en Salud Mental,
- Universidad San Ignacio de Loyola
Open access
Publication Information
Output type
Original language
EnglishArticle number
101179Journal (Volume, Issue Number)
Educacion Medica (Volume 27, Issue 3)Publication milestones
- Published - 01/05/2026
Publication status
ISSN
1575-1813Publication IDs
- Scopus: 105034635250
Abstract
IntroductionTo evaluate the performance of four state-of-the-art AI models (GPT-5, Claude 4.5 Sonnet, Gemini 2.5-Flash, DeepSeek V3) on psychiatry subspecialty assessments, an area that remains relatively unexplored.Materials and methodsThe AI models were benchmarked against Peru's National Psychiatry Subspecialty Examination (2022–2025, n = 400 questions) using a zero-shot prompting strategy. The comparison group consisted of 42 licensed psychiatrists.ResultsAll models exceeded 90% accuracy (range: 91.0%–94.2%), with no statistically significant differences between them (p = 0.32). AI models consistently outperformed human psychiatrists, with mean accuracy gaps ranging from 10.8 to 20.8 percentage points. Diagnostic questions yielded the highest accuracy (95.9%), while treatment items showed lower performance (88.2%–91.2%). Among the 12 concurrent model failures, 83% (10/12) were attributed to item construction issues: six to defective or ambiguous design, and four to conflicts with current medical consensus.ConclusionAI performance matches or exceeds that of psychiatrists on knowledge-based multiple-choice assessments. These findings suggest that medical education assessments should be reoriented toward clinical judgment and therapeutic reasoning competencies.
