Skip to search boxSkip to navigationSkip to main content

Benchmarking state-of-the-art artificial intelligence models against human psychiatrists on the national subspecialty examination in Peru: A cross-sectional study

  • ,
  • Hospital Víctor Larco Herrera
    ,
  • Universidad Privada San Juan Bautista
    ,
  • Yachay - Centro de Excelencia en Salud Mental
    ,
  • Universidad San Ignacio de Loyola
Research Output:
Contribution to journal
Article
Peer-review

Open access

Publication Information

Output type

Research Output:
Contribution to journal
Article
Peer-review

Original language

English

Article number

101179

Journal (Volume, Issue Number)

Educacion Medica (Volume 27, Issue 3)

Publication milestones

  • Published - 01/05/2026

Publication status

Published - 01/05/2026

ISSN

1575-1813

Publication IDs

  • Scopus: 105034635250

Abstract

IntroductionTo evaluate the performance of four state-of-the-art AI models (GPT-5, Claude 4.5 Sonnet, Gemini 2.5-Flash, DeepSeek V3) on psychiatry subspecialty assessments, an area that remains relatively unexplored.Materials and methodsThe AI models were benchmarked against Peru's National Psychiatry Subspecialty Examination (2022–2025, n = 400 questions) using a zero-shot prompting strategy. The comparison group consisted of 42 licensed psychiatrists.ResultsAll models exceeded 90% accuracy (range: 91.0%–94.2%), with no statistically significant differences between them (p = 0.32). AI models consistently outperformed human psychiatrists, with mean accuracy gaps ranging from 10.8 to 20.8 percentage points. Diagnostic questions yielded the highest accuracy (95.9%), while treatment items showed lower performance (88.2%–91.2%). Among the 12 concurrent model failures, 83% (10/12) were attributed to item construction issues: six to defective or ambiguous design, and four to conflicts with current medical consensus.ConclusionAI performance matches or exceeds that of psychiatrists on knowledge-based multiple-choice assessments. These findings suggest that medical education assessments should be reoriented toward clinical judgment and therapeutic reasoning competencies.