Skip to search boxSkip to navigationSkip to main content

ChatGPT as a global doctor: a rapid review of its performance on national licensing medical examination

*Corresponding author for this work
  • ,
  • Hospital Víctor Larco Herrera
    ,
  • Universidad Privada de Tacna
Research Output:
Contribution to journal
Review article
Peer-review

Open access

Publication Information

Output type

Research Output:
Contribution to journal
Review article
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 284-293 (10 pages)

Journal (Volume, Issue Number)

Acta Medica Peruana (Volume 42, Issue 4)

Publication milestones

  • Published - 01/10/2025

Publication status

Published - 01/10/2025

ISSN

1018-8800

Publication IDs

  • Scopus: 105030997668

Abstract

Objective: To evaluate ChatGPT's performance on NLMEs worldwide and determine whether it could achieve licensure to practice medicine across different countries. Methods: We searched PubMed, Scopus, and Google Scholar for studies evaluating ChatGPT's performance on NLMEs. Reference lists of included studies were also reviewed. Two reviewers independently screened studies and extracted the accuracy rates (performance) of GPT-3.5 and GPT-4, including those that passed thresholds, human examinee scores, and other study characteristics. The risk of bias was assessed using the JBI Critical Appraisal Checklist for Prevalence Studies. Results: We identified 37 studies evaluating ChatGPT's performance across 18 NLMEs. Most studies assessed the United States, Chinese, and Japanese examinations. While most studies used official datasets, others relied on unofficial third-party sources, and few employed advanced prompting techniques. GPT-4 was superior to GPT-3.5 in all NLMEs, with accuracy rates ranging from 67% to 89%. GPT-4 passed all 18 NLMEs (100%), while GPT-3.5 passed 10 of 15 (67%). Compared to human examinees, GPT-4 outperformed the average score in 6 of 7 NLMEs (86%); the sole exception was Japan, where examinees achieved 84.9% versus 81.5% for GPT-4. Conclusion: Current evidence demonstrates that GPT-4 can pass all 18 NLMEs evaluated, surpassing human examinees in most cases. However, this finding likely reflects low passing thresholds rather than AI superiority over physicians.

Sustainable Development Goals

  • SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well