JMIR Med Educ. 2026 Aug 14;12:e92826. doi: 10.2196/92826.
ABSTRACT
BACKGROUND: Strengthening the global health workforce is central to achieving universal health coverage, but health systems cannot improve what they cannot measure. Valid and scalable assessment of clinical competency is essential for monitoring workforce readiness and ensuring that expanded service coverage translates into high-quality care. Traditional standardized patients, however, remain resource-intensive, difficult to scale, and vulnerable to evaluator-related bias. Recent advances in AI have enabled AI-led simulated standardized patients (SSPs) that may address these limitations.
OBJECTIVE: Although existing reviews have examined AI-SSPs broadly within medical education, their use as clinical competency assessment systems remains less well characterized. This scoping review aimed to address this gap by systematically mapping the scope, design features, and validation approaches of AI-SSP tools for clinical competency assessment.
METHODS: We registered the protocol prospectively with the Open Science Framework and conducted a scoping review following the Joanna Briggs Institute Manual for Evidence Synthesis. We searched MEDLINE, Embase, CINAHL Complete, Education Source, Web of Science Core Collection, Inspec, ERIC, and Cochrane CENTRAL, supplemented by manual searching. Eligible primary studies involved an AI-based SSP interaction that yielded a quantitative score of respondents' clinical competency. Two reviewers independently screened records and resolved conflicts through discussion. Data were charted on study characteristics and populations, frontend platform and interface features, backend AI architectures, performance scoring mechanisms, and tool evaluation methods and outcomes.
RESULTS: Of 5987 database records and 4 identified through other methods, 21 studies describing 20 unique AI-SSP systems were included. Studies were published between 2008 and 2026, with 15 (71%) published in 2024 or later. Most studies were conducted in high-income settings and involved medical students. Systems shifted over time from rule-based virtual patient to large language model-enabled conversational platforms after 2022. History-taking was the most frequently assessed competency, and checklist coverage was the most common scoring approach. Among studies comparing AI-generated scores with human ratings, findings were mixed: some reported moderate to strong agreement, while others showed score inflation and inconsistent performance. Most studies focused on feasibility; fewer evaluated large-scale use or sustained real-world implementation.
CONCLUSIONS: This review is innovative in shifting attention from AI-SSPs as educational simulations to AI-SSPs as emerging infrastructures for clinical competency assessment. Unlike prior reviews of virtual patients or AI in medical education, this review maps not only system design and learner interaction but also which competencies are assessed, how scores are generated, and how the assessment evidence has been validated. AI-SSPs could make clinical competency assessment more scalable and potentially more standardized, although more evidence is still needed to determine whether they achieve validity and reliability across diverse learner or practitioner groups.
PMID:42605156 | DOI:10.2196/92826

