Background: Large language models (LLMs) have achieved physician-comparable performance on multiple-choice medical knowledge examinations, but their capabilities in clinical history taking, urgency assessment, and safety remain insufficiently evaluated.
We proposed Japanese Stroke LLM Evaluation, a multi-turn conversational benchmark for stroke care in Japanese, and evaluated LLM performance and safety under practice-oriented conditions.
Methods:
We created 10 stroke and related-condition cases and evaluated LLMs in multi-turn Japanese conversations. The LLM acted as physician, while a board-certified neurosurgeon acted as simulated patient and evaluator.
Each case comprised history-taking and action phases scored using pre-specified criteria. Errors that could directly threaten life were defined as critical mistakes.
The safety threshold was at least 80% overall with zero critical mistakes.
Eighteen models were evaluated in October 2025 and June 2026.
Results:
Claude Fable 5 achieved the highest score (87.4%) with zero critical mistakes, followed by Claude Opus 4.7 (80.3%) and GLM-5.2 (75.6%).
Two leaders met the safety threshold.
Eleven models made 17 critical mistakes, including failure to confirm laboratory results or blood glucose before t-PA, surgery before airway stabilization, omission of cervical vascular evaluation, and t-PA outside its indication.
History-taking question count correlated with history-taking score (r = 0.648, p = 0.007).
Conclusions:
Japanese Stroke LLM Evaluation provides a benchmark for LLM performance under practice-oriented conditions, including a cap on history-taking questions.
Cases and evaluations were created by neurosurgical specialists rather than using an LLM-as-judge approach.
Performance improved across cloud-based and on-premise models in 2026, with some exceeding the safety threshold.
Further evaluation using real-world cases is required.