首页 > AI前沿 > Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models

Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models

arXiv自然语言 2026-09-15 15:13 5 阅读 查看原文

Background: Large language models (LLMs) have achieved physician-comparable performance on multiple-choice medical knowledge examinations, but their capabilities in clinical history taking, urgency assessment, and safety remain insufficiently evaluated.

We proposed Japanese Stroke LLM Evaluation, a multi-turn conversational benchmark for stroke care in Japanese, and evaluated LLM performance and safety under practice-oriented conditions.

Methods:

We created 10 stroke and related-condition cases and evaluated LLMs in multi-turn Japanese conversations. The LLM acted as physician, while a board-certified neurosurgeon acted as simulated patient and evaluator.

Each case comprised history-taking and action phases scored using pre-specified criteria. Errors that could directly threaten life were defined as critical mistakes.

The safety threshold was at least 80% overall with zero critical mistakes.

Eighteen models were evaluated in October 2025 and June 2026.

Results:

Claude Fable 5 achieved the highest score (87.4%) with zero critical mistakes, followed by Claude Opus 4.7 (80.3%) and GLM-5.2 (75.6%).

Two leaders met the safety threshold.

Eleven models made 17 critical mistakes, including failure to confirm laboratory results or blood glucose before t-PA, surgery before airway stabilization, omission of cervical vascular evaluation, and t-PA outside its indication.

History-taking question count correlated with history-taking score (r = 0.648, p = 0.007).

Conclusions:

Japanese Stroke LLM Evaluation provides a benchmark for LLM performance under practice-oriented conditions, including a cap on history-taking questions.

Cases and evaluations were created by neurosurgical specialists rather than using an LLM-as-judge approach.

Performance improved across cloud-based and on-premise models in 2026, with some exceeding the safety threshold.

Further evaluation using real-world cases is required.