A cross-sectional evaluation of large language model chatbot interfaces for patient-facing herpes zoster information: safety, information quality, and readability

Scritto il 01/10/2026
da Dalian Liang

Front Public Health. 2026 Sep 16;14:1955701. doi: 10.3389/fpubh.2026.1955701. eCollection 2026.

ABSTRACT

BACKGROUND: Large language model (LLM) chatbots are increasingly used to obtain health information. However, fluent and clinically plausible responses may still contain safety-relevant omissions, inadequate source attribution and disclosure, or difficult-to-read text.

OBJECTIVE: To evaluate the safety, accuracy, empathy, information quality, response-level source attribution and disclosure, overall quality, and readability of five LLM chatbot interfaces answering lay-oriented questions about herpes zoster.

METHODS: This cross-sectional evaluation used 46 standardised English-language questions across seven herpes zoster domains. Each question was submitted once to ChatGPT-5.5 Instant, DeepSeek-V4-Pro, Doubao-Seed-2.0-Pro, Gemini 3.5 Flash, and Qwen3.7-Plus under consumer-access conditions on June 2-3, 2026. Five dermatologists independently assessed safety, accuracy, empathy, DISCERN, Ensuring Quality Information for Patients (EQIP), Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Score (GQS). Six readability indices were calculated. Paired comparisons and inter-rater agreement were evaluated using prespecified statistical methods.

RESULTS: The five interfaces generated 230 complete responses. Under the study conditions, 66 responses (28.7%) were classified as potentially unsafe using the prespecified ≥3/5-rater majority threshold, with no significant difference between interfaces (Cochran's Q = 1.661, df = 4, p = 0.798). Sensitivity analyses using alternative thresholds showed consistent results. Accuracy and empathy differed across interfaces (both p < 0.001; Kendall's W = 0.127 and 0.211, respectively), although effect sizes were small. Information-quality and overall-quality measures also showed outcome-specific differences. All six readability indices differed across interfaces (all p < 0.001); DeepSeek-V4-Pro generally produced text estimated to be easier to read, whereas Gemini 3.5 Flash produced text estimated to be more difficult to read. Safety agreement was substantial (Fleiss' κ = 0.682), and ICC (2,1) values for other manually rated outcomes ranged from 0.832 to 0.889.

CONCLUSION: In this standardised benchmark, LLM chatbot interfaces provided generally favourable accuracy and information-quality scores but showed clinically relevant safety limitations, sparse source attribution and disclosure, and readability challenges. No interface consistently performed best across all outcomes. These findings represent single first responses obtained under the tested conditions and do not establish response stability across repeated queries. LLM-generated herpes zoster information should therefore be interpreted cautiously and should not replace individualised professional assessment or professionally reviewed patient information.

PMID:42819370 | PMC:PMC13623941 | DOI:10.3389/fpubh.2026.1955701