别只看用理想病人测AI的论文,这篇用2053个真实对话发现,交流风格直接改变分诊结果,准确率才55%!
该论文分析了2053个真实患者与AI聊天机器人的对话,发现用户交流模式和情绪表达差异显著。研究开发了患者模拟器,分别建模临床内容、情感状态、对话策略和交流风格。在15位人类评分员的图灵测试中,模拟对话与真实对话几乎无法区分,准确率仅55%。使用5种患者角色和1164个临床医生评分案例,评估了4个LLM的紧急评估性能,发现交流风格显著改变分诊结果。
The complexities of patient-centred conversational artificial intelligence
Consumer-facing health chatbots powered by large language models (LLMs) are increasingly used for symptom assessment. However, chatbot development and evaluation often rely on cooperative, articulate, simulated patients. We analysed 2,053 real patient-chatbot conversations and found that communication patterns and expression of emotions vary widely across users. We developed a patient simulator that separately models clinical content, emotional state, conversational strategy, and communication style. In a Turing-inspired evaluation of realism with 15 human graders, simulated conversations were nearly indistinguishable from real ones, with human graders achieving an accuracy of 55%. We used five distinct patient personae, across 1,164 clinician-graded cases, to evaluate the performance of four LLMs in urgency assessment. We found that communication style can significantly alter triage outcomes. Patient-centred conversational artificial intelligence must accommodate communication diversity: systems designed for idealised, rather than realistic, interactions risk underperforming and amplifying health disparities when deployed in the real world.