Large Language Models Not Yet Safe for Autonomous Clinical Decisions in Real-World Healthcare
ai-safety autonomous reasoning
| Source: ArXiv | Original article
Large language models pass medical exams, but aren't yet safe for autonomous clinical decisions. They can rival physicians in diagnostic reasoning, but have limitations.
Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support
Large language models have made significant strides in medical licensing examinations and diagnostic reasoning, rivaling physicians in certain cases. However, their use in autonomous clinical decision support is not yet safe. The issue lies not in medical knowledge, but in the fidelity of clinical evaluation. Models optimized for probable text continuation are not optimized for safe decision-making, particularly in cases where the correct answer is an improbable but crucial diagnosis.
This limitation is crucial, as large language models face challenges in real-world clinical settings due to the safety-critical and context-dependent nature of medical decision-making. While they perform well on medical exams, their performance in actual clinical settings, such as emergency departments, is less reliable. Furthermore, biases in large language models, including sex bias, can manifest in diagnoses, making their deployment in AI-enabled care problematic.
As researchers continue to develop and refine large language models for medical applications, it is essential to address these safety concerns. The development of models like DrAgent, which aims to empower large language models as medical agents, is a step in the right direction. However, more work is needed to ensure that these models can safely and effectively support clinical decision-making in real-world settings.
Sources
Back to AIPULSEN