RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation

May 13, 20262605.13542

Chengzhi Shen, Weixiang Shen, Tobias Susetzky, Chen, Chen + 6 more

cs.AIcs.CLcs.LGcs.MA

TLDR

RealICU is a new benchmark for evaluating LLM agents on long-context ICU data, revealing recall-safety tradeoffs and anchoring biases in existing models.

Key contributions

Introduces RealICU, a hindsight-annotated benchmark for LLMs evaluating realistic ICU decision support.
Formulates four physician-motivated tasks: Patient Status, Acute Problems, Recommended, and Red Flag Actions.
Releases RealICU-Gold (930 windows) and RealICU-Scale (11,862 windows via Oracle LLM) datasets.
Exposes LLM failures like recall-safety tradeoffs and anchoring bias in long-context ICU data.

Why it matters

This paper introduces a critical benchmark, RealICU, for evaluating LLMs in high-stakes ICU settings, moving beyond suboptimal historical clinician actions. It highlights significant limitations in current LLMs, such as safety tradeoffs and biases, pushing for more reliable AI decision support in critical care.

Original Abstract

Intensive care units (ICU) generate long, dense and evolving streams of clinical information, where physicians must repeatedly reassess patient states under time pressure, underscoring a clear need for reliable AI decision support. Existing ICU benchmarks typically treat historical clinician actions as ground truth. However, these actions are made under incomplete information and limited temporal context of the underlying patient state, and may therefore be suboptimal, making it difficult to assess the true reasoning capabilities of AI systems. We introduce RealICU, a hindsight-annotated benchmark for evaluating large language models (LLMs) under realistic ICU conditions, where labels are created after senior physicians review the full patient trajectory. We formulate four physician-motivated tasks: assess Patient Status, Acute Problems, Recommended Actions, and Red Flag actions that risk unsafe outcomes. We partition each trajectory with 30-min windows and release two datasets: RealICU-Gold with 930-window annotations from 94 MIMIC-IV patients, and RealICU-Scale with 11,862 windows extended by Oracle, a physician-validated LLM hindsight labeler. Existing LLMs including memory-augmented ones performed poorly on RealICU, exposing two failure modes: a recall-safety tradeoff for clinical recommendations, and an anchoring bias to early interpretations of the patient. We further introduce ICU-Evo to study structured-memory agents that improves long-horizon reasoning but does not fully eliminate safety failures. Together, RealICU provides a clinically grounded testbed for measuring and improving AI sequential decision-support in high-stakes care. Project page: https://chengzhi-leo.github.io/RealICU-Bench/

View on arXiv Download PDF

📬 Weekly AI Paper Digest

Get the top 10 AI/ML arXiv papers from the week — summarized, scored, and delivered to your inbox every Monday.

TLDR

Key contributions

Why it matters

Original Abstract

📬 Weekly AI Paper Digest

Related papers