cs.CVSep 26, 2026

AnesTRACE: Benchmarking Intraoperative Anesthesia from Multimodal Perception to Multi-step Decision-Making

Authors: Ziwei Huang, Qi Gao, Zhe Ji, Yuanyuan Yao, Fengjiang Zhang, Min Yan, Zhongle Xie, Gang Chen

Organizations: School of Software Technology, Zhejiang University · Department of Anesthesiology, The Second Affiliated Hospital of Zhejiang University School of Medicine · School of Software, Central South University · Zhejiang University

Abstract

Intraoperative anesthesia requires systems to interpret evolving multimodal evidence, recommend timely management, and revise decisions as patient states change, yet existing benchmarks usually isolate perception or single-point reasoning. We introduce AnesTRACE, an evaluation suite comprising AnesTRACE-Bench and AnesTRACE-Eval. Built from public perioperative datasets with anesthesiologist annotation, AnesTRACE-Bench evaluates Intraoperative Perception, Single-point Anesthesia Decision-Making, and Multi-step Anesthesia Decision-Making. AnesTRACE-Eval assesses open-ended responses through anesthesiologist-defined criteria for Clinical Correctness, Evidence Grounding, Task Completeness, and Safety, with Temporal Consistency for multi-step decisions; its domain-specific evaluator is trained by supervised fine-tuning and preference alignment on expert-reviewed judgments. Across more than 30 models, fine-grained visual grounding and intervention selection remain difficult: the leading model reaches only 32.2 mIoU for TEE visual grounding and retains a 17.5% Major/Critical Safety Error Rate in multi-step management. Evaluator alignment with anesthesiologists improves across both training stages, while the best decision quality is accompanied by a 74.3-second P95 Latency. These results show that aggregate performance alone does not establish safe, timely longitudinal decision-making. We release our code at https://zjudbxai.github.io/AnesTRACE/.

Explore similar work

CardsList
  1. MedTRACE: Tool-Augmented Multimodal Clinical Reasoning Agents for Evidence-Grounded Decision-Making

    Sep 13, 2026Ji Lu, Lifei Liu, Haoran Yu +5Multimodal Clinical DataMultimodal Reasoning

  2. Evaluation Awareness Is Not One Capability: Evidence from Open Language Models

    Jun 22, 2026Nilesh Nayan, Aishwarya Sampath Kumar, Rishiraj Girmal +5Safety Benchmarks

  3. TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos

    May 8, 2026Hengyi Feng, Hao Liang, Mingrui Chen +6Audio-Visual ReasoningMultimodal Reasoning