HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding
Abstract
Recent advances in Vision-Language Models (VLMs) have led to rapid progress in video understanding across a wide range of benchmark tasks. However, existing evaluations largely focus on short-term reasoning, failing to assess a critical capability: maintaining cumulative temporal consistency over extended time horizons. To close this evaluation gap, we introduce HeiCo-FOCUS, a clinically grounded dataset for evaluating long-context video understanding through the task of Foreign Object Contextual Understanding in Surgery. Built on a dataset of Heidelberg Colorectal surgeries, this task requires models to continuously track multiple objects as they are inserted, manipulated, occluded, and removed over procedures lasting up to hours. HeiCo-FOCUS comprises 30,000 visual question answering (VQA) pairs covering five core capabilities: object recognition, temporal grounding, aggregation, event and procedural understanding, and complex reasoning. The dataset was constructed through a rigorous multi-stage annotation pipeline involving large-scale crowd annotation and 39 surgical domain experts to ensure high quality and clinical relevance. To systematically probe model behavior, we introduce a multi-track evaluation framework that progressively increases temporal and contextual demands from single frames to full procedures. Experiments with ten frontier VLMs show that HeiCo-FOCUS tasks are far from solved: only around half of the models clearly outperform a text-only baseline. Across the video tracks, models perform best on event and procedural understanding (mean Accuracy: 56.5% across all models), while temporal grounding remains particularly challenging for all evaluated models (mean Accuracy: 19.7%). We therefore expect HeiCo-FOCUS to serve as a catalyst for the development of models capable of reliable, temporally consistent reasoning over hours-long videos.
Figures & tables
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | #Videos | #VQAs | Context Length | Source | |||
| Frame | Segment | Procedure | |||||
| (30 sec t 5 min) | t 5 min | ||||||
| Cholec80-VQA ( Seenivasan et al., 2022 ) | 40 | 43,182 | 43,182 | 0 | 0 | Class. labels | |
| EndoVis-18-VQA ( Seenivasan et al., 2022 ) | 14 | 11,783 | 11,783 | 0 | 0 | Class. labels | |
| EndoVis-VQLA ( Bai et al., 2023 ) | 24 | 12,255 | 12,255 | 0 | 0 | Class. labels | |
| PSI-AVA-VQA ( Seenivasan et al., 2023 ) | 8 | 10,291 | 10,291 | 0 | 0 | Class. labels | |
| Stage | Purpose | Input | Output | #annotations (total across 30 videos) | Annotators / Method |
|---|---|---|---|---|---|
| Stage 1 Foreign object presence screening | Identification of video segments with foreign objects | 5 s long videos | Presence flag for each class of foreign objects on each clip | 69,469 annotated 5 s videos; 16% contained a foreign object | Quality Match GmbH, Häusserstraße 36, 69115 Heidelberg, Germany |
| Stage 2 Frame-level localization (bounding boxes at 1 fps) | Identification and localization of foreign objects in individual frames | Frames (sampled at 1 fps) corresponding to clips with foreign objects | Frames (sampled at 1 fps) with bounding boxes and class assignment for all foreign objects | 53,001 bounding boxes on a total of 54,528 frames | Quality Match GmbH, Häusserstraße 36, 69115 Heidelberg, Germany; iMerit, 160 West Santa Clara St., Suite 600, San Jose, CA 95113, USA |
| Stage 3 Automated bounding box completion / correction | Automatic refinement of crowd-based foreign object identification and localization | Frames (sampled at 1 fps) with bounding boxes and class assignment for all foreign objects | Automatically corrected frames (sampled at 1 fps) with bounding boxes and class assignment for all foreign objects | 4,832 corrected frames | Automated method |
| Stage 4 Instance assignment | Instance assignment to individual objects (clips only first three frames), final correction of bounding boxes, and class assignments. | Frames (sampled at 1 fps) with bounding boxes and class assignment for all foreign objects | Frames (sampled at 1 fps) with class and instance assignments and bounding boxes around all foreign objects | 79,682 bounding boxes (instance identities for all classes except clips); 324,273 curated frames | 18 Medical students |
| Stage 5 VQA generation and clinical question design | Generation of questions and answers based on foreign object information (auto-generated + expert-added) | Frames (sampled at 1 fps) with class and instance assignments, bounding boxes around all foreign objects, and anchor moments | VQA pairs balanced across the 5 core capabilities, including 15 sub-capabilities | 30,000 VQA pairs balanced across the 5 capability groups. FRAME track: 12,000 VQA pairs SEGMENT track: 12,000 VQA pairs PROCEDURE track: 6,000 VQA pairs | 18 Medical students and 7 expert surgeons; automatic method based on instance information and expert annotations |
| Stage 6 End-to-end VQA verification | Verification of the question generation pipeline | A spreadsheet (via link) containing question IDs, corresponding timestamps, and the associated surgical procedure video | The same spreadsheet completed with answers, a confidence score (1–10), optional comments, and, where necessary, a “Needs revision” flag | 500 verified questions | 14 Surgical AI researchers + 3 Medical students + 5 expert surgeons |
| Format | Counted as correct if |
| Number | the integer equals the reference |
| Yes/no | the answer equals the reference (case-insensitive) |
| Foreign-object class | the set of named classes equals the reference set (order- and case-insensitive; “none” only on its own) |
| Percentage | within 5 percentage points of the reference |
| Time point, duration | within s of the reference: 1.3, 2.3, and 4.3 s for 30 s, 2 min, and 5 min windows; 5 s for windows of 6 min or longer |
| Open-ended, multiple choice, matching | judged semantically equivalent to the reference by gpt-oss-120b; answers longer than 300 characters count as incorrect |
| FRAME | SEGMENT | PROCEDURE | |
| Human raters | 62.0 | 69.8 | 57.4 |
| Gemini 3.8 Flash | 46.0 | 57.7 | 50.5 |
| GPT-5.6 Sol | 40.0 | 55.0 | 45.5 |
| Gemini 3 Flash | 50.0 | 49.0 | 39.6 |
| Meta Muse Spark 1.3 | 30.0 | 49.0 | 42.6 |
| Gemini 3.5 Flash Lite | 44.0 | 46.3 | 42.6 |
| Annotation issue/challenge | Sample image |
|---|---|
| Sponges tucked deep into the surgical site are challenging to identify and link. | |
| Staples can be confused with clips and are infeasible to count. | |
| Specimen and clips attached to them should not be annotated as foreign objects as soon as they are placed in the specimen bag. | |
| Due to inadequate lighting at the image borders, distinguishing between shadows and blood-soaked sponges is challenging. | |
| For several sponges, it was difficult to decide whether they had been removed from the surgical site or remained in situ only to reappear later, particularly when the video was interrupted by a prolonged blue screen. |