cs.CVSep 29, 2026

EGSD: Event-Grounded Self-Distillation for Streaming Video Understanding

Authors: Yuwei Miao, Xuesheng Zhang, Wenhao Zou, Jixia Zhang, Jianwei Lv, Bo Yuan, Junfeng Wang, Shiao Xie

Organizations: Baidu · University of Chinese Academy of Sciences · China University of Petroleum, Beijing

Abstract

Real-time video understanding requires incrementally maintaining a memory of streaming content, and optimizing this requires dense process signals. On-Policy Self-Distillation (OPSD), which lets one model serve as both teacher and student with the teacher receiving additional privileged information such as the question and ground-truth (GT) answer, can supply such token-level signals. However, applying it directly to streaming video raises two problems. (1) The student cannot be optimized end-to-end, where memory is written before the question arrives, yet the teacher scores it with the question-and-GT privilege, misaligning their preferences. (2) Effective-entity memory collapses, where the question-and-GT privilege makes the teacher favor only question-relevant entities, and token-mean averaging over a memory renders its signal invariant to how many entities that memory covers, both driving memory against the streaming need for diversity. To address these issues, we propose Event-Grounded Self-Distillation (EGSD), which characterizes streaming memory as an incremental update over verifiable Events (key visual entities, actions, and details) and targets the two problems on this basis. For problem (1), we adapt the OPSD signal into a multiplicative weight combined with the outcome reward; for problem (2), we re-weight the teacher with Events as privileged information to counter its question-relevance bias, and add an entity-coverage reward to supply the coverage preference the token-mean teacher lacks. Extensive experiments on mainstream online and offline benchmarks show EGSD achieves strong performance, reaching 79.8% on StreamingBench and 73.4% on the OVO-Bench Real-Time track, while memory analysis shows effective-entity recall rises 17.4% at only 6.8% more memory length.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. StreamScout: Learning When to Look Deeper for Streaming Video Understanding

    Aug 31, 2026Ce Zhang, Jing Bi, Jinxi He +9Streaming Video UnderstandingBounded-Memory

  2. MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding

    Sep 30, 2026Yinying Li, Yuqian Fu, Yulin Dai +3Streaming Video UnderstandingStreaming Video

  3. What Should a Streaming Video Model Remember?

    Jun 15, 2026Haonan Ge, Yiwei Wang, Hang Wu +1Streaming Video UnderstandingStreaming