cs.CVOct 5, 2026

Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning

Authors: Lingyu Shen, Wei Tang, Fakhri Karray, Min-Ling Zhang

Organizations: School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, United Arab Emirates

Abstract

Multi-instance partial-label learning (MIPL) addresses inexact supervision in both the instance and label spaces, which can be applied to video classification. However, bag-level labels do not explicitly supervise the correspondence between candidate classes and temporal evidence. We propose {\ours}, which couples label disambiguation with temporal evidence allocation through a joint class--time assignment. Occupancy-regularized spherical matching associates contextualized video features while learning nonuniform temporal mass and discouraging excessive concentration. During training, candidate-restricted inference recomputes the assignment within the candidate label set. A dual-marginal KL projection then constructs a structured teacher that incorporates momentum-refined class beliefs while preserving the proposal's temporal occupancy. A single plan-level KL objective aligns the full-space predictor with this teacher. Our analysis characterizes when candidate re-solving differs from masking and shows that, under the stated construction, the joint objective decomposes into class-marginal and class-conditional temporal supervision. We construct VCMIPL benchmarks from Breakfast, DoTA, and FineAction using model-generated candidate labels and evaluate the method across four feature representations. Extensive experimental results demonstrate that PIVOTMIPL outperforms existing MIPL algorithms in both effectiveness and efficiency.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Calibratable Disambiguation Loss for Multi-Instance Partial-Label Learning

    Dec 19, 2025Wei Tang, Yin-Fang Yang, Weijia Zhang +1Multiple Instance LearningDisambiguation

  2. Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence

    May 5, 2026Zhiyuan Li, Rongzhen Zhao, Wenyan Yang +3Object-Centric Representations

  3. Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval

    May 7, 2026Jun Li, Peifeng Lai, Xuhang Lou +5Fine-Grained Video UnderstandingMultimodal Retrieval