Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning
Organizations: School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, United Arab Emirates
Abstract
Multi-instance partial-label learning (MIPL) addresses inexact supervision in both the instance and label spaces, which can be applied to video classification. However, bag-level labels do not explicitly supervise the correspondence between candidate classes and temporal evidence. We propose {\ours}, which couples label disambiguation with temporal evidence allocation through a joint class--time assignment. Occupancy-regularized spherical matching associates contextualized video features while learning nonuniform temporal mass and discouraging excessive concentration. During training, candidate-restricted inference recomputes the assignment within the candidate label set. A dual-marginal KL projection then constructs a structured teacher that incorporates momentum-refined class beliefs while preserving the proposal's temporal occupancy. A single plan-level KL objective aligns the full-space predictor with this teacher. Our analysis characterizes when candidate re-solving differs from masking and shows that, under the stated construction, the joint objective decomposes into class-marginal and class-conditional temporal supervision. We construct VCMIPL benchmarks from Breakfast, DoTA, and FineAction using model-generated candidate labels and evaluate the method across four feature representations. Extensive experimental results demonstrate that PIVOTMIPL outperforms existing MIPL algorithms in both effectiveness and efficiency.
Figures & tables
| 1 | Initialize parameters and . |
|---|---|
| 2 | For each epoch and each training bag : |
| 3 | Encode valid ordered positions and compute . |
| 4 | Solve full-space actor . |
| 5 | Solve proposal . |
| 6 | Update class belief using Eq. ( 16 ). |
| 7 | Scale the proposal to using Eq. ( 18 ). |
| Dataset | Method | DINOv3 | VideoMAEv2 | ResNet | SlowFast |
|---|---|---|---|---|---|
| Breakfast -MIPL | DeMipl | ||||
| EliMipl | |||||
| MiplMa | |||||
| ProMipl | |||||
| PsMipl | |||||
| DualG |
| Dataset | Feature | PivotMipl | A1 | A2 | A3 | A4 | A5 | A6 |
|---|---|---|---|---|---|---|---|---|
| Breakfast -MIPL | DINOv3 | |||||||
| VideoMAEv2 | ||||||||
| ResNet | ||||||||
| SlowFast | ||||||||
| DoTA -MIPL | DINOv3 | |||||||
| VideoMAEv2 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Videos | Classes | Mean | Range |
|---|---|---|---|---|
| Breakfast | 1,989 | 10 | 3.3896 | 1–9 |
| DoTA | 4,577 | 9 | 3.2733 | 1–8 |
| FineAction | 10,675 | 78 | 11.5592 | 2–64 |
| Feature | Total instances | Maximum | Minimum | Mean | Dimension |
| Breakfast: 1,989 videos, 10 classes, 100 fine-tuning epochs | |||||
| Candidate size: 1–9, mean 3.3896 | |||||
| ResNet | 4,103,316 | 9,745 | 186 | 2,063.00 | 2,048 |
| DINOv3 | 4,103,316 | 9,745 | 186 | 2,063.00 | 4,096 |
| SlowFast | 1,026,577 | 2,437 | 47 | 516.13 | 2,304 |
| VideoMAEv2 | 1,026,577 | 2,437 | 47 | 516.13 | 768 |
| Setting | Value or implementation |
| Framework / hardware | PyTorch / NVIDIA H100 |
| Optimizer | AdamW; ; weight decay |
| Learning rates | Prototypes: ; temporal encoder: |
| Gradient clipping | Global norm clipped to |
| Training budget | epochs |
| Motion module | Average-filter kernel ; convolution kernel ; dilations ; motion-residual dropout |
| Dataset | Method | DINOv3 | VideoMAEv2 | ResNet | SlowFast |
|---|---|---|---|---|---|
| Breakfast -MIPL | DeMipl | ||||
| EliMipl | |||||
| MiplMa | |||||
| ProMipl | |||||
| PsMipl | |||||
| DualG |
| Variant | Modification | Full-model wins | Mean gain |
|---|---|---|---|
| A1 | No occupancy regularization | 10/12 | |
| A2 | Uniform temporal occupancy | 9/12 | |
| A3 | Class-level KL | 12/12 | |
| A4 | Class-marginal-only teacher | 12/12 | |
| A5 | Single prototype per class | 11/12 | |
| A6 | No temporal encoder | 11/12 |
| Dataset | Feature | AgopMipl | AgopMipl Temporal Encoder | PivotMipl |
|---|---|---|---|---|
| Breakfast | DINOv3 | |||
| VideoMAEv2 | ||||
| ResNet | ||||
| SlowFast | ||||
| DoTA | DINOv3 | |||
| VideoMAEv2 |