Continuous video counting requires distinguishing new observations from new objects or completed events. We introduce StaMina (State Maintenance), which learns to maintain counting state through state-conditioned updates. Recurrent visual context supports recognition; learned transitions maintain visibility, persistent identities, and completed-event records. A differentiable recurrence trains event transitions over legal paths constrained by count endpoints; visibility and association objectives train the object branch. A multi-source pipeline organizes 39.8K spatial queries and complementary event annotations into counting trajectories. On SVCBench, we evaluate counting adaptation with partial video overlap and held-out groups of linked annotations. Under prefix replay (Full) and persistent streaming (Stream), 4B and 8B models reach 41.9/36.4 and 44.9/38.2 Gaussian Precision Accuracy, respectively. The 8B model gains 10.9/3.2 points over Counting-SFT on the same queries. Matched-graph comparisons isolate phase conditioning and trajectory supervision, assessing training objectives alongside hard decisions. Online video benchmarks and count-conditioned decisions assess online understanding and task eligibility. Project Page: https://PLACEHOLDER.github.io/StaMina/
Figures & tables
Figure 2 : StaMina architecture and state-transition learning. (a) Short-window attention and delta memory supply causal features for event transitions and independent visibility/identity updates. DEFER postpones identity association while preserving visible candidates. Accepted updates maintain counting state and an evidence log, with separate readouts for general QA and counting. Previous state conditions transition decisions. (b) The event graph defines legal changes. Training sums the probabilities of paths matching supported count endpoints; online inference selects a legal outgoing edge at each observation.
Figure 3 : Counting-data construction. Source-specific evidence checks produce canonical trajectories and causal queries, with inherited provenance and supervision-family exclusion.
Model
Access
O1
O2
E1
E2
Overall
Snap
Delta
Unique
Gain
Action
Transit
Episode
Periodic
GPA
MoC
UDA
Human reference
Human
Human
96.7
100.0
94.5
100.0
94.9
98.3
97.0
93.2
96.1
100.0
99.3
Text-only reference
GPT-4-Turbo
Text
15.7
19.4
15.8
50.0
13.1
23.6
21.9
0.0
18.7
95.7
4.3
Proprietary VLMs
Table 1 : SVCBench counting (%, ↑ ). StaMina adapts with partial source-video overlap and held-out supervision families; references retain their protocols. Groups denote weight availability; Access denotes Full/Stream. Bold/underline mark best/second-best per access mode, separately for reference and adapted rows; ties share rank. Protocols: Appendix C ; shared-query controls: Table 18 .
Phase input
Endpoint objective
GPA ↑
Hard exact ↑
False updates ↓
Off
Mean
40.8
40.1
1.1
On
Mean
46.3
45.7
0.5
Off
Path
50.7
49.9
0.8
On
Path
65.1
64.4
0.4
Table 2: Event-head controls on frozen 8B Stream features. The diagnostic inputs and legal graph are fixed. GPA, hard prefix exact match, and false-update rate are percentages; a lower false-update rate is better.
Figure 4 : Hard-decision diagnostics at 8B. Event errors distinguish repeated or missed completions from stable-interval updates. Object endpoints compare the differentiable cumulative count with the accepted identity count under the same rollout. Detailed denominators and supervision strata are given in Appendix G .
Model
Access
OVO-Bench Avg.
OVO-S L-Avg.
Qwen3-VL-4B
Full
50.81
46.93
Qwen3-VL-8B
Full
55.02
48.58
StaMina-4B (Ours)
Full
51.66
48.57
StaMina-8B (Ours)
Full
55.46
50.32
StreamForest-7B
Stream
55.57
44.13
Flash-VStream-7B
Stream
33.15
24.94
Table 3: Online video understanding ( ↑ ). OVO-Bench Avg. averages three task groups; OVO-S L-Avg. averages four spatial levels. Full and Stream are ranked separately; evaluation protocols are specified in Appendix K .
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Memory heads × width; block stride
8×64 ; 4
Transition/association hidden width
256 (GELU)
Pooling queries; entity-event slots
32; 4
Identity shortlist; visibility threshold
8; 0.5
Deduplication/track cosine thresholds
0.95/0.5
Token chunk; continuous truncation
128; 8 chunks
Appendix
Table 4: Default complete-system configuration for both model scales. Pretrained dimensions follow Qwen3-VL; the frozen-input event-head controls use the training protocol in Appendix E .
Operator
Queries
O1-Snap
16379
O1-Delta
6746
O2-Unique
10627
O2-Gain
6077
E1-Action
4606
E1-Transit
885
Appendix
Table 5: Composition of the multi-source pool by counting operator. Entries count query–answer pairs, not independent source videos.
Phase
Objective
GPA ↑
Exact ↑
NLL ↓
Duplicate ↓
Miss ↓
Delay ↓
Off
Mean
40.8
40.1
0.93
2.8
17.4
1.11
On
Mean
46.3
45.7
0.94
1.2
16.6
0.93
Off
Path
50.7
49.9
0.67
1.8
10.7
1.15
On
Path
65.1
64.4
0.58
1.1
7.0
0.89
Appendix
Table 6: Event hard decisions and trajectory fit. GPA, exact match, duplicates, and misses are percentages. NLL is normalized by the number of supervised endpoints and delay is in seconds.
Phase
Objective
Source stratum
GPA ↑
Exact ↑
NLL ↓
Off
Mean
Endpoint only
38.1
37.4
0.95
Off
Mean
Boundary supported
43.5
42.8
0.90
On
Mean
Endpoint only
41.6
41.0
0.98
On
Mean
Boundary supported
51.0
50.3
0.90
Off
Path
Endpoint only
45.3
44.5
0.71
Off
Path
Boundary supported
56.2
55.3
0.62
Appendix
Table 7: Event diagnostics by source supervision. Evaluation annotations are held out; each source stratum is ranked separately.
Identity loss
Source stratum
Soft U
Hard U
Soft Gain
Hard Gain
λi=0
Count only
0.56
0.89
0.55
0.68
λi=0
Identity verified
0.56
0.94
0.56
0.69
Complete
Count only
0.40
0.69
0.39
0.50
Complete
Identity verified
0.23
0.43
0.25
0.31
Appendix
Table 8: Soft and hard object errors. Each column reports MAE ( ↓ ) on the same endpoint or interval queries; ranks compare the two training conditions within each source stratum.
Identity loss
NEW precision ↑
NEW recall ↑
Re-entry ↓
Unresolved ↓
Delay ↓
λi=0
88.1
87.9
11.8
18.3
3.04
Complete
96.3
95.1
3.7
7.1
1.08
Appendix
Table 9: Identity decisions on verified tracks. Precision, recall, re-entry double-counting and unresolved rates are percentages; delay is in seconds.
Access
GPA ↑
Hard exact ↑
Full
63.2
62.4
Stream
60.9
60.1
Appendix
Table 10: Matched-timestamp access on the event diagnostic collection. Both conditions use the complete 8B model.
Selection statistic
Count
Source train pool
47882
OVO media exclusion
2710
Dependency-family exclusion
2316
Graph-incompatible exclusion
694
Other unsupported records
550
Selected training queries
41612
Appendix
Table 11: Training selection from the source pool. Counts refer to queries unless a row specifies parents.
Supervision
Selected queries
Endpoint only
26068
Verified identity
11328
Verified boundary
4216
Verified stable interval
12894
Appendix
Table 12: Supervision coverage in the selected training set. The first three strata are mutually exclusive; verified stability overlaps them.
Source
Audited records
Endpoint correct (%)
Increment correct (%)
Objects
320
93.1
89.4
Natural events
240
88.8
85.4
Periodic
240
96.7
93.8
Appendix
Table 13: Stratified visual-label audit. Each record contributes one endpoint and one interval judgment; both rates use the audited-record count as denominator.
Operator
Endpoint queries
Full upper coverage
Stream upper coverage
Action
1281
100.0
100.0
Transit
205
100.0
100.0
Episode
513
100.0
100.0
Periodic
280
44.3
58.2
Appendix
Table 14: Observation-budget capacity coverage on official event endpoints (%). This is an upper bound on reachable coverage; all endpoints remain in the official score.
Access / population
Configuration
GPA ↑
Full / 3706 queries
First-stage adaptation
42.01
Full / 3706 queries
Two-stage adaptation
55.61
Full / 3706 queries
Two-stage, branch gates zero
30.65
Stream / 866 queries
Recurrent initialization
31.91
Stream / 866 queries
Fixed fast weights
26.12
Stream / 866 queries
Full-adapted initialization
36.80
Appendix
Table 15: Continuous-component diagnostics. Full rows evaluate the adaptation split; Stream rows evaluate the development split. Each block retains its precision and training setup and is ranked separately.
Condition
Full ↑
Stream ↑
Complete model
44.9
38.2
Threshold updates in place of FSM
38.2
34.4
Clear count projection
6.1
6.1
Cross-video count projection
12.7
15.6
Previous-endpoint projection
16.1
15.8
Appendix
Table 16: Structured updates and count-state readout on SVCBench (GPA). The threshold variant is retrained; the remaining controls modify only the count projection read from a fixed complete-model rollout.
OVO-Bench
OVO-S
Model
Access
RT
BT
FA
L1
L2
L3
L4
Qwen3-VL-4B
Full
57.32
46.38
48.75
42.01
49.01
53.67
43.05
Qwen3-VL-8B
Full
62.16
50.56
52.35
42.48
50.68
55.20
45.98
StaMina-4B (Ours)
Full
56.11
49.63
49.23
41.91
54.67
53.72
43.96
StaMina-8B (Ours)
Full
60.71
52.97
52.68
42.68
56.93
54.91
46.77
StreamForest-7B
Stream
61.20
52.02
53.49
46.64
45.20
49.75
34.92
Appendix
Table 17 : Task-group and spatial-level breakdown ( ↑ ). RT: real-time perception; BT: backward tracing; FA: forward active responding. L1–L4 follow the spatial benchmark hierarchy. Group and level aggregates reproduce Table 3 . Results are ranked within each access mode.
Figure 5 : System and supervision ablations at 8B on SVCBench. State only removes delta memory; without FSM replaces the finite-state machine with independent thresholded increments.
Model
Full ↑
Stream ↑
Qwen3-VL-8B
31.0
28.0
Counting-SFT
34.0
35.0
Counting-SFT + delta memory
37.1
36.5
Counting-SFT + explicit state
41.3
36.3
StaMina-8B (Ours)
44.9
38.2
Appendix
Table 18: SVCBench system variants at 8B. Adapted rows share counting queries and update budgets; auxiliary objectives follow their components. Full and Stream share a checkpoint within each row. The two single-component variants are independent controls.
Table 23
Figure 6 : Model scale and access on SVCBench. Each bar is the corresponding official-question aggregate in Table 1 .
Figure 7 : Action-predicate accuracy on recorded task videos. Both 8B conditions use Stream with identical predicates and decision times.
Video understanding requires models to continuously track and update world state during playback. Although existing benchmarks have advanced video understanding evaluation across multiple dimensions, they provide limited visibility into how models maintain world state over time. We propose SVCBench, a Streaming Video Counting Benchmark that repositions counting as a minimal, controlled probe for diagnosing models' world-state maintenance capability. We decompose this capability into object counting and event counting, forming 8 fine-grained subcategories. Object counting covers tracking currently visible objects and cumulative unique identities, while event counting covers detecting instantaneous actions and tracking complete activity cycles. SVCBench contains 406 videos with frame-by-frame annotations of 10,071 event occurrences and object state changes, yielding 1,000 streaming QA pairs with 4,576 query points distributed along video timelines. By observing state maintenance trajectories through streaming multi-point queries, we design three complementary metrics to diagnose numerical precision, trajectory consistency, and temporal awareness. Evaluations of mainstream video-language models show that current models still exhibit significant deficiencies in spatial-temporal state maintenance, with especially poor performance on periodic event counting. SVCBench provides a diagnostic framework for measuring and improving state maintenance in video understanding systems. Our code and data are available at https://buaa-colalab.github.io/SVCBench.
Pengyiang Liu, Zhongyue Shi, Hongye Hao +7
Institute of Artificial Intelligence, Beihang University, Beijing, China · Hangzhou International Innovation Institute of Beihang University, Hangzhou, China
Large vision-language models (VLMs) can recognize \textit{what} happens in video but fail to count \textit{how many} times. We introduce \textbf{PushupBench}, 446 long-form clips (avg. 36.7s) for evaluating repetition counting. The best frontier model achieves 42.1% exact accuracy; open-source 4B models score ∼6%, matching supervised baselines. We show that accuracy alone misleads -- weaker models exploit the modal count rather than reason temporally. Fine-tuning on counting with 1k samples transfers to general video understanding: MVBench (+2.15), PerceptionTest (+1.88), TVBench (+4.54), suggesting counting is a proxy for broader temporal reasoning.PushupBench incorporated in \texttt{lmms-eval} (https://github.com/EvolvingLMMs-Lab/lmms-eval/pull/1262) and hosted on (pushupbench.com/)
We introduce S3T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking. Our method treats temporal sampling density as privileged information, based on the hypothesis that a denser view of the same clip recovers the running state more accurately. This view serves as the teacher, while a sparse-view student with the same weights learns to match its next-token distribution. The model generates its own target, so training requires no labels, separate teacher, or reward signal, and adds no inference cost. On LLaVA-OneVision-2-8B, S3T improves VSTAT accuracy by +1.74 as a single model, +2.38 with souping, and +2.70 with additional vision-encoder adaptation, while prior self-evolving methods leave state tracking largely unchanged. The capability learned from unlabeled synthetic clips transfers to real videos, improving performance by +7.95 on VSTAT-YouTube state-tracking questions and +4.50 on MVBench Action Count.
Shravan Venkatraman, Wenshuai Zhao, Mohammad Hassan Vali +1
Mohamed bin Zayed University of Artificial Intelligence, UAE · Aalto University, Finland · ELLIS Institute Finland & Department of Computer Science, Aalto University, Finland