Video-language models are ranked by multiple-choice accuracy on frames from a uniform grid. The grid has two parameters, a rate and a phase, and benchmarks report only the rate. The phase moves answers: two deployed samplers differing only by a half-step phase offset answer 23.6% of questions differently while scoring within a point, and across four releases from two families shifting only the phase changes roughly one answer in five after controlling option order. PHASEFUSION decodes three offset grids and averages the option posteriors. The grids are the polyphase components of the dense grid. Fusion matches a 32-frame single pass in accuracy within a prespecified margin (logit-scored) and cuts the answers a half-step shift of all three grids changes from 18.2% to 10.1%. Option order, which changes only the presentation, is flagged instead by a one-pass answer margin. Report the phase convention with the budget, or marginalize it.
Figures & tables
Figure 1: The sampling phase changes answers, not scores. (a) One EgoSchema clip decoded at two phases of the same uniform 8-frame grid, whose sample times are drawn on a common axis (four of eight frames shown per phase). Each phase is unanimous across all ten option orderings, yet they pick different options, one of them wrong. (b) Four releases from two families. Dark bars: consensus answers that differ between two phases, with each family’s measured re-run floor printed beside the bar (InternVL3.5 has no re-run corpus). Light bars: answers that change under any of the three phases.
Figure 2: The disagreement needs no perturbation from us. (a) Four uniform samplers taken from deployed harnesses, on a common time axis: left endpoint ⌊iT/b⌋ , midpoint ⌊(i+21)T/b⌋ , endpoint-inclusive round(ib−1T−1) , and decimation to 1 fps followed by linspace . Left endpoint and midpoint have the same rate and differ only by a half-step phase offset. (b) All six pairs disagree far above the re-run floor, and the pure-phase pair disagrees most , on 23.6% of consensus answers, while all four score within 1.2 points of one another.
Figure 3: PhaseFusion . Top: the S phase grids are decoded separately and their content-space posteriors averaged. The margin of the fused posterior is available for flagging order fragility at no extra pass (Section 4.3 ). Bottom: the three offset grids interleave into the dense Sb -frame grid, sample i of phase s being sample k=iS+s (Eq. 1 ).
A. Cost and accuracy , Qwen3-VL, n=442 paired items
arm
attn.
acc. %
Δ vs. 32f
p
PhaseFusion 3×4 f
48
58.8
-1.13
.603
single pass, 8f
64
55.0
-4.98
.004
PhaseFusion 3×8 f
192
60.4
+0.45
.883
single pass, 16f
256
57.5
-2.49
.109
single pass, 32f
1024
60.0
+0.00
–
Table 1: What averaging buys, on one table. (A) At S=3 , b=8 fusion sits on the 32-frame accuracy at 3/16 of its visual-prefix attention. (B) Under a shift of all three grids by half their spacing, seeing the same 24 frames as one sequence buys the accuracy but not the stability: the dense pass still changes 13.3% of its answers, fusion 10.1%. Attention units are Sb2 for the visual prefix, and end-to-end latency is reported in Section 3 . p : McNemar, paired. Shaded rows are PhaseFusion . All blocks are Qwen3-VL at the identity ordering, and each uses the items that carry all of its arms, hence the differing n and accuracies.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
pool
used
options
perms
EgoSchema
500
500
5 (all)
10
MVBench
1420
500
3 / 4
6
( 369 / 131 )
Video-MME
2700
500
4 (all)
10
Appendix
Table S1: Evaluated subsets. “perms” is the number of distinct option orderings available given the smallest option count present, capped at the ten we evaluate.
Confirmatory test
p
crit.
Margin vs. MSP, flip prediction
.005
.0063
S=3 vs. S=2 phase count
.016
.0125
Fusion vs. 32f, clustered TOST
.022
.0188
S=3 vs. its own single views
.022
.0250
Fusion vs. union pass, origin shift
.032
.0312
MiMo-VL fusion gain
.037
.0375
Appendix
Table S2: Benjamini–Hochberg over the confirmatory family ( m=8 , q=.05 ). The largest passing rank is 8 , so the step-up procedure rejects all eight; intermediate ranks need not clear their own thresholds.
#
Intervention (cost in passes)
Level
Outcome on the failure it targets
Verdict
1
k -ordering plurality vote (10 × )
inference
+2.08 points Qwen-pooled ( +1.74 points over all 9 cells); two disjoint 4-votes disagree on 18.2% of items
lottery persists
2
SPR-gated adaptive re-ask (2–5 × )
inference
oracle gate 55.0% at its best budget vs. best uniform 55.4% ( −0.4 points)
no headroom
3
Per-option content scoring (cloze, 2nopt× )
inference
raw 35.0/token-mean 37.0/PMI 42.0 % vs. letter 53.0%, same items
−11 points at best
4
Margin-routed channel switch ( ∼ 1.5 × )
inference
+1.9 points on flipped items, −25 points on stable ones
54.7 vs. 54.5% (random pair) on exact-reverse subset
pairing worth 0.2 points
6
Phase-ensemble, argmax majority (3 × )
inference
3×8 f vote 57.8% ≈ 16f single 58.1% ( p=.79 ); logit mixture +1.5 points over the vote ( p=.020 )
counting loses to pooling
Appendix
Table S3: The repair ladder, tested rung by rung. Every intervention level available to a practitioner (inference-time aggregation and debiasing, prompt re-design, measurement reallocation, training-time consistency objectives, and calibration transfer) fails to repair per-question instability, several worsening it. Two results bound the space we test: an oracle re-ask allocator shows no headroom over the best uniform budget (row 2; bootstrap CI [−1.9,+0.7] points), and no LoRA consistency objective reduces flips at p<.05 (rows 11–12), because forcing agreement at genuine near-ties trades instability for arbitrary commitment. The instability is epistemic signal, not a removable artifact, which is exactly why it can be screened (Section 4.3 ) even though it cannot be fixed . SPR is the single-pass answer margin of Section 4.3 of the main paper.
Model
Benchmark
Flip%
Swing (points)
BH p
Q2.5
EgoSchema
52.8
5.8
<10−15
Q2.5
MVBench
66.4
4.6
<10−15
Q2.5
Video-MME
47.8
2.6
<10−15
Q3
EgoSchema
40.8
3.2
<10−15
Q3
MVBench
58.2
4.0
<10−15
Q3
Video-MME
51.8
3.8
<10−15
Appendix
Table S4: Per-cell flip rate vs. accuracy swing. The rightmost column is the Benjamini–Hochberg-adjusted p for a one-sided binomial test of H0 : flip ≤15% ; all nine fall below 10−15 , and the least extreme is 3.4×10−44 (Q3/EgoSchema).
Pair
Benchmark
Shared
Model-sp.
Lift
Q2.5 × Q3
EgoSchema
32.0
29.6
1.49
Q2.5 × Q3
MVBench
43.2
38.2
1.12
Q2.5 × Q3
Video-MME
35.0
29.6
1.41
IV2.5 × Q3
EgoSchema
33.0
36.4
1.31
IV2.5 × Q3
MVBench
40.6
31.8
1.27
IV2.5 × Q3
Video-MME
36.4
30.6
1.36
Appendix
Table S5: Flip-state decomposition (% of shared items). Lift is the shared-unstable mass over its independence null.
Model
Benchmark
c-fix
l-fix
pos-mass
Q2.5
EgoSchema
0.82
0.40
0.31
Q2.5
MVBench
0.78
0.54
0.38
Q2.5
Video-MME
0.82
0.44
0.34
Q3
EgoSchema
0.86
0.40
0.38
Q3
MVBench
0.78
0.53
0.47
Q3
Video-MME
0.81
0.44
0.40
Appendix
Table S6: Content-, letter- and position-fixation per cell.
8f
16f
24f
32f
PhaseFusion 3×8 f
Q3 ( n=448 )
54.9
57.4
59.2
59.6
60.3
Q2.5 ( n=435 )
49.4
52.9
54.5
54.5
53.3
Attn. units
64
256
576
1024
192
Appendix
Table S7: Union control, generation-scored pipeline, identity ordering. Fusion vs. 24f: Q3 +1.12 points (McNemar p=.51 ), Q2.5 −1.15 points ( p=.60 ). Fusion vs. 32f: +0.67 points ( p=.77 ) / −1.15 points ( p=.61 ). Fusion vs. 8f: +5.36 points ( p=4×10−4 ) / +3.91 points ( p=.043 ). The gain over the single pass is attributable to the additional frame dose, delivered at 1/3 the attention.
Figure S1: The union control, visually. Single-pass accuracy against the number of distinct frames seen (8/16/24/32; whiskers ±1 SE), with PhaseFusion 3×8 f plotted at its own dose of 24 distinct frames (star). On the generation-scored path, the star differs from the union pass by +1.12/−1.15 points; McNemar p=.51/.60 tests difference, not equivalence. The same-logit TOST ( p=.035 / .046 ; “Same-pipeline replication” below) establishes equivalence within the prespecified ±2.5 -point margin. Fusion uses 192 attention units instead of 576 and reduces origin sensitivity.
Configuration
Attn. units
Median s/q
Peak GB
single 4f
16
0.55
17.67
single 8f
64
0.99
17.83
single 16f
256
1.92
18.29
single 24f
576
2.86
18.93
single 32f
1024
3.92
19.77
fusion 3 × 4f
48
1.72
17.67
Appendix
Table S8: Measured cost of each configuration in the matched run. Means carry heavy decode tails on long videos, so medians are reported; peak memory is flat because 8B bfloat16 weights dominate activations at these lengths.
Configuration
Attn.
Acc. (%)
vs. 32f
Single pass, 4f
16
53.6
−6.3 ∗
PhaseFusion , 3×4 f
48
58.8
−1.1
Single pass, 8f
64
55.0
−5.0 ∗
Argmax vote, 3×8 f
192
59.3
−0.7
PhaseFusion , 2×8 f
128
55.9
−4.1 ∗
PhaseFusion , 3×8 f
192
60.4
+0.5
Appendix
Table S9: The (S,b) frontier (Qwen3-VL, paired n=442 items carrying every configuration and baseline; attention cost Sb2 in frame-token 2 ; ∗ below the 32-frame pass at McNemar p<.05 ). Every S=3 configuration sits above the single-pass curve at its cost. PhaseFusion at 3×4 f sits 1.1 points below the 32-frame pass and runs at 1.72 s/question against 3.92 s in the matched run, and costs no more than a 16-frame pass (1.92 s; +1.4 points, p=.51 ). S=2 is the exception (Section G ); S=5 buys nothing. Argmax vote is shown at this subset’s own value; its pooled gap is in Section 4 of the main paper.
Figure S2: Risk–coverage over the pooled three-model logit subset ( n=1350 ): accuracy on the answered subset as low-confidence items are abstained on, for the single-pass margin vs. the ten-pass vote-share.
α=.30
α=.25
Signal
viol.
cov.
viol.
cov.
Margin ( SPR )
6.0
11.6
4.7
6.7
Top-1 prob. (MSP)
6.7
12.3
3.0
9.5
Answer entropy
6.3
13.2
5.3
9.9
10-pass vote-share
1.7
5.9
1.7
2.6
Random
0.0
0.1
0.0
0.1
Appendix
Table S10: In-cell certified mode (%; 300 splits, 75 calibration labels per cell). Violations stay inside the δ=10% budget for every signal. At matched validity the one-pass margin certifies twice the coverage of the ten-pass vote-share, at a tenth of the cost. At α=.20 every signal certifies zero coverage. The fixed-sequence walk starts at the ten highest-scoring calibration items, and the Clopper–Pearson upper bound for ten items with no errors at δ=10% is 0.206 , above α , so the first hypothesis in the sequence fails and the walk stops before any threshold can be accepted. The binding constraint there is the start of the sequence under this δ , not the signal ( results/derived/conformal_selective.json ).
Signal
Passes
flip
correct
Random
0
0.484
0.496
Option count only
0
0.450
0.423
Surface features only †
0
0.619
—
Answer entropy
1
0.803
0.693
Top-1 prob. (MSP)
1
0.823
0.700
Learned, 4 features
1
0.829
0.696
Appendix
Table S11: Reliability-signal ablation (AUROC; n=1,350 pooled, three models; fused-margin row per model on its own fused answers, n=446/445 ). No single-pass signal separates on correctness; the margin wins on flip prediction (vs. MSP p=.005 ), learning adds nothing ( p=.26 ), and the 10-pass vote-share loses to one pass on the shared target ( p=.042 ). † Qwen-only 5-fold CV ( n=3000 ); transfers to InternVL2.5 ( n=1500 ) at 0.551.
Transfer setting
Passes
flip AUROC
in-cell, per cell (range of 9)
1
0.793–0.925
cross-cell (train 8, test 1; mean)
1
0.87
cross-model Q2.5 → Q3
1
0.900
cross-model Q3 → Q2.5
1
0.837
held-out family (worst/best of 3)
1
0.842/0.898
held-out RL family (MiMo-VL)
1
0.882
Appendix
Table S12: Transfer of the learned four-feature flip predictor when fit on other cells, models, or entire families ( n=450 per held-out family). The training-free margin itself needs no fit; its untouched-family results are 0.867 (InternVL2.5) and 0.876 (MiMo-VL), as reported in Section 4.3 of the main paper.
When asked which of two events came first, video large language models can fail in two opposite ways: cave to a false claim, or reject a true one. Prior video sycophancy work measures only the first and mitigates it by teaching the model to trust the user less, a fix known in text and image models to worsen the second. In video, both failures come from two causes the literature treats as one: availability, whether the sparse sampled frames contain the two events, and weighting, whether that evidence is trusted over the user. We separate them with two interventions that keep the claim fixed: a frame-preserving reorder that flips the claim's truth, and a sampling-offset shift that captures or misses both events at a fixed frame budget. When the events are missed, the two twins present identical frames, so each of the nine models we evaluate accepts a true and a false claim at the same rate, making Youden's J=0 by construction. Availability is necessary but not sufficient. Five of the nine read the order, yet four of those five still cave to the false claim, so their deference hits a weighting ceiling. Since trust cannot be calibrated over evidence that was never sampled, we propose a reversal test that cancels the model's order prior by scoring the sampled frames forward and reversed, then answers, resamples, or abstains without reading the claim. The test raises the order accuracy to 0.92-1.00 on the models that read the order and abstains rather than guesses on those that cannot.
Yuxin Cao, Wei Song, Jingling Xue +1
National University of Singapore, Singapore · University of New South Wales, Australia
Recent advances in large audio language models (LALMs) have primarily been assessed using a multiple-choice question answering (MCQA) framework. However, subtle changes, such as shifting the order of choices, result in substantially different results. Existing MCQA frameworks do not account for this variability and report a single accuracy number per benchmark or category. We dive into the MCQA evaluation framework and conduct a systematic study spanning three benchmarks (MMAU, MMAR and MMSU) and four models: Audio Flamingo 2, Audio Flamingo 3, Qwen2.5-Omni-7B-Instruct, and Kimi-Audio-7B-Instruct. Our findings indicate that models are sensitive not only to the ordering of choices, but also to the paraphrasing of the question and the choices. Finally, we propose a simpler evaluation protocol and metric that account for subtle variations and provide a more detailed evaluation report of LALMs within the MCQA framework.
Fernando López, Santosh Kesiraju, Jordi Luque
Scientific Research, Telef´onica Innovaci´on Digital, Spain · Brno University of Technology, Czech Republic
Vision-language models (VLMs) can ingest only a limited number of video frames, making frame selection a practical necessity. But do current Video QA benchmarks genuinely require temporal frame selection, or can most questions be answered regardless of which frames are shown? We introduce Frame Selection Sensitivity (FSS), a per-sample diagnostic that measures how much VLM accuracy changes when the most relevant frames are replaced with the least relevant ones. Across six benchmarks and eight VLMs, we find that a large majority of samples are frame-agnostic: only a minority are genuinely sensitive to frame choice. Combining FSS with a Language Independence Score (LIS) reveals that merely 5.5--31% of samples are Temporally Sensitive. We construct TempCore, compact evaluation subsets that isolate these temporal samples from existing benchmarks, and will release code and per-sample annotations upon publication.