Vision-Language Models (VLMs) enable video annotation at scale, but costs accumulate quickly: processing a typical 60-second short-form video at one frame per second requires millions of tokens. To reduce costs, researchers rely on heuristics such as sampling a subset of frames, compressing videos into image grids, or using only a single modality. However, it remains unclear which heuristics save cost, and whether they preserve the downstream conclusions these annotations enable. To address this gap, we conduct a systematic evaluation of these heuristics using short-form videos, on two computational social science (CSS) tasks: sentiment and topic classification. We evaluate each configuration along three axes the literature typically treats separately: classification accuracy, validity of downstream inference, and per-video token cost. First, we find that accuracy and validity diverge: the highest-accuracy configuration can produce wrong conclusions. Second, modality value is not guaranteed: text alone can yield strong performance, indicating that adding modalities can add cost without adding signal. Finally, we find that cost can be decoupled from video length when annotating short-form videos: a single 2×8 image grid built via simple shot-transition detection approaches full-video understanding (κ within~.05), at ∼15% of the token cost. Based on these findings, we derive guidelines that can enable cost-aware VLM annotation in CSS.
Figures & tables
Figure 1: Evaluation framework for cost-saving VLM annotation heuristics. We evaluate cost-saving heuristics for VLM-based video annotation along three configurable choices (frame selection, grid composition, modality), evaluating each configuration on classification accuracy, downstream inferential validity, and per-video token cost. Bias correction via PPI is applied to recover valid confidence intervals from a small set of human labels.
Table 1: High classification accuracy does not guarantee valid downstream inference, but bias correction helps recover it. Downstream estimation accuracy on (a) MELD and (b) TikTok. RF = random forest on Vertex AI video embeddings; RF V+T = random forest on Vertex AI video + transcript embeddings
Method
Avg. Tokens
Avg. Cost (USD)
Min Tokens
Max Tokens
3 × 3 grid (9 frames)
36,596
$0.00549
25,501
48,169
2 × 8 grid (16 frames)
47,665
$0.00715
25,501
48,169
1 FPS ( ∼ 60 frames)
2,275,411
$0.34131
110,505
25,821,335
Table 2: Grid sampling decouples token cost from video length. Average GPT-4o-mini token cost per video by frame input method on the TikTok dataset (n=900, avg. duration 62.3s). Packing frames into a single grid costs ∼ 37–48k tokens regardless of video length, whereas 1 FPS sampling scales with duration and is ∼ 50 × more expensive on average (and over 500 × in the worst case). The two grids share the same max-token value ( ∼ 48k), reflecting the API’s internal per-image token cap, which bounds cost regardless of grid resolution.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
VLM-based (GPT-4o-mini)
Random Forest
MELD: Sentiment
Frames-only
Text-only
ShotT 16 (4 × 4)
ShotT 9 (3 × 3)
Uni 16 (4 × 4)
Uni 9 (3 × 3)
RF
RF V+T
Negative
.309
.572
.516
.583
.575
.581
.361
.288
Neutral
.514
.628
.592
.502
.507
.470
.770
.874
Positive
.482
.627
.689
.684
.705
.586
.145
.275
Overall
.442
.610
.589
.567
.571
.530
.507
.560
κ
.139
.398
.371
.350
.360
.365
.157
.250
Appendix
Table 3: The value of the visual modality is task-dependent. Per-class accuracy across annotation configurations on MELD (sentiment, 3 classes) and TikTok (hashtag, 18 classes). On MELD, Text-only achieves the highest overall accuracy ( κ =0.398), while on TikTok all multimodal grid configurations substantially outperform both unimodal baselines ( κ 0.808–0.825 vs. 0.695–0.696), illustrating that modality value is task-dependent. Green highlights the best-performing configuration overall.
Figure 2: Accuracy of TikTok video annotations by video duration (Shot Transition Detection 2×8 ).
Figure 3: Annotation cost vs. quality on TikTok. The x-axis shows frames per video (in units of tokens-per-frame, TPF): 9 ( 3×3 grid), 16 ( 2×8 grid), and 62 (Gemini at 1 FPS, since TikTok videos average 62.3s). The dashed line at κ=.80 marks the conventional “almost perfect agreement” threshold ( Landis and Koch, 1977 ) . The ShotT 3×3 grid method exceeds this threshold at a fraction of the per-video token cost.
Method
Accuracy
κ
Gemini (video)
.870
.860
ShotT 2 × 8 (Gemini)
.870
.860
ShotT 2 × 8 (GPT-4o-mini)
.820
.806
Appendix
Table 4: Accuracy and Cohen’s κ on 100 common videos ( N=100 ).
Dataset
Method
Accuracy [95% CI]
Cohen’s κ [95% CI]
MELD
Frames-only
0.442 [0.410, 0.474]
0.139 [0.090, 0.187]
Text-only
0.610 [0.579, 0.641]
0.398 [0.351, 0.446]
ShotT 16 (4 × 4)
0.589 [0.557, 0.621]
0.371 [0.323, 0.419]
Uni 16 (4 × 4)
0.571 [0.539, 0.603]
0.360 [0.313, 0.406]
TikTok ( n=900 )
Frames-only
0.712 [0.683, 0.742]
0.696 [0.665, 0.727]
Text-only
0.713 [0.683, 0.743]
0.696 [0.664, 0.728]
Appendix
Table 5: Accuracy and Cohen’s κ across modality and frame-sampling conditions, with bootstrap CIs
Figure 4: Per-hashtag PPI 95% confidence intervals vs. ground truth.
Model
Text-only
Img-only (4 × 4 s.)
Img+Text (3 × 3 s.)
Img+Text (4 × 4 s.)
Img+Text (4 × 4 u.)
Claude Sonnet 4.6
.61 / .39
.52 / .18
.59 / .37
.61 / .41
.60 / .39
GPT-5.1
.60 / .37
.48 / .16
.42 / .17
.48 / .24
.51 / .29
Grok 4.3
.64 / .40
.41 / .09
.40 / .16
.59 / .33
.62 / .39
GPT-4o-mini
.56 / .32
.48 / .13
.56 / .33
.60 / .38
.60 / .37
Appendix
Table 6: MELD sentiment ( N=100 ). Each cell: accuracy / Cohen’s κ . Bold = best per model.
Model
Text-only
Img-only (2 × 8 s.)
Img+Text (3 × 3 s.)
Img+Text (2 × 8 s.)
Img+Text (2 × 8 u.)
Claude Sonnet 4.6
.78 / .76
.80 / .78
.88 / .87
.84 / .83
.85 / .84
GPT-5.1
.77 / .75
.73 / .71
.87 / .86
.81 / .79
.82 / .81
Grok 4.3
.75 / .73
.69 / .67
.87 / .86
.83 / .82
.85 / .84
GPT-4o-mini
.76 / .74
.72 / .70
.85 / .84
.82 / .81
.84 / .83
Appendix
Table 7: TikTok hashtag ( N=100 , 18 classes). Each cell: accuracy / Cohen’s κ . Bold = best per model.
Positive
Neutral
Negative
nh
β
Δ
SD
Coverage
β
Δ
SD
Coverage
β
Δ
SD
Coverage
80
+ 0.50
+ 0.37
3.64
91%
− 0.30
+ 0.01
0.71
96%
+ 1.23
+ 0.95
3.73
91%
100
+ 0.10
− 0.03
0.67
95%
− 0.32
− 0.01
0.58
93%
+ 0.60
+ 0.33
0.70
97%
200
+ 0.17
+ 0.04
0.51
96%
− 0.35
− 0.04
0.36
95%
+ 0.43
+ 0.16
0.41
95%
250
+ 0.12
− 0.01
0.41
98%
− 0.38
− 0.07
0.28
99%
+ 0.46
+ 0.18
0.36
98%
450
+ 0.07
− 0.06
0.22
100%
− 0.37
− 0.06
0.20
99%
+ 0.42
+ 0.15
0.25
97%
Appendix
Table 8: PPI sensitivity analysis on MELD. For each nh , we sample nh human-labeled examples from a pool of 1,000 videos, averaged over 100 random draws. Δ=β−βGT . GT coefficients: Pos. β=+0.13 , Neu. β=−0.31 , Neg. β=+0.27 . Coverage is the fraction of the 100 per-draw 95% PPI intervals that contain βGT .
#funnyvideos
#travel
#anime
nh
θ
Δ
SD
Cov.
θ
Δ
SD
Cov.
θ
Δ
SD
Cov.
GT
5.11%
—
—
—
6.01%
—
—
—
4.60%
—
—
—
80
4.86%
− 0.24
2.43
89%
5.19%
− 0.81
1.72
80%
4.14%
− 0.46
1.51
92%
100
5.19%
+ 0.09
1.76
95%
5.35%
− 0.66
1.63
87%
4.16%
− 0.45
1.25
92%
200
5.11%
+ 0.00
1.06
99%
5.87%
− 0.14
1.01
97%
4.43%
− 0.17
0.80
95%
250
4.88%
− 0.23
0.97
96%
5.79%
− 0.21
1.00
98%
4.43%
− 0.18
0.74
96%
Appendix
Table 9: PPI sensitivity to human-annotation budget ( nhuman ) on TikTok.
School of Biomedical Engineering, University of Science and Technology of China · Data Darkness Lab, MIRACLE Center, Suzhou Institute for Advanced Research · China +1