Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristics
Organizations: Johns Hopkins University
Abstract
Vision-Language Models (VLMs) enable video annotation at scale, but costs accumulate quickly: processing a typical 60-second short-form video at one frame per second requires millions of tokens. To reduce costs, researchers rely on heuristics such as sampling a subset of frames, compressing videos into image grids, or using only a single modality. However, it remains unclear which heuristics save cost, and whether they preserve the downstream conclusions these annotations enable. To address this gap, we conduct a systematic evaluation of these heuristics using short-form videos, on two computational social science (CSS) tasks: sentiment and topic classification. We evaluate each configuration along three axes the literature typically treats separately: classification accuracy, validity of downstream inference, and per-video token cost. First, we find that accuracy and validity diverge: the highest-accuracy configuration can produce wrong conclusions. Second, modality value is not guaranteed: text alone can yield strong performance, indicating that adding modalities can add cost without adding signal. Finally, we find that cost can be decoupled from video length when annotating short-form videos: a single image grid built via simple shot-transition detection approaches full-video understanding ( within~.05), at of the token cost. Based on these findings, we derive guidelines that can enable cost-aware VLM annotation in CSS.
Figures & tables
| Method | Avg. Tokens | Avg. Cost (USD) | Min Tokens | Max Tokens |
| 3 3 grid (9 frames) | 36,596 | $0.00549 | 25,501 | 48,169 |
| 2 8 grid (16 frames) | 47,665 | $0.00715 | 25,501 | 48,169 |
| 1 FPS ( 60 frames) | 2,275,411 | $0.34131 | 110,505 | 25,821,335 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| VLM-based (GPT-4o-mini) | Random Forest | |||||||
| MELD: Sentiment | Frames-only | Text-only | ShotT 16 (4 4) | ShotT 9 (3 3) | Uni 16 (4 4) | Uni 9 (3 3) | RF | RF V+T |
| Negative | .309 | .572 | .516 | .583 | .575 | .581 | .361 | .288 |
| Neutral | .514 | .628 | .592 | .502 | .507 | .470 | .770 | .874 |
| Positive | .482 | .627 | .689 | .684 | .705 | .586 | .145 | .275 |
| Overall | .442 | .610 | .589 | .567 | .571 | .530 | .507 | .560 |
| .139 | .398 | .371 | .350 | .360 | .365 | .157 | .250 | |
| Method | Accuracy | |
| Gemini (video) | .870 | .860 |
| ShotT 2 8 (Gemini) | .870 | .860 |
| ShotT 2 8 (GPT-4o-mini) | .820 | .806 |
| Dataset | Method | Accuracy [95% CI] | Cohen’s [95% CI] |
| MELD | Frames-only | 0.442 [0.410, 0.474] | 0.139 [0.090, 0.187] |
| Text-only | 0.610 [0.579, 0.641] | 0.398 [0.351, 0.446] | |
| ShotT 16 (4 4) | 0.589 [0.557, 0.621] | 0.371 [0.323, 0.419] | |
| Uni 16 (4 4) | 0.571 [0.539, 0.603] | 0.360 [0.313, 0.406] | |
| TikTok ( ) | Frames-only | 0.712 [0.683, 0.742] | 0.696 [0.665, 0.727] |
| Text-only | 0.713 [0.683, 0.743] | 0.696 [0.664, 0.728] |
| Model | Text-only | Img-only (4 4 s.) | Img+Text (3 3 s.) | Img+Text (4 4 s.) | Img+Text (4 4 u.) |
| Claude Sonnet 4.6 | .61 / .39 | .52 / .18 | .59 / .37 | .61 / .41 | .60 / .39 |
| GPT-5.1 | .60 / .37 | .48 / .16 | .42 / .17 | .48 / .24 | .51 / .29 |
| Grok 4.3 | .64 / .40 | .41 / .09 | .40 / .16 | .59 / .33 | .62 / .39 |
| GPT-4o-mini | .56 / .32 | .48 / .13 | .56 / .33 | .60 / .38 | .60 / .37 |
| Model | Text-only | Img-only (2 8 s.) | Img+Text (3 3 s.) | Img+Text (2 8 s.) | Img+Text (2 8 u.) |
| Claude Sonnet 4.6 | .78 / .76 | .80 / .78 | .88 / .87 | .84 / .83 | .85 / .84 |
| GPT-5.1 | .77 / .75 | .73 / .71 | .87 / .86 | .81 / .79 | .82 / .81 |
| Grok 4.3 | .75 / .73 | .69 / .67 | .87 / .86 | .83 / .82 | .85 / .84 |
| GPT-4o-mini | .76 / .74 | .72 / .70 | .85 / .84 | .82 / .81 | .84 / .83 |
| Positive | Neutral | Negative | ||||||||||
| SD | Coverage | SD | Coverage | SD | Coverage | |||||||
| 80 | 0.50 | 0.37 | 3.64 | 91% | 0.30 | 0.01 | 0.71 | 96% | 1.23 | 0.95 | 3.73 | 91% |
| 100 | 0.10 | 0.03 | 0.67 | 95% | 0.32 | 0.01 | 0.58 | 93% | 0.60 | 0.33 | 0.70 | 97% |
| 200 | 0.17 | 0.04 | 0.51 | 96% | 0.35 | 0.04 | 0.36 | 95% | 0.43 | 0.16 | 0.41 | 95% |
| 250 | 0.12 | 0.01 | 0.41 | 98% | 0.38 | 0.07 | 0.28 | 99% | 0.46 | 0.18 | 0.36 | 98% |
| 450 | 0.07 | 0.06 | 0.22 | 100% | 0.37 | 0.06 | 0.20 | 99% | 0.42 | 0.15 | 0.25 | 97% |
| #funnyvideos | #travel | #anime | ||||||||||
| SD | Cov. | SD | Cov. | SD | Cov. | |||||||
| GT | 5.11% | — | — | — | 6.01% | — | — | — | 4.60% | — | — | — |
| 80 | 4.86% | 0.24 | 2.43 | 89% | 5.19% | 0.81 | 1.72 | 80% | 4.14% | 0.46 | 1.51 | 92% |
| 100 | 5.19% | 0.09 | 1.76 | 95% | 5.35% | 0.66 | 1.63 | 87% | 4.16% | 0.45 | 1.25 | 92% |
| 200 | 5.11% | 0.00 | 1.06 | 99% | 5.87% | 0.14 | 1.01 | 97% | 4.43% | 0.17 | 0.80 | 95% |
| 250 | 4.88% | 0.23 | 0.97 | 96% | 5.79% | 0.21 | 1.00 | 98% | 4.43% | 0.18 | 0.74 | 96% |