Rapid progress in video generation has led to a plethora of models that differ substantially in capability and generation cost. This raises a natural question: can each request be efficiently routed to an appropriate model? We find that even when the consensus of the other annotators is used as an oracle, it agrees with each annotator's own favorite only 34-55% of the time. Motivated by this observation, we introduce TasteRoute, a personalized video-generation router that selects a generator jointly based on the input request, user preferences, and available generation budget. Across text-to-video and image-to-video settings, TasteRoute is competitive with strong simple baselines on preference routing while reducing average generation cost. The cost saving increases under higher budget caps. Finally, we release TasteRoute-3k, a human-annotated dataset containing multi-model video comparisons, quality judgments, preference rankings, and user-profile signals to facilitate future research on personalized and cost-aware video routing.
Figures & tables
Quantity
Count
Definition
Distinct prompts
277
unique prompt texts
Pools
394
one per request, i.e. prompt × setting (274 T2V, 120 I2V); 393 evaluable
Groups
812
2.06 per pool, at most six videos each
Annotated videos
3,128
distinct (pool, generator) generations
Generator endpoints
30
catalog of Appendix E ; 28 win at least one pool
Annotators
9
four general annotators, five professional video creators
Table 1: TasteRoute-3K at a glance. A pool is one request in one setting together with every generator labeled for it; a group is the unit one annotator ranks at a time. A comparison is one ordered pair of videos within a group; the consensus outcome of a pair is the label chosen by the most annotators, and pairs with no strict plurality are ambiguous and excluded from the consensus targets.
Figure 1: The best generator changes with the prompt, in both settings. Each cell is how much better or worse a generator scores on prompts with a given attribute than on prompts without it (%); the bottom strip is the mean over generators. Left (T2V) and right (I2V) use the same twelve attributes and the same estimator. Different generators lead on different attributes, and the leader changes on 5 of the 12 attributes in T2V and 8 of the 12 in I2V.
Figure 2: Annotators agree on bad videos but barely on good ones. How often two annotators who judged the same pair chose the same video, for pairs where one video was rejected, where one carried a defect, and where both were clean.
Figure 3: The consensus winner is nobody’s favorite. One T2V pool in which no video was rejected; the first column is the pool’s consensus rank, the rest are the individual annotators’.
Figure 4: How far consensus gets each annotator. (A) Pool level: how often the favorite of the other annotators on the same pool is also this annotator’s own top choice (tie-aware). Gray marks a random pick; the dashed line at 100% ( personal oracle ) is a router that always matches the annotator. (B) Pair level: how often the annotator prefers the same video as another annotator in a single comparison. Bars in both panels are pool-level bootstrap 95% intervals.
Figure 5: Routing quality by limiting the cost available to the router in two settings : generalized routing (A) and personalized routing (B). In both settings, TasteRoute (Generic, Personal) performs competitively against the strong kNN baseline. (C) We also compare training dataset scaling across all routers and find that TasteRoute-Personal scales better as data continues to grow. All performance is measured using NDCG.
Figure 6: TasteRoute-Generic by prompt attribute. Held-out outer-fold T2V pools grouped by the keyword-detected attributes of Section 5 , counts in parentheses, all graded against the pool’s consensus ranking. The router is compared with two prompt-independent references: the single best generator on the training folds ( global best ) and the best generator for each attribute on the training folds ( attribute lookup ). Markers are offset vertically within a row and each dotted line is that series’ rate over all attributes. Note that Figure 8 uses a different ground truth (each annotator’s own ranking) and a different unit (annotator–pool pairs), so the two are not comparable.
Figure 8
$0.50 cap
$1.00 cap
$1.50 cap
$1.70 cap
λ
Top-1% ↑
Saving % ↑
Top-1% ↑
Saving % ↑
Top-1% ↑
Saving % ↑
Top-1% ↑
Saving % ↑
0
47.01
0.00
42.51
0.00
39.85
0.00
39.85
0.00
0.003
47.01
0.00
42.51
0.03
39.85
0.03
39.85
0.03
0.01
47.01
0.00
42.51
0.03
39.85
0.03
39.85
0.03
0.03
46.92
0.13
42.33
0.84
39.69
0.73
39.69
0.73
0.1
46.25
0.91
40.95
2.86
38.79
3.82
38.79
3.81
Table 2: Cost-preference trade-off for the personalized TasteRoute. Top-1 accuracy(%) and cost savings(%) are reported, averaged over nine annotators and three split seeds on a fixed T2V/I2V subset of the personalization test splits. Savings are relative to unchanged training ( λ=0 ) at the same budget cap. The utility-score formula is unchanged; λ weights only the additional training cost loss.
Figure 9: Recovering opposing preferences. Accuracy on held-out video pairs where two annotators disagree. A router without user information scores 50% by construction.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Generator
Fractional pool wins
Win rate
ltx-2.3
39.7
10.10%
veo-3.1-fast
34.8
8.84%
veo-3.1-lite-generate-001
31.2
7.95%
seedance-2.0-mini
28.8
7.32%
veo-3.1-lite
28.2
7.18%
Appendix
Table 3: Distribution of pool-level winning generators by consensus utility. Each of the 393 evaluable pools contributes one vote, with ties divided fractionally among co-winners. The five most frequent winners are shown; nine of the 30 generators each win more than 5% of the pools and 28 win at least one.
Figure 10: Prompt attributes barely move agreement. Among clean pairs (T2V and I2V pooled), the change in preference agreement when a keyword-detected prompt attribute is present versus absent, with pool-level bootstrap 95% intervals, since comparisons within a pool are correlated. Only cartoon or 3D prompts have an interval that excludes zero.
Figure 11: Cyclic favorites. A T2V pool in which no annotator rated any video Not acceptable, laid out as in Figure 3 , where ⋆ marks a rank shared by more than one generator. A2 prefers Grok Imagine, A3 LTX 2.5 Fast, and A8 FLUX 3; the majority relation is cyclic (Grok ≻ FLUX ≻ LTX 2.5 Fast ≻ Grok), and the consensus ties all three for first.
Figure 12: NDCG across the three routing experiments, all methods , laid out as Figure 5 .
Figure 13: Top-1 across the three routing experiments , laid out as Figure 5 . (A) General routing by budget. (B) Personalized routing by budget. (C) Matched-budget label scaling.
Figure 14: Top-3 across the three routing experiments , laid out as Figure 5 . (A) General routing by budget. (B) Personalized routing by budget. (C) Matched-budget label scaling.
Figure 15: Onboarding step 1: the five-item preference questionnaire. Responses initialize the annotator’s preference profile and do not affect the separate quality-evaluation labels.
Figure 16: Onboarding step 2: comprehension check confirming that annotators understand the preference examples ask for the result they personally would rather receive, not an objectively correct one.
Figure 17: Onboarding steps 3–6: a reference-video comparison. Annotators watch both videos and choose which result they would rather receive.
Figure 18: Follow-up questions shown after each choice in a reference-video comparison.
Figure 19: The ranking interface. Each video card shows its quality ratings; videos are ordered into rank groups, with ties allowed. Annotator-identifying information is redacted.
Model
Endpoint ID
Tasks
$/s
LongCat Video Distilled
longcat-video-distilled-720p
T2V
0.010
Wan 2.2 5B
wan-v2.2-5b
T2V, I2V
0.025 ‡
Veo 3.1 Lite
veo-3.1-lite
T2V, I2V
0.030
Veo 3.1 Lite
veo-3.1-lite-generate-001
T2V
0.030
LTX Video 13B Distilled
ltx-video-13b-distilled
T2V, I2V
0.040
Pika 2.2
pika-v2.2
T2V, I2V
0.040
Appendix
Table 4: Video generators in the candidate pool. Tasks gives the settings in which each generator appears in TasteRoute-3K. Prices are in USD per second of 720p video without audio; Veo and gemini-omni-flash prices follow Google’s API catalogs, and all others follow fal.ai. † No 720p tier; billed at the 1080p rate. ‡ Flat $0.15 per video, converted at 6 s. § Base price covers 5 s plus a per-second rate; shown as the 6 s equivalent.