Rapid progress in video generation has led to a plethora of models that differ substantially in capability and generation cost. This raises a natural question: can each request be efficiently routed to an appropriate model? We find that even when the consensus of the other annotators is used as an oracle, it agrees with each annotator's own favorite only 34-55% of the time. Motivated by this observation, we introduce TasteRoute, a personalized video-generation router that selects a generator jointly based on the input request, user preferences, and available generation budget. Across text-to-video and image-to-video settings, TasteRoute is competitive with strong simple baselines on preference routing while reducing average generation cost. The cost saving increases under higher budget caps. Finally, we release TasteRoute-3k, a human-annotated dataset containing multi-model video comparisons, quality judgments, preference rankings, and user-profile signals to facilitate future research on personalized and cost-aware video routing.
Figures & tables
Quantity
Count
Definition
Distinct prompts
277
unique prompt texts
Pools
394
one per request, i.e. prompt × setting (274 T2V, 120 I2V); 393 evaluable
Groups
812
2.06 per pool, at most six videos each
Annotated videos
3,128
distinct (pool, generator) generations
Generator endpoints
30
catalog of Appendix E ; 28 win at least one pool
Annotators
9
four general annotators, five professional video creators
Table 1: TasteRoute-3K at a glance. A pool is one request in one setting together with every generator labeled for it; a group is the unit one annotator ranks at a time. A comparison is one ordered pair of videos within a group; the consensus outcome of a pair is the label chosen by the most annotators, and pairs with no strict plurality are ambiguous and excluded from the consensus targets.
Figure 1: The best generator changes with the prompt, in both settings. Each cell is how much better or worse a generator scores on prompts with a given attribute than on prompts without it (%); the bottom strip is the mean over generators. Left (T2V) and right (I2V) use the same twelve attributes and the same estimator. Different generators lead on different attributes, and the leader changes on 5 of the 12 attributes in T2V and 8 of the 12 in I2V.
Figure 2: Annotators agree on bad videos but barely on good ones. How often two annotators who judged the same pair chose the same video, for pairs where one video was rejected, where one carried a defect, and where both were clean.
Figure 3: The consensus winner is nobody’s favorite. One T2V pool in which no video was rejected; the first column is the pool’s consensus rank, the rest are the individual annotators’.
Figure 4: How far consensus gets each annotator. (A) Pool level: how often the favorite of the other annotators on the same pool is also this annotator’s own top choice (tie-aware). Gray marks a random pick; the dashed line at 100% ( personal oracle ) is a router that always matches the annotator. (B) Pair level: how often the annotator prefers the same video as another annotator in a single comparison. Bars in both panels are pool-level bootstrap 95% intervals.
Figure 5: Routing quality by limiting the cost available to the router in two settings : generalized routing (A) and personalized routing (B). In both settings, TasteRoute (Generic, Personal) performs competitively against the strong kNN baseline. (C) We also compare training dataset scaling across all routers and find that TasteRoute-Personal scales better as data continues to grow. All performance is measured using NDCG.
Figure 6: TasteRoute-Generic by prompt attribute. Held-out outer-fold T2V pools grouped by the keyword-detected attributes of Section 5 , counts in parentheses, all graded against the pool’s consensus ranking. The router is compared with two prompt-independent references: the single best generator on the training folds ( global best ) and the best generator for each attribute on the training folds ( attribute lookup ). Markers are offset vertically within a row and each dotted line is that series’ rate over all attributes. Note that Figure 8 uses a different ground truth (each annotator’s own ranking) and a different unit (annotator–pool pairs), so the two are not comparable.
Figure 8
$0.50 cap
$1.00 cap
$1.50 cap
$1.70 cap
λ
Top-1% ↑
Saving % ↑
Top-1% ↑
Saving % ↑
Top-1% ↑
Saving % ↑
Top-1% ↑
Saving % ↑
0
47.01
0.00
42.51
0.00
39.85
0.00
39.85
0.00
0.003
47.01
0.00
42.51
0.03
39.85
0.03
39.85
0.03
0.01
47.01
0.00
42.51
0.03
39.85
0.03
39.85
0.03
0.03
46.92
0.13
42.33
0.84
39.69
0.73
39.69
0.73
0.1
46.25
0.91
40.95
2.86
38.79
3.82
38.79
3.81
Table 2: Cost-preference trade-off for the personalized TasteRoute. Top-1 accuracy(%) and cost savings(%) are reported, averaged over nine annotators and three split seeds on a fixed T2V/I2V subset of the personalization test splits. Savings are relative to unchanged training ( λ=0 ) at the same budget cap. The utility-score formula is unchanged; λ weights only the additional training cost loss.
Figure 9: Recovering opposing preferences. Accuracy on held-out video pairs where two annotators disagree. A router without user information scores 50% by construction.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Generator
Fractional pool wins
Win rate
ltx-2.3
39.7
10.10%
veo-3.1-fast
34.8
8.84%
veo-3.1-lite-generate-001
31.2
7.95%
seedance-2.0-mini
28.8
7.32%
veo-3.1-lite
28.2
7.18%
Appendix
Table 3: Distribution of pool-level winning generators by consensus utility. Each of the 393 evaluable pools contributes one vote, with ties divided fractionally among co-winners. The five most frequent winners are shown; nine of the 30 generators each win more than 5% of the pools and 28 win at least one.
Figure 10: Prompt attributes barely move agreement. Among clean pairs (T2V and I2V pooled), the change in preference agreement when a keyword-detected prompt attribute is present versus absent, with pool-level bootstrap 95% intervals, since comparisons within a pool are correlated. Only cartoon or 3D prompts have an interval that excludes zero.
Figure 11: Cyclic favorites. A T2V pool in which no annotator rated any video Not acceptable, laid out as in Figure 3 , where ⋆ marks a rank shared by more than one generator. A2 prefers Grok Imagine, A3 LTX 2.5 Fast, and A8 FLUX 3; the majority relation is cyclic (Grok ≻ FLUX ≻ LTX 2.5 Fast ≻ Grok), and the consensus ties all three for first.
Figure 12: NDCG across the three routing experiments, all methods , laid out as Figure 5 .
Figure 13: Top-1 across the three routing experiments , laid out as Figure 5 . (A) General routing by budget. (B) Personalized routing by budget. (C) Matched-budget label scaling.
Figure 14: Top-3 across the three routing experiments , laid out as Figure 5 . (A) General routing by budget. (B) Personalized routing by budget. (C) Matched-budget label scaling.
Figure 15: Onboarding step 1: the five-item preference questionnaire. Responses initialize the annotator’s preference profile and do not affect the separate quality-evaluation labels.
Figure 16: Onboarding step 2: comprehension check confirming that annotators understand the preference examples ask for the result they personally would rather receive, not an objectively correct one.
Figure 17: Onboarding steps 3–6: a reference-video comparison. Annotators watch both videos and choose which result they would rather receive.
Figure 18: Follow-up questions shown after each choice in a reference-video comparison.
Figure 19: The ranking interface. Each video card shows its quality ratings; videos are ordered into rank groups, with ties allowed. Annotator-identifying information is redacted.
Model
Endpoint ID
Tasks
$/s
LongCat Video Distilled
longcat-video-distilled-720p
T2V
0.010
Wan 2.2 5B
wan-v2.2-5b
T2V, I2V
0.025 ‡
Veo 3.1 Lite
veo-3.1-lite
T2V, I2V
0.030
Veo 3.1 Lite
veo-3.1-lite-generate-001
T2V
0.030
LTX Video 13B Distilled
ltx-video-13b-distilled
T2V, I2V
0.040
Pika 2.2
pika-v2.2
T2V, I2V
0.040
Appendix
Table 4: Video generators in the candidate pool. Tasks gives the settings in which each generator appears in TasteRoute-3K. Prices are in USD per second of 720p video without audio; Veo and gemini-omni-flash prices follow Google’s API catalogs, and all others follow fal.ai. † No 720p tier; billed at the 1080p rate. ‡ Flat $0.15 per video, converted at 6 s. § Base price covers 5 s plus a per-second rate; shown as the 6 s equivalent.
Wepresent Alice v1, a 14-billion parameter open-source video generation model that achieves state-of-the-art quality through consistency distillation with score regularization (rCM). Contrary to conventional distillation-which trades quality for speed-we demonstrate that rCM-based distillation can exceed teacher model quality. We attribute this to three mechanisms: (1) the score regularization term acts as a mode-seeking objective that concentrates probability mass on high-quality outputs rather than covering the full teacher distribution, (2) our targeted synthetic data pipeline with hard example mining provides training signal specifically for failure modes (physics, hands, faces) that the teacher handles inconsistently, and (3) consistency enforcement acts as implicit regularization, eliminating "lucky path" dependence on specific noise samples. Alice v1 generates 5-second 720p videos at 24fps in 4 denoising steps (~8 seconds on H100), a 7x speedup over the 50-step teacher while improving VBench score from 84.0 (Wan2.2) to 91.2. This surpasses both the teacher and closed-source systems including Veo3 (~90) and Sora2 (~88) on automated benchmarks, with competitive results in human preference studies. We release all model weights, training code, synthetic data pipelines, and evaluation scripts to advance open research in video generation.
Video-to-video (V2V) generation is difficult to evaluate because outputs must both follow editing instructions and preserve frame-level correspondence with the source video, which existing T2V and I2V metrics do not capture. We introduce V2V-Bench, a 11-dimension benchmark organized into five categories: temporal alignment, structural fidelity, transformation quality, video quality, and semantic alignment. V2V-Bench pairs diverse source videos with challenging editing tasks and evaluates two commercial models, Grok Imagine and Gemini Veo3, and one open-source model, Open Sora 2. Results show complementary model strengths: Grok performs better on editing fidelity, while Veo3 achieves stronger visual quality. On six V2V-specific dimensions, V2V-Bench reaches a Spearman correlation of 0.905 with human judgments.
We present Paris 2.0, the first video generation model pre-trained through decentralized computation. Its training recipe builds upon Paris 1.0 (arXiv:2510.03434), the first ever open-weight Decentralized Diffusion Model (DDM), which showed that image generation can be trained without a monolithic GPU cluster. However, temporally coherent video generation had remained an open problem under decentralized training, and Paris 2.0 closes it. In low-resolution text-to-video training, against a monolithic model trained on the same data under a matched total compute budget, Paris 2.0 cuts Frechet Video Distance (FVD) from 561.04 to 279.01, a ~2.0x improvement, and lifts CLIP text-video similarity and aesthetic score.