Under the AI Flow framework, communication networks distribute intelligence across devices, edge servers, and clouds, and computation at the receiver becomes a resource that can substitute for transmitted bits. Generative video compression (GVC) embodies this exchange by sending compact tokens with ultra-low bitrate and letting a generative decoder synthesize the video, yet how much bandwidth savings a unit of decoder compute actually achieves has never been quantified. To fill this vacancy, we model reconstruction quality as a two-factor power law in data rate and decoder compute, which fits measured DISTS of two GVC decoders with a mean error below 3%, and define the information capacity (IC) as the negative logarithmic slope along an iso-quality contour, namely the fraction of rate saved per fractional increase in compute at identical quality. IC is dimensionless and unit-invariant, thus enabling an architecture-agnostic comparison. It forms a field over the operating plane, locating where additional denoising steps are worth their cost. Across five datasets, the 14B decoder trades more compute for fewer rate about ten times more efficiently than the 1.3B decoder. IC also varies significantly across datasets, indicating imbalanced performance on the rate-compute trade-off in GVC methods.
Figures & tables
Figure 1: Fitted DISTS surface from Eq. ( 1 ) on MCL-JCV for GVC-1.3B (left) and GVC-14B (right) on logarithmic rate–compute axes with iso-quality contours. Circles are the measured configurations, colored by their measured DISTS. Contours of the 1.3B decoder are steep, so quality is governed mainly by rate, whereas those of the 14B decoder are nearly horizontal at few denoising steps and bend toward the rate axis only at higher steps.
Figure 2: Information capacity field IC(R,C) from Eq. ( 3 ) for GVC-1.3B (top) and GVC-14B (bottom) on five datasets with a shared logarithmic color scale. Black curves are analytical contours of IC and white dots denote the measured configurations. IC increases with rate and decreases with compute on every panel, and the 14B decoder exhibits IC about one order of magnitude higher than the 1.3B decoder.
Dataset
Model
b
d
e
Err (%)
IC0
IC range
η2 (%)
HOIGen-1M
1.3B
3.13
0.06
0.000
0.9
0.051
0.022–0.12
3.5
14B
2.51
2.02
0.052
0.7
0.462
0.021–9.95
27.4
MCL-JCV
1.3B
3.06
0.41
0.098
0.5
0.108
0.033–0.35
7.2
14B
1.32
1.26
0.103
0.7
1.322
0.198–8.84
60.0
SA-V
1.3B
1.32
0.54
0.097
0.5
0.662
0.205–2.13
36.8
14B
1.23
0.57
0.057
2.7
4.155
1.391–12.41
94.4
Table 1: Fitted surface parameters and information capacity for DISTS. b and d are the rate and compute exponents, e the distortion floor, and Err the mean relative fit error. IC0=IC(R0,C0) is the grid-center value, the range spans the measured grid, and η2 is the rate saved by doubling compute at (R0,C0) according to Eq. ( 4 ).
Figure 3: Information capacity at the grid center IC(R0,C0) per dataset and decoder in logarithmic scale. Whiskers indicate the minimum and maximum of IC over the measured grid.