cs.GROct 8, 2026
SaveTabula Rasa: Monte Carlo estimation of unit-variance noise with controlled spatio-temporal correlation
Organizations: Karlsruhe Institute of Technology, Germany
Abstract
We suggest a method to generate time-varying Gaussian noise with controlled variance and controlled temporal correlation. This noise is used in several downstream tasks for temporal control and temporal coherence. The core technical idea is to phrase this problem as joint Monte-Carlo estimation of both a classic pixel reconstruction and estimation of variance using the concept of "sketching" from the database literature. We demonstrate that our method allows temporal control for downstream tasks with simpler and faster code than previous methods.
Figures & tables
Figure 1. A sample video generated by our method driven by geometry of bunnies. Maintaining temporal correlation of noise patterns is critical for controlling video generation, but maintaining unit variance of the noise is necessary to produce high-quality video. Generating a frame takes on an Nvidia A100 and adds three lines of code to an off-the-shelf Monte Carlo path tracer.
Figure 2. Footprint of a single output pixel (blue) in the random field for 2D (left) and 3D (right). For 3D, two surfaces fall onto one pixel, resulting in a discontinuous (pink) 3D correlation a 2D flow cannot express.
Figure 3. Convergence analysis: Different colors mean different tasks. In the top row, horizontal axis is sample count , while it is bin count in the second row. The vertical axis is, first, MSE of the estimator, where less is better, second, the whiteness of the spectrum i.e., how similar the noise is to white noise, measured as one minus the -distance to the flat spectrum of white noise (less is better), third the variance of the estimated random variable (ideally, unit variance, so, 1), and fourth and finally the Gaussianity measured by the Shapiro-Wilk test (less is better).
Figure 4. Performance instrumentation. In every subplot, the vertical axis is time. In the first row, the horizontal axis is sample count , the second one is image resolution, and the third one is the bin count of the unique count estimator. All horizontal axes are in log scale. Different colors are different tasks, while different line styles mean different underlying noise fields (with prefiltering is solid while dashed is the prefiltered grid). The top row of plots is a different implementation (JAX), while the bottom row is the one in WebGPU. Please see text, for an interpretation.
| Method | LPIPS | SSIM | PSNR | WRP-E | 3D WRP-E |
|---|---|---|---|---|---|
| Random | 0.24 | 0.62 | 23.41 | 369.73 | 790.03 |
| Fixed | 0.23 | 0.62 | 23.48 | 316.07 | 784.45 |
| GWTF | 0.24 | 0.62 | 23.44 | 276.67 | 766.99 |
| Ours | 0.24 | 0.62 | 23.67 | 274.21 | 706.25 |
Table 1. Comparison of different methods across various metrics. Arrows indicate whether higher ( ) or lower ( ) values are better.
Figure 5. Still frame of our animated texture synthesis at the left, an epipolar slice of the animation with our noise in the middle and using random noise at the right.
Figure 6. Epipolar slices (horizontal is space, vertical is time) from super-resolution result of different methods (rows).
Figure 7. Light transport results: the horizontal axis are prompts and the vertical axis is the correlations of the underlying noise.
Figure 8. Using our noise as an NPR primitive in two scenes (columns) . For previous methods (top) , primitive density varies in 3D space to be consistent. For ours (bottom) it is kept constant and consistent.
Figure 9. Producing correlated noise from many views can be used as input to generate lightfields (top, 4 samples of 48 shown) or stereo pairs (bottom, shown in anaglyph).
Figure 10. Video generation results. The top row shows the noise, composed on top of the normal channel (not visible to the generation, only to visualize the 3D configuration). All other rows show four frames from an animation generated from this noise using different prompts. Note how all videos share the 3D structure, where the only guidance comes from the noise correlation.
Figure 11. Correlation of two falling spheres (seen in Fig. 10 ) in conflict with a prompt asking for ice cubes. The generative model produces motion that complies with the one contained in our noise, but despite the correlation with falling spheres, cubes with the right motion are produced. The supplemental video has more results for other conflicting prompts, too.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1. Variance of multiple variates.
Figure 2. 3D and 2D correlation.
Figure 3. Illustration of our estimation procedure for two pixels, 0 and 1 in a 1D-image using a ray-tracing image formation model with 8 samples in each pixel. For each sample, a ray is traced into the 2D world (red lines) to find the intersection with 2D surfaces (blue). Each hit coordinate is hashed to two codes: one to estimate variance (upright Greek, to ) and one to estimate the variate (Roman, A to P). First, the mean of all values for the first (“large”) hash codes is computed. Next, we build a histogram of all the repeating (“small”) hash codes within the pixel. This allows us to recompute the variance and adjust the estimate as per the last row. The hashing can have collisions, as shown in the pixel where large code and both map to code , hence overestimating the variance.
Figure 4. Residual spatial correlation: any method that resamples a pixel grid (two pixels U and V shown as big squares) to a non-aligned texel grid (colored areas) is subject to some form of correlation. Consider the different values at each texel, visualized as colors. As the texel and pixel grid do not align, there are always some texels that are shared by two pixels, in this case , , and . The resulting pixel value is a combination of colors, which all are, without loss of generality, distinct, besides the shared ones.
Figure 5. An example of 1D spatial correlation.
Figure 6. A texel grid of doubling resolution resp. magnification (columns), falling onto a single pixel (yellow square). The first row shows the configuration where texels that would induce correlation with neighbouring pixels are marked orange. The second row shows the respective area of unique vs. correlated texels. We see that with typical resolution the correlated pixel areas vanishes and so does the correlation.
Figure 7. Comparisons of non-photorealistic stippling and hatching renders based on world-space noise (“Baseline”) versus our multi-level, variance-corrected noise (“Ours”). For each scene, two views are shown to highlight the different noise behaviors at different zoom levels. Notice how ours maintains a consistent scale of patterns. Readers are encouraged to zoom in for a closer inspection.
| Modality of Optical Flows | Used by | Supported by | ||
|---|---|---|---|---|
| Type | Direction | Dim. | ||
| Absolute | Forward | 2D | — | Ours |
| Backward | 2D | — | Ours | |
| Relative | Forward | 2D | GWTF | GWTF |
| Backward | 2D | GWTF | GWTF | |
| Absolute | Forward | 3D | — | Ours |
Table 1. Optical Flow Modality comparison.