cs.GROct 8, 2026
SaveTabula Rasa: Monte Carlo estimation of unit-variance noise with controlled spatio-temporal correlation
Organizations: Karlsruhe Institute of Technology, Germany
Abstract
We suggest a method to generate time-varying Gaussian noise with controlled variance and controlled temporal correlation. This noise is used in several downstream tasks for temporal control and temporal coherence. The core technical idea is to phrase this problem as joint Monte-Carlo estimation of both a classic pixel reconstruction and estimation of variance using the concept of "sketching" from the database literature. We demonstrate that our method allows temporal control for downstream tasks with simpler and faster code than previous methods.
Figures & tables
Figure 1. A sample video generated by our method driven by geometry of bunnies. Maintaining temporal correlation of noise patterns is critical for controlling video generation, but maintaining unit variance of the noise is necessary to produce high-quality video. Generating a frame takes on an Nvidia A100 and adds three lines of code to an off-the-shelf Monte Carlo path tracer.
Figure 2. Footprint of a single output pixel (blue) in the random field for 2D (left) and 3D (right). For 3D, two surfaces fall onto one pixel, resulting in a discontinuous (pink) 3D correlation a 2D flow cannot express.
Figure 3. Convergence analysis: Different colors mean different tasks. In the top row, horizontal axis is sample count , while it is bin count in the second row. The vertical axis is, first, MSE of the estimator, where less is better, second, the whiteness of the spectrum i.e., how similar the noise is to white noise, measured as one minus the -distance to the flat spectrum of white noise (less is better), third the variance of the estimated random variable (ideally, unit variance, so, 1), and fourth and finally the Gaussianity measured by the Shapiro-Wilk test (less is better).
Figure 4. Performance instrumentation. In every subplot, the vertical axis is time. In the first row, the horizontal axis is sample count , the second one is image resolution, and the third one is the bin count of the unique count estimator. All horizontal axes are in log scale. Different colors are different tasks, while different line styles mean different underlying noise fields (with prefiltering is solid while dashed is the prefiltered grid). The top row of plots is a different implementation (JAX), while the bottom row is the one in WebGPU. Please see text, for an interpretation.
| Method | LPIPS | SSIM | PSNR | WRP-E | 3D WRP-E |
|---|---|---|---|---|---|
| Random | 0.24 | 0.62 | 23.41 | 369.73 | 790.03 |
| Fixed | 0.23 | 0.62 | 23.48 | 316.07 | 784.45 |
| GWTF | 0.24 | 0.62 | 23.44 | 276.67 | 766.99 |
| Ours | 0.24 | 0.62 | 23.67 | 274.21 | 706.25 |
Table 1. Comparison of different methods across various metrics. Arrows indicate whether higher ( ) or lower ( ) values are better.
Figure 5. Still frame of our animated texture synthesis at the left, an epipolar slice of the animation with our noise in the middle and using random noise at the right.
Figure 6. Epipolar slices (horizontal is space, vertical is time) from super-resolution result of different methods (rows).
Figure 7. Light transport results: the horizontal axis are prompts and the vertical axis is the correlations of the underlying noise.
Figure 8. Using our noise as an NPR primitive in two scenes (columns) . For previous methods (top) , primitive density varies in 3D space to be consistent. For ours (bottom) it is kept constant and consistent.
Figure 9. Producing correlated noise from many views can be used as input to generate lightfields (top, 4 samples of 48 shown) or stereo pairs (bottom, shown in anaglyph).
Figure 10. Video generation results. The top row shows the noise, composed on top of the normal channel (not visible to the generation, only to visualize the 3D configuration). All other rows show four frames from an animation generated from this noise using different prompts. Note how all videos share the 3D structure, where the only guidance comes from the noise correlation.
Figure 11. Correlation of two falling spheres (seen in Fig. 10 ) in conflict with a prompt asking for ice cubes. The generative model produces motion that complies with the one contained in our noise, but despite the correlation with falling spheres, cubes with the right motion are produced. The supplemental video has more results for other conflicting prompts, too.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1. Variance of multiple variates.
Figure 2. 3D and 2D correlation.
Figure 3. Illustration of our estimation procedure for two pixels, 0 and 1 in a 1D-image using a ray-tracing image formation model with 8 samples in each pixel. For each sample, a ray is traced into the 2D world (red lines) to find the intersection with 2D surfaces (blue). Each hit coordinate is hashed to two codes: one to estimate variance (upright Greek, to ) and one to estimate the variate (Roman, A to P). First, the mean of all values for the first (“large”) hash codes is computed. Next, we build a histogram of all the repeating (“small”) hash codes within the pixel. This allows us to recompute the variance and adjust the estimate as per the last row. The hashing can have collisions, as shown in the pixel where large code and both map to code , hence overestimating the variance.
Figure 4. Residual spatial correlation: any method that resamples a pixel grid (two pixels U and V shown as big squares) to a non-aligned texel grid (colored areas) is subject to some form of correlation. Consider the different values at each texel, visualized as colors. As the texel and pixel grid do not align, there are always some texels that are shared by two pixels, in this case , , and . The resulting pixel value is a combination of colors, which all are, without loss of generality, distinct, besides the shared ones.
Figure 5. An example of 1D spatial correlation.
Figure 6. A texel grid of doubling resolution resp. magnification (columns), falling onto a single pixel (yellow square). The first row shows the configuration where texels that would induce correlation with neighbouring pixels are marked orange. The second row shows the respective area of unique vs. correlated texels. We see that with typical resolution the correlated pixel areas vanishes and so does the correlation.
Figure 7. Comparisons of non-photorealistic stippling and hatching renders based on world-space noise (“Baseline”) versus our multi-level, variance-corrected noise (“Ours”). For each scene, two views are shown to highlight the different noise behaviors at different zoom levels. Notice how ours maintains a consistent scale of patterns. Readers are encouraged to zoom in for a closer inspection.
| Modality of Optical Flows | Used by | Supported by | ||
|---|---|---|---|---|
| Type | Direction | Dim. | ||
| Absolute | Forward | 2D | — | Ours |
| Backward | 2D | — | Ours | |
| Relative | Forward | 2D | GWTF | GWTF |
| Backward | 2D | GWTF | GWTF | |
| Absolute | Forward | 3D | — | Ours |
Table 1. Optical Flow Modality comparison.
Explore similar work
We analyze the variance of temporal difference (TD) learning using the phased setting with tabular representation, and show that one of the mechanisms behind its ability to reduce variance is by effectively aggregating over a larger number of independent trajectories. Based on this insight, we demonstrate that (1) the variance of TD is asymptotically bounded from above by Monte Carlo (MC) estimators, and (2) shorter horizon updates incurs less variance for a fixed number of samples. Beyond TD, we show that Direct Advantage Estimation (DAE), a method for estimating the advantage function, can be seen as a type of regression-adjusted control variate, which achieves a tighter bound on the variance compared to TD in the large-sample limit. Finally, we numerically illustrate the behaviors of these estimators with carefully designed environments.
Cross-View Variance Correlation in Path-Traced Stereo:A Hidden Shortcut in Synthetic Training Data
Path-traced synthetic stereo data underlie a large fraction of modern disparity-estimation training pipelines. We report a previously unrecognised property of such data: while the Monte Carlo (MC) noise streams of the two cameras are statistically independent, the underlying \emph{variance fields} -- deterministic per-pixel functions of the rendering integrand -- are highly correlated once aligned by the ground-truth disparity warp. Across 20 scenes rendered with Mitsuba~3, the warped Pearson correlation reaches across 20 scenes at , and on a representative scene remains essentially invariant () over a range of samples per pixel. The effect is strongest in Lambertian regions () and substantially weaker in glass (), as predicted by an integrand decomposition into view-independent and view-dependent components. A residual-shuffle intervention that breaks the cross-view alignment while preserving the clean image degrades the GT cost margin by on non-glass and the variance-based winner-take-all accuracy on glass by , confirming the structure functions as a matching cue. This signal is unique to MC-rendered data and constitutes a candidate sim-to-real shortcut whose impact on trained networks remains to be quantified.
SteadySplats: Resampling of Low-Variance Gaussians for High-Fidelity Stochastic Rendering
Stochastic order-independent transparency enables efficient and elegant rendering of primitive-based radiance fields like 3D Gaussian Splatting models, but remains impractical due to the inherent visible noise in the output. We propose a principled approach to minimize high-frequency noise, addressing its sources at the representation and image synthesis level. During stochastic rendering, our history-based spatial resampling scheme drastically accelerates image convergence, while temporal importance resampling ensures coherence under camera movement. During training, a color regularizer implicitly reduces the variance along view rays in the 3DGS models. With these properties, our optimized, Vulkan-based renderer effectively mitigates output noise at low and high sample counts, achieving a substantial 13~dB PSNR increase in quality over previous stochastic methods at 1 sample per pixel and quickly converging to sorted 3DGS with an average L1 error of less than .