stat.MLJun 21, 2026

Data Evolution by Wittgenstein's Rule Following

Authors: Aydin GhojoghBenyamin Ghojogh

Abstract

This paper introduces Wittgenstein's Rule Following (WRF) data evolution, a framework in philomatics for evolving or generating a new dataset from a sequence of previously observed datasets. The method is inspired by Ludwig Wittgenstein's rule-following considerations and his notion of family resemblance in Philosophical Investigations. Unlike standard synthetic data generation, where the goal is usually to sample from or augment a fixed distribution, WRF aims to continue the implicit rule expressed by a historical sequence of datasets while preserving resemblance to the previous datasets. WRF represents each dataset by structural descriptors rather than pointwise correspondences. These descriptors summarize geometric, distributional, clustering, and, in the supervised case, label-based properties of the data. The method predicts a rule-following target by extrapolating descriptor trajectories and a family-resemblance target by averaging historical descriptors. Candidate datasets are then generated from the observed history through balanced or bounded mixture recombination, scored according to these targets, and optionally refined through differentiable optimization in descriptor space. The proposed framework allows both sample size and feature dimension to vary over time and does not assume that the next dataset is a direct transformation of the last one. Simulations on synthetic and image datasets show that WRF can generate meaningful continuations of evolving datasets in both unsupervised and supervised settings.

Explore similar work

May 11, 2026cs.LG

The two clocks and the innovation window: When and how generative models learn rules

Generative models trained on finite data face a fundamental tension: their score-matching or next-token objective converges to the empirical training distribution rather than the population distribution we seek to learn. Using rule-valid synthetic tasks, we trace this tension across two training timescales: τruleτ_{\mathrm{rule}}, the step at which generations first become rule-valid, and τmemτ_{\mathrm{mem}}, the step at which models begin reproducing training samples. Focusing on parity and extending to other binary rules and combinatorial puzzles, we characterize how these two clocks, τruleτ_{\mathrm{rule}} and τmemτ_{\mathrm{mem}}, depend on key aspects of the learning setup. Specifically, we show that τruleτ_{\mathrm{rule}} increases with rule complexity and decreases with model capacity, while τmemτ_{\mathrm{mem}} is approximately invariant to the rule and scales nearly linearly with dataset size NN. We define the \emph{innovation window} as the interval [τrule,τmem][τ_{\mathrm{rule}}, τ_{\mathrm{mem}}]. This window widens with increasing NN and narrows with rule complexity, and may vanish entirely when τruleτmemτ_{\mathrm{rule}} \geq τ_{\mathrm{mem}}. The same two-clock structure arises in both diffusion (DiT) and autoregressive (GPT) models, with architecture-dependent offsets. Dissecting the learned score of DiT models reveals a corresponding evolution of the optimization landscapes, where rule-valid samples' basins expand substantially around τruleτ_{\mathrm{rule}}, while training samples' basins begin to dominate around τmemτ_{\mathrm{mem}}. Together, these results yield a unified and predictive account of when and how generative models exhibit genuine innovation.
Binxu Wang, Emma Lucia Byrnes Finn, Bingbin Liu
Jan 21, 2026cs.LG

Ambient Dataloops: Generative Models for Dataset Refinement

We propose Ambient Dataloops, an iterative framework for refining datasets that makes it easier for diffusion models to learn the underlying data distribution. Modern datasets contain samples of highly varying quality, and training directly on such heterogeneous data often yields suboptimal models. We propose a dataset-model co-evolution process; at each iteration of our method, the dataset becomes progressively higher quality, and the model improves accordingly. To avoid destructive self-consuming loops, at each generation, we treat the synthetically improved samples as noisy, but at a slightly lower noisy level than the previous iteration, and we use Ambient Diffusion techniques for learning under corruption. Empirically, Ambient Dataloops achieve state-of-the-art performance in unconditional and text-conditional image generation and de novo protein design. We further provide a theoretical justification for the proposed framework that captures the benefits of the data looping procedure.
Adrián Rodríguez-Muñoz, William Daspit, Adam Klivans +3
Jun 4, 2026cs.LG

Synthics: Synthetic Physics-like Datasets for Machine Learning

Representative data is fundamental in machine learning, as limited data hinders generalisation. Collecting sufficient real-world samples is often infeasible. Synthetic data generation offers a practical solution, but only if the generated data faithfully reflects the structure of real observations. In this paper, a method for generating synthetic regression datasets that structurally resemble physics equations from a given equation corpus is presented. The approach uses a Bayesian Probabilistic Context-Free Grammar to capture the underlying algebraic structure of the corpus, from which novel equations are sampled. To ensure the generated inputs lie within a physically meaningful domain, the applicability domain is characterised for each equation through non-intrusive probing, also recovering inter-variable constraints. Input sampling further mimics realistic experimental conditions by drawing from random sub-ranges of the valid domain with mixed uniform and truncated normal distributions. The generated data is statistically validated against the Feynman equation corpus using Kolmogorov-Smirnov tests. The generated equations match the corpus on all of the eight studied structural features, compared to only two for an unsmoothed purely probabilistic grammar, demonstrating that the Bayesian prior is essential for structural fidelity given the size of the corpus. In a downstream hyperparameter-tuning task, a gradient-boosted regressor tuned on the synthetic data picks, on average, the 6th-best configuration out of 20 on real data, matching the result of tuning on real data itself and substantially outperforming random expression trees (10th) and noise (19th).
Jari Vepsäläinen