Novelty Search

Momentum

2 papers in the last four weeks, with none the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 10

Sep 27, 2026cond-mat.mtrl-sci

Let CSP Be Your ANCHOR: Adaptive Crystal Search over Frozen Structure Priors

De novo crystal generation (DNG) models decide where to search in composition space and how to generate structures with one set of weights. We argue that discovery is better served by separating the two. A crystal structure prediction (CSP) model is a physical prior that should be improved by likelihood training, while rewards, including novelty measured against the search's own history, should act on a search over compositions. We introduce ANCHOR, a GRPO composition policy trained with multi-objective rewards around a frozen CSP model, and continuous adaptive novelty (CAN), a graded novelty score against known structures and a growing discovery history. Using the frozen CSP model as a fixed ruler under one evaluator, we test where adaptation should act. Replacing DNG compositions with ANCHOR's policy on the same CSP backbone raises MSUN from 11.4% to 47.6% and SUN from 1.1% to 22.1% at 99.9% formula uniqueness. Fine-tuning DNG models directly on the same rewards instead moves their composition marginal without raising their on-hull fraction. We show that KL-regularized fine-tuning of a DNG model can only reweight chemistry the pretrained model already supports by a bounded factor, while unregularized DNG fine-tunes move toward known or less stable chemistry. Even a stability-only reward routed into ANCHOR's CSP backbone roughly halves SUN relative to the frozen backbone, whereas likelihood training on structures found during search can improve a CSP backbone. Under MatterGen's evaluation pipeline, ANCHOR raises state-of-the-art MSUN from 29.2% to 41.3%, transfers without retraining to two further CSP backbones, and reaches 47.1% after distillation into Crystalite-CSP. As with any model optimised against a potential, its on-hull rate depends on that potential.
Sep 27, 2026cs.AI

RINI: Seeing the Prior Is Not Enough

A research proposal can describe an established mechanism correctly while claiming to introduce it. We study whether providing the earlier paper corrects such contribution claims. Three controlled experiments compare proposals generated with a contribution-bearing prior and a same-topic control. Providing the prior yields no clear aggregate reduction in unsupported novelty. Human analysis of 175 interpretable exposed proposals finds that 137 recognize the prior's relevance, but 61 correctly attribute the established contribution. Of 71 proposed remaining distinctions, 37 are covered by the same prior. We introduce Research Idea Novelty Inspection (RINI), which audits contribution claims against evidence, checks the remaining distinction, and applies local revisions. Five human annotators evaluate 1,080 original-revision pairs across three methods. On the same 240 originals judged to require correction, successful repair is 11.7% for Self-Revision, 39.1% for Retrieve-and-Revise, and 72.2% for RINI, with research tasks weighted equally. The improvement over same-evidence direct revision is 33.0 percentage points. The revised proposals retain their research questions and technical methods. These results motivate explicit contribution attribution when using literature to generate and revise research proposals.
Jul 20, 2026cs.LG

Intelligence from Learnable Novelty

Intelligence appears under different names in different fields: as data compression in statistics and machine learning, as universal computation in dynamical systems, and as adaptive behavior in agents. Each field carries its own objective, and the two most influential drives often fail in mirror image: novelty search, which seeks surprise, is transfixed by a noisy television screen, while the free-energy principle, which avoids surprise, is most content in a dark room. Both failures have a single cause: each objective treats as one quantity the surprise a learner can convert into knowledge and the surprise it never can. Here we show that the learnable part of that information, which we call learnable novelty, yields the seemingly disparate projections of intelligence, and we give a closed-form estimator of it built on a cheap and differentiable reservoir computer. Used as a measure, with no supervision of any kind, the estimator recovers decades of complexity classification, ranking the Turing-complete rule110 highest among the elementary cellular automata. Used as an objective, its gradient carries a neural cellular automaton from simple dynamics into a regime of solitons, the traveling, colliding structures by which rule110 computes, as well as organizes the representation of an image encoder around the ten digit classes of MNIST, fully unsupervised: no label ever enters training. Handed to a reinforcement-learning agent as an intrinsic reward, it supplies the exploration that task rewards lack, improving on the task baseline in nine of ten environments and collapsing in none. Complexity generation, abstraction, and exploration, ordinarily pursued with unrelated objectives in separate fields, thus emerge from ascent on one differentiable quantity, and the projections of intelligence gain a common quantitative footing.
Jul 1, 2026cs.NE

Novelty Search with Cross-Task Collaborative Discovery

Novelty Search (NS) promotes exploration by rewarding behaviorally novel solutions rather than directly optimizing predefined objectives, making it particularly useful when objective guidance is sparse or deceptive. However, existing NS methods are typically designed for a single search task. When multiple related NS tasks are considered independently, their search processes may repeatedly explore similar regions or rediscover solutions that could be useful across tasks, leading to redundant use of the evaluation budget. To address this issue, we formulate a multitask novelty search setting and propose Multitask Novelty Search with Cross-Task Collaborative Discovery (MTNS-CoD). The central idea is to coordinate discovery across tasks so that different task-specific searches explore complementary regions while useful discoveries can still be reused across tasks. Specifically, MTNS-CoD introduces a multitask repulsion mechanism to discourage redundant exploration in similar genotype regions, together with an adaptive inter-task transfer mechanism that adjusts transfer probabilities according to the observed utility of cross-task exchanges during evolution. Their joint use promotes complementary exploration while selectively reusing beneficial discoveries. We further extend MTNS-CoD to novelty-augmented optimization, where behavioral novelty and objective information are jointly considered to support exploration under deceptive objective landscapes. Experiments on synthetic benchmarks, deceptive maze navigation, MuJoCo policy optimization, and generative novelty search show that MTNS-CoD can improve behavioral discovery and search coverage over the considered single-task and multitask baselines, with additional benefits observed on problems containing deceptive objectives.
Jun 23, 2026cs.AI

Heuresis: Search Strategies for Autonomous AI Research Agents Across Quality, Diversity and Novelty

Autonomous AI Research promises to accelerate the scientific progress of machine learning. To realise this goal, current Large Language Model (LLM)-based agents need to go beyond just writing code, to mastering the exploration of simultaneously performant, diverse and novel ideas. To this end, we introduce Heuresis, a framework that abstracts the research pipeline into a set of general and composable primitives, enabling open-ended scientific exploration in machine learning research. We implement six search strategies: a greedy baseline, two archive-based (MAP-Elites, Go-Explore), one evolutionary (Islands), and two divergent (Curiosity, Omni), and evaluate them across three axes (Quality, Diversity, and Novelty) on three domains (LLM Pretraining, On-Policy RL, and Model Unlearning), totalling 3,222 scored runs. We find that completely novel ideas are rare. No idea across our scored runs is rated as "Original", and only a few achieve only "Minor Similarity" to prior work. Moreover, novel ideas never approach the highest-performing known-recipe scores. Across all six strategies and three domains, only one such idea lands in the top-10 by quality. We also observed agents resorting to a variety of reward-hacking techniques during execution (40 confirmed fabrications across 1,628 scored runs), and detecting them was necessary to keep the search faithful to the task. Our results show that while current search and Quality-Diversity strategies enable us to steer where the generated ideas land on the quality, diversity, and novelty axes, they do not expand the quality-novelty frontier. Bridging this gap is the open challenge towards the ultimate goal of perpetual, autonomous scientific progress. Code is available at github.com/a-antoniades/Heuresis.
May 7, 2026cs.AI

Novelty-based Tree-of-Thought Search for LLM Reasoning and Planning

Although advances such as chain-of-thought, tree-of-thought or reinforcement learning have improved the performance of LLMs in reasoning and planning tasks, they are still brittle and have not achieved human-level performance in many domains, and often suffer from high time and token costs. Inspired by the success of width-based search in planning, we explore how the concept of novelty can be transferred to language domains and how it can improve tree-of-thought reasoning. A tree of thoughts relies on building possible "paths" of consecutive ideas or thoughts. These are generated by repeatedly prompting an LLM. In our paper, a measurable concept of novelty is proposed that describes the uniqueness of a new node (thought) in comparison to nodes previously seen in the search tree. Novelty is estimated by prompting an LLM and making use of embedded general knowledge from pre-training. This metric can then be used to prune branches and reduce the scope of the search. Although this method introduces more prompts per state, the overall token cost can be reduced by pruning and reducing the overall tree size. This procedure is tested and compared using several benchmarks in language-based planning and general reasoning.
Apr 23, 2026cs.NE

Novelty-Based Generation of Continuous Landscapes with Diverse Local Optima Networks

Local Optima Networks (LONs) represent the global structure of search spaces as graphs, but their construction requires iterative execution of a search algorithm to find local optima and approximate transitions between Basins of Attraction (BoAs). In continuous optimization, this high computational cost prevents systematic investigation of the relationship between LON features and evolutionary algorithm performance. To address this issue, we propose an alternative definition of BoAs for Max-Set of Gaussians (MSG) landscapes with explicitly tunable multimodality. This bypasses search-based BoA identification, enabling low-cost LON construction. Moreover, we leverage Novelty Search (NS) to explore the parameter space of the MSG landscape generator, producing instances with diverse graph topologies. Our experiments show that the proposed BoAs closely align with gradient-based BoAs, and that NS successfully generates instances with varied search difficulty and connectivity patterns among optima. Finally, over the instances generated by NS, we predict the success rate of two well-established evolutionary algorithms from LON features. While our LON construction is specific to MSG landscapes, the proposed framework provides a dataset that serves as a foundation for landscape-aware optimization.
Apr 13, 2026cs.NE

Evolving Many Worlds: Towards Open-Ended Discovery in Petri Dish NCA via Population-Based Training

The generation of sustained, open-ended complexity from local interactions remains a fundamental challenge in artificial life. Differentiable multi-agent systems, such as Petri Dish Neural Cellular Automata (PD-NCA), exhibit rich self-organization driven purely by spatial competition; however, they are highly sensitive to hyperparameters and frequently collapse into uninteresting patterns and dynamics, such as frozen equilibria or structureless noise. In this paper, we introduce PBT-NCA, a meta-evolutionary algorithm that evolves a population of PD-NCAs subject to a composite objective that rewards both historical behavioral novelty and contemporary visual diversity. Driven by this continuous evolutionary pressure, PBT-NCA spontaneously generates a plethora of emergent lifelike phenomena over extended horizons-a hallmark of true open-endedness. Strikingly, the substrate autonomously discovers diverse morphological survival and self-organization strategies. We observe highly regular, coordinated periodic waves; spore-like scattering where homogeneous groups eject cell-like clusters to colonize distant territories; and fluid, shape-shifting macro-structures that migrate across the substrate, maintaining stable outer boundaries that enclose highly active interiors. By actively penalizing monocultures and dead states, PBT-NCA sustains a state of effective complexity that is neither globally ordered nor globally random, operating persistently at the "edge of chaos".
Jul 15, 2025cs.AI

Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light

Artificial learning systems are graduating from passive learners to increasingly autonomous agents, lending pragmatic urgency to the question of what constitutes agency. Reinforcement learning (RL) offers arguably the most explicit formulation of agent-environment interaction, built on three core tenets: the environment as a Markov decision process, learning as policy optimization, and the agent as a maximizer of scalar reward. Recent work has called to revise these tenets: reconceptualizing learning as adaptation rather than optimization, broadening goals beyond scalar reward, and noting the absence of a formal theory of the agent in a formalism that so heavily emphasizes the environment. We argue that the artificial life community is uniquely positioned to illuminate this critique and concretize an alternative. We draw on open-ended novelty search as a complementary model of adaptation and goal-directed behavior beyond reward optimization, and ground such evolutionary dynamics in thermodynamic theories of origin-of-life and agency, toward a more biologically faithful and formally grounded account of what it is to be an adaptive agent.
Jun 5, 2024stat.ML

BEACON: A Bayesian Optimization Inspired Strategy for Efficient Novelty Search

Novelty search (NS) aims to uncover diverse system behaviors through simulation or experiment without requiring a pre-specified scalar objective. This capability is especially relevant to modern discovery problems in chemistry, materials science, and molecular design, where researchers often seek broad coverage of attainable property space rather than a single optimum and where each evaluation may require a costly computation or experiment. For such expensive black-box settings, we propose BEACON, a sample-efficient NS strategy inspired by Bayesian optimization principles. BEACON models the input-to-outcome mapping using multi-output Gaussian processes and selects new inputs by scoring how far plausible posterior outcomes lie from a denoised archive of previously observed outcomes. This gives a distance-based novelty acquisition that accounts for predictive uncertainty and observational noise while operating directly in continuous outcome space, rather than requiring direct optimization over a discretized partition of behaviors. By leveraging efficient posterior sampling together with scalable high-dimensional Gaussian process models, the proposed framework can be extended to settings with large data sets and high-dimensional design variables. We demonstrate BEACON on established benchmark problems together with real-world case studies in materials and molecular discovery. Across these settings, BEACON consistently discovers broader sets of distinct behaviors than several competing baselines under limited evaluation budgets.