Search-Based Software Engineering and AI Foundation Models: Current Landscape and Future Roadmap
Authors: Hassan Sartaj, Shaukat Ali, Paolo Arcaini, Andrea Arcuri
Organizations: Simula Research Laboratory Oslo, Norway · National Institute of Informatics Tokyo, Japan · Kristiania University of Applied Sciences and Oslo Metropolitan University Oslo, Norway
Search-based software engineering (SBSE), which integrates metaheuristic search techniques with software engineering, has been an active area of research for about 25 years. It has been applied to solve numerous problems across the entire software engineering lifecycle and has demonstrated its versatility in multiple domains. With recent advances in Artificial Intelligence (AI), particularly the emergence of foundation models (FMs) such as large language models (LLMs), the evolution of SBSE alongside these models remains undetermined. In this window of opportunity, we present a research roadmap that articulates the current landscape of SBSE in relation to FMs, identifies open challenges, and outlines potential research directions to advance SBSE through its synergy with FMs. Specifically, we analyze three core aspects: utilizing FMs to enhance SBSE, applying SBSE to advance FMs, and exploring the integration of SBSE and FMs. Furthermore, we present a forward-thinking perspective that envisions the future of SBSE in the era of FMs, highlighting promising research opportunities to address challenges in emerging domains.
Figures & tables
Figure 1: Roadmap overview showing the discussion flow: strengths and weaknesses ( Section 3 ), SBSE–FM synergies ( Sections 4 , 5 and 6 ), empirical evaluation challenges ( Section 7 ), and the 2030 research horizon ( Section 8 ).
Figure 2: Overview of the methodology used to develop the proposed roadmap.
Figure 3: Key aspects of the potential synergy between SBSE and FMs. The abbreviations used are: FMs (Foundation Models), SBSE (Search-Based Software Engineering), and SE (Software Engineering).
Figure 4: Key aspects of employing FMs to enhance SBSE. The abbreviations used are: FMs (Foundation Models), LLMs (Large Language Models), VLMs (Vision-Language Models), MMs (Multimodal Models), SBSE (Search-Based Software Engineering), and SE (Software Engineering).
FMs for SBSE Design
✦Automating the design and implementation of fitness functions for SBSE problems.
✦Designing search operators for complex domain problems involving multimodal data.
✦Automating solution encoding for the varying nature of problems.
✦Defining appropriate search spaces for SBSE to balance efficiency and solution quality.
✦Generating partial or complete SBSE implementations to enhance automation and efficiency.
✦Automating the testing, debugging, and repair of SBSE implementations.
Table 1: Key challenges – FMs for SBSE
Figure 5: Key aspects of applying SBSE to enhance FMs. The abbreviations used are: FMs (Foundation Models), SBSE (Search-Based Software Engineering), and SE (Software Engineering).
SBSE for FM Limitations
✦Managing non-determinism in FMs to ensure consistent and reliable outputs.
✦Effectively handling data-related uncertainties in FMs to improve model performance.
✦Detecting and mitigating hallucinations in FMs.
✦Optimizing the computationally expensive fine-tuning process of FMs, particularly for evolving SE requirements.
✦Balancing trade-offs in automated prompt optimization.
SBSE for SE Artifacts Generated with FMs
Table 2: Key challenges – SBSE for FMs
Figure 6: Integration between FMs and SBSE (two-way interactions). The abbreviations used are: FMs (Foundation Models), SBSE (Search-Based Software Engineering), SE (Software Engineering), SDLC (Software Development Lifecycle), and ADS (Autonomous Driving Systems).
Integrating SBSE and FMs
✦Developing FM-inspired search algorithms to enhance initialization, fitness calculation, and solution repair during the search process.
✦Using FMs to support SBSE in traditional SE problems, such as requirement prioritization, fault identification, and repair for SE tasks, including GUI testing.
✦Effectively integrating FMs with SBSE to tackle complex problems in emerging domains.
Table 3: Key challenges – Integration of SBSE and FMs
Empirical Evaluations involving both SBSE and FMs
✦Ensuring fair comparisons between FM-based and SBSE techniques, considering differences in computational resources and energy consumption.
✦Determining appropriate scaling factors for resource allocation to achieve fairness in experiments.
✦Addressing the replicability of experiments with remote, commercial FMs, which can change unpredictably over time.
✦Balancing the use of commercial FMs for best results with local, open-source LLMs for replicability in scientific research.
✦Managing non-determinism in FMs and SBSE techniques through repeated experiments and proper statistical analysis.
✦Establishing detailed guidelines for fair and reliable comparisons of hybrid techniques involving both FMs and SBSE.
Table 4: Key challenges – Empirical Evaluations
Figure 7: An overview of the tetrad illustrating the disruptive effects of FMs on SBSE.
Large language models (LLMs) embedded in multi-turn agentic harnesses are reshaping software engineering (SWE), but routing every task to a frontier model is wasteful when many issues admit cheap fixes. Existing LLM routers operate on the task description alone, which inherits an information-theoretic Bayes-error floor in agentic settings: a similar issue can hide either a localized typo or a multi-module refactor, and the prompt does not separate the two. We introduce SWE-Router, a value-based temporal approach that lets a cheap model run for a few exploratory turns and reads the resulting partial trajectory before deciding whether to continue cheaply or to escalate to an expensive model. We provide a Bayes-optimality theorem showing that conditioning on the partial trajectory never harms routing and is strictly better whenever exploration is informative. Across the LLM pairs of weak and strong models spanning the contemporary cost--capability frontier, we show that SWE-Router greatly improves the cost efficiency of SWE tasks, while maintaining the majority of the performances of the stronger model. We additionally release a multi-LLM trajectory dataset which allows reproduction of our trajectory-level routing.
Seongho Son, Sangwoong Yoon, Jiahua Tang +3
University College London, United Kingdom · Ulsan National Institute of Science and Technology, South Korea · PSL Research University, France +1
Search has been proposed as an effective method for self-improving language models and agentic systems, both for post-training sample generation and for inference. However, widely used methods such as best-of-N sampling and tree search face two fundamental limitations: they are guided by sparse verification signals, and they construct candidates primarily through autoregressive expansion, restricting exploration to regions with substantial model probability mass. To address these, we propose Bidirectional Evolutionary Search (BES), a search framework that couples forward candidate evolution with backward goal decomposition. In the forward search, BES augments standard expansion with evolution operators that recombine partial trajectories to generate candidates that are difficult to obtain from a single model rollout. In the backward search, BES recursively decomposes the original task into checkable subgoals, producing dense intermediate feedback that guides forward search. We provide theoretical motivation showing that candidates generated by expansion-only search are confined to a narrow entropy shell while evolutionary operators can escape it, and that backward search can exponentially reduce the number of required samples to find a correct answer. Experiments show that on challenging post-training tasks where mainstream post-training algorithms fail to improve, BES enables consistent gains, and on three open problem solving benchmarks at inference time, BES outperforms existing open-source frameworks in both average and best-case performance. Code and trained models are available at https://github.com/Embodied-Minds-Lab/BES.
Deep search capabilities have become an indispensable competency for frontier Large Language Model (LLM) agents, yet their development remains dominated by industrial giants. The typical industry recipe involves a highly resource-intensive pipeline spanning pre-training, continual pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). In this report, we show that when fueled with informative and high-difficulty trajectories, a simple SFT approach could be surprisingly powerful for training frontier search agents. By introducing three simple data synthesis modifications: scaling knowledge graph size for richer exploration, expanding the tool set size for broader functionality, and strict low-step filtering, we establish a stronger baseline. Trained on merely 10.6k data points, our OpenSeeker-v2 achieves state-of-the-art performance across 4 benchmarks (30B-sized agents with ReAct paradigm): 46.0% on BrowseComp, 58.1% on BrowseComp-ZH, 34.6% on Humanity's Last Exam, and 78.0% on xbench, surpassing even Tongyi DeepResearch trained with heavy CPT+SFT+RL pipeline, which achieves 43.4%, 46.7%, 32.9%, and 75.0%, respectively. Notably, OpenSeeker-v2 represents the first state-of-the-art search agent within its model scale and paradigm to be developed by a purely academic team using only SFT. We are excited to open-source the OpenSeeker-v2 model weights and share our simple yet effective findings to make frontier search agent research more accessible to the community.