OptiCom : A Unified Framework for State-Conditioned Composition in LLM-Driven Optimization
Organizations: School of Computing and Data Science, The University of Hong Kong · Shenzhen Loop Area Institute · ByteDance · Huazhong University of Science and Technology · School of Information Technology, Carleton University · Hong Kong University of Science and Technology (Guangzhou)
Abstract
Large language models (LLMs) are increasingly deployed to solve complex scientific and practical problems via iterative optimization. However, dynamically coordinating diverse search mechanisms as candidate quality, failure modes, and resource budgets evolve remains a critical open challenge. Targeted empirical diagnostics reveal that mechanism effectiveness is highly state-dependent. Motivated by this, we analyze how individual decisions drive final outcomes, decomposing the expected terminal improvement under a shared budget into cumulative decision opportunities minus cumulative selection losses. Guided by this opportunity-loss theoretical foundation, we propose OptiCom, a unified framework that represents LLM-driven optimizers within a shared configuration space: C=(A,Q,O,E,M,S), corresponding to artifact, query, operator, evaluation, memory, and strategy. Operating within this space, a fast LLM-based Optimization Controller dynamically composes immediate mechanisms through structured Action Packages, while a slower Strategy Adapter refines long-term selection preferences, operator weights, and templates based on accumulated trajectory feedback. Comprehensive evaluations across 32 benchmark groups demonstrate the superiority of framework: OptiCom achieves an average Max-score rank of 1.72 among 14 evaluated configurations, securing the top score in 23 groups. Ultimately, these results highlight the broad applicability and high extensibility of OptiCom as a general-purpose paradigm for robust LLM test-time scaling.
Figures & tables
| Feedback-guided revision and experience reuse | Evolutionary search and adaptive control | Shared abstractions and meta-optimization | |
|---|---|---|---|
| TextGrad Yuksekgonul et al. (2025) | AdaEvolve Cemri et al. (2026) | DSPy Khattab et al. (2024) | |
| Optimizable graph variables | Candidate programs | LM program instructions and demonstrations | |
| Graph context and propagated feedback | Selected programs and search evidence | Training examples and execution traces | |
| Textual-feedback-guided updates | LLM-generated program variations | Instruction and demonstration updates | |
| Objective feedback | Program fitness | User-specified task metric | |
| Graph and update state | Populations and improvement statistics | Optimizer-dependent candidates and search state |
| Domain | Math | Systems | GPU | Algorithms | Reasoning | Prompts | Quantum | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Benchmark | Heilbronn | LLM-SQL | TriMul | Frontier-CS | ARC | HotpotQA | QNN | |||||||
| Method | Max | Mean | Max | Mean | Max | Mean | Max | Mean | Max | Mean | Max | Mean | Max | Mean |
| ProTeGi | 0.828 | 0.813 | 0.630 | 0.622 | 2.541 | 2.531 | 0.632 | 0.619 | 0.473 | 0.461 | 0.441 | 0.427 | 0.8833 | 0.8367 |
| LATS | 0.887 | 0.867 | 0.686 | 0.682 | 2.576 | 2.542 | 0.627 | 0.609 | 0.457 | 0.432 | 0.413 | 0.409 | 0.8500 | 0.8300 |
| EvoX | 0.943 | 0.931 | 0.684 | 0.677 | 2.593 | 2.583 | 0.626 | 0.617 | 0.461 | 0.443 | 0.386 | 0.357 | 0.8167 | 0.8067 |
| EvoX (official) | 0.956 | 0.941 | 0.696 | 0.681 | 2.592 | 2.579 | 0.613 | 0.597 | 0.473 | 0.451 | 0.397 | 0.363 | 0.8333 | 0.8133 |
| Config. | Heilbronn | HotpotQA |
|---|---|---|
| Full OptiCom | ||
| Fixed composition | ||
| Random composition | ||
| w/o Strat. Adapter | ||
| Operator-only adapt. | ||
| w/o Exp. Memory |
| Heilbronn Triangle | HotpotQA | |||||
|---|---|---|---|---|---|---|
| Backbone | w/o Adapter | Full OptiCom | Gain | w/o Adapter | Full OptiCom | Gain |
| Doubao-Seed-2.0-pro | ||||||
| GPT-5.5 | ||||||
| Claude Opus 4.6 | ||||||
| GLM-5.3 | ||||||
| Kimi-K3 | ||||||
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Dimension | Optimization concept | Functional responsibility | Illustrative realizations |
|---|---|---|---|
| : Artifact | Solution representation | What is optimized, and how candidates are represented. | Source code, numerical parameters, mathematical constructions, prompts, or images. |
| : Query | Information acquisition | What evidence and context are acquired or assembled to inform an update. | Inspect selected failures, request profiling evidence, retrieve past attempts, or use only the current candidate and feedback. |
| : Operator | Search and update operations | How candidates are generated or modified. | Generate a new candidate, repair an error, apply a local revision, or recombine existing candidates. |
| : Evaluation | Objective and constraint assessment | How candidate quality and feasibility are assessed, and what feedback is returned. | Correctness checks, runtime measurements, diagnostic tests, or higher-fidelity verification. |
| : Memory | Search history and experience | What information persists across iterations, and how it is maintained. | Retain an incumbent, a candidate archive, evaluation records, or reusable lessons; omit persistent experience storage when unused. |
| : Strategy | Search control | How the other functions are coordinated and resources are allocated as optimization proceeds. | Follow a prescribed schedule, apply state-dependent rules, or adapt selection preferences using accumulated outcomes. |
| Method / status | Artifact, query, and operator | Evaluation, memory, and strategy |
|---|---|---|
| OPRO Yang et al. (2024) B | A: Candidate solutions, including prompts. Q: Task description and selected scored solutions. O: Generate proposals conditioned on that history. | E: Task objective values. M: Evaluated candidate–score records. S: Repeated generation and evaluation with history selection and stopping rules. |
| TextGrad Yuksekgonul et al. (2025) B | A: Optimizable variables in a computational graph. Q: Relevant graph context and propagated feedback. O: Variable updates guided by textual feedback. | E: Objective evaluation and associated criticism. M: Graph state, current variables, and optimizer-dependent update history. S: Feedback propagation and variable-update procedure. |
| ProTeGi Pryzant et al. (2023) B | A: Task prompts. Q: Minibatch examples and observed errors. O: Generate textual critiques and corresponding prompt revisions. | E: Predictive performance on evaluation examples. M: Beam candidates and evaluation statistics. S: Beam expansion and bandit-based candidate selection. |
| Reflexion Shinn et al. (2023) B | A: Task attempts, including programs or action trajectories. Q: Current feedback and previous verbal reflections. O: Produce a subsequent attempt informed by reflection. | E: Task-dependent external or internal feedback. M: Episodic verbal reflections. S: Attempt, evaluate, reflect, and retry. |
| GEPA Agrawal et al. (2025) B | A: Prompts in an AI system. Q: Execution trajectories, diagnostics, and selected candidate context. O: Reflective mutation and combination of complementary candidates. | E: Task evaluations, including performance across examples. M: Candidate pool, scores, and ancestry. S: Pareto-aware selection and evolutionary updates under an evaluation budget. |
| FunSearch Romera-Paredes et al. (2023) B | A: Functions within a program specification. Q: Selected prior functions assembled into a prompt. O: LLM-generated function variants. | E: Automated execution and scoring. M: An island-structured database of evaluated programs. S: Program sampling, insertion, and island reset rules. |
| Aspect | Mechanism to preserve | Alignment check |
|---|---|---|
| Candidate generation | The information and update rule used to produce new candidates. | Check that generation receives the intended parents, examples, feedback, or revision instructions. |
| Selection and search | The method’s candidate selection, branching, population, or tree-search rule. | Check that these rules affect executed actions, rather than appearing only in configuration descriptions. |
| Feedback use | The feedback representation and its role in subsequent optimization. | Check that evaluation records, reflections, or textual critiques reach the decisions they are intended to guide. |
| Memory and reuse | The information retained across iterations and the rule for retrieving it. | Check that retained candidates or experiences remain available and are retrieved according to the configured mechanism. |
| Strategy adaptation | The trigger, evidence, and scope of changes to the optimization procedure. | Check that strategy updates affect later decisions; distinguish adaptive rules within a configuration from changes to the configuration itself. |
| Evaluation and budget | Task validity requirements, scoring semantics, and resource constraints. | Check that candidates use the shared benchmark evaluator and that reported limits include the relevant optimization operations. |
| State | Initial condition | Success criterion |
|---|---|---|
| s1 | Execution failure or hard-constraint violation | Recover a candidate that executes and satisfies the required hard constraints. |
| s2 | Failed correctness tests | Preserve feasibility and pass all required correctness checks. |
| s3 | Objective stagnation | Preserve feasibility and correctness and exceed the predefined objective-improvement criterion. |
| s4 | Limited remaining budget | Reach the s3 objective criterion within the declared remaining budget. |
| s5 | High performance variability | Reduce variability beyond the predefined criterion without exceeding the allowed degradation in mean quality, feasibility, or correctness. |
| ID | Situation | Starting condition |
|---|---|---|
| s1-1 | Runtime exception | Point generation or updating raises an exception, such as an out-of-range index or incompatible array dimensions, and terminates without returning a point set. |
| s1-2 | Construction timeout | Excessive search, too many internal iterations, or an ineffective termination condition prevents the construction from finishing within the declared execution limit. |
| s1-3 | Invalid output interface | The program terminates but returns the wrong number of points, an array other than shape , or coordinates containing NaN or infinity. |
| s2-1 | Incorrect boundary handling | The program returns a well-formed point set, but rectangular clipping or an incorrect projection leaves points outside the sloping boundaries of the equilateral triangle. |
| s2-2 | Incorrect area objective | The search uses an incorrect internal area calculation, such as omitting the absolute determinant or using an incorrect normalization factor, so its internal objective disagrees with the specified geometric objective. |
| s2-3 | Incomplete triple enumeration | The internal objective evaluates only a subset of triples, such as consecutive or locally selected points, and can miss the triple determining the true minimum area. |
| ID | Mechanism | s1-1 | s1-2 | s1-3 | s2-1 | s2-2 | s2-3 | s3-1 | s3-2 | s3-3 | s4-1 | s4-2 | s4-3 | s5-1 | s5-2 | s5-3 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Q1 | none | 3 | 3 | 4 | 3 | 3 | 3 | 3 | 3 | 3 | 4 | 3 | 4 | 2 | 2 | 2 |
| Q2 | self-only (R) | 5 | 4 | 6 | 5 | 5 | 4 | 4 | 3 | 4 | 4 | 3 | 3 | 3 | 3 | 3 |
| Q3 | env-probe | 9 | 7 | 8 | 8 | 8 | 7 | 4 | 4 | 4 | 2 | 3 | 3 | 3 | 3 | 3 |
| Q4 | memory-lookup | 3 | 2 | 4 | 3 | 3 | 3 | 3 | 3 | 3 | 4 | 3 | 4 | 2 | 2 | 2 |
| O1 | local-revision (R) | 5 | 4 | 6 | 5 | 5 | 4 | 4 | 3 | 4 | 4 | 3 | 3 | 3 | 3 | 3 |
| O2 | repair | 8 | 6 | 8 | 8 | 8 | 7 | 2 | 2 | 2 | 5 | 4 | 6 | 2 | 2 | 2 |
| ID | Mechanism | s1-1 | s1-2 | s1-3 | s2-1 | s2-2 | s2-3 | s3-1 | s3-2 | s3-3 | s4-1 | s4-2 | s4-3 | s5-1 | s5-2 | s5-3 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| M2 | episodic + Q4 | 6 | 5 | 7 | 6 | 6 | 5 | 5 | 3 | 5 | 5 | 4 | 4 | 3 | 3 | 4 |
| M3 | structured + Q4 | 6 | 5 | 7 | 7 | 6 | 6 | 6 | 4 | 6 | 7 | 7 | 7 | 5 | 5 | 6 |
| M4 | adaptive + Q4 | 7 | 6 | 8 | 7 | 6 | 7 | 6 | 5 | 7 | 7 | 7 | 7 | 5 | 5 | 7 |
| Configuration | Query | Operator | Successes | Avg. iterations |
|---|---|---|---|---|
| R | self-only | local-revision | 14/30 | 5.2 |
| Q-only | env-probe | local-revision | 23/30 | 2.9 |
| O-only | self-only | repair | 23/30 | 3.6 |
| Q+O | env-probe | repair | 25/30 | 2.1 |
| Field | Decision | Consumer and constraints |
|---|---|---|
| Query plan | Query executor: permitted tools, evidence requests, and retrieval operations; may be empty. | |
| Context scope | Context builder: candidate identifiers, diagnostics, and experience to expose within the context limit. | |
| Update operator | Operator executor: a registered mechanism with compatible artifact inputs and required parent candidates. | |
| Branch width | Scheduler: a positive bounded number of candidate proposals, subject to available resources. | |
| Evaluation plan | Evaluator: registered stages, effort limits, and feedback form; final required checks remain unchanged. | |
| Retention policy | Archive and experience stores: which evaluated candidates, records, and lessons to retain or expose. |
| Domain | Benchmark family | Evaluation target |
|---|---|---|
| Math | Mathematical optimization | Numerical objectives subject to task-specific feasibility constraints. |
| Systems | ADRS | Workload performance, cost, and composite system objectives. |
| GPU | GPU Mode; KernelBench | Numerical correctness and kernel execution performance; scaled inverse runtime or eager-baseline speedup, respectively. |
| Algorithms | Frontier-CS; ALE-Bench-Lite | Bounded algorithmic-problem scores or private contest-performance scores over the stated problem sets. |
| Reasoning | ARC | Correctness of candidate transformations on benchmark inputs. |
| Creative | Sky Festival | Satisfaction of semantic and compositional image requirements. |
| Heilbronn Triangle | HotpotQA | |||
|---|---|---|---|---|
| Backbone | Full OptiCom | w/o Adapter | Full OptiCom | w/o Adapter |
| Doubao-Seed-2.0-pro | ||||
| GPT-5.5 | ||||
| Claude Opus 4.6 | ||||
| GLM-5.3 | ||||
| Kimi-K3 | ||||
| Backbone | Heilbronn Triangle | HotpotQA |
|---|---|---|
| Doubao-Seed-2.0-pro | ||
| GPT-5.5 | ||
| Claude Opus 4.6 | ||
| GLM-5.3 | ||
| Kimi-K3 |
| Configuration | Final score | Adapter calls | Total tokens |
|---|---|---|---|
| No adaptation | 0 | 853K | |
| Event-triggered adaptation | 10 | 865K | |
| Periodic adaptation ( ) | 5 | 861K | |
| Every-iteration adaptation | 29 | 877K |
| Final score | Total tokens | |||
|---|---|---|---|---|
| Feedback | Full OptiCom | w/o Adapter | Full OptiCom | w/o Adapter |
| Basic | 809K | 799K | ||
| Rich | 865K | 853K | ||
| Initial configuration | Fixed | Adaptive OptiCom |
|---|---|---|
| Local-revision-first | ||
| Diagnosis-and-repair-first | ||
| Exploration-and-recombination-first |