ConflictGuide: AutoResearch Improves When Competing Behaviors Are Made Visible
Organizations: National University of Singapore · Nanjing University of Science and Technology
Abstract
When designing machine learning models, desirable properties are often in tension: improving one behavior can impair another, so task progress can depend on alleviating the conflict. LLM-based AutoResearch systems, which iteratively edit model code and retain edits based on scalar task-performance feedback, have largely ignored this trade-off. We find that scalar feedback supports broad exploration early in search, but it does not reveal how edits affect competing behaviors. In matched-budget experiments, introducing competing-behavior feedback as task gains diminish increases the share of proposals that improve both behaviors and sustains progress beyond scalar-only plateaus. Obtaining this feedback for a given model requires identifying its competing behaviors and designing probes to measure them. To make competing-behavior feedback actionable, we introduce ConflictGuide. Its reusable ConflictGuide-Skill combines a literature-grounded taxonomy with model-specific evidence to identify competing behaviors and specify probes for a code agent to implement as metrics. Evolution proceeds in two stages: Stage I explores with task feedback; Stage II uses probe feedback to steer proposals toward conflict alleviation and retains marginal-gain edits only when probes indicate sufficient alleviation. Across five diverse model families, ConflictGuide reduces task and conflict-related errors by up to 28% and 14%, respectively, relative to scalar-only AutoResearch, with gains extending to other code agents.
Figures & tables
| Model | Domain / Task | Conflict mechanism (Axis) | Primary Probe metric(s) |
|---|---|---|---|
| SpecB–FNO | Scientific ML / PDE operator learning | R2.M2 ( space ) Frequency-structure mismatch: dominant-mode fit vs. non-dominant-mode utilization | |
| SNGP | Uncertainty-aware classification / OOD detection | R3.M4 ( objective ) Compression vs. uncertainty: predictive fit vs. distance-aware feature geometry | |
| ESN | Reservoir computing / time-series prediction | R2.M4 ( time ) Temporal-component mismatch: memory retention vs. nonlinear processing | |
| TCM–Lite | Learned image compression / rate–distortion coding | R3.M10 ( space ) Compact representation damaging fine structure: latent rate vs. detail preservation | |
| GCNII | Graph learning / node classification | R2.M5 ( space ) Relational-structure mismatch: useful aggregation vs. incompatible messages |
| Round | Method | Main Cumulative Edit | NRMSE | ND-NMSE | Params (M) |
|---|---|---|---|---|---|
| 0 | Reference | None (Original) | 328.5 | ||
| Autoresearch | Spatial gate | 328.9 | |||
| ConflictGuide | Mode refinement | 333.5 | |||
| 1 | Reference | None (Original) | 328.5 | ||
| Autoresearch | Energy match | 403.9 | |||
| ConflictGuide | Multi-scale residual | 407.2 |
| Round | Method | Main Cumulative Edit | NLL | OOD AUPR | |
|---|---|---|---|---|---|
| SVHN | CIFAR-10 | ||||
| 0 | Reference | None (Original) | |||
| Autoresearch | Multi-scale RFF | ||||
| ConflictGuide | LN-residual RFF | ||||
| 1 | Reference | None (Original) | |||
| Autoresearch | Adaptive normalized RFF | ||||
| Round | Method | Main Cumulative Edit | Accuracy | NLL | |
|---|---|---|---|---|---|
| Clean | Contaminated | ||||
| 0 | Reference | None (Original) | |||
| Autoresearch | Disagreement gate | ||||
| ConflictGuide | Residual–neighbor mixer | ||||
| 1 | Reference | None (Original) | |||
| Autoresearch | Disagreement self-gate | ||||
| Round | Method | Main Cumulative Edit | NRMSE | MSE | |
|---|---|---|---|---|---|
| 0 | Reference | None (Original) | |||
| AutoResearch | Gated leak | ||||
| ConflictGuide | State–drive coupling | ||||
| 1 | Reference | None (Original) | |||
| AutoResearch | State–dependent leak | ||||
| ConflictGuide | Dual state coupling |
| Round | Method | Main Cumulative Edit | Actual bpp | MS-SSIM | |
|---|---|---|---|---|---|
| 0 | Reference | None (Original) | |||
| Autoresearch | Cross-branch gating | ||||
| ConflictGuide | Intra-encoder skip | ||||
| 1 | Reference | None (Original) | |||
| Autoresearch | Cross-scale gating | ||||
| ConflictGuide | Channel recalibration |
| Configuration | Guidance | Retention | SpecB-FNO | SNGP | ||
|---|---|---|---|---|---|---|
| NRMSE | ND-NMSE | NLL | OOD AUPR | |||
| Baseline | ✗ | ✗ | ||||
| Guidance only | ✓ | ✗ | ||||
| ConflictGuide | ✓ | ✓ | ||||
| Configuration | GCNII | SNGP | ||
|---|---|---|---|---|
| Clean NLL | Contaminated NLL | NLL | OOD AUPR | |
| Multi-metric feedback | ||||
| Trade-off prompt only | ||||
| ConflictGuide | ||||
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Mechanism label | Included code changes |
|---|---|
| Frequency-specific spectral | Fourier-mode weighting, selection, filtering, or spectral operators. |
| Spatial/local pathway | Spatial or pointwise feature extraction and local helper paths alongside the spectral path. |
| Normalization/activation/channel | Normalization, nonlinear activation, channel width, or channel projection. |
| Stage fusion/residual scale | Fusion between stages or branches and scaling of residual corrections. |
| Optimization/regularization | Optimizer, schedule, loss regularization, or training-time stabilization. |
| Configuration | Test NRMSE | ND-NMSE |
|---|---|---|
| Reference | ||
| Scalar-only | ||
| Probe from outset |
| Configuration | Test NRMSE | ND-NMSE |
|---|---|---|
| Reference | ||
| Scalar-only | ||
| ConflictGuide |
| Stage | Operation | Main evidence | Output |
|---|---|---|---|
| 0 | Model context | Architecture, task, data, objective, , editable scope, and runtime budget | Normalized Model Context Card |
| 1 | Taxonomy screening | Target-model code and configuration, interpreted through the fixed Root–Axis–Mechanism taxonomy | Taxonomy selection or explicit abstention |
| 2 | Conflict instantiation | Selected mechanism, task relevance, shared component or resource, and bidirectional interference | Model-specific competing behaviors and an observability analysis of |
| 3 | Probe design | Existing outputs, deterministic diagnostics, and lightweight read-only observables | Fixed Conflict Probe Card and implementation specification |
| 4 | Probe qualification | Null replicates and pre-specified parent–candidate replay pairs | Qualified, rejected, or inconclusive Probes |
| Condition | Skill response |
|---|---|
| Insufficient model, task, data, objective, or scalar-metric context | Abstain. The taxonomy selection cannot be grounded in target-specific evidence. |
| No two independently desirable behaviors | Reject. The proposed pair does not constitute a task-relevant behavioral conflict. |
| No identifiable shared component, representation, parameter path, or budget | Reject. The proposed behaviors lack a defensible coupling mechanism. |
| Only one interference direction is plausible | Reject. The case describes a one-sided failure rather than competing behaviors. |
| Behavior-specific observables are unavailable or incomparable across candidates | Abstain or reject. The conflict cannot be operationalized reliably under the current evaluation setting. |
| A Probe duplicates the scalar task metric, leaks unavailable information, or depends on post-evolution tuning | Reject. The proposed measurement does not provide valid fixed feedback. |
| Field | Description |
|---|---|
| Target behavior | The declared competing behavior measured by the Probe. |
| Observable and computation | Source value, formula or algorithm, sampling, and aggregation. |
| Direction and normalization | Whether an increase or decrease is desirable and how values are made comparable across candidates. |
| Evaluation setting | Frozen dataset, split, references, perturbations, and randomness policy. |
| Additional requirements | Whether extra training, labels, hooks, checkpoints, or model outputs are required. |
| Relation to scalar task metric | What behavior-specific information the Probe adds beyond the scalar task metric. |
| Card component | Recorded information |
|---|---|
| Model context | Architecture, task, data, objective, scalar task metric, evolution boundary, budget, and evidence sources. |
| Taxonomy selection | Selection status, Root, axis, mechanism, causal links, optimization signatures, and abstention reason when applicable. |
| Behavioral conflict | Competing behaviors, desirable directions, shared coupling, interference directions, confidence, and rejection reason when applicable. |
| Feedback observability | Definition and direction of the scalar task metric, conflict visibility, and behavior-specific information it aggregates away. |
| Probes | Fixed Probe specifications, information added beyond the scalar task metric, computational costs, and failure modes. |
| Qualification | Null calibration, scalar-tied informativeness test, direction test, decision state, and rejection conditions. |
| Root | Conflict is against | Compute test | Primary discriminator | Typical causal role |
|---|---|---|---|---|
| R1 | Compute, memory, latency, parameter count, or bandwidth | Partially resolvable | The conflict substantially eases with greater capacity, resolution, or evaluation budget. | May induce downstream information loss or structural compromise. |
| R2 | Mismatch between an inductive bias or structural constraint and the signal | Not resolvable | The imposed prior is too weak, too strong, or incompatible with the signal structure; architectural change is required. | Often upstream of shortcut learning or failure under distribution shift. |
| R3 | Information loss imposed by compression or abstraction | Not resolvable | Required information has been discarded or incompatible information demands must share a limited representation. | Often downstream of resource limits and upstream of uncertainty or generalization failures. |
| R4 | The gap between the training distribution and deployment conditions | Not resolvable | Performance relies on non-causal or distribution-specific evidence that fails under domain, group, temporal, or intervention shifts. | Usually a downstream symptom; upstream causes should be identified. |
| R5 | New adaptation competing with retained knowledge or behavior | Partially resolvable | The conflict arises sequentially across training stages rather than simultaneously among objectives. | May follow distribution, label-process, or normative changes. |
| R6 | A fixed search budget divided between coverage and solution quality | Partially resolvable | The conflict concerns exploration of solutions, actions, or hypotheses, not coverage of demographic populations. | May expose resource limits or aggravate structural and normative violations. |
| ID | Mechanism | Core diagnostic definition |
|---|---|---|
| R1: Resource allocation vs fidelity | ||
| R1.M1 | Insufficient fidelity | Does the model systematically underfit the target because the allocated resolution, step count, or capacity is too low? |
| R1.M2 | Uniform allocation mismatch | Does a uniform allocation waste resources on easy cases while remaining insufficient for hard cases, long-tail cases, or complex samples? |
| R1.M3 | Insufficient expressive bandwidth | Is the current positional encoding, basis, feature map, latent grid, channel rank, codebook, recurrent width, or local representation insufficient to express the required rate of variation? - Primarily space and time . May be module when a specific branch’s rank is the limit, or modality when a specific modality’s encoder is the limit. - If the loss is informational rather than budgetary (i.e. more capacity would not help because the representation deliberately discards the factor), use R3 instead. |
| R1.M4 | Added freedom causing overfitting | After adding local parameters, high-frequency encoding, extra steps, extra depth, or extra experts, does the model become more likely to fit noise, label errors, spurious patterns, or training-set-specific details? |
| R1.M5 | Cross-partition consistency issue | Do outputs become inconsistent across the partitions created by the chosen axis? |
| ID | Artifact | Diagnostic definition |
|---|---|---|
| OPT.M1 | Gradient scale imbalance | Gradients from some objectives are persistently larger, so shared training primarily serves them although the substantive demands are compatible. |
| OPT.M2 | Convergence rhythm conflict | Objectives converge at different speeds, so later updates damage early-converged objectives or slow objectives drag down shared training. |
| OPT.M3 | Weighted-sum masking | Aggregation in the training objective hides degradation of individual objectives while the weighted sum appears normal. |
| Model | Evolution training and auxiliary data | Online keep-validation | Post-search retraining and formal evaluation |
|---|---|---|---|
| SpecB–FNO | Navier–Stokes official training pool: 200 training examples and 48 disjoint mask-construction examples; the remaining 704 examples are unused online. | 48 fixed examples, disjoint from training and mask construction. | Retrain: 800/100/100 train/mask/validation partition of the official training pool. Test: 200 official examples (indices 1000–1199). |
| SNGP | CIFAR-100 training set, stratified per class into 400 training and 20 temperature-calibration images. A further 50 images/class are reserved for search-inaccessible audit. | 30 fixed images/class, disjoint from training, calibration, and audit. | Retrain: all 50,000 CIFAR-100 training images. Test: CIFAR-100; balanced OOD comparisons with SVHN, and CIFAR-10. |
| GCNII | Chameleon public split 0: its 60% per-class training partition; five frozen, label-free disagreement graphs are constructed without test nodes. | The corresponding 20% validation partition; the 20% test partition is inaccessible online. | Retrain/test: the frozen source is trained independently on each of ten public 60/20/20 splits of Cora, Citeseer, Pubmed, Chameleon, Cornell, Texas, and Wisconsin. |
| ESN | One frozen NARMA-30 realization: 200 washout and 5,000 readout-fit samples, evaluated with reservoir seeds 1103, 2207, and 3301. | The subsequent 2,500 samples from the same realization. | Refit/test: independent NARMA-30 and Mackey–Glass sequences crossed with ten independent reservoir seeds; a fresh ridge readout is fitted for every pair. |
| TCM–Lite | An image-disjoint DIV2K+Flickr2K patch bank: 32,768 training patches and four separate 2,048-patch panels for mechanism development, hidden audit, qualification discovery, and qualification confirmation. | 2,048 fixed RGB patches from a fifth source-image-disjoint panel. | Retrain: all 3,450 DIV2K+Flickr2K source images. Test: full-resolution Kodak images using actual arithmetic coding. |
| Dataset | Layers | Hidden | Dropout | Weight decay | ||
|---|---|---|---|---|---|---|
| Chameleon | 8 | 64 | 0.5 | 0.2 | 1.5 | |
| Cornell | 16 | 64 | 0.5 | 0.5 | 1.0 | |
| Texas | 32 | 64 | 0.5 | 0.5 | 1.5 | |
| Wisconsin | 16 | 64 | 0.5 | 0.5 | 1.0 | |
| Cora | 64 | 64 | 0.5 | 0.2 | 0.5 | |
| Citeseer | 64 | 64 | 0.5 | 0.5 | 0.5 |
| Model | Task-gain condition | Required Probe improvement | Guard condition |
|---|---|---|---|
| SpecB–FNO | Direct: . Probe route: . | and | |
| SNGP | Direct: . Probe route: . | and | |
| GCNII | Direct: . Probe route: , where . | No additional model-specific guard; the non-negative task-gain requirement prevents clean-NLL regression. | |
| ESN | Direct: . Probe route: , where . | Either or | and , frozen from the 95th percentiles of paired no-edit variation. |
| TCM–Lite | Direct: . Probe route: . | Either or | and |
| Model | Search seeds | Stage I | Stage-II AutoResearch | Stage-II ConflictGuide |
|---|---|---|---|---|
| SpecB–FNO | 0, 1, 2 | 100 | 100 | 100 |
| SNGP | 0, 1, 2 | 100 | 100 | 100 |
| GCNII | 0, 1, 2 | 100 | 100 | 100 |
| ESN | 0, 1, 2 | 20 | 50 | 50 |
| TCM–Lite | 0, 1, 2 | 100 | 100 | 100 |
| Aspect | Upstream AutoResearch | Our controlled harness | Reason for the modification |
|---|---|---|---|
| Loop ownership | The coding agent edits the program, invokes training, interprets the output, and performs Git-based keep or reset operations. | The agent only proposes source edits. A protected external controller owns validation, training, evaluation, retention, rollback, and accounting. | Prevents evaluator modification and information leakage, and applies the same frozen decision rule to every candidate. |
| Session and history | The agent follows an iterative experiment loop and records a compact results.tsv . | Every proposal uses a fresh session supplied with a deterministic, source-hash-bound structured history view; a separate complete ledger archives all attempts. | Removes unrecorded conversational state and permits exact reconstruction of the information available at each iteration. |
| Task substrate | The public implementation optimizes one language-model training program under a fixed wall-clock budget and a scalar validation metric. | We instantiate the loop for five model families with frozen model-specific training budgets and scalar task metrics . | A single training duration and metric are not meaningful across PDE, classification, graph, reservoir, and compression tasks. |
| Feedback | Each experiment is summarized primarily by the scalar task metric. | The scalar-only arm receives only , whereas ConflictGuide additionally receives registered competing-behavior Probes and Probe-aware decision information. | This is the experimental intervention whose effect we study. |
| Search design | The public loop follows a single sequential trajectory. | Each search seed runs its own Stage-I scalar trajectory, after which matched AutoResearch and ConflictGuide continuations begin from the same branch point. | Creates a paired comparison that controls initialization and all pre-branch search outcomes. |
| Editable boundary | train.py is editable while supporting infrastructure is fixed. | We retain the one-file editable boundary, but impose model-specific API, shape, determinism, runtime, memory, and parameter constraints. | Preserves evaluator semantics and prevents edits from changing the experimental task or resource budget. |
| Model | Proposal generation | Candidate evaluation |
|---|---|---|
| SpecB–FNO | 1,800 s | 1,500 s |
| SNGP | 1,800 s | 1,024 s |
| GCNII | 1,800 s | 168 s |
| ESN | 1,800 s | 1,800 s |
| TCM–Lite | 660 s | 300 s |
| Model | Scalar-only evaluator | ConflictGuide evaluator | Incremental-cost accounting |
|---|---|---|---|
| SpecB–FNO | Task metric and hidden spectral Probes | Identical computation, with Probes exposed to the agent | The realized difference is s. Across eight matched calibration evaluations, the spectral Probe pass required a median of s, corresponding to of the complete candidate-evaluation time. |
| SNGP | Training, GP-state reconstruction, temperature fitting, and task metrics | The same computation plus representation collection and geometry Probes | Across three matched no-edit branch anchors, the median paired increase was s, or of the scalar-only evaluation time. This is a descriptive wall-clock estimate because the paired jobs were not executed simultaneously. |
| GCNII | Median candidate-evaluation time of s | Probe-normalized median candidate-evaluation time of s | The estimated incremental cost is s per candidate, corresponding to a increase over the scalar-only evaluator. |
| ESN | Median total candidate-evaluation time of s, including hidden Jacobian Probes | The same median total time of s, with Probe measurements exposed to the agent | The three reservoir-seed Jacobian Probe passes together required approximately s, corresponding to of the complete candidate-evaluation time. |
| TCM–Lite | Median total candidate-evaluation time of s, including hidden rate and detail measurements | The same median total time of s, with the measurements exposed to the agent | The separately timed multiscale-detail pass required approximately s, or less than of the total. The rate measurement adds no separate pass because it is already required by the rate–distortion task metric. |
| Model | Multi-metric feedback | Trade-off-only prompting | ConflictGuide |
|---|---|---|---|
| SNGP | Fifteen-bin calibration error and development OOD Dempster–Shafer AUPR. | A generic qualitative reminder, without additional metrics or model-specific conflict information. | Low-quantile centroid margin , input–feature distance distortion , their desirable directions, and Probe-aware retention feedback. |
| GCNII | Mean NLL and accuracy under five frozen, randomly rewired degree-matched graphs. | A generic qualitative reminder, without additional metrics or model-specific conflict information. | Disagreement-contamination sensitivity , its desirable direction, and Probe-aware retention feedback. |
| Setting | Code generation | Post-execution bug review | Search structure |
|---|---|---|---|
| Claude Code | Claude Opus 4.6 | Not used | Sequential retained-parent search |
| Kimi Code | Kimi K2.7 Code | Not used | Sequential retained-parent search |
| AIDE (OpenAI) | o4-mini | GPT-4.1-mini | Solution tree with improve and debug operations |
| AIDE (GLM) | GLM-5.3 (low reasoning effort) | GLM-5.3-Flash | Solution tree with improve and debug operations |