Sustaining industrial recommendation research requires using the results of one experiment to decide what to investigate next. We present AgentX-Model, the next generation of AgentX's model research framework, which connects proposal development and model experimentation within sandboxes defined by business inputs and prediction tasks. AgentX-Model adopts a dual-agent architecture comprising a Research Agent and a Model Agent. The Research Agent develops independently reviewed proposals from papers and experimental findings, while the Model Agent conducts multi-round investigations and returns code, measurements, and unresolved questions. Using the returned results, the Research Agent selects a starting implementation and formulates the next research question, allowing subsequent experiments to build on earlier findings. We organize this continuing research around four actions: Reproduce, Follow-up, Composition, and Diagnose. The first three actions drive routine research, while Diagnose acquires the evidence needed to choose a repair, including for issues raised by business feedback and online evaluation, such as prediction bias measured by PCOC. Across the production evaluation, 560 of 636 completed model-changing experiments recorded AUC above their business baselines. As research continued, some experiments recorded AUC above every comparable ancestor in their lineages. The five latest online A/B evaluations across different business settings reported gains including 10-15% in acquisition efficiency, 15-20% in target-segment advertising spend, and 0.3-0.8% in watch time; the watch-time model used approximately 10% fewer FLOPs and parameters. A dependency-aware historical-replay benchmark further evaluates research allocation, with initial results showing no consistent efficiency gain from more complex scheduling when agents already analyze and select concrete candidates.
Figures & tables
Information carried forward
Decision in the next experiment
Round-specific code and its baseline
Choose the implementation to resume from
Reference model, evaluation conditions, and measurements
Choose comparable reference results
Diagnostic observations and unresolved hypotheses
Design a test that distinguishes explanations
Complete effective changes from each source
Resolve overlap and design the joint modification
Table 1: Information retained from an experiment and its use in subsequent research.
Dataset
Size
Use
Production Records
Proposal production (10-day sample)
886 attempts; 262 deliveries
Throughput
Offline results ( ∼ 25 days)
636 model-changing experiments
Baseline gains
Continuation archive
1,218 experiments; 923 valid results
Research continuity
Agent Evaluation
Execution subset
200 delivered proposals
Outputs and timing
Table 2: Production records and agent evaluation datasets.
Measure
Observation
Proposal-production attempts
886
Proposals delivered to the Model Agent
262
Delivered proposals assessed for execution
200
Proposals with Model Agent outputs
198/200
Proposals with verifiable experimental measurements
177/200
Median proposal-production time
33.6 min
Table 3: Proposal production in the sampled 10-day window. Execution outcomes and production times are reported for a subset of 200 delivered proposals.
Evaluation setting
Baseline AUC
Best AUC
Δ AUC
AUC above business baseline / Valid (%)
Scenario A
0.773683
0.797352
+0.023669
500/527 (94.9%)
Scenario B
0.815403
0.823145
+0.007742
21/29 (72.4%)
Scenario C
0.816849
0.820246
+0.003397
20/33 (60.6%)
Scenario D
0.809656
0.811432
+0.001776
19/47 (40.4%)
Table 4: Offline model gains over approximately 25 days. Valid results are completed experiments whose published AUC is verified against round-level measurements. Best AUC is the highest measured AUC among these experiments; Δ AUC is its gain over the corresponding baseline. Results counted as above business baseline exceed it by more than 10−6 .
Evaluation setting
Experiment category
Valid results
AUC above business baseline
Share (%)
Scenario A
Reproduce
51
41
80.4%
Composition
308
302
98.1%
Follow-up
168
157
93.5%
Scenario B
Reproduce
20
15
75.0%
Follow-up
9
6
66.7%
Scenario C
Reproduce
24
15
62.5%
Table 5: The same 636 completed model-changing experiments, grouped by category. Definitions follow Table 4 ; percentages use valid results within each category as the denominator.
LR
Model change
Relative improvement
1
Adaptive attention temperature
Acquisition efficiency: 10–15%
2
Attention-module combination with normalization
Target-segment advertising spend: 15–20%
3
Multi-task objective pruning
Watch time across two clients: 0.3–0.8%; FLOPs and parameter count: approximately 10% lower
4
Hierarchical aggregation with calibration
Overall daily active users: 0.5–1%; target-page follow actions: 3–4%
5
Cross-setting mechanism transfer
Target-segment advertising spend: 5–10%
Table 6: Online A/B results from the five latest Launch Reviews. Relative gains use coarse ranges.
Research action
AUC above strongest direct parent
AUC above all ancestors
Follow-up
46/91 (50.5%)
7/77 (9.1%)
Composition
6/139 (4.3%)
5/120 (4.2%)
Table 7: Scenario A continuation results. Each cell gives experiments with a higher best-round AUC over the number with the required comparable evidence. An increase must exceed 10−6 to avoid rounding ties. Complete ancestry requires more evidence, so the column denominators differ.
Implementation
AUC
PCOC before
PCOC after
Error reduction
Business baseline
0.773683
–
–
–
SMES best
0.778940
–
–
–
Diagnose
0.775103
0.953836
–
–
Follow-up R1
0.500000
–
–
–
Follow-up R2
0.775632
0.951715
0.962735
22.8%
Follow-up R3
0.775318
0.948190
0.959965
22.7%
Table 8: The Scenario A calibration path. Before/after values compare outputs of the same run, retaining the inherited calibration and then adding the new factor. Error reduction is relative to ∣PCOC−1∣ before correction. All Follow-up rounds use the same code patch; R1 collapsed during training.
Proposal source
Valid results
Δ AUC >0
Paper-derived
7
0 (0%)
Knowledge Transfer
15
5 (33.3%)
Table 9: Historical Scenario D results by proposal source. Δ AUC is the published AUC minus the common baseline AUC of 0.809656.
Evaluation setting
Baseline
Random
Greedy
Fixed
Bandit
Joint
Scenario A
0.773683
0.782285
0.779145
0.780657
0.782500
0.780012
Scenario B
0.815403
0.823145
0.823141
0.823141
0.823141
0.823141
Scenario C
0.816849
0.819780
0.819780
0.819780
0.819780
0.819780
Scenario E
0.815403
0.821067
0.820931
0.821011
0.821614
0.821011
Table 10: Mean best AUC after 20 selections on the four graphs for Scenarios A, B, C, and E. Each policy completes three runs per graph; baselines and evaluation contexts are row-specific. Fixed and Bandit allocate the type before agent candidate selection; Joint selects both.
Policy
Mean
Median
Range
Uniform Random
18.00
18
16–20
Parent-result Greedy
13.33
15
8–17
Fixed routing + Agent
14.67
16
12–16
Bandit routing + Agent
16.67
17
15–18
Joint selection
10.67
9
7–16
Table 11: Selections to first reach the highest recorded AUC (0.819780) on the Scenario C graph. Each policy has three repeats; all 15 runs reach this result within the 20-experiment budget.
Above-baseline selections
Summaries
Evaluation setting
Raw
+ Summary
Paired Δ Best@20
Initial
Scenario A
37/40
37/40
+0.000998,−0.000950
Initial
Scenario E
51/60
55/60
0,0,−0.001809
Comparative
Scenario A
54/60
52/60
−0.001052,+0.000329,+0.017284
Comparative
Scenario E
49/60
51/60
0,0,−0.001602
Table 12: Summary comparisons in historical replay. Above-baseline selections count experiments whose best valid AUC exceeds their baseline, summed across repeats. Paired Δ Best@20 is with-summary minus raw-evidence best AUC. The initial Scenario A row includes two complete pairs; each other row includes three pairs.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Metric
Before
After
First-draft structural pass rate
37.5%
95.0%
Semantic fidelity
72.5%
95.0%
Plans with unacknowledged mechanism omissions
4
0
Appendix
Table 13: First-draft implementation plans before and after instruction revision in a regression evaluation on the same 40 paper-reproduction inputs.