PEAR: Progressive Evidence-Based AutoResearch for Industrial Search Systems
Organizations: Global E-Commerce Agentic Search Team
Abstract
AutoResearch improves systems through iterative experimentation: agents propose candidate modifications, evaluate them, and use the results to guide subsequent exploration. Applying this paradigm to industrial search presents two challenges. (1) Common AutoResearch approaches follow a keep-if-better rule, retaining the highest-scoring candidate for subsequent experiments. Under non-stationary traffic, transient gains may be mistaken for persistent improvements, impairing reliable accumulation of search knowledge. (2) Candidate modifications can be evaluated at multiple fidelity levels, from low-cost proxies to online validation, differing in cost, objective alignment, and statistical reliability. Existing methods rely on individual signals or task-specific procedures, lacking a unified basis for using evidence across levels to guide search. We introduce Progressive Evidence-Based AutoResearch (PEAR) with two complementary components. Evidence-driven AutoResearch maintains an independent, hypothesis-guided research state for each strategy task within a predefined objective and intervention scope. Each state evolves through a Plan-Execute-Evaluate-Update transition that links experimentation to context-aware evidence interpretation and hypothesis revision. Confidence-Gated Verifier Ladder organizes evaluation into four levels of increasing fidelity: Offline Replay, Shadow-Traffic Evaluation, Rapid Online Evaluation, and Decision-Grade Online Evaluation. A unified confidence-based gate promotes candidates only when evidence supports a statistically significant positive effect, enabling broad low-cost exploration while reserving costly online experiments for promoted candidates. In a real-world industrial search system, strategies optimized with PEAR significantly increased Main Order/DAU by 2.7336% and 3.2957% relative to their respective baselines in two A/B experiments.
Figures & tables
| Task phase | Experimental rounds | Groups | Unique configs. | Promoted from L2 | Promotion rate |
| A | 19 | 133 | 20 | 2 | 1.50% |
| B–I | 12 | 84 | 23 | 5 | 5.95% |
| B–II | 2 | 14 | 6 | 1 | 7.14% |
| B–III | 21 | 147 | 53 | 2 | 1.36% |
| C | 16 | 76 | 54 | 3 | 3.95% |
| Experimental round / research question | Controlled interventions | Evidence-driven update |
|---|---|---|
| 1. Which constraint axis accounts for the observed gain? | Apply one-axis changes to global and format-specific constraints; also test the coupled global change . | Only yields a significant positive order signal. The effects of increasing , increasing , and experimental variation remain entangled. |
| 2. Can be reproduced, and which variable explains its gain? | Replicate ; compare window-only , capacity-only , and boundary settings including . | produces a stronger signal. Replications of yield mixed directional effects and indicate potential degradation in the exit metric. The next experimental round prioritizes capacity and replication. |
| 3. Is the gain driven by the coupled change or by capacity alone? | Replicate and ; test capacity-only and wider-window boundary settings. | achieves the strongest observed result. The revised hypothesis attributes the gain primarily to capacity relaxation without requiring a wider window. |
| 4. Are the marginal effects of window and capacity stable? | Concentrate replications on , , and ; retain window-only as a counterfactual. | remains directionally favorable on order and exit signals but is not significant. The capacity-centered hypothesis is retained for further replication. |
| Phase | Experimental rounds | Groups | Unique configs. | Promoted from L2 | L4 outcome |
|---|---|---|---|---|---|
| I | 12 | 84 | 23 | 5 | No significant gain |
| II | 2 | 14 | 6 | 1 | Positive, not significant |
| III | 21 | 147 | 53 | 2 | Significant order gain |
| Click | Order | GMV | |
|---|---|---|---|
| Signed relative error |
| Candidate | L2: Shadow-Traffic Evaluation | L3: Rapid Online Evaluation | L4: Decision-Grade Online Evaluation |
|---|---|---|---|
| A | Order proxy | Observed order | Main Order/DAU |
| B | Order proxy | Observed order | Main Order/DAU |
| Task | SearchPV/DAU | ASN | GMV/DAU | Main Order/DAU | SKU Order/DAU | PayPV/PV | Main OPMS | SKU OPMS |
|---|---|---|---|---|---|---|---|---|
| A | ||||||||
| B |