From Weak Task Specifications to Scientific Extraction Agents: Optimizing Task Construction
Organizations: School of Computer Science and Technology, University of Science and Technology of China · State Key Laboratory of Cognitive Intelligence · University of Science and Technology of China · CCCC Second Highway Consultants Co., Ltd.
Abstract
Most methods that optimize LLM prompts and agent workflows assume that task-specific output schemas, extraction instructions, and evaluation criteria are predefined. For scientific extraction agents, however, a short task goal may not fully determine these components, while specifying them manually is costly. We study the upstream problem of constructing the task-specific configuration from a weak specification containing only a short goal and unannotated reference documents. Rather than treating automatic construction as a fixed preprocessing step, our framework constructs a task-specific schema, extraction instructions, and base training rubrics, then keeps schema construction and extraction instructions editable during optimization. Failure-focused updates concentrate textual-gradient feedback on lower-scoring documents, while training-time evaluation criteria adapt to recurring failures. On a heterogeneous-catalysis literature corpus, automatic construction remains improvable, and optimizing both schema construction and extraction instructions performs best across all four judge-rubric settings, with ablations and blinded human evaluation supporting the proposed formulation.
Figures & tables
| Method | T-R1 | T-R2 | 5.5-R1 | 5.5-R2 |
|---|---|---|---|---|
| Auto-init | 79.63 | 77.42 | 79.02 | 77.40 |
| AgentCAT-short [ 17 ] | 81.55 | 78.93 | 79.05 | 77.20 |
| AgentCAT [ 17 ] | 81.04 | 80.49 | 79.52 | 78.17 |
| Schema-TG (best) | 81.68 | 81.63 | 80.95 | 79.02 |
| Extract-TG (best) | 83.63 | 82.90 | 81.30 | 80.75 |
| Joint-TG (best/best) | 84.50 | 84.38 | 83.67 | 81.05 |
| Rubric | Feedback + checkpoint | R1 | R2 |
|---|---|---|---|
| Adaptive | All + best/best | 81.91 | 80.29 |
| Fixed | Bottom + best/best | 82.15 | 81.98 |
| Adaptive | Bottom + final/final | 82.43 | 81.10 |
| Adaptive | Bottom + best/best | 84.50 | 84.38 |