Adapt Semantics, Not Structure: Few-Instance Schema Calibration for Scientific PDF Extraction
Authors: Zixiao Dong, Wei Yang, Zihao Liu, Chenshu Li, Longzhang Liu, Tao Tan, Hong Xie
Organizations: School of Computer Science and Technology, University of Science and Technology of China · State Key Laboratory of Cognitive Intelligence · University of Science and Technology of China · CCCC Second Highway Consultants Co., Ltd.
A well-designed extraction schema is not necessarily ready for reliable LLM execution. When only limited verified extractions are available, manually tuning hundreds of field definitions through trial and error is costly. We frame this problem as few-instance schema calibration: adapting the operational semantics of an existing schema from a few annotated documents while preserving its structural contract. We introduce CPSE, a contract-preserving semantic extraction framework that jointly calibrates extraction prompts and field-level semantic descriptions from a few gold annotations. CPSE decomposes the schema into an invariant structural contract and mutable field semantics, and further separates identity discovery from record completion using manifest-conditioned resolution. On expert-annotated polymer-science documents, CPSE improves extraction by 9.93 points over an execution-matched baseline, with consistent gains under an independent judge and in a blinded expert audit. These results show that CPSE enables low-resource schema execution while preserving the output structure required downstream.
Figures & tables
Figure 1: Overview of CPSE for few-instance schema calibration. Schema induction produces S0=(C,d0) , after which the structural contract C remains fixed while textual feedback calibrates prompts and field descriptions. At inference, an ordered material manifest is resolved in bounded batches before deterministic merging and validation.
Method
Sol
Terra
Single-pass baseline
76.45
75.84
Prompt adaptation
83.26
82.44
Description adaptation
82.09
84.28
Prompt + description
86.74
86.64
Manifest-conditioned baseline
81.02
83.88
Full CPSE
90.95
91.09
Table 1: Internal-ablation scores on 17 held-out PDFs.
Method
Sol
Terra
GEPA prompt only [ 1 ]
82.12
83.03
GEPA [ 1 ] + manifest
85.52
85.50
MIPROv2 full [ 10 ]
79.44
77.86
MIPROv2 instruction only [ 10 ]
74.21
72.44
OPRO [ 16 ]
71.35
71.24
Fixed 3-shot ICL
76.50
74.37
Table 2: Mean Sol and Terra scores for external optimizers and non-optimized references.
Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and measured cost together. The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating their challenge scenarios. The scalable schema and ground-truth curation pipeline combines independent-system agreement for real documents, known values for synthetic lists, and human verification for forms. We report order-insensitive value F1 for value accuracy, plus two grounding metrics for source traceability: word- and page-level F1. Commercial VLMs perform well on short documents but often truncate record lists on long ones, while coding agents retain higher accuracy at much higher cost. LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost. Dataset and evaluation code are available on \href{https://huggingface.co/datasets/llamaindex/ExtractBench}{HuggingFace} and \href{https://github.com/run-llama/ExtractBench}{GitHub}.
Extracting structured data from unstructured text using large language models (LLMs) becomes challenging when target schemas are large and complex. In such cases, including the full schema in the prompt increases cost and latency, risks lost-in-the-middle performance degradation, and can exceed context length limits. We propose SchemaRAG, a retrieval-augmented generation (RAG) framework that dynamically prunes the output schema space for schema-conditioned information extraction tasks by leveraging schema metadata and few-shot examples when available. We evaluate SchemaRAG on real-world healthcare and e-commerce datasets. Our results show that SchemaRAG can achieve up to an 8.8% increase in micro-F1, a 47% reduction in latency, and a 48% reduction in token costs, demonstrating its practicality for large-schema extraction.
Most methods that optimize LLM prompts and agent workflows assume that task-specific output schemas, extraction instructions, and evaluation criteria are predefined. For scientific extraction agents, however, a short task goal may not fully determine these components, while specifying them manually is costly. We study the upstream problem of constructing the task-specific configuration from a weak specification containing only a short goal and unannotated reference documents. Rather than treating automatic construction as a fixed preprocessing step, our framework constructs a task-specific schema, extraction instructions, and base training rubrics, then keeps schema construction and extraction instructions editable during optimization. Failure-focused updates concentrate textual-gradient feedback on lower-scoring documents, while training-time evaluation criteria adapt to recurring failures. On a heterogeneous-catalysis literature corpus, automatic construction remains improvable, and optimizing both schema construction and extraction instructions performs best across all four judge-rubric settings, with ablations and blinded human evaluation supporting the proposed formulation.
Zixiao Dong, Wei Yang, Zihao Liu +4
School of Computer Science and Technology, University of Science and Technology of China · State Key Laboratory of Cognitive Intelligence · University of Science and Technology of China +1