Document-level Information Extraction (DocIE) aims to produce an output template with the entities, relations, and events of interest occurring in the given document. Standard practices include prompting decoder-only LLMs using greedy decoding to avoid output variability. Rather than treating this variability as a limitation, we show that sampling can produce substantially better solutions than greedy decoding, especially when using reasoning models. We thus propose ThinkTwice, a sampling and selection framework in which the LLM generates multiple candidate templates for a given document, and a selection module chooses the most suitable one. We introduce both an unsupervised method that exploits agreement across generated outputs, and a supervised selection method using reward models trained on labeled DocIE data. To address the scarcity of golden reasoning trajectories for DocIE, we propose a rejection-sampling-based method to generate silver training data that pairs output templates with reasoning traces. Our experiments show the validity of unsupervised and supervised ThinkTwice, consistently outperforming greedy baselines and the supervised state-of-the-art.
Figures & tables
Figure 1: Sampling results on MUC-4, maximum reports the results of oracle selection among generated samples.
Figure 2: ThinkTwice architecture, with the inference process at the bottom. The supervised option includes two steps: \raisebox{-.9pt} {1}⃝ The iterative procedure to generate the silver dataset with trajectories and to fine-tune the reasoning model; \raisebox{-.9pt} {2}⃝ Training the selector with silver preference data that include trajectories.
Figure 3: Zero-shot greedy decoding results for the reasoning and non-reasoning baselines of two LLMs on MUC (English), MultiMUC and BETTER.
Method
Selector
MultiMUC
BETTER
GPT 5.5
✗
19.84
27.95
Greedy Llama R1
✗
12.67
14.78
ThinkTwice Llama R1
Majority
13.83 ± 0.66
10.51 ± 2.30
F1 Voting
14.41 ± 0.33
15.29 ± 0.88
(oracle)
31.77
34.08
Greedy Qwen3
✗
14.65
16.12
Table 1: Zero-shot results on MultiMUC and BETTER. Greedy results are compared to two unsupervised selectors. Underline and bold indicate the best result among greedy and the selectors per model and across all models, respectively.
Method
Selector
MUC-4
TempGen BART large
✗
28.30
GTT BERT base
✗
32.30
IterX T5-enc large
✗
35.20
Greedy Llama R1
✗
28.52
ThinkTwice Llama R1
Majority
28.41 ± 1.49
F1 Voting
36.56 ± 0.22
Table 2: Supervised results for ThinkTwice , greedy baseline and the state-of-the-art on MUC-4.
Method
Selector
English
Arabic
Farsi
Korean
Russian
Chinese
Average
GPT 5.5
✗
29.50
24.60
22.03
07.59
17.70
17.59
17.90
Gantt et al. (2024) Supervised
✗
35.20
21.46
20.66
23.91
23.77
21.93
22.35
ThinkTwice Zero-shot
F1 Voting
24.30 ± 0.57
17.66 ± 0.32
19.62 ± 0.44
07.22 ± 0.37
14.34 ± 0.16
16.20 ± 0.16
15.01 ± 0.31
Reward
33.46 ± 0.70
26.05 ± 0.24
26.80 ± 0.21
08.71 ± 0.53
20.66 ± 0.25
25.47 ± 0.74
21.54 ± 0.45
ThinkTwice English FT
F1 Voting
40.77 ± 0.53
22.66 ± 0.44
23.61 ± 0.77
08.73 ± 0.75
20.53 ± 1.18
22.15 ± 0.92
19.54 ± 0.85
Reward
42.51 ± 0.72
30.27 ± 0.49
30.16 ± 0.66
09.62 ± 0.14
29.95 ± 1.23
29.75 ± 0.75
25.95 ± 0.74
Table 3: Cross-lingual transfer performance of ThinkTwice . The Qwen3-based LLM and reward model are trained exclusively on English data (MUC-4) and evaluated on multiple languages. Results are compared against state-of-the-art models trained on the target language ( Gantt et al., 2024 ) , as well as GPT 5.5 under a zero-shot setting. The averages are calculated without using the English scores.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Maximum, mean, and 95% confidence interval results in train split across different numbers of rejection sampling fine-tuning iterations.
Reasoning
Non-Reasoning
Hyperparameter
Llama
Qwen
Llama
Qwen
Temperature
0.7
0.6
0.6
0.7
Top-p
1
0.95
1
0.8
Top-k
-1
20
-1
20
Min-p
0
0
0
0
Appendix
Table 4: Inference hyperparameters.
Hyperparameter
Reasoning
Reward
Batch Size
128
512
Learning Rate
5e−5
2e−5
Epochs
5
4
Weight Decay
5e−5
5e−5
Appendix
Table 5: Fine-tuning hyperparameters.
Model
F1 similarity
GPT 5.5
71.76 ± 12.65
Llama R1
23.20 ± 10.03
Qwen3
26.72 ± 11.93
Appendix
Table 6: Mean pairwise F1 similarity between N=64 sampled templates on the BETTER dataset for each model. Higher values indicate lower output diversity across samples.
MultiMUC
BETTER
Model
English
Arabic
Farsi
Korean
Russian
Chinese
English
Llama R1 Zero-shot
364 ± 146
368 ± 127
379 ± 125
364 ± 132
560 ± 221
485 ± 140
447 ± 138
Llama R1 Fine-tuned
466 ± 505
456 ± 554
452 ± 531
433 ± 501
454 ± 504
433 ± 487
–
Qwen3 Zero-shot
792 ± 578
749 ± 519
750 ± 505
826 ± 563
809 ± 561
709 ± 466
701 ± 364
Qwen3 Fine-tuned
995 ± 848
978 ± 863
1005 ± 868
1009 ± 867
1013 ± 855
994 ± 873
–
Appendix
Table 7: Mean and standard deviation of reasoning budget (in tokens) for each model and dataset, computed over all instances and traces.
Method
Selector
BETTER
GPT 5.5
✗
27.95
Greedy Llama 3.3
✗
3.20
ThinkTwice Llama 3.3
Majority
1.72 ± 0.42
F1
1.60 ± 0.22
(oracle)
5.71
Greedy Llama R1
✗
14.78
Appendix
Table 8: Zero-shot scenario results in the BETTER Granular dataset. We compare our approach using unsupervised selectors.
Method
Selector
English
Arabic
Farsi
Korean
Russian
Chinese
Average
Reference
✗
35.2
21.46
20.66
23.91
23.77
21.93
24.49
Greedy Llama R1
✗
28.52
3.57
1.30
3.07
0.51
4.21
6.86
ThinkTwice Llama R1
Majority
28.41 ± 1.49
2.69 ± 0.50
0.53 ± 0.57
2.29 ± 0.15
0.06 ± 0.11
4.72 ± 0.84
6.45 ± 0.77
F1 Voting
36.56 ± 0.22
3.79 ± 0.30
1.43 ± 0.39
3.88 ± 0.31
0.24 ± 0.10
5.34 ± 0.46
8.54 ± 0.32
Reward
41.11 ± 0.48
5.65 ± 0.15
2.43 ± 0.34
5.08 ± 0.21
1.36 ± 0.21
10.98 ± 0.27
11.10 ± 0.30
(oracle)
56.60
12.79
4.66
11.82
2.96
19.44
18.04
Appendix
Table 9: Supervised and cross-lingual results on MUC-4 (English) and MultiMUC comparing greedy decoding against different selection strategies in reasoning fine-tuned models on English data. We also report the results of the state-of-the-art method, IterX T5-enc large ( Chen et al., 2023 ) fine-tuned using both English and target-language data for training (Reference) reported in the work of ( Gantt et al., 2024 ) .
Method
Selector
English
Arabic
Farsi
Korean
Russian
Chinese
Average
GPT 5.5
✗
29.50
24.60
22.03
7.59
17.70
17.59
19.84
Greedy Llama 3.3
✗
18.27
14.52
14.57
5.55
10.84
11.02
12.46
ThinkTwice Llama 3.3
Majority
18.69 ± 0.19
14.62 ± 0.63
15.09 ± 0.18
5.69 ± 0.12
11.27 ± 0.43
11.68 ± 0.27
12.84 ± 0.35
F1 Voting
18.60 ± 0.14
14.80 ± 0.11
15.02 ± 0.06
5.79 ± 0.17
10.99 ± 0.18
11.62 ± 0.18
12.80 ± 0.15
Reward
22.84 ± 0.36
18.83 ± 0.13
18.60 ± 0.62
7.26 ± 0.20
15.04 ± 0.28
17.01 ± 0.03
16.60 ± 0.33
(oracle)
26.76
25.71
26.07
10.87
19.37
20.51
21.55
Appendix
Table 10: Zero-shot scenario results. We compare our approach using unsupervised selectors (F1 and Majority Voting) and supervised selectors (Reward trained in English data) in MUC-4 (English) and MultiMUC.
Document-level relation extraction (DocRE) aims to extract relations among multiple entities across extended contexts while maintaining consistency across predicted triples. Although large language models (LLMs) show remarkable reasoning capabilities in information extraction, their predictions are typically generated independently for each candidate triple and may violate fundamental relational constraints such as transitivity, symmetry, and functional uniqueness, leading to contradictory and unreliable outputs. We propose CONSISTRE, a unified consistency-aware framework for DocRE that addresses this limitation through two complementary tracks. The first operates at inference time for black-box LLMs, combining constraint-aware prompting, constraint-based verification, and iterative self-reflection to refine predictions without task-specific fine-tuning. The second injects consistency knowledge into smaller open-source models via a knowledge distillation and reinforcement learning pipeline: reasoning traces from a powerful teacher are distilled into a student via supervised fine-tuning, followed by GRPO alignment using a composite reward that jointly optimizes extraction performance and relational consistency. Together, the two tracks cover both API-accessible and locally deployable scenarios under a unified consistency formulation. Experiments on DocRED show that both tracks outperform their baselines, with the inference-time track achieving competitive F1 using off-the-shelf black-box LLMs and the training-time track substantially narrowing the gap between 7--8B open-source models and state-of-the-art proprietary LLMs at a fraction of their inference cost. Ablation studies confirm that explicit consistency modeling mitigates relational contradictions and enhances the reliability of LLM-based DocRE across both deployment paradigms.
Mingxuan Sun
Department of Computer Science, Universit´e de Sherbrooke, Sherbrooke, Canada
Information Extraction aims to distill structured, decision-relevant information from unstructured text, serving as a foundation for downstream understanding and reasoning. However, it is traditionally treated merely as a terminal objective: once extracted, the resulting structure is often consumed in isolation rather than maintained and reused during multi-step inference. Moving beyond this, we propose \textit{IE-as-Cache}, a framework that repurposes IE as a cognitive cache to enhance agentic reasoning. Drawing inspiration from hierarchical computer memory, our approach combines query-driven extraction with cache-aware reasoning to dynamically maintain compact intermediate information and filter noise. Experiments on challenging benchmarks across diverse LLMs demonstrate significant improvements in reasoning accuracy, indicating that IE can be effectively repurposed as a reusable cognitive resource and offering a promising direction for future research on downstream uses of IE.
Hang Lv, Sheng Liang, Hongchao Gu +5
University of Science and Technology of China · Huawei Technologies Co., Ltd.
Modern reasoning language models generate dense, sequential chain-of-thought traces implicitly assuming that every token contributes and that steps must be consumed in order. We challenge both assumptions through a systematic intervention pipeline--removal, masking, shuffling, and noise injection--applied to model-generated reasoning chains across three models and three benchmarks. Our findings are counterintuitive on three dimensions. Order: Does the sequential order of a reasoning chain matter for answer extraction? No--line-level shuffling reduces accuracy by less than 0.5 pp; word-level shuffling retains 62%-89% accuracy; only token-level shuffling collapses to near zero. Pretrained-only and instruction-tuned variants exhibit near-identical tolerance (78.67% vs. 78.00% under line shuffling), indicating order-independence originates from pretraining rather than reasoning-specific fine-tuning. Dense: Is all the information in a reasoning chain important for answer extraction? No--masking numeric digits collapses accuracy to exactly 0%, while masking alphabetic prose improves accuracy by 4.7 pp. Robustness: Is a reasoning chain that is both order-shuffling and non-dense still robust? Yes--the most aggressively reduced representation (all natural language removed, lines arbitrarily shuffled) still achieves 83% accuracy, and injecting false answers at 3x true-answer frequency leaves accuracy unchanged (83.3%->83.3%), falsifying a frequency-based extraction account. These results establish that answer extraction operates on a sparse, order-insensitive, and structurally robust informational substrate, opening paths toward parallelized and token-efficient reasoning generation.
Yi-Chang Chen, Feng-Ting Liao, Da-shan Shiu +1
MediaTek Research · Artificial Intelligence Center of Research Excellence, National Taiwan University