Scientific research often begins by synthesizing ideas from a set of related papers to identify gaps and formulate new directions. However, training language models to perform this form of literature-grounded ideation remains challenging, as existing approaches based on prompting or feedback lack structured supervision for how papers should be synthesized. We introduce IdeaAnchor, a paradigm for training LLMs to perform research ideation using structured specifications as privileged signals. Each IdeaAnchor instance encodes how each input paper should be synthesized into a successful idea, including their functional roles, relationships, and target synthesis criteria. We build this paradigm by mining instances from published papers, capturing how real ideas emerge from prior literature. We then train models via demonstration, self-distillation, and reinforcement learning, and further enhance generation with retrieval at inference time. Experiments show consistent improvements in ideation quality. Our analysis reveals a functional decomposition: anchor-based training strengthens creative synthesis, retrieval enhances detail elaboration, and combining both yields the best performance.
Figures & tables
Figure 1: Overview of IdeaAnchor . Starting from a set of papers, we extract IdeaAnchor instances from published papers by identifying the functional roles of prior work, their cross-paper relationships, and the target synthesis criteria. These structured anchors make otherwise implicit synthesis logic explicit, providing privileged supervision for demonstration learning, self-distillation, and reinforcement learning. At inference time, the framework uses either related papers or a topic query as input. It can also incorporate retrieval augmentation, which complements the learned synthesis capability by supplying factual and methodological details for elaborating generated research ideas.
Figure 2: A sample IdeaAnchor instance derived from a paper on Byzantine-resilient distributed learning. The anchor encodes per-paper role assignments with learned insights ( top ), a cross-paper relationship analysis identifying the synthesized gap ( bottom-left ), and literature-grounded checkable criteria ( bottom-right ).
Figure 3: Main automated evaluation on the ICLR 2026 benchmark with 924 instances. Bars report CSR (%) for the base, self-distillation (SSD), SFT, and RL variants of Qwen3-8B and Qwen3.5-9B . Horizontal lines show proprietary model references.
Figure 4
Variant
RAG-Full
RAG-Sum.
SFT
73.6
74.2
SSD
68.7
70.3
RL
76.8
75.5
Table 1: Pairwise preference of RAG over abstract-only outputs. Each cell reports the win rate of the retrieval variant against the corresponding abstract-only output under a GPT-5.4 judge.
Ideation System
Novelty
Grounding
Specificity
Feasibility
Overall
Qwen3-8B
3.1
2.8
3.2
3.0
2.9
Qwen3-8B-RL w/ RAG-Full
3.5
3.9
3.8
3.2
3.7
Table 2: Human evaluation of topic-driven ideation. Scores are averaged over proposals from six open ML topics.
Table 7: Per-category criteria distribution across training domains.
Criteria count
ML
Natural Science
Evaluation
6 criteria
1,042
1,572
0
7 criteria
6,452
5,117
725
8 criteria
0
0
199
Appendix
Table 8: Criteria count distribution. Number of instances with each rubric count.
Strategy
Leakage ↓
Rubric Compl. ↑
ROUGE-L ↑
Rank
Naive (baseline)
100% (50/50)
0.836
0.147
5
Persona knowledge
0% (0/50)
1.000
0.193
2
Bottleneck PI
0% (0/50)
0.522
0.164
7
Quality characteristics
4% (2/50)
1.000
0.192
1
Critique–edit
4% † (2/50)
1.000
0.193
3
Uncertainty steering
4% † (2/50)
0.955
0.186
4
Appendix
Table 9: Comparison of anti-leak strategies for self-distillation. Leakage = fraction of outputs that reference anchor content; Rubric Compliance = LLM-as-judge pass rate over motivation/method/overall criteria; ROUGE-L = overlap with ground-truth reference. † False positives: the word “criterion” used in normal academic context.
Judge
Base
SFT
SSD
RL
GPT-5.4
10.9
21.7
16.3
24.6
GPT-5.4-mini
10.4
20.9
15.8
23.7
GPT-5.3-chat
8.7
18.2
13.9
20.4
Appendix
Table 10: CSR evaluated by different LLM judges. All results on the 924-instance ICLR 2026 benchmark (abstract-only setting).