Scientific ideation is increasingly mediated by large language models, but current ideation systems are usually trained and evaluated on immediately judgeable proxies such as novelty, clarity, and feasibility. This leaves open whether delayed signals of scientific uptake can be used as feedback for steering models toward research directions with higher expected \emph{impact}. We study this question using citation-normalized impact as a noisy but scalable proxy for scholarly uptake. We construct a large-scale dataset from over 100K computer science papers by extracting goal-conditioned idea descriptions and assigning each paper an ordinal, year-normalized citation label. We then train a goal-conditioned reward model to predict citation-impact labels from research goal and idea pairs, and use this reward to align an idea generator through supervised fine-tuning followed by reinforcement learning. To reduce circularity, we evaluate generated ideas with a held-out, reference-grounded protocol that compares model outputs against historical ideas under the same research goal and weights judgments by the reference idea's citation-impact label. Experiments show that our RL-tuned model consistently produces ideas with higher estimated impact than both the base model and supervised fine-tuning baselines. Our findings position scientific impact as a practical, outcome-grounded feedback signal for aligning LLMs in open-ended scientific discovery.
Figures & tables
Figure 1 : Dataset construction pipeline.
Dataset
Split
Zero/VL
Low
Med.
High
V.High
Total
Reward
Train
7,231
7,232
7,241
7,236
6,033
34,973
Reward
Val
1,552
1,550
1,549
1,552
1,292
7,495
Reward
Test
1,552
1,552
1,548
1,551
1,293
7,496
SFT/RL
Train
7,239
7,238
7,236
7,239
6,038
34,990
SFT/RL
Val
1,552
1,552
1,552
1,551
1,292
7,499
SFT/RL
Test
1,549
1,552
1,551
1,550
1,295
7,497
Table 1 : Paper distribution across ordinal classes for distinct dataset subsets and splits. Counts remain approximately uniform across train, validation, and test sets, indicating successful stratification.
Figure 2 : Impact Aligned Ideation Training formulation
Model
Acc
MAE
±1 Acc
Spear
QWK
Zero-shot
Qwen3-8B
20.7
1.410
58.7
0.065
0.004
GPT-4o
23.9
1.155
68.1
0.354
0.235
GPT-5
27.7
1.027
75.3
0.402
0.338
Fine-tuned
RM
48.0 (+20.3)
0.632 (-0.39)
90.6 (+15.3)
0.756 (+0.35)
0.754 (+0.41)
Table 2 : Trained Reward Model (RM: Qwen3-8B + LoRA + CORN) Performance comparison with zero-shot models on the ordinal classification task. Improvement of RM over GPT-5. Acc: Accuracy, Spear: Spearman Correlation
Model
V.Low
Low
Med.
High
V.High
GPT-4o
1.95
1.41
1.07
0.96
0.38
GPT-5
0.66
0.86
0.94
1.02
1.79
Qwen3-8B
3.01
2.01
1.00
0.00
1.00
RM
0.50
0.71
0.77
0.74
0.43
Table 3 : Class-wise Mean Absolute Error (MAE) for RM and baseline models. Best , Second Best Performance
Figure 3 : Confusion matrices on the ordinal classification benchmark. Darker cells indicate larger sample counts.
Eval.
Method
IR
I’Ideas
RGCW
nRGCW
GPT-4o
Base
32.66
2440
0.432
0.147
SFT
27.30
2040
0.320
0.109
RL
50.58
3779
0.709
0.242
GPT-4.1
Base
52.60
3930
0.965
0.329
SFT
59.53
4448
1.171
0.399
RL
81.80
6112
1.753
0.597
Table 4 : Reference-grounded citation-weighted idea impact evaluation results. IR: Impact Rate, I’Ideas: Impactful Ideas RGCW: Reference-Grounded Citation-Weighted score, nRGCW: Normalized Reference-Grounded Citation-Weighted score, Eval.: Evaluator, Maj. Vote: Majority Vote. Number of evaluation test samples ( N ): 7472. The upper bound of RGCW is approximately 2.93 based on ground-truth citation-impact labels.
Pattern
GPT-5.1
GPT-4.1
GPT-4o
Cnt
%
Cnt
%
Cnt
%
RL only
1252
76.3
1130
30.2
1206
37.6
RL+SFT
149
9.1
1427
38.1
653
20.4
Base+RL
74
4.5
821
21.9
829
25.9
Base only
156
9.5
251
6.7
397
12.4
Base+SFT
11
0.7
116
3.1
119
3.7
Table 5 : Distribution of evaluator-confirmed impact patterns across generation methods. Cnt: Count, Base: Baseline. Percentages are normalized within each evaluator.
Table 6 : Qualitative Comparison of Baseline, SFT, and RL Generated Research Ideas
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Item
Value
Number of samples
100,000
Input tokens per sample
8,000
Output tokens per sample
500
Total input tokens
800,000,000
Total output tokens
50,000,000
Input price per 1M tokens
$0.15
Appendix
Table 7 : Estimated cost for processing 100,000 samples with GPT-4o-mini
Dataset
Train
Validation
Test
Total
Reward
34,973
7,495
7,496
49,964
SFT/RL
34,990
7,499
7,497
49,986
Appendix
Table 8 : Dataset statistics after Batch API generation. The original target size was 50K instances per dataset; differences indicate failed or missing generations.
Figure 4 : Distribution of research goal lengths across the reward and SFT dataset splits. The research goal lengths approximately follow a normal distribution with a peak around 100 words, indicating consistent problem-context formulation across samples.
Figure 5 : Distribution of idea description lengths across the reward and SFT dataset splits. The idea descriptions exhibit a normal distribution with a peak around 200 words, reflecting richer methodological and implementation-oriented detail in the proposed research ideas.
Figure 6 : Publication year distribution of the reward dataset across citation-impact bins. The dataset spans a broad temporal range and contains papers from diverse citation-impact categories, demonstrating balanced impact coverage across publication years.
Figure 7 : Publication year distribution of the SFT dataset across citation-impact bins. The dataset exhibits broad temporal coverage with citation-impact diversity distributed consistently across years.
Training / Evaluation Parameter
Value
Base Model
Qwen3-8B
RL Algorithm
DAPO
Reward Model
Qwen3 8B with CORN
Maximum Sequence Length
2048
Maximum Completion Length
256
Evaluation Generation Length
500
Appendix
Table 10 : Training and evaluation configuration used for GRPO-based impact-aligned research idea generation.
Evaluation Parameter
Value
Base Model
Qwen3-8B
Maximum Sequence Length
2048
Maximum New Tokens
500
Batch Size
128
Sampling Strategy
Stochastic Sampling
Temperature
0.9
Appendix
Table 11: Generation and evaluation configuration used for idea generation experiments. All models were evaluated using identical decoding and inference settings to ensure fair comparison across baseline, SFT, and RL models. Experiments were conducted on a single NVIDIA A100 80GB PCIe GPU.
Table 12 : Qualitative Comparison of Baseline, SFT, and RL Generated Research Ideas
Table 13 : Qualitative Comparison of Baseline, SFT, and RL Generated Research Ideas
Table 14 : Qualitative Comparison of Baseline, SFT, and RL Generated Research Ideas
Table 15: Qualitative Comparison of Baseline, SFT, and RL Generated Research Ideas