Scientific ideation is increasingly mediated by large language models, but current ideation systems are usually trained and evaluated on immediately judgeable proxies such as novelty, clarity, and feasibility. This leaves open whether delayed signals of scientific uptake can be used as feedback for steering models toward research directions with higher expected \emph{impact}. We study this question using citation-normalized impact as a noisy but scalable proxy for scholarly uptake. We construct a large-scale dataset from over 100K computer science papers by extracting goal-conditioned idea descriptions and assigning each paper an ordinal, year-normalized citation label. We then train a goal-conditioned reward model to predict citation-impact labels from research goal and idea pairs, and use this reward to align an idea generator through supervised fine-tuning followed by reinforcement learning. To reduce circularity, we evaluate generated ideas with a held-out, reference-grounded protocol that compares model outputs against historical ideas under the same research goal and weights judgments by the reference idea's citation-impact label. Experiments show that our RL-tuned model consistently produces ideas with higher estimated impact than both the base model and supervised fine-tuning baselines. Our findings position scientific impact as a practical, outcome-grounded feedback signal for aligning LLMs in open-ended scientific discovery.
Figures & tables
Figure 1 : Dataset construction pipeline.
Dataset
Split
Zero/VL
Low
Med.
High
V.High
Total
Reward
Train
7,231
7,232
7,241
7,236
6,033
34,973
Reward
Val
1,552
1,550
1,549
1,552
1,292
7,495
Reward
Test
1,552
1,552
1,548
1,551
1,293
7,496
SFT/RL
Train
7,239
7,238
7,236
7,239
6,038
34,990
SFT/RL
Val
1,552
1,552
1,552
1,551
1,292
7,499
SFT/RL
Test
1,549
1,552
1,551
1,550
1,295
7,497
Table 1 : Paper distribution across ordinal classes for distinct dataset subsets and splits. Counts remain approximately uniform across train, validation, and test sets, indicating successful stratification.
Figure 2 : Impact Aligned Ideation Training formulation
Model
Acc
MAE
±1 Acc
Spear
QWK
Zero-shot
Qwen3-8B
20.7
1.410
58.7
0.065
0.004
GPT-4o
23.9
1.155
68.1
0.354
0.235
GPT-5
27.7
1.027
75.3
0.402
0.338
Fine-tuned
RM
48.0 (+20.3)
0.632 (-0.39)
90.6 (+15.3)
0.756 (+0.35)
0.754 (+0.41)
Table 2 : Trained Reward Model (RM: Qwen3-8B + LoRA + CORN) Performance comparison with zero-shot models on the ordinal classification task. Improvement of RM over GPT-5. Acc: Accuracy, Spear: Spearman Correlation
Model
V.Low
Low
Med.
High
V.High
GPT-4o
1.95
1.41
1.07
0.96
0.38
GPT-5
0.66
0.86
0.94
1.02
1.79
Qwen3-8B
3.01
2.01
1.00
0.00
1.00
RM
0.50
0.71
0.77
0.74
0.43
Table 3 : Class-wise Mean Absolute Error (MAE) for RM and baseline models. Best , Second Best Performance
Figure 3 : Confusion matrices on the ordinal classification benchmark. Darker cells indicate larger sample counts.
Eval.
Method
IR
I’Ideas
RGCW
nRGCW
GPT-4o
Base
32.66
2440
0.432
0.147
SFT
27.30
2040
0.320
0.109
RL
50.58
3779
0.709
0.242
GPT-4.1
Base
52.60
3930
0.965
0.329
SFT
59.53
4448
1.171
0.399
RL
81.80
6112
1.753
0.597
Table 4 : Reference-grounded citation-weighted idea impact evaluation results. IR: Impact Rate, I’Ideas: Impactful Ideas RGCW: Reference-Grounded Citation-Weighted score, nRGCW: Normalized Reference-Grounded Citation-Weighted score, Eval.: Evaluator, Maj. Vote: Majority Vote. Number of evaluation test samples ( N ): 7472. The upper bound of RGCW is approximately 2.93 based on ground-truth citation-impact labels.
Pattern
GPT-5.1
GPT-4.1
GPT-4o
Cnt
%
Cnt
%
Cnt
%
RL only
1252
76.3
1130
30.2
1206
37.6
RL+SFT
149
9.1
1427
38.1
653
20.4
Base+RL
74
4.5
821
21.9
829
25.9
Base only
156
9.5
251
6.7
397
12.4
Base+SFT
11
0.7
116
3.1
119
3.7
Table 5 : Distribution of evaluator-confirmed impact patterns across generation methods. Cnt: Count, Base: Baseline. Percentages are normalized within each evaluator.
Table 6 : Qualitative Comparison of Baseline, SFT, and RL Generated Research Ideas
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Item
Value
Number of samples
100,000
Input tokens per sample
8,000
Output tokens per sample
500
Total input tokens
800,000,000
Total output tokens
50,000,000
Input price per 1M tokens
$0.15
Appendix
Table 7 : Estimated cost for processing 100,000 samples with GPT-4o-mini
Dataset
Train
Validation
Test
Total
Reward
34,973
7,495
7,496
49,964
SFT/RL
34,990
7,499
7,497
49,986
Appendix
Table 8 : Dataset statistics after Batch API generation. The original target size was 50K instances per dataset; differences indicate failed or missing generations.
Figure 4 : Distribution of research goal lengths across the reward and SFT dataset splits. The research goal lengths approximately follow a normal distribution with a peak around 100 words, indicating consistent problem-context formulation across samples.
Figure 5 : Distribution of idea description lengths across the reward and SFT dataset splits. The idea descriptions exhibit a normal distribution with a peak around 200 words, reflecting richer methodological and implementation-oriented detail in the proposed research ideas.
Figure 6 : Publication year distribution of the reward dataset across citation-impact bins. The dataset spans a broad temporal range and contains papers from diverse citation-impact categories, demonstrating balanced impact coverage across publication years.
Figure 7 : Publication year distribution of the SFT dataset across citation-impact bins. The dataset exhibits broad temporal coverage with citation-impact diversity distributed consistently across years.
Training / Evaluation Parameter
Value
Base Model
Qwen3-8B
RL Algorithm
DAPO
Reward Model
Qwen3 8B with CORN
Maximum Sequence Length
2048
Maximum Completion Length
256
Evaluation Generation Length
500
Appendix
Table 10 : Training and evaluation configuration used for GRPO-based impact-aligned research idea generation.
Evaluation Parameter
Value
Base Model
Qwen3-8B
Maximum Sequence Length
2048
Maximum New Tokens
500
Batch Size
128
Sampling Strategy
Stochastic Sampling
Temperature
0.9
Appendix
Table 11: Generation and evaluation configuration used for idea generation experiments. All models were evaluated using identical decoding and inference settings to ensure fair comparison across baseline, SFT, and RL models. Experiments were conducted on a single NVIDIA A100 80GB PCIe GPU.
Table 12 : Qualitative Comparison of Baseline, SFT, and RL Generated Research Ideas
Table 13 : Qualitative Comparison of Baseline, SFT, and RL Generated Research Ideas
Table 14 : Qualitative Comparison of Baseline, SFT, and RL Generated Research Ideas
Table 15: Qualitative Comparison of Baseline, SFT, and RL Generated Research Ideas
Scientific discovery depends on expert judgement and foresight, which we call scientific taste: the ability to judge and propose research ideas with potential for long-term scientific impact. Whether AI can learn this ability remains an open question. Here we provide evidence that artificial intelligence can learn judgement and ideation. We introduce Reinforcement Learning from Community Feedback (RLCF), a training paradigm that uses large-scale signals from scientific community as supervision. We first train Scientific Judge on field- and time-matched pairs of high- vs. low-citation papers to judge ideas. We then train a Scientific Thinker, to propose research ideas with high potential impact. Experiments show that the 30B Scientific Judge variant outperforms strong LLM baselines (e.g., GPT-5.4 Thinking), while Scientific Judge generalizes across future-year papers, unseen fields, and other community metrics. Furthermore, Scientific Thinker proposes research ideas with higher potential impact than baselines. These results suggest that AI can learn scientific taste, marking an important step towards AI systems that could help accelerate scientific discovery.
Jingqi Tong, Mingzhe Li, Hangcheng Li +20
Fudan University · Shanghai Innovation Institute · OpenMOSS Team +2
Large Language Models (LLMs) have demonstrated potential in automating scientific ideation, yet current approaches relying on iterative prompting or complex multi-agent architectures often suffer from hallucination or computational inefficiency. A critical bottleneck in applying Reinforcement Learning (RL) to this open-ended domain is reward hacking -- where models exploit imperfect evaluation proxies to maximize scores without producing genuine scientific innovation. To address these limitations, we propose an RL framework explicitly tailored for high-quality scientific idea generation. We propose the first multi-agent reward function designed to serve as a judge, decoupling methodological validation from implementation details while providing strict binary rewards that are robust to reward hacking. To effectively optimize against this sparse signal, we utilize an unbiased variant of Group Relative Policy Optimization to mitigate artificial length bias. We grounded our training in ICLR-320, a curated dataset of problem-solution pairs extracted from ICLR 2024 proceedings. Experiments demonstrate that our framework significantly outperforms state-of-the-art baselines across expert-evaluated metrics of novelty, feasibility, and effectiveness.
Scientific discovery is an extended process of ideation--surveying prior work, forming hypotheses, and refining reasoning--yet existing approaches treat this phase as a brief preamble despite its central role in research. We introduce SCISENSE, a sensemaking-grounded framework that operationalizes ideation as a structured sequence of eight cognitive stages (Pirolli & Card, 2005). We construct SCISENSE-Traj, a 100K-scale dataset of citation-conditioned research trajectories in two modes: Target, where an LLM reconstructs the ideation path leading to a known paper from its cited works, and Infer, where the LLM proposes novel directions from the same citations. We distill these into SCISENSE-LM, a family of sensemaking LLMs spanning 3B to 70B parameters. Contrary to the assumption that looser supervision promotes greater exploration, Target-trained models achieve a 2.0% improvement in trajectory quality over Infer-trained models while also producing more novel and diverse outputs. This advantage propagates downstream: coding agents conditioned on Target trajectories produce research artifacts with higher executability and quality than those conditioned on Infer trajectories. This suggests that targeted ideation reduces cognitive burden on downstream agents, freeing them to explore more creatively. SCISENSE offers both a practical tool for augmenting LLM-driven research workflows and a principled testbed for studying how planning shapes scientific discovery.