Large language models (LLMs) are increasingly used to assist ideation in research, but evaluating the quality of LLM-generated research proposals remains difficult: novelty and soundness are hard to measure automatically, and large-scale human evaluation is costly. We propose a verifiable alternative by reframing proposal generation as a time-sliced scientific forecasting problem. Given a research question and inspiring papers available before a cutoff time, the model generates a structured proposal and is evaluated by whether it anticipates research directions that appear in papers published after the time. We operationalize this objective with the Future Alignment Score (FAS), computed via retrieval and LLM-based semantic scoring against a held-out future corpus. To train models, we build a time-consistent dataset of 21,835 paper occurrences across 3,642 instances from targets and their pre-cutoff citations, and synthesize reasoning traces that teach gap identification and inspiration borrowing. Across Llama-3.1 and Qwen2.5 models, future-aligned tuning improves future alignment over unaligned baselines (up to +10.6% overall FAS), and domain-expert human evaluation corroborates improved proposal quality. Finally, we demonstrate practical impact by implementing two model-generated proposals with a code agent, obtaining 4.17% accuracy gain on MATH from a new prompting strategy and consistent improvements for a novel model-merging method. Our code and data are publicly available at https://github.com/Arthur-Heng/future-aligned-proposals.
Figures & tables
Figure 1: Given inspiring papers S and a research question q available before a cutoff time tC , the model generates a proposal P~ . We evaluate whether the proposal anticipates future human research directions by comparing it against papers published after tC using retrieval and LLM-based semantic alignment.
Figure 2: Overview of the proposed future-aligned learning framework. Time-consistent supervision constructs training data from historical papers without future leakage, and citation-grounded stepwise reasoning decomposes proposal generation into staged scientific planning. Together, these enable LoRA-based supervised fine-tuning of a proposal generator, which is evaluated by Future Alignment Score (FAS) against a held-out future corpus and further validated through human evaluation and execution-based case studies.
Method
Hypothesis
Proposed Method
Novelty Claims
Exp. Details
Overall
Llama-3.1-8B-Instruct
RQ only
63.0
52.8
51.1
52.4
60.0
Paper only
55.2
49.4
46.9
48.0
52.4
Prompting
64.5
56.8
54.4
55.0
62.1
AI-Researcher
57.9
46.0
45.1
44.4
53.7
Chain-of-Ideas
63.2
54.1
51.6
45.9
59.3
Table 1: Main results (FAS; higher is better) for future-aligned proposal prediction. We report component-wise FAS along with the overall FAS. Future-aligned SFT improves the FAS substantially over the unaligned baselines, while synthetic structured reasoning trace supervision provides additional gains, improving the overall score by 10.6%.
Figure 3: Pairwise human evaluation results (win/tie/lose). Each stacked bar shows the fraction of instances where Stepwise CoT is preferred (win), the two proposals are judged equivalent (tie), or Stepwise CoT is not preferred (lose), aggregated by majority vote across three annotators.
Figure 4: Two proposals generated by Qwen2.5-14B- Instruct (stepwise CoT tuned). The content is summarized for readability. The proposals are textually sound and are turned into reasonable experimental results and findings with the implementation and execution of code agents.
Table 3: Multi-dimensional LLM judge evaluation (1–5 scale). Future-aligned tuned models consistently outperform the baselines, while the Stepwise-CoT model achieves the highest scores on all three dimensions.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Method
LLM (Mistral-7B)
Vision (ViT-B/16)
GSM8K
ARC-E
ARC-C
HSwag
PPL ↓
Acc.
F1
Simple Average
0.530
0.430
0.290
0.620
13.06
0.473
0.453
Uniform Sparsity
0.480
0.750
0.605
0.605
12.21
0.425
0.399
TIES-Merging
0.545
0.125
0.090
0.655
14.57
0.530
0.515
TIES (w/ sign)
0.545
0.455
0.310
0.630
12.97
–
–
MALS (ours)
0.525
0.755
0.605
0.625
12.22
0.493
0.472
Appendix
Table 4: Model merging results on LLM (Mistral-7B) and Vision (ViT-B/16) tasks. Best results per column in bold . ↓ indicates lower is better.
Method
GSM8K
MATH
BBH
Direct
93.0
46.0
84.0
CoT
89.0
42.0
90.0
Self-Consistency
95.0
48.0
92.0
Strategy Search (ours)
95.0
50.0
88.0
Appendix
Table 5: Accuracy of the proposed Strategy Search method on three reasoning benchmarks. Bold means the best performance.
Comparison
Mean diff.
95% bootstrap CI
Bootstrap p
Wilcoxon p
Llama-3.1-8B: Ours vs. prompting
+3.17
[2.31,4.04]
<10−4
1.54×10−13
Qwen2.5-7B: Ours vs. prompting
+2.89
[2.25,3.55]
<10−4
3.35×10−17
Qwen2.5-14B: Ours vs. prompting
+6.69
[6.13,7.25]
<10−4
5.10×10−82
Qwen2.5-14B: Ours vs. w/o stepwise
+3.60
[3.06,4.14]
<10−4
3.84×10−34
Appendix
Table 6: Paired significance tests for overall FAS on all 819 evaluation instances. The bootstrap uses 10,000 paired resamples. All confidence intervals exclude zero, and both tests reject the null hypothesis for every comparison shown.
k
Embed
Judge
Ours
Untuned
CoI
r
ρ
10
3-large
4.1-mini
6.94
6.46
6.07
–
–
5
3-large
4.1-mini
6.93
6.46
6.05
0.946
0.931
10
3-small
4.1-mini
6.74
6.36
5.93
0.712
0.771
10
3-large
4o-mini
6.72
6.57
6.25
0.705
0.664
Appendix
Table 7: Robustness of FAS under evaluation-pipeline variations (n=300). We vary retrieval depth ( k ), embedding model, and judge model from the default setting (shaded). Rankings remain consistent ( Ours > Untuned > CoI). Avg. Pearson ( r ) and Spearman ( ρ ) are instance-level correlations with the baseline.
Judge
Model
Pearson r
Spearman ρ
Within 1 pt
GPT-4o-mini
Stepwise CoT
0.681
0.614
98.7%
Untuned
0.694
0.684
97.7%
CoI
0.740
0.693
94.3%
GPT-5-mini
Stepwise CoT
0.803
0.796
86.0%
Untuned
0.796
0.781
88.6%
CoI
0.824
0.765
89.0%
Appendix
Table 8: Inter-judge agreement with the primary judge (GPT-4.1-mini). Both judges yield the identical model ordering ( Ours > CoI > Untuned).
Dim.
Model
Human-Derived Reference
Generated Proposal
Hypothesis
Ours (7/10)
Incorporating explicit selection and demonstration of diverse reasoning types (deductive, inductive, abductive, analogical) for each problem instance will enable LLMs to solve a broader range of problems more effectively than approaches relying on a single or undifferentiated reasoning strategy.
Explicitly biasing and coordinating the use of high-level reasoning strategies within LLMs, particularly by introducing strategy biasing and modular reasoning operations , will significantly enhance their performance on complex logical and mathematical tasks compared to standard prompting methods.
Untuned (4/10)
Integrating diverse reasoning strategies within a unified framework will significantly enhance the problem-solving capabilities of large language models, particularly in complex logical and mathematical tasks, compared to models relying solely on deductive reasoning.
CoI (4/10)
By integrating a Hybrid Reasoning Framework that includes abductive, inductive, deductive, and analogical reasoning modalities, along with dynamic strategy selection and enhanced interactive mechanisms , LLMs can significantly enhance their problem-solving capabilities, particularly in tasks involving incomplete data, pattern recognition, and prediction.
Method
Ours (7/10)
Three main components: (1) a meta-thinker , fine-tuned to predict the effectiveness of each reasoning type based on empirical success rates; (2) an explicit demonstration collection , storing successful solutions for each type, enabling retrieval of relevant few-shot examples ; (3) a reasoner applying the selected type, guided by retrieved demonstrations. The system aggregates answers via weighted voting .
Two key components: Strategy Biasing —explicitly biases the LLM toward specific strategies (inductive, abductive, analogical) by conditioning the prompt on the desired strategy; Modular Reasoning Operations —task-adaptive reasoning functions that can be flexibly combined and executed in parallel. A Strategy Coordination mechanism schedules and combines the outputs of different reasoning strategies to optimize task performance, leveraging the strengths of each strategy for a given problem.
Untuned (3/10)
Reasoning Strategy Module : dynamically switch between abductive, inductive, and analogical reasoning … Graph Representation : graph-based structure where nodes represent reasoning steps and edges represent dependencies. Multiagent Debate : multiple LLM instances generate and critique reasoning paths using different strategies. World Model Integration : incorporate a world model to simulate state transitions and integrate planning algorithms to guide the reasoning process.
CoI (4/10)
Three core components: Integrated Reasoning Modalities, Dynamic Strategy Selection, and Enhanced Interactive Mechanisms. Each reasoning type is a distinct module, designed to operate independently yet collaboratively. A decision-making layer evaluates task requirements and selects the most suitable strategy. Leveraging the multiagent debate framework , the method facilitates collaboration among multiple model instances, refining and validating reasoning outcomes through iterative feedback.
Appendix
Table 9: Case Study 1— TypedThinker: Typed Thinking Improves Large Language Model Reasoning . The human-derived proposal designs a meta-thinker that selects among reasoning types with demonstration memory and weighted voting. Our model captures this core design, while the baselines propose generic frameworks with unrelated components.
Dim.
Model
Human-Derived Reference
Generated Proposal
Hypothesis
Ours (9/10)
A small, specialized subset of attention heads, termed retrieval heads , are primarily responsible for retrieval from long contexts. These are universal, sparse, intrinsic to pretrained models, dynamically activated, and causally linked to factuality and complex reasoning.
A subset of attention heads in long-context LLMs act as retrieval heads , selectively copying and retrieving information from arbitrary positions in the input sequence, which is causally responsible for improved factuality and reasoning performance.
Untuned (6/10)
Specific attention head circuits in LLMs are responsible for long-context retrieval, and these mechanisms significantly enhance the model’s factuality and reasoning abilities in downstream tasks.
CoI (6/10)
It is hypothesized that induction heads significantly contribute to long-context retrieval through specific interaction patterns with other attention heads, thereby enhancing the model’s ability to maintain and leverage long-term dependencies. Furthermore … novel prompting strategies and architectural designs can improve long-context retrieval and reasoning.
Method
Ours (7/10)
Define a retrieval score for each attention head, quantifying copy-paste behavior during autoregressive decoding. Needle-in-a-Haystack tests with unique QA pairs embedded at random positions … Retrieval scores computed across diverse contexts and model variants. Examine universality, sparsity, intrinsic nature , and dynamic activation across model families, scales, and fine-tuning types.
Identify and characterize retrieval heads —attention heads that selectively copy and retrieve from arbitrary positions. Cluster heads based on copying behavior using PCA of per-token loss vectors ; retrieval heads identified as those exhibiting long-range copying across multiple training snapshots. Causal role validated through ablation studies , where retrieval heads are removed or replaced and impact on retrieval and downstream reasoning is measured. Retention patterns in the KV cache analyzed for memory usage and efficiency.
Untuned (5/10)
Three separate tracks: (1) Mechanistic Analysis —per-token loss PCA, identify attention heads via sequence copying tasks, architectural perturbations and direct ablations; (2) Chain-of-Thought Prompting —evaluate reasoning with and without identified retrieval heads, compare to standard prompting; (3) KV Cache Compression —adaptive techniques (e.g., FastGen), use retrieval heads to inform compression policies, ensure critical context retained.
CoI (5/10)
Multi-faceted approach: mechanistic analysis of attention mechanisms, focusing on induction heads and their interactions … New prompting strategies based on chain-of-thought prompting to encourage engagement with long-range dependencies. Architectural designs to prioritize long-term context, including modifications to the attention mechanism and novel training objectives.
Appendix
Table 10: Case Study 2— Retrieval Head Mechanistically Explains Long-Context Factuality . The human-derived proposal identifies sparse “retrieval heads” responsible for copy-paste retrieval, validated via Needle-in-a-Haystack tests. Our model nearly matches this design, correctly naming the concept and proposing causal ablation. The baselines dilute across loosely connected tracks (Untuned) or misidentify the mechanism as “induction heads” (CoI).
Setting
Value
Epochs
2
Max sequence length
8000
Batch size (per device)
1
Gradient accumulation
4
Effective batch size
4
Learning rate
2×10−5
Appendix
Table 11: SFT training configuration for all the fine-tuning. Effective batch size equals batch size × gradient accumulation steps.
Figure 5: System prompts used for different proposal-generation variants.
Figure 6: Additional user-side instructions used for CoT-based proposal generation.
Figure 7: Prompt used to convert a paper into a structured proposal target for supervision.
Figure 8: Prompt used to identify the most directly inspiring citations for each target paper.
Figure 9: Prompt used to synthesize direct chain-of-thought reasoning traces from inspiring papers and a target paper outcome.
Figure 10: Prompt used to synthesize stepwise chain-of-thought reasoning traces interleaved with proposal construction.
Split
Raw
Proposal
Training completions (NeurIPS’24 + ICLR’24)
Stepwise-CoT
909
460
CoT
900
460
No-CoT
460
460
Test reference (NeurIPS’25 + ICML’25 + ICLR’25)
Human-derived proposal
—
460
Appendix
Table 12: Mean proposal length in words. For training completions, we report both the full output (including reasoning) and the proposal-only portion. For generated proposals, we report the proposal after stripping reasoning steps.
Figure 11: The annotation interface of the human evaluation.
Comparison
Dimension
Win
Tie
Loss
Win Rate (95% CI)
Unanimity
Stepwise CoT vs. Human
Overall
25
10
25
50.0% [37.7, 62.3]
11.7%
Soundness
21
11
28
44.2% [32.3, 56.7]
11.7%
Excitement
20
18
22
48.3% [36.2, 60.7]
6.7%
Stepwise CoT vs. Prompting
Overall
31
3
26
54.2% [41.7, 66.1]
26.7%
Soundness
29
7
24
54.2% [41.7, 66.1]
31.7%
Excitement
30
10
20
58.3% [45.7, 69.9]
26.7%
Appendix
Table 13: Detailed human evaluation results. Win/Tie/Loss counts reflect the majority vote across three annotators per pair. Win rate treats ties as 0.5 wins. CI: Wilson score 95% confidence interval. Unanimity: percentage of pairs where all three annotators agree.
Figure 12: Prompt used for LLM-based future-alignment scoring between a generated proposal and a candidate future paper.
Figure 13: Prompt used for LLM-based proposal quality evaluation.
Large language models increasingly generate research ideas, yet judging their novelty or feasibility at generation time does not establish whether they anticipate subsequent work. We introduce IdeaForecastBench to evaluate research idea forecasting. Given a community's literature up to a cutoff, a system produces up to five ranked ideas, which are evaluated against later papers. The benchmark comprises 624 rolling episodes across 52 topics, with a fixed retrieve-then-judge protocol and separately reported results from two judges. We compare five history-compression strategies across GPT-4.1, Qwen2.5-7B/14B, and Qwen3.5-9B, together with a learned Mode-Decomposition Forecaster (MDF). Under the primary GPT-4.1-mini judge, Summary improves on Direct in Hit@5 and Precision@5 across all four backbones. Qwen2.5 scores above GPT-4.1, whereas Qwen3.5 scores below it. An outcome-blind assessment finds that Qwen2.5 produces broader forecasts, but does not identify how much breadth contributes to its advantage. Threshold and judge diagnostics further clarify the limits of interpreting realization as precise anticipation. IdeaForecastBench provides a common task for studying which research ideas a community subsequently pursues and how reliably this outcome can be measured.
We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.
Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate because reviews and research ideas lack uniquely verifiable outcomes. We introduce Research Attention Prediction (RAP), a rolling benchmark covering 278 AI/ML fields and 1,390 episodes. At each cut-off, an LLM agent searches a temporally restricted arXiv corpus and predicts the next six months' paper shares across eight frozen research directions. Search generally helps, but all four diagnostic models perform worse than an exact-count exponentially weighted moving average (EWMA) baseline in compositional accuracy. We identify two linked bottlenecks. Under cumulative-history access, State carry-forward outperforms direct Forecast for all four diagnostic models; frozen-evidence replay links a shared component of this reversal to Forecast-oriented policies retrieving a smaller share of recent evidence. Even with exact historical activity, future-specific updating remains limited, with only GPT-5.5 plus reopened Search slightly surpassing EWMA. Fine-tuning on realised outcomes improves Qwen3-4B's forecast Spearman correlation by 0.105 on held-out fields at later origins, with gains also on change-rich episodes.
Yingqian Wu, Jingcong Liang, Siyuan Wang +4
Fudan University · The Chinese University of Hong Kong · University of Oxford +1