Learning to Predict Future-Aligned Research Proposals with Language Models
Organizations: University of Illinois Urbana-Champaign
Abstract
Large language models (LLMs) are increasingly used to assist ideation in research, but evaluating the quality of LLM-generated research proposals remains difficult: novelty and soundness are hard to measure automatically, and large-scale human evaluation is costly. We propose a verifiable alternative by reframing proposal generation as a time-sliced scientific forecasting problem. Given a research question and inspiring papers available before a cutoff time, the model generates a structured proposal and is evaluated by whether it anticipates research directions that appear in papers published after the time. We operationalize this objective with the Future Alignment Score (FAS), computed via retrieval and LLM-based semantic scoring against a held-out future corpus. To train models, we build a time-consistent dataset of 21,835 paper occurrences across 3,642 instances from targets and their pre-cutoff citations, and synthesize reasoning traces that teach gap identification and inspiration borrowing. Across Llama-3.1 and Qwen2.5 models, future-aligned tuning improves future alignment over unaligned baselines (up to +10.6% overall FAS), and domain-expert human evaluation corroborates improved proposal quality. Finally, we demonstrate practical impact by implementing two model-generated proposals with a code agent, obtaining 4.17% accuracy gain on MATH from a new prompting strategy and consistent improvements for a novel model-merging method. Our code and data are publicly available at https://github.com/Arthur-Heng/future-aligned-proposals.
Figures & tables
| Method | Hypothesis | Proposed Method | Novelty Claims | Exp. Details | Overall |
| Llama-3.1-8B-Instruct | |||||
| RQ only | 63.0 | 52.8 | 51.1 | 52.4 | 60.0 |
| Paper only | 55.2 | 49.4 | 46.9 | 48.0 | 52.4 |
| Prompting | 64.5 | 56.8 | 54.4 | 55.0 | 62.1 |
| AI-Researcher | 57.9 | 46.0 | 45.1 | 44.4 | 53.7 |
| Chain-of-Ideas | 63.2 | 54.1 | 51.6 | 45.9 | 59.3 |
| Ablation Type | Hyp. | Method | Novelty | Overall |
| background ( =775) | 6.75 | 7.21 | 8.31 | 6.59 |
| method ( =765) | 6.93 | 7.37 | 8.35 | 6.63 |
| benchmark ( =248) | 6.41 | 6.85 | 7.30 | 6.13 |
| Model | Resource | Task–Method | Task–Exp. | Avg. |
| Prompting | 3.26 | 3.20 | 2.99 | 3.15 |
| CoI | 3.52 | 3.20 | 3.12 | 3.28 |
| AI-Researcher | 3.30 | 3.13 | 3.03 | 3.16 |
| Future-aligned SFT | 3.63 | 3.42 | 3.21 | 3.42 |
| w/o reasoning traces | 3.54 | 3.41 | 3.16 | 3.37 |
| w/o stepwise | 3.39 | 3.30 | 3.07 | 3.25 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | LLM (Mistral-7B) | Vision (ViT-B/16) | |||||
| GSM8K | ARC-E | ARC-C | HSwag | PPL | Acc. | F1 | |
| Simple Average | 0.530 | 0.430 | 0.290 | 0.620 | 13.06 | 0.473 | 0.453 |
| Uniform Sparsity | 0.480 | 0.750 | 0.605 | 0.605 | 12.21 | 0.425 | 0.399 |
| TIES-Merging | 0.545 | 0.125 | 0.090 | 0.655 | 14.57 | 0.530 | 0.515 |
| TIES (w/ sign) | 0.545 | 0.455 | 0.310 | 0.630 | 12.97 | – | – |
| MALS (ours) | 0.525 | 0.755 | 0.605 | 0.625 | 12.22 | 0.493 | 0.472 |
| Method | GSM8K | MATH | BBH |
| Direct | 93.0 | 46.0 | 84.0 |
| CoT | 89.0 | 42.0 | 90.0 |
| Self-Consistency | 95.0 | 48.0 | 92.0 |
| Strategy Search (ours) | 95.0 | 50.0 | 88.0 |
| Comparison | Mean diff. | 95% bootstrap CI | Bootstrap | Wilcoxon |
| Llama-3.1-8B: Ours vs. prompting | ||||
| Qwen2.5-7B: Ours vs. prompting | ||||
| Qwen2.5-14B: Ours vs. prompting | ||||
| Qwen2.5-14B: Ours vs. w/o stepwise |
| Embed | Judge | Ours | Untuned | CoI | |||
| 10 | 3-large | 4.1-mini | 6.94 | 6.46 | 6.07 | – | – |
| 5 | 3-large | 4.1-mini | 6.93 | 6.46 | 6.05 | 0.946 | 0.931 |
| 10 | 3-small | 4.1-mini | 6.74 | 6.36 | 5.93 | 0.712 | 0.771 |
| 10 | 3-large | 4o-mini | 6.72 | 6.57 | 6.25 | 0.705 | 0.664 |
| Judge | Model | Pearson | Spearman | Within 1 pt |
| GPT-4o-mini | Stepwise CoT | 0.681 | 0.614 | 98.7% |
| Untuned | 0.694 | 0.684 | 97.7% | |
| CoI | 0.740 | 0.693 | 94.3% | |
| GPT-5-mini | Stepwise CoT | 0.803 | 0.796 | 86.0% |
| Untuned | 0.796 | 0.781 | 88.6% | |
| CoI | 0.824 | 0.765 | 89.0% |
| Dim. | Model | Human-Derived Reference | Generated Proposal |
| Hypothesis | Ours (7/10) | Incorporating explicit selection and demonstration of diverse reasoning types (deductive, inductive, abductive, analogical) for each problem instance will enable LLMs to solve a broader range of problems more effectively than approaches relying on a single or undifferentiated reasoning strategy. | Explicitly biasing and coordinating the use of high-level reasoning strategies within LLMs, particularly by introducing strategy biasing and modular reasoning operations , will significantly enhance their performance on complex logical and mathematical tasks compared to standard prompting methods. |
| Untuned (4/10) | Integrating diverse reasoning strategies within a unified framework will significantly enhance the problem-solving capabilities of large language models, particularly in complex logical and mathematical tasks, compared to models relying solely on deductive reasoning. | ||
| CoI (4/10) | By integrating a Hybrid Reasoning Framework that includes abductive, inductive, deductive, and analogical reasoning modalities, along with dynamic strategy selection and enhanced interactive mechanisms , LLMs can significantly enhance their problem-solving capabilities, particularly in tasks involving incomplete data, pattern recognition, and prediction. | ||
| Method | Ours (7/10) | Three main components: (1) a meta-thinker , fine-tuned to predict the effectiveness of each reasoning type based on empirical success rates; (2) an explicit demonstration collection , storing successful solutions for each type, enabling retrieval of relevant few-shot examples ; (3) a reasoner applying the selected type, guided by retrieved demonstrations. The system aggregates answers via weighted voting . | Two key components: Strategy Biasing —explicitly biases the LLM toward specific strategies (inductive, abductive, analogical) by conditioning the prompt on the desired strategy; Modular Reasoning Operations —task-adaptive reasoning functions that can be flexibly combined and executed in parallel. A Strategy Coordination mechanism schedules and combines the outputs of different reasoning strategies to optimize task performance, leveraging the strengths of each strategy for a given problem. |
| Untuned (3/10) | Reasoning Strategy Module : dynamically switch between abductive, inductive, and analogical reasoning … Graph Representation : graph-based structure where nodes represent reasoning steps and edges represent dependencies. Multiagent Debate : multiple LLM instances generate and critique reasoning paths using different strategies. World Model Integration : incorporate a world model to simulate state transitions and integrate planning algorithms to guide the reasoning process. | ||
| CoI (4/10) | Three core components: Integrated Reasoning Modalities, Dynamic Strategy Selection, and Enhanced Interactive Mechanisms. Each reasoning type is a distinct module, designed to operate independently yet collaboratively. A decision-making layer evaluates task requirements and selects the most suitable strategy. Leveraging the multiagent debate framework , the method facilitates collaboration among multiple model instances, refining and validating reasoning outcomes through iterative feedback. |
| Dim. | Model | Human-Derived Reference | Generated Proposal |
| Hypothesis | Ours (9/10) | A small, specialized subset of attention heads, termed retrieval heads , are primarily responsible for retrieval from long contexts. These are universal, sparse, intrinsic to pretrained models, dynamically activated, and causally linked to factuality and complex reasoning. | A subset of attention heads in long-context LLMs act as retrieval heads , selectively copying and retrieving information from arbitrary positions in the input sequence, which is causally responsible for improved factuality and reasoning performance. |
| Untuned (6/10) | Specific attention head circuits in LLMs are responsible for long-context retrieval, and these mechanisms significantly enhance the model’s factuality and reasoning abilities in downstream tasks. | ||
| CoI (6/10) | It is hypothesized that induction heads significantly contribute to long-context retrieval through specific interaction patterns with other attention heads, thereby enhancing the model’s ability to maintain and leverage long-term dependencies. Furthermore … novel prompting strategies and architectural designs can improve long-context retrieval and reasoning. | ||
| Method | Ours (7/10) | Define a retrieval score for each attention head, quantifying copy-paste behavior during autoregressive decoding. Needle-in-a-Haystack tests with unique QA pairs embedded at random positions … Retrieval scores computed across diverse contexts and model variants. Examine universality, sparsity, intrinsic nature , and dynamic activation across model families, scales, and fine-tuning types. | Identify and characterize retrieval heads —attention heads that selectively copy and retrieve from arbitrary positions. Cluster heads based on copying behavior using PCA of per-token loss vectors ; retrieval heads identified as those exhibiting long-range copying across multiple training snapshots. Causal role validated through ablation studies , where retrieval heads are removed or replaced and impact on retrieval and downstream reasoning is measured. Retention patterns in the KV cache analyzed for memory usage and efficiency. |
| Untuned (5/10) | Three separate tracks: (1) Mechanistic Analysis —per-token loss PCA, identify attention heads via sequence copying tasks, architectural perturbations and direct ablations; (2) Chain-of-Thought Prompting —evaluate reasoning with and without identified retrieval heads, compare to standard prompting; (3) KV Cache Compression —adaptive techniques (e.g., FastGen), use retrieval heads to inform compression policies, ensure critical context retained. | ||
| CoI (5/10) | Multi-faceted approach: mechanistic analysis of attention mechanisms, focusing on induction heads and their interactions … New prompting strategies based on chain-of-thought prompting to encourage engagement with long-range dependencies. Architectural designs to prioritize long-term context, including modifications to the attention mechanism and novel training objectives. |
| Setting | Value |
| Epochs | 2 |
| Max sequence length | 8000 |
| Batch size (per device) | 1 |
| Gradient accumulation | 4 |
| Effective batch size | 4 |
| Learning rate |
| Split | Raw | Proposal |
| Training completions (NeurIPS’24 + ICLR’24) | ||
| Stepwise-CoT | 909 | 460 |
| CoT | 900 | 460 |
| No-CoT | 460 | 460 |
| Test reference (NeurIPS’25 + ICML’25 + ICLR’25) | ||
| Human-derived proposal | — | 460 |
| Comparison | Dimension | Win | Tie | Loss | Win Rate (95% CI) | Unanimity |
| Stepwise CoT vs. Human | Overall | 25 | 10 | 25 | 50.0% [37.7, 62.3] | 11.7% |
| Soundness | 21 | 11 | 28 | 44.2% [32.3, 56.7] | 11.7% | |
| Excitement | 20 | 18 | 22 | 48.3% [36.2, 60.7] | 6.7% | |
| Stepwise CoT vs. Prompting | Overall | 31 | 3 | 26 | 54.2% [41.7, 66.1] | 26.7% |
| Soundness | 29 | 7 | 24 | 54.2% [41.7, 66.1] | 31.7% | |
| Excitement | 30 | 10 | 20 | 58.3% [45.7, 69.9] | 26.7% |