Organizations: State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, China · University of Chinese Academy of Sciences, China · Peking University · DeepWisdom · Shanghai Artificial Intelligence Laboratory · Renmin University of China
Research idea innovation is a fundamental engine of scientific progress, yet it remains difficult to generate and evaluate in a scalable and controllable way. This challenge lies in its inherently open-ended and multi-objective nature, where ideas should balance novelty, plausibility and feasibility. While recent LLM-based approaches have made progress through carefully designed prompts or agent pipelines, they are constrained by predefined, static ideation workflows. To address this limitation, we propose MindFlow, a framework that explicitly formulates ideation as a graph-structured Flow in Mind, which is composed of modular thinking operators and modeled by a probabilistic mind supernet. Given a research topic, a controller dynamically samples thinking flows to generate candidate ideas. This open-ended problem is optimized using a tournament-based relative ranking, enabling the controller to progressively favor higher-quality thinking flows. We further introduce an evaluation protocol that jointly assesses problem finding and problem solving, going beyond title- or abstractonly judgments. Across diverse topics, MindFlow shows its superiority as an explicit, controllable and optimizable research idea innovator.
Figures & tables
Figure 1 : Overview of challenges in research idea generation and comparison between prior pipelines and MindFlow.
Figure 2 : The overall framework of our proposed MindFlow. MindFlow formulates research idea innovation as a graph-structured thinking flow composed of atomic thinking operators and modeled by a probabilistic mind supernet. Given a research topic, a controller dynamically samples a tailored thinking flow to generate a idea, and the controller is optimized using a tournament-based relative ranking.
Method
Problem Finding (Motivation)
Problem Solving (Method)
Overall
Novelty
Significance
Timeliness
MOScore
Novelty
Effectiveness
Feasibility
MOScore
Generate
0.294
0.602
0.455
0.441
0.071
0.171
0.308
0.169
0.305
GenerateCoT
0.687
0.379
0.469
0.504
0.308
0.076
0.213
0.185
0.344
AI Scientist
0.199
0.768
0.498
0.456
0.028
0.123
0.455
0.159
0.308
AI-Researcher
0.692
0.351
0.450
0.488
0.265
0.081
0.175
0.165
0.326
VIRSCI
0.711
0.645
0.645
0.667
0.445
0.332
0.109
0.274
0.470
Table 1 : Win-rate evaluation under our LLM-judged protocol. We report dimension-wise win rates for problem finding (novelty, significance, and timeliness) and problem solving (novelty, effectiveness, and feasibility), along with their aggregated multi-objective scores (MOScore) and overall average. The best and second-best results are labeled with bold and underline.
Method
Novelty
Diversity
Effectiveness
Feasibility
Generate
0.370
0.184
0.642
0.539
GenerateCoT
0.422
0.275
0.626
0.191
AI Scientist
0.352
0.167
0.663
0.172
AI-Researcher
0.474
0.293
0.589
0.672
VIRSCI
0.457
0.281
0.630
0.583
MindFlow
0.541
0.322
0.665
0.735
Table 2 : Computable evaluation on novelty, diversity, effectiveness, and feasibility, reflecting multi-objective beyond LLM judgments. The best and second-best are labeled with bold and underline.
Method
Problem Finding (Motivation)
Problem Solving (Method)
Overall
Novelty
Significance
Timeliness
MOScore
Novelty
Effectiveness
Feasibility
MOScore
AIScientist
0.175
0.688
0.425
0.403
0.100
0.213
0.638
0.271
0.337
AI-Researcher
0.600
0.313
0.388
0.428
0.325
0.163
0.275
0.248
0.338
VIRSCI
0.650
0.563
0.575
0.595
0.613
0.450
0.138
0.356
0.475
MindFlow
0.688
0.663
0.650
0.667
0.550
0.538
0.400
0.494
0.580
Table 3 : Human expert evaluation on 50 randomly sampled queries from IdeaBench. We report dimension-wise win rates for problem finding (motivation) and problem solving (method), along with their aggregated multi-objective scores (MOScore) and overall average. The best and second-best results are labeled with bold and underline.
Agreement Type
PF-Novelty
PF-Significance
PF-Timeliness
PS-Novelty
PS-Effectiveness
PS-Feasibility
Average
Inter-judge Agreement Rate
79.3%
68.7%
76.0%
80.7%
72.0%
72.7%
74.9%
Human–LLM Agreement Rate
82.1%
71.4%
78.8%
83.5%
75.3%
75.6%
77.8%
Table 4 : Agreement analysis of the LLM-judged evaluation protocol.
Figure 3 : The visualization of MindFlow’s operator sampling process and visualization of the sampling flow.
Operator
Problem Finding (Motivation)
Problem Solving (Method)
Overall
Novelty
Significance
Timeliness
MOScore
Novelty
Effectiveness
Feasibility
MOScore
Divergent + Convergent Thinking
0.569
0.564
0.531
0.554
0.265
0.194
0.180
0.210
0.383
Counterfactual Thinking
0.901
0.412
0.569
0.611
0.569
0.123
0.128
0.240
0.426
Analogical Thinking
0.578
0.716
0.635
0.641
0.271
0.271
0.171
0.234
0.438
Constraint-Driven Thinking
0.791
0.526
0.706
0.669
0.370
0.346
0.142
0.275
0.472
Critical Thinking
0.275
0.796
0.531
0.511
0.062
0.192
0.275
0.162
0.336
Table 5 : Operator-wise win-rate evaluation under our LLM-judged protocol. We report dimension-wise win rates for problem finding (motivation) and problem solving (method), along with their aggregated multi-objective scores (MOScore) and overall average. The best and second-best results are labeled with bold and underline.
Method
CV
NLP
Multimodal
Audio & Speech
Robotics
Science
General ML
Theory
Generate
0.302
0.289
0.334
0.257
0.284
0.327
0.281
0.365
GenerateCoT
0.355
0.366
0.307
0.259
0.419
0.348
0.352
0.372
AI Scientist
0.281
0.318
0.348
0.254
0.344
0.388
0.305
0.315
AI-Researcher
0.348
0.293
0.264
0.317
0.335
0.314
0.342
0.422
VIRSCI
0.502
0.450
0.436
0.324
0.533
0.473
0.468
0.568
MindFlow (Ours)
0.503
0.498
0.480
0.398
0.694
0.412
0.576
0.660
Table 6 : Performance across topic domains ( problem finding and problem solving ). We report per-domain scores under our LLM-judged protocol. The best and second-best results are labeled with bold and underline.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
CV
NLP
Multimodal
Audio & Speech
Robotics
Science
General ML
Theory
Divergent +
Convergent Thinking
0.363
0.394
0.374
0.268
0.392
0.359
0.416
0.391
Counterfactual Thinking
0.423
0.400
0.402
0.406
0.490
0.334
0.425
0.539
Analogical Thinking
0.457
0.449
0.434
0.331
0.402
0.409
0.388
0.564
Critical Thinking
0.338
0.350
0.359
0.276
0.345
0.335
0.313
0.433
Constraint-Driven Thinking
0.439
0.473
0.445
0.237
0.628
0.316
0.532
0.601
Appendix
Table 7 : Operator-wise performance across topic domains ( problem finding and problem solving ) under our LLM-judged protocol. We report domain-wise scores for each thinking operator. The best and second-best results are highlighted in bold and underlined, respectively.
Problem Finding (Motivation)
Problem Solving (Method)
Overall
Method
Novelty
Significance
Timeliness
MOScore
Novelty
Effectiveness
Feasibility
MOScore
AIScientist
0.175
0.688
0.425
0.403
0.100
0.213
0.638
0.271
0.337
AI-Researcher
0.600
0.313
0.388
0.428
0.325
0.163
0.275
0.248
0.338
VIRSCI
0.650
0.563
0.575
0.595
0.613
0.450
0.138
0.356
0.475
MindFlow
0.688
0.663
0.650
0.667
0.550
0.538
0.400
0.494
0.580
Appendix
Table 8 : Human expert evaluation results.
Problem Finding (Motivation)
Problem Solving (Method)
Average
Novelty
Significance
Timeliness
Novelty
Effectiveness
Feasibility
Inter-judge Agreement
79.3%
68.7%
76.0%
80.7%
72.0%
72.7%
74.9%
Human-LLM Agreement
82.1%
71.4%
78.8%
83.5%
75.3%
75.6%
77.8%
Appendix
Table 9 : Agreement analysis across inter-judge agreement among three LLM judges and between human experts and LLM judges .
Method
Novelty
Significance
Timeliness
MOScore
VIRSCI
0.711±0.018
0.645±0.021
0.645±0.016
0.667±0.015
MindFlow
0.744±0.015
0.782±0.019
0.702±0.017
0.742±0.012
Appendix
Table 10 : Problem Finding (Motivation) performance (mean ± std over three runs).
Method
Novelty
Effectiveness
Feasibility
MOScore
VIRSCI
0.445±0.023
0.332±0.019
0.109±0.014
0.274±0.017
MindFlow
0.361±0.020
0.408±0.022
0.257±0.018
0.339±0.014
Appendix
Table 11 : Problem Solving (Method) performance (mean ± std over three runs).
Recent advancements in large language models (LLMs) have demonstrated their potential in automating the scientific research ideation. Existing approaches primarily focus on prompting techniques, often producing ideas misaligned with expert standards - novelty, feasibility, and effectiveness, which are widely recognized by the research community as the three key subdimensions of high-quality ideas. Also, balancing these dimensions remains challenging due to their inherent trade-offs. To address these limitations, we propose the first framework that employs a two-stage approach combining Supervised Fine-Tuning (SFT) and controllable Reinforcement Learning (RL) for the task. In the SFT stage, the model learns foundational patterns from pairs of research papers and their corresponding follow-up ideas. In the RL stage, multi-dimensional reward models guided by fine-grained feedback evaluate and optimize the model across key dimensions. During inference, dimensional controllers coordinated by a sentence-level decoder enable dynamic context-aware steering of the idea generation process. Our framework provides a balanced approach to research idea generation, achieving high-quality outcomes in the experiment by dynamically navigating the trade-offs among novelty, feasibility, and effectiveness.
Ruochen Li, Liqiang Jing, Chi Han +2
University of Texas at Dallas · UIUC · Stony Brook University
Generating novel research ideas is fundamental to scientific progress. While Large Language Models (LLMs) show promise in assisting this process, existing approaches often exhibit semantic convergence, resulting in limited diversity and novelty. To address this, we introduce EvoGens, an evolution-inspired framework that recasts scientific idea generation as an evolutionary search over a population of ideas. EvoGens iteratively applies rank-based mutation with differentiated retrieval planning to incorporate external knowledge, and semantic-aware crossover to fuse complementary concepts for conceptual reorganization. A lightweight evaluation signal guides the selection process, encouraging sustained exploration while mitigating premature convergence. Extensive experiments demonstrate that EvoGens substantially enhances exploration capabilities compared to state-of-the-art baselines. Specifically, it improves the Novelty from 0.1 to 0.4 and the Diversity from 0.24 to 0.55, while maintaining comparable idea quality under the current automatic evaluation protocol. These findings suggest that evolutionary mechanisms can serve as a useful framework for exploration-oriented research ideation, especially for broadening the novelty and diversity of candidate ideas under a shared automatic evaluation setting.
Xu Li, Hanzhe Tu, Xinyi Li +3
Southwest Petroleum University, Chengdu, China · Sichuan Police College, Luzhou, China
Scientific progress depends on the continual generation of innovative re-search ideas. However, the rapid growth of scientific literature has greatly increased the cost of knowledge filtering, making it harder for researchers to identify novel directions. Although existing large language model (LLM)-based methods show promise in research idea generation, the ideas they produce are often repetitive and lack depth. To address this issue, this study proposes a multi-agent iterative planning search strategy inspired by com-binatorial innovation theory. The framework combines iterative knowledge search with an LLM-based multi-agent system to generate, evaluate, and re-fine research ideas through repeated interaction, with the goal of improving idea diversity and novelty. Experiments in the natural language processing domain show that the proposed method outperforms state-of-the-art base-lines in both diversity and novelty. Further comparison with ideas derived from top-tier machine learning conference papers indicates that the quality of the generated ideas falls between that of accepted and rejected papers. These results suggest that the proposed framework is a promising approach for supporting high-quality research idea generation. The source code and dataset used in this paper are publicly available on Github repository: https://github.com/ChenShuai00/MAGenIdeas. The demo is available at https://huggingface.co/spaces/cshuai20/MAGenIdeas.
Shuai Chen, Chengzhi Zhang
Department of Information Management, Nanjing University of Science and Technology, Nanjing, 210094 China