Organizations: State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, China · University of Chinese Academy of Sciences, China · Peking University · DeepWisdom · Shanghai Artificial Intelligence Laboratory · Renmin University of China
Research idea innovation is a fundamental engine of scientific progress, yet it remains difficult to generate and evaluate in a scalable and controllable way. This challenge lies in its inherently open-ended and multi-objective nature, where ideas should balance novelty, plausibility and feasibility. While recent LLM-based approaches have made progress through carefully designed prompts or agent pipelines, they are constrained by predefined, static ideation workflows. To address this limitation, we propose MindFlow, a framework that explicitly formulates ideation as a graph-structured Flow in Mind, which is composed of modular thinking operators and modeled by a probabilistic mind supernet. Given a research topic, a controller dynamically samples thinking flows to generate candidate ideas. This open-ended problem is optimized using a tournament-based relative ranking, enabling the controller to progressively favor higher-quality thinking flows. We further introduce an evaluation protocol that jointly assesses problem finding and problem solving, going beyond title- or abstractonly judgments. Across diverse topics, MindFlow shows its superiority as an explicit, controllable and optimizable research idea innovator.
Figures & tables
Figure 1 : Overview of challenges in research idea generation and comparison between prior pipelines and MindFlow.
Figure 2 : The overall framework of our proposed MindFlow. MindFlow formulates research idea innovation as a graph-structured thinking flow composed of atomic thinking operators and modeled by a probabilistic mind supernet. Given a research topic, a controller dynamically samples a tailored thinking flow to generate a idea, and the controller is optimized using a tournament-based relative ranking.
Method
Problem Finding (Motivation)
Problem Solving (Method)
Overall
Novelty
Significance
Timeliness
MOScore
Novelty
Effectiveness
Feasibility
MOScore
Generate
0.294
0.602
0.455
0.441
0.071
0.171
0.308
0.169
0.305
GenerateCoT
0.687
0.379
0.469
0.504
0.308
0.076
0.213
0.185
0.344
AI Scientist
0.199
0.768
0.498
0.456
0.028
0.123
0.455
0.159
0.308
AI-Researcher
0.692
0.351
0.450
0.488
0.265
0.081
0.175
0.165
0.326
VIRSCI
0.711
0.645
0.645
0.667
0.445
0.332
0.109
0.274
0.470
Table 1 : Win-rate evaluation under our LLM-judged protocol. We report dimension-wise win rates for problem finding (novelty, significance, and timeliness) and problem solving (novelty, effectiveness, and feasibility), along with their aggregated multi-objective scores (MOScore) and overall average. The best and second-best results are labeled with bold and underline.
Method
Novelty
Diversity
Effectiveness
Feasibility
Generate
0.370
0.184
0.642
0.539
GenerateCoT
0.422
0.275
0.626
0.191
AI Scientist
0.352
0.167
0.663
0.172
AI-Researcher
0.474
0.293
0.589
0.672
VIRSCI
0.457
0.281
0.630
0.583
MindFlow
0.541
0.322
0.665
0.735
Table 2 : Computable evaluation on novelty, diversity, effectiveness, and feasibility, reflecting multi-objective beyond LLM judgments. The best and second-best are labeled with bold and underline.
Method
Problem Finding (Motivation)
Problem Solving (Method)
Overall
Novelty
Significance
Timeliness
MOScore
Novelty
Effectiveness
Feasibility
MOScore
AIScientist
0.175
0.688
0.425
0.403
0.100
0.213
0.638
0.271
0.337
AI-Researcher
0.600
0.313
0.388
0.428
0.325
0.163
0.275
0.248
0.338
VIRSCI
0.650
0.563
0.575
0.595
0.613
0.450
0.138
0.356
0.475
MindFlow
0.688
0.663
0.650
0.667
0.550
0.538
0.400
0.494
0.580
Table 3 : Human expert evaluation on 50 randomly sampled queries from IdeaBench. We report dimension-wise win rates for problem finding (motivation) and problem solving (method), along with their aggregated multi-objective scores (MOScore) and overall average. The best and second-best results are labeled with bold and underline.
Agreement Type
PF-Novelty
PF-Significance
PF-Timeliness
PS-Novelty
PS-Effectiveness
PS-Feasibility
Average
Inter-judge Agreement Rate
79.3%
68.7%
76.0%
80.7%
72.0%
72.7%
74.9%
Human–LLM Agreement Rate
82.1%
71.4%
78.8%
83.5%
75.3%
75.6%
77.8%
Table 4 : Agreement analysis of the LLM-judged evaluation protocol.
Figure 3 : The visualization of MindFlow’s operator sampling process and visualization of the sampling flow.
Operator
Problem Finding (Motivation)
Problem Solving (Method)
Overall
Novelty
Significance
Timeliness
MOScore
Novelty
Effectiveness
Feasibility
MOScore
Divergent + Convergent Thinking
0.569
0.564
0.531
0.554
0.265
0.194
0.180
0.210
0.383
Counterfactual Thinking
0.901
0.412
0.569
0.611
0.569
0.123
0.128
0.240
0.426
Analogical Thinking
0.578
0.716
0.635
0.641
0.271
0.271
0.171
0.234
0.438
Constraint-Driven Thinking
0.791
0.526
0.706
0.669
0.370
0.346
0.142
0.275
0.472
Critical Thinking
0.275
0.796
0.531
0.511
0.062
0.192
0.275
0.162
0.336
Table 5 : Operator-wise win-rate evaluation under our LLM-judged protocol. We report dimension-wise win rates for problem finding (motivation) and problem solving (method), along with their aggregated multi-objective scores (MOScore) and overall average. The best and second-best results are labeled with bold and underline.
Method
CV
NLP
Multimodal
Audio & Speech
Robotics
Science
General ML
Theory
Generate
0.302
0.289
0.334
0.257
0.284
0.327
0.281
0.365
GenerateCoT
0.355
0.366
0.307
0.259
0.419
0.348
0.352
0.372
AI Scientist
0.281
0.318
0.348
0.254
0.344
0.388
0.305
0.315
AI-Researcher
0.348
0.293
0.264
0.317
0.335
0.314
0.342
0.422
VIRSCI
0.502
0.450
0.436
0.324
0.533
0.473
0.468
0.568
MindFlow (Ours)
0.503
0.498
0.480
0.398
0.694
0.412
0.576
0.660
Table 6 : Performance across topic domains ( problem finding and problem solving ). We report per-domain scores under our LLM-judged protocol. The best and second-best results are labeled with bold and underline.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
CV
NLP
Multimodal
Audio & Speech
Robotics
Science
General ML
Theory
Divergent +
Convergent Thinking
0.363
0.394
0.374
0.268
0.392
0.359
0.416
0.391
Counterfactual Thinking
0.423
0.400
0.402
0.406
0.490
0.334
0.425
0.539
Analogical Thinking
0.457
0.449
0.434
0.331
0.402
0.409
0.388
0.564
Critical Thinking
0.338
0.350
0.359
0.276
0.345
0.335
0.313
0.433
Constraint-Driven Thinking
0.439
0.473
0.445
0.237
0.628
0.316
0.532
0.601
Appendix
Table 7 : Operator-wise performance across topic domains ( problem finding and problem solving ) under our LLM-judged protocol. We report domain-wise scores for each thinking operator. The best and second-best results are highlighted in bold and underlined, respectively.
Problem Finding (Motivation)
Problem Solving (Method)
Overall
Method
Novelty
Significance
Timeliness
MOScore
Novelty
Effectiveness
Feasibility
MOScore
AIScientist
0.175
0.688
0.425
0.403
0.100
0.213
0.638
0.271
0.337
AI-Researcher
0.600
0.313
0.388
0.428
0.325
0.163
0.275
0.248
0.338
VIRSCI
0.650
0.563
0.575
0.595
0.613
0.450
0.138
0.356
0.475
MindFlow
0.688
0.663
0.650
0.667
0.550
0.538
0.400
0.494
0.580
Appendix
Table 8 : Human expert evaluation results.
Problem Finding (Motivation)
Problem Solving (Method)
Average
Novelty
Significance
Timeliness
Novelty
Effectiveness
Feasibility
Inter-judge Agreement
79.3%
68.7%
76.0%
80.7%
72.0%
72.7%
74.9%
Human-LLM Agreement
82.1%
71.4%
78.8%
83.5%
75.3%
75.6%
77.8%
Appendix
Table 9 : Agreement analysis across inter-judge agreement among three LLM judges and between human experts and LLM judges .
Method
Novelty
Significance
Timeliness
MOScore
VIRSCI
0.711±0.018
0.645±0.021
0.645±0.016
0.667±0.015
MindFlow
0.744±0.015
0.782±0.019
0.702±0.017
0.742±0.012
Appendix
Table 10 : Problem Finding (Motivation) performance (mean ± std over three runs).
Method
Novelty
Effectiveness
Feasibility
MOScore
VIRSCI
0.445±0.023
0.332±0.019
0.109±0.014
0.274±0.017
MindFlow
0.361±0.020
0.408±0.022
0.257±0.018
0.339±0.014
Appendix
Table 11 : Problem Solving (Method) performance (mean ± std over three runs).