Reinforcing Agentic Creativity in Scientific Ideation with Night Science
Organizations: University of Illinois Urbana-Champaign · Microsoft · Microsoft Research
Abstract
Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to loosely structured, serendipitous night science that reaches ideas beyond those typically considered. We introduce AI Night-Scientist, an agentic framework that uses reinforcement learning to teach models when and how to depart from predictable reasoning. Grounded in cognitive science, we model creativity along three axes: action (what to do and how creatively), process (when to explore versus exploit), and outcome (the novelty and usefulness of the resulting idea). We use these axes to train models with GRPO, exposing them to varying degrees and forms of creativity throughout training. This produces substantially more diverse scientific proposals, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model. It also improves predicted citation impact by up to 32.0 percentage points and originality by 66.2 points. These gains cannot be reproduced by simply increasing decoding temperature; instead, we find that semantic guidance specifying what kind of creativity to pursue is critical. Overall, our results suggest that creativity is a learnable, multi-level ability that can be shaped to help researchers reach ideas beyond those typically explored by LLMs.
Figures & tables
| Action | Mechanism & Output | Creativity Levels |
| a Search | Generate query retrieve arXiv papers based on embedding similarity | L1: Search proposal-specific background using title and core terms. L3: Search tangentially related background for broader context. L5: Search distant domains, alternate perspectives, or broader questions. |
| b Debate | Select participants and topic retrieve relevant papers simulate discussion | L1: Discuss proposal specifics with a close-domain colleague. L3: Debate with a peer from the same field but a different topic. L5: Explore with experts from distant disciplines in an open-ended debate. |
| c Spark | Identify assumption ( Bit ) invert it ( Flip ) reframe it ( Spark ) | L1: Challenge a narrow assumption specific to the current proposal. L3: Challenge a meaningful assumption underlying the approach. L5: Challenge a broad, field-level assumption through radical reframing. |
| d Write | Synthesize prior trajectory into a proposal draft or revision | Fixed: consolidate prior trajectory; preserve credit assignment |
| e Stop | End trajectory & return final proposal | Judge final proposal by (precedence), (feasibility), and (relevance) |
| Dimension | Definition | Positive Example (Reward ) | Negative Example (Reward ) |
| Precedence | Distance from prior work and existing approaches. | Core idea opens genuinely new technical directions. | Recombines familiar components without new insight. |
| Feasibility | Credibility & specificity of the execution plan. | Clear methods with precedent or scoped evaluation strategy. | Broad, underspecified plan or implausible scope. |
| Relevance | Alignment of the proposal with problem . | Addresses a critical bottleneck in the target domain. | Drifts off-topic or fails to engage with the core challenges of . |
| Category | Method | Citation (%) | Originality (%) |
| Zero-shot | Llama-3.1-8B | 1.61 | 1.15 |
| Qwen3-8B | 0.46 | 2.53 | |
| Qwen3-14B | 1.15 | 2.53 | |
| GPT-4.1 | 3.90 | 12.18 | |
| Temp. | Qwen3-8B + Temp | 2.07 | 0.69 |
| GPT-4.1 + Temp | 2.99 | 9.89 |
| Method | Key Proposal Direction | Main Limitation |
| GPT-4.1 | Uses AI to segment prerecorded lectures and represent the material through conversational agents , with adaptive pacing, clarification, and multimodal accessibility support. | The proposal is coherent and inclusive, but largely combines established lecture-segmentation and chatbot mechanisms rather than introducing a distinct technical idea. |
| ReAct | Introduces a multimodal engagement score that combines visual, auditory, and physiological signals and uses it to adapt lecture pacing and content granularity in real time. | The score is a concrete artifact, but the proposed RL-based adaptation is underspecified : the state, action, and reward spaces are left unclear for how engagement signals translate into adaptation decisions. |
| Night-8B ( ) | Proposes Agentify , which turns a prerecorded lecture into a live, agent-mediated session where an AI interleaves questions, hints, and dialogue based on learner behavior. | The proposal substantially reframes the interaction, but some components are only loosely motivated or underdeveloped , including the cognitive-tier co-evolution mechanism and forced speed controls. |
| Night-8B ( ) | Proposes Progressive Content Enactors (PCEs), which deliberately inject controlled errors and ambiguities for learners to detect and resolve , shifting the lecture from passive delivery toward active problem solving. | The central idea is distinctive, but the proposal also contains unverifiable quantitative claims and weakly grounded evaluation measures , reducing methodological credibility. |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Proposal Quality | Research-Idea Diversity | |||
| Method | Citation (%) | Originality (%) | Paradigm | Contribution |
| Zero-shot | ||||
| Qwen3-8B | 0.46 [0.13, 1.66] | 2.53 [1.42, 4.47] | 0.738 [0.682, 0.785] | 0.370 [0.340, 0.399] |
| GPT-4.1 | 3.91 [2.45, 6.17] | 12.18 [9.44, 15.59] | 0.785 [0.733, 0.826] | 0.439 [0.401, 0.472] |
| RL Baselines | ||||
| Qwen3-8B + Temp | 11.52 [8.85, 14.87] | 19.82 [16.34, 23.82] | 0.792 [0.742, 0.831] | 0.339 [0.313, 0.364] |
| Domain | Count | % of Labels |
| Computing & AI | 6,418 | 58.0% |
| Artificial Intelligence & Machine Learning | 2558 | 23.1 |
| Networks & Communications | 657 | 5.9 |
| Computer Systems & Architecture | 626 | 5.7 |
| Data Science & Analytics | 611 | 5.5 |
| Cybersecurity & Privacy | 547 | 4.9 |
| Domain | Citation Gain (pp) | Originality Gain (pp) |
| Computing & AI | ||
| Networks & Communications | +40.0 7.3 | +59.6 7.2 |
| Robotics & Autonomous Systems | +34.0 6.9 | +57.4 7.7 |
| Artificial Intelligence & Machine Learning | +29.6 3.4 | +57.0 3.7 |
| Cybersecurity & Privacy | +43.2 7.5 | +55.6 7.7 |
| Computer Systems & Architecture | +30.6 7.0 | +54.0 7.5 |
| Method | Key Proposal Excerpt | Assessment |
| GPT-4.1 | “Develop AI algorithms to automatically segment video lectures into coherent topics and extract key instructional elements. Design and implement conversational agents that re-present segmented lecture content interactively , supporting adaptive pacing, clarifications, and multimodal delivery (text, visuals, sign language, audio descriptions) . Empirically evaluate effectiveness and inclusivity against standard video lectures. Assess scalability and real-world deployment challenges within existing platforms.” | Pros: Well-structured four-phase plan; covers accessibility and inclusivity. Cons: Restates existing segmentation and chatbot approaches without introducing a new mechanism; no novel technical contribution beyond integration. |
| ReAct | “Develop a multimodal AI framework for real-time content adaptation leveraging visual, auditory, and physiological engagement metrics to dynamically adjust lecture content. Introduces a ‘multimodal engagement score’ synthesizing data across modalities . Paired with a content adaptation engine using reinforcement learning to optimize lecture pacing and content granularity [no state, action, or reward space defined]. Validated via controlled A/B study with 120 participants across three engagement conditions.” | Pros: Concrete novel artifact (multimodal engagement score); controlled experimental design. Cons: RL formulation is underspecified; experimental design is elaborate relative to the degree of technical novelty. |
| AI Night-Scientist (Outcome) | “ ‘Agentify’: a framework that transforms passive lectures into live, agent-mediated sessions where AI agents interweave structured questions, personalized hints, and interactive dialogues based on real-time participant behavior. Proposes co-evolution of agent templates and lecture cognitive tiers. Randomizes 200 learners across 40 lectures (STEM, Social Science, Humanities) into four groups comparing passive, over-moderated, fixed-tier, and adaptive-tier conditions. Includes a forced-speed-slider component with 8 granular speed settings (1x–8x) to calibrate attention span variability.” | Pros: Novel reframing of lectures as co-creative sessions rather than content delivery artifacts. Central hypothesis is testable. Ambitious but sensible multi-group experimental design. Cons: Some design elements appear contrived (speed-slider rationale). Cognitive tier co-evolution mechanism is underdeveloped. |
| AI Night-Scientist (Process + Outcome) | “ Progressive Content Enactors (PCEs): virtual agents that shift from automated fidelity to intentional pedagogical maladaptation —deliberately embedding controlled errors and ambiguity so that learners are forced to detect and resolve them, promoting active problem-solving over passive reception . Uses a Layered Semantic Transformative Stack (LISS) to anchor distortions to content structure. Reports cognitive engagement score (NCIR) of , retention improvement (TAR-film, 50:1 decay), and a 2.78- increase in credibility-based efficacy —none of which are independently verifiable or grounded in standard evaluation frameworks.” | Pros: Most novel direction: purposeful errors as a pedagogical tool mirrors well-established ideas (e.g., error-based learning, productive failure). PCE is a fresh, actionable concept that is clearly differentiated from prior work. Cons: Background section relies on unverifiable metrics and internally inconsistent citations. Evaluation framework is not grounded in standard methodology. |
| Description | |
| search | |
| 1 | Extremely relevant to the proposal, focusing on background information and related works based on terms extracted from the proposal title. |
| 2 | Very relevant, focusing on background and related works closely related to the proposal idea. |
| 3 | Somewhat relevant, focusing on background that is tangentially related to the proposal idea. |
| 4 | Not very relevant; explores broader topics or concepts that may not be directly related but still provide useful context. |
| 5 | Very distantly relevant; explores specific alternate perspectives, domains, or philosophical questions of general interest. |
| Metric | Raw Agreement | Cohen’s | Interpretation |
| Impact | 80.0% (12/15) | 0.625 | Substantial |
| Originality | 86.7% (13/15) | 0.766 | Substantial |
| Metric | LLM Judge | Human–LLM Agreement |
| Impact | SciJudge-30B | 76.7% (46/60) |
| Originality | GPT-5.1 | 72.4% (43/60) |
| Hyperparameter | Value |
| Data | |
| Train batch size (prompts) | 32 |
| Max prompt length | 16,384 tokens |
| Max response length | 25,000 tokens |
| Dynamic batching | Enabled |
| Rollout (SGLang) | |