cs.CLSep 27, 2026

The Effects of Incremental Instruction Delivery on Language-Model Creative Writing

Authors: Anshuman Singh, Abrar Eyasir, Haseeb Yaqoob, John Manavalan

Organizations: SGT UNIVERSITY · CSE, University of Dhaka · NED University of Engineering and Technology · Metea Valley High School, Illinois

Abstract

Large language models are increasingly used as interactive writing tools, where users develop stories, revise ideas, and introduce new requirements across multiple turns rather than specifying a complete brief upfront. Yet most evidence on multi-turn instruction degradation comes from tasks with objectively verifiable outcomes, leaving unclear whether incremental interaction harms creative artifacts in ways that explicit requirement checks cannot capture. We study this question using 160 human-authored creative-writing tasks across six genres, presenting each intended specification either upfront or progressively over 5-9 turns to six distinct open-weight model families, yielding 960 matched pairs. Progressive delivery reduces explicit constraint adherence and produces its largest writing-quality degradation in structure/coherence. The structural gap persists among outputs with equal observed adherence, suggesting that measured requirement loss alone does not explain the observed structural difference. We define Creative Integrity as a compact measure of joint adherence and narrative structure; under incremental delivery, models retain 71.2% of FULL Creative Integrity (95% CI [68.2%, 74.3%]). A three-rater human study over 50 matched pairs independently recovers FULL advantages in structure/coherence, craft, and genre effectiveness, while automated scores remain positively associated with aggregated human ratings. These findings show that interactive creative-writing systems should be evaluated not only on whether requirements survive conversation, but also on whether evolving requirements remain coherently integrated into the final artifact. Our dataset, benchmarks, and source code are available at: https://github.com/solusops/SISTER-2026-Team19

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Aug 7, 2026cs.CL

CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity

While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.g., story generation) as well as those that require it implicitly, e.g., reinforcement learning (RL). We instead propose CreativeInstruct, a scalable instruction-tuning method that teaches LLMs to balance creative, base-model-like generations with the quality of post-trained models, by learning to inject special [StartCreativity] spans that bias generation toward creativity. Furthermore, we introduce a structural diversity metric based on graph edit distance, which captures narrative level variation missed by purely lexical and semantic metrics. On narrative generation, CreativeInstruct matches or exceeds the diversity of both multi-model baselines and distilled variants of their outputs, without sacrificing quality or requiring multiple models at inference time. These results are mirrored in our human evaluation, where we find that annotators rate CreativeInstruct generations as more creative than the post-trained LLMs' generations in 70.3% of cases. We also show the benefits of creative models as a substrate for RL: GRPO applied to a CreativeInstruct checkpoint improves by ~4% on AMC and ~5% points on MATH over the same training applied to the post-trained checkpoint.
May 27, 2026cs.CL

Playing with Words, Improving with Rewards: Training Language Models for Creative Association

Large Language Models (LLMs) are being applied to increasingly difficult problems and use cases. To navigate their vast solution spaces effectively, LLMs need to be creative. Yet the subjective nature of creativity and the limits of human judgment make training LLMs for creativity especially challenging. As a solution, we train LLMs on Codenames, a word-association game that exercises the two central axes of creativity, divergent and convergent thinking, while yielding objectively verifiable outcomes. This verifiability lets us bypass human judgment and train with Reinforcement Learning with Verifiable Rewards (RLVR). We train Qwen3-1.7B, 4B, and 8B models and evaluate them on ten creativity and four reasoning benchmarks. We find that the precision-diversity trade-off is scale-dependent: the 8B model prioritizes creativity over precision, while the 1.7B and 4B models gain reasoning precision at the cost of creativity. Concretely, the 8B model shows modest but consistent creativity gains (8 of 10 benchmarks) with only minor reasoning degradation, whereas the smaller models achieve substantial gains on reasoning tasks. Our study presents a scalable and effective solution to train LLMs for creativity.
May 16, 2026cs.AI

Towards Human-Level Book-Writing Capability

Large language models are optimized for instruction following and agentic tasks remain poorly aligned with the requirements of high-quality creative writing. We show that a purpose-built creative writing model can outperform both GPT-5.5 and Claude Opus 4.8 on writing quality evaluation. Fiction frequently depends on behaviors that assistant-tuned models are explicitly trained to avoid, particularly deception, moral ambiguity, and unreliable narration. As a result, generated stories often appear structurally correct while remaining stylistically generic, overly explanatory, or weakly grounded in human literary behavior. We present a dataset construction and training framework for book-scale creative writing that reframes supervised fine-tuning as a prompt-to-book generation task grounded in human-authored fiction. Starting from public-domain novels, we derive a multi-resolution Planning Scaffold by summarizing each book at progressively finer levels, from a high-level premise to chapter- and scene-level structure. We then invert this hierarchy during training: the model learns to expand a prompt into increasingly detailed plans and finally into the original human-authored book text. This formulation preserves human prose as the final supervised target while using intermediate summaries to make book-scale generation learnable. We train a long-context language model on these prompt-to-book trajectories and show that this objective shifts generation away from assistant-style prose and toward human literary writing.