Large language models are increasingly used as interactive writing tools, where users develop stories, revise ideas, and introduce new requirements across multiple turns rather than specifying a complete brief upfront. Yet most evidence on multi-turn instruction degradation comes from tasks with objectively verifiable outcomes, leaving unclear whether incremental interaction harms creative artifacts in ways that explicit requirement checks cannot capture. We study this question using 160 human-authored creative-writing tasks across six genres, presenting each intended specification either upfront or progressively over 5-9 turns to six distinct open-weight model families, yielding 960 matched pairs. Progressive delivery reduces explicit constraint adherence and produces its largest writing-quality degradation in structure/coherence. The structural gap persists among outputs with equal observed adherence, suggesting that measured requirement loss alone does not explain the observed structural difference. We define Creative Integrity as a compact measure of joint adherence and narrative structure; under incremental delivery, models retain 71.2% of FULL Creative Integrity (95% CI [68.2%, 74.3%]). A three-rater human study over 50 matched pairs independently recovers FULL advantages in structure/coherence, craft, and genre effectiveness, while automated scores remain positively associated with aggregated human ratings. These findings show that interactive creative-writing systems should be evaluated not only on whether requirements survive conversation, but also on whether evolving requirements remain coherently integrated into the final artifact. Our dataset, benchmarks, and source code are available at: https://github.com/solusops/SISTER-2026-Team19
Figures & tables
Genre
Tasks
Delivery variants
Final outputs
Matched pairs
Fantasy
20
40
240
120
Historical fiction
20
40
240
120
Mystery
20
40
240
120
Romance
20
40
240
120
Science fiction
20
40
240
120
Comedy
60
120
720
360
Table 1: Benchmark composition and resulting evaluation scale. Each task has Full and Sharded delivery variants and is run on all six models.
Model
FULL CI
SHARDED CI
ICIR
95% CI
Gemma 4 12B
0.647
0.568
87.8%
[82.8, 92.8]
GPT-OSS 20B
0.512
0.384
75.0%
[68.0, 82.4]
Llama 3.1 8B
0.488
0.366
74.9%
[68.8, 81.3]
Qwen 3.5 9B
0.546
0.392
71.7%
[65.7, 78.2]
Mistral 7B
0.408
0.227
55.7%
[50.7, 61.2]
Granite 4 H Tiny
0.325
0.148
45.4%
[39.1, 52.5]
Table 2: Model-wise Creative Integrity (CI) under each delivery condition and Incremental Creative Integrity Retention (ICIR), with 95% story-clustered bootstrap confidence intervals for ICIR.
Dimension
N
FULL
SHARDED
Δ (FULL − SHARDED)
95% CI
Constraint adherence
960
0.852
0.742
+ 0.111
[0.096, 0.126]
Craft
960
3.044
2.796
+ 0.248
[0.200, 0.295]
Structure/coherence
960
3.231
2.755
+ 0.476
[0.413, 0.542]
Originality
960
2.447
2.525
− 0.078
[ − 0.123, − 0.032]
Genre effectiveness
960
3.125
2.900
+ 0.225
[0.168, 0.280]
Characterization
945
2.634
2.580
+ 0.054
[0.004, 0.104]
Table 3: Raw paired Full / Sharded means, the raw paired difference Δ , and its 95% story-clustered bootstrap confidence interval, over all matched pairs (constraint adherence has a 0–1 scale; the five quality dimensions are on a 1–5 scale). Characterization is scored for fewer pairs as it is not applicable for all items ( section 3.3 ).
Dimension
All pairs
Equal adherence
Craft
+ 0.248
0.000
Structure/coherence
+ 0.476
+ 0.239
Originality
− 0.078
− 0.361
Genre effectiveness
+ 0.225
− 0.133
Characterization
+ 0.054
− 0.233
Table 4: Paired Full − Sharded quality effects before and after restricting to pairs with equal observed constraint adherence (N=180). Positive values favor Full .
Dimension
Automated Δ
Human Δ
Human 95% CI
Craft
+ 0.248
+ 0.460
[0.240, 0.673]
Structure/coherence
+ 0.476
+ 0.527
[0.313, 0.740]
Originality
− 0.078
+ 0.187
[ − 0.007, 0.380]
Genre effectiveness
+ 0.225
+ 0.347
[0.080, 0.607]
Characterization
+ 0.054
+ 0.193
[ − 0.013, 0.397]
Table 5: Automated and human paired Full − Sharded effects across writing-quality dimensions. Human effects use 50 matched pairs. Positive values favor Full .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Dimension
Ordinal α
Exact agreement
Mean abs. disagreement
Craft
0.232
0.390
0.840
Structure/coherence
0.201
0.340
0.927
Originality
0.036
0.303
1.053
Genre effectiveness
0.213
0.327
0.927
Characterization
0.077
0.313
1.056
Appendix
Table 6: Inter-rater statistics for pointwise human annotations. Each dimension covers 100 responses and three raters; characterization has 292 valid ratings because eight ratings are N/A, compared with 300 for the other dimensions.
Dimension
Spearman ρ
95% CI
MAE
Craft
0.624
[0.505, 0.723]
0.497
Structure/coherence
0.633
[0.500, 0.746]
0.530
Originality
0.437
[0.269, 0.584]
0.810
Genre effectiveness
0.497
[0.329, 0.635]
0.627
Characterization
0.506
[0.319, 0.663]
0.633
Appendix
Table 7: Automated-pointwise alignment with mean human ratings. Confidence intervals resample the 50 matched pairs.
Dimension
FULL
SHARDED
Δ (95% CI)
F/T/S
Craft
3.560
3.100
+ 0.460 [0.240, 0.673]
34/6/10
Structure/coherence
3.513
2.987
+ 0.527 [0.313, 0.740]
35/5/10
Originality
3.260
3.073
+ 0.187 [ − 0.007, 0.380]
27/5/18
Genre effectiveness
3.640
3.293
+ 0.347 [0.080, 0.607]
30/5/15
Characterization
3.197
3.003
+ 0.193 [ − 0.013, 0.397]
27/7/16
Appendix
Table 8: Human paired effects over 50 matched pairs. F/T/S reports pairs favoring FULL, tied, and favoring SHARDED.
Dimension
Median Δ (95% CI)
Reviewer-centered Δ (95% CI)
Craft
+ 0.460 [0.220, 0.700]
+ 0.460 [0.240, 0.673]
Structure/coherence
+ 0.600 [0.320, 0.880]
+ 0.527 [0.313, 0.740]
Originality
+ 0.240 [0.000, 0.480]
+ 0.187 [ − 0.007, 0.380]
Genre effectiveness
+ 0.420 [0.160, 0.680]
+ 0.347 [0.080, 0.607]
Characterization
+ 0.300 [0.090, 0.520]
+ 0.208 [0.001, 0.412]
Appendix
Table 9: Sensitivity of human paired effects to response-level aggregation. Positive values favor Full ; confidence intervals resample the 50 matched pairs.
Outcome
Equal-genre FULL − SHARDED
95% CI
Adherence
+ 0.119
[0.103, 0.136]
Craft
+ 0.276
[0.226, 0.327]
Structure/coherence
+ 0.493
[0.424, 0.561]
Originality
− 0.056
[ − 0.102, − 0.007]
Genre effectiveness
+ 0.277
[0.223, 0.331]
Characterization
+ 0.090
[0.038, 0.140]
Appendix
Table 10: Equal-genre sensitivity analysis. Positive values favor Full .
Figure 2: Model-wise standardized FULL − SHARDED effect dav,m ( eq. 5 , computed per model over 160 story pairs) for constraint adherence and each writing-quality dimension. Blue cells favor FULL and orange cells favor SHARDED.
Figure 3: Final-generation diagnostics by model and condition (N=160 per box). Boxes show interquartile ranges, horizontal lines show medians, and whiskers extend to 1.5 times the interquartile range.
Dimension
Directional agreement
Quadratic-weighted κ
Constraint following
63.3%
0.544
Creative-writing quality
40.0%
0.078
Appendix
Table 11: Human–model agreement for the simple pairwise LLM evaluator over 30 blinded pairs.
Figure 4: Pairwise evaluator diagnostics over the fixed 30-pair blinded human sample. (A) Directional agreement with human judgments. (B) Order-swap directional consistency; the starred annotation gives evidence-first atomic constraint-status stability.
Model
Publisher
Params
Quant.
Context (tok.)
GPT-OSS 20B
OpenAI
20B
MXFP4
16,384
Gemma 4 12B QAT
Google
12B
Q4_0
16,384
Llama 3.1 8B Instruct
Meta
8B
Q4_K_M
16,384
Granite 4 H Tiny
IBM
7B
Q4_K_M
16,384
Mistral 7B Instruct v0.3
Mistral AI
7B
Q4_K_M
16,384
Qwen3.5 9B
Qwen
9B
Q4_K_M
16,384 / 22,528 †
Appendix
Table 12: Models evaluated in the benchmark. All models were run locally with the listed 4-bit quantized variants.
Model
Official source model card
Selected local variant
License or terms
GPT-OSS 20B
openai/gpt-oss-20b
openai/gpt-oss-20b@mxfp4
Apache-2.0
Gemma 4 12B QAT
Google Gemma 4
google/gemma-4-12b-qat@q4_0
Apache-2.0
Llama 3.1 8B Instruct
meta-llama/Llama-3.1-8B-Instruct
llama-3.1-8b-instruct@q4_k_m
Llama 3.1 Community License
Granite 4 H Tiny
ibm-granite/granite-4.0-h-tiny
ibm/granite-4-h-tiny@q4_k_m
Apache-2.0
Mistral 7B Instruct v0.3
mistralai/Mistral-7B-Instruct-v0.3
mistralai/mistral-7b-instruct-v0.3@q4_k_m
Apache-2.0
Qwen3.5 9B
Qwen/Qwen3.5-9B
qwen/qwen3.5-9b@q4_k_m
Apache-2.0
Appendix
Table 13: Source-model and license documentation. “Selected variant” is the exact local identifier recorded in the run index.
While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.g., story generation) as well as those that require it implicitly, e.g., reinforcement learning (RL). We instead propose CreativeInstruct, a scalable instruction-tuning method that teaches LLMs to balance creative, base-model-like generations with the quality of post-trained models, by learning to inject special [StartCreativity] spans that bias generation toward creativity. Furthermore, we introduce a structural diversity metric based on graph edit distance, which captures narrative level variation missed by purely lexical and semantic metrics. On narrative generation, CreativeInstruct matches or exceeds the diversity of both multi-model baselines and distilled variants of their outputs, without sacrificing quality or requiring multiple models at inference time. These results are mirrored in our human evaluation, where we find that annotators rate CreativeInstruct generations as more creative than the post-trained LLMs' generations in 70.3% of cases. We also show the benefits of creative models as a substrate for RL: GRPO applied to a CreativeInstruct checkpoint improves by ~4% on AMC and ~5% points on MATH over the same training applied to the post-trained checkpoint.
Ananya Sahu, Mohit Bansal, Elias Stengel-Eskin
Columbia · UNC Chapel Hill · University of Texas at Austin
Large Language Models (LLMs) are being applied to increasingly difficult problems and use cases. To navigate their vast solution spaces effectively, LLMs need to be creative. Yet the subjective nature of creativity and the limits of human judgment make training LLMs for creativity especially challenging. As a solution, we train LLMs on Codenames, a word-association game that exercises the two central axes of creativity, divergent and convergent thinking, while yielding objectively verifiable outcomes. This verifiability lets us bypass human judgment and train with Reinforcement Learning with Verifiable Rewards (RLVR). We train Qwen3-1.7B, 4B, and 8B models and evaluate them on ten creativity and four reasoning benchmarks. We find that the precision-diversity trade-off is scale-dependent: the 8B model prioritizes creativity over precision, while the 1.7B and 4B models gain reasoning precision at the cost of creativity. Concretely, the 8B model shows modest but consistent creativity gains (8 of 10 benchmarks) with only minor reasoning degradation, whereas the smaller models achieve substantial gains on reasoning tasks. Our study presents a scalable and effective solution to train LLMs for creativity.
Large language models are optimized for instruction following and agentic tasks remain poorly aligned with the requirements of high-quality creative writing. We show that a purpose-built creative writing model can outperform both GPT-5.5 and Claude Opus 4.8 on writing quality evaluation. Fiction frequently depends on behaviors that assistant-tuned models are explicitly trained to avoid, particularly deception, moral ambiguity, and unreliable narration. As a result, generated stories often appear structurally correct while remaining stylistically generic, overly explanatory, or weakly grounded in human literary behavior. We present a dataset construction and training framework for book-scale creative writing that reframes supervised fine-tuning as a prompt-to-book generation task grounded in human-authored fiction. Starting from public-domain novels, we derive a multi-resolution Planning Scaffold by summarizing each book at progressively finer levels, from a high-level premise to chapter- and scene-level structure. We then invert this hierarchy during training: the model learns to expand a prompt into increasingly detailed plans and finally into the original human-authored book text. This formulation preserves human prose as the final supervised target while using intermediate summaries to make book-scale generation learnable. We train a long-context language model on these prompt-to-book trajectories and show that this objective shifts generation away from assistant-style prose and toward human literary writing.