Large language models are increasingly used as interactive writing tools, where users develop stories, revise ideas, and introduce new requirements across multiple turns rather than specifying a complete brief upfront. Yet most evidence on multi-turn instruction degradation comes from tasks with objectively verifiable outcomes, leaving unclear whether incremental interaction harms creative artifacts in ways that explicit requirement checks cannot capture. We study this question using 160 human-authored creative-writing tasks across six genres, presenting each intended specification either upfront or progressively over 5-9 turns to six distinct open-weight model families, yielding 960 matched pairs. Progressive delivery reduces explicit constraint adherence and produces its largest writing-quality degradation in structure/coherence. The structural gap persists among outputs with equal observed adherence, suggesting that measured requirement loss alone does not explain the observed structural difference. We define Creative Integrity as a compact measure of joint adherence and narrative structure; under incremental delivery, models retain 71.2% of FULL Creative Integrity (95% CI [68.2%, 74.3%]). A three-rater human study over 50 matched pairs independently recovers FULL advantages in structure/coherence, craft, and genre effectiveness, while automated scores remain positively associated with aggregated human ratings. These findings show that interactive creative-writing systems should be evaluated not only on whether requirements survive conversation, but also on whether evolving requirements remain coherently integrated into the final artifact. Our dataset, benchmarks, and source code are available at: https://github.com/solusops/SISTER-2026-Team19
Figures & tables
Genre
Tasks
Delivery variants
Final outputs
Matched pairs
Fantasy
20
40
240
120
Historical fiction
20
40
240
120
Mystery
20
40
240
120
Romance
20
40
240
120
Science fiction
20
40
240
120
Comedy
60
120
720
360
Table 1: Benchmark composition and resulting evaluation scale. Each task has Full and Sharded delivery variants and is run on all six models.
Model
FULL CI
SHARDED CI
ICIR
95% CI
Gemma 4 12B
0.647
0.568
87.8%
[82.8, 92.8]
GPT-OSS 20B
0.512
0.384
75.0%
[68.0, 82.4]
Llama 3.1 8B
0.488
0.366
74.9%
[68.8, 81.3]
Qwen 3.5 9B
0.546
0.392
71.7%
[65.7, 78.2]
Mistral 7B
0.408
0.227
55.7%
[50.7, 61.2]
Granite 4 H Tiny
0.325
0.148
45.4%
[39.1, 52.5]
Table 2: Model-wise Creative Integrity (CI) under each delivery condition and Incremental Creative Integrity Retention (ICIR), with 95% story-clustered bootstrap confidence intervals for ICIR.
Dimension
N
FULL
SHARDED
Δ (FULL − SHARDED)
95% CI
Constraint adherence
960
0.852
0.742
+ 0.111
[0.096, 0.126]
Craft
960
3.044
2.796
+ 0.248
[0.200, 0.295]
Structure/coherence
960
3.231
2.755
+ 0.476
[0.413, 0.542]
Originality
960
2.447
2.525
− 0.078
[ − 0.123, − 0.032]
Genre effectiveness
960
3.125
2.900
+ 0.225
[0.168, 0.280]
Characterization
945
2.634
2.580
+ 0.054
[0.004, 0.104]
Table 3: Raw paired Full / Sharded means, the raw paired difference Δ , and its 95% story-clustered bootstrap confidence interval, over all matched pairs (constraint adherence has a 0–1 scale; the five quality dimensions are on a 1–5 scale). Characterization is scored for fewer pairs as it is not applicable for all items ( section 3.3 ).
Dimension
All pairs
Equal adherence
Craft
+ 0.248
0.000
Structure/coherence
+ 0.476
+ 0.239
Originality
− 0.078
− 0.361
Genre effectiveness
+ 0.225
− 0.133
Characterization
+ 0.054
− 0.233
Table 4: Paired Full − Sharded quality effects before and after restricting to pairs with equal observed constraint adherence (N=180). Positive values favor Full .
Dimension
Automated Δ
Human Δ
Human 95% CI
Craft
+ 0.248
+ 0.460
[0.240, 0.673]
Structure/coherence
+ 0.476
+ 0.527
[0.313, 0.740]
Originality
− 0.078
+ 0.187
[ − 0.007, 0.380]
Genre effectiveness
+ 0.225
+ 0.347
[0.080, 0.607]
Characterization
+ 0.054
+ 0.193
[ − 0.013, 0.397]
Table 5: Automated and human paired Full − Sharded effects across writing-quality dimensions. Human effects use 50 matched pairs. Positive values favor Full .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Dimension
Ordinal α
Exact agreement
Mean abs. disagreement
Craft
0.232
0.390
0.840
Structure/coherence
0.201
0.340
0.927
Originality
0.036
0.303
1.053
Genre effectiveness
0.213
0.327
0.927
Characterization
0.077
0.313
1.056
Appendix
Table 6: Inter-rater statistics for pointwise human annotations. Each dimension covers 100 responses and three raters; characterization has 292 valid ratings because eight ratings are N/A, compared with 300 for the other dimensions.
Dimension
Spearman ρ
95% CI
MAE
Craft
0.624
[0.505, 0.723]
0.497
Structure/coherence
0.633
[0.500, 0.746]
0.530
Originality
0.437
[0.269, 0.584]
0.810
Genre effectiveness
0.497
[0.329, 0.635]
0.627
Characterization
0.506
[0.319, 0.663]
0.633
Appendix
Table 7: Automated-pointwise alignment with mean human ratings. Confidence intervals resample the 50 matched pairs.
Dimension
FULL
SHARDED
Δ (95% CI)
F/T/S
Craft
3.560
3.100
+ 0.460 [0.240, 0.673]
34/6/10
Structure/coherence
3.513
2.987
+ 0.527 [0.313, 0.740]
35/5/10
Originality
3.260
3.073
+ 0.187 [ − 0.007, 0.380]
27/5/18
Genre effectiveness
3.640
3.293
+ 0.347 [0.080, 0.607]
30/5/15
Characterization
3.197
3.003
+ 0.193 [ − 0.013, 0.397]
27/7/16
Appendix
Table 8: Human paired effects over 50 matched pairs. F/T/S reports pairs favoring FULL, tied, and favoring SHARDED.
Dimension
Median Δ (95% CI)
Reviewer-centered Δ (95% CI)
Craft
+ 0.460 [0.220, 0.700]
+ 0.460 [0.240, 0.673]
Structure/coherence
+ 0.600 [0.320, 0.880]
+ 0.527 [0.313, 0.740]
Originality
+ 0.240 [0.000, 0.480]
+ 0.187 [ − 0.007, 0.380]
Genre effectiveness
+ 0.420 [0.160, 0.680]
+ 0.347 [0.080, 0.607]
Characterization
+ 0.300 [0.090, 0.520]
+ 0.208 [0.001, 0.412]
Appendix
Table 9: Sensitivity of human paired effects to response-level aggregation. Positive values favor Full ; confidence intervals resample the 50 matched pairs.
Outcome
Equal-genre FULL − SHARDED
95% CI
Adherence
+ 0.119
[0.103, 0.136]
Craft
+ 0.276
[0.226, 0.327]
Structure/coherence
+ 0.493
[0.424, 0.561]
Originality
− 0.056
[ − 0.102, − 0.007]
Genre effectiveness
+ 0.277
[0.223, 0.331]
Characterization
+ 0.090
[0.038, 0.140]
Appendix
Table 10: Equal-genre sensitivity analysis. Positive values favor Full .
Figure 2: Model-wise standardized FULL − SHARDED effect dav,m ( eq. 5 , computed per model over 160 story pairs) for constraint adherence and each writing-quality dimension. Blue cells favor FULL and orange cells favor SHARDED.
Figure 3: Final-generation diagnostics by model and condition (N=160 per box). Boxes show interquartile ranges, horizontal lines show medians, and whiskers extend to 1.5 times the interquartile range.
Dimension
Directional agreement
Quadratic-weighted κ
Constraint following
63.3%
0.544
Creative-writing quality
40.0%
0.078
Appendix
Table 11: Human–model agreement for the simple pairwise LLM evaluator over 30 blinded pairs.
Figure 4: Pairwise evaluator diagnostics over the fixed 30-pair blinded human sample. (A) Directional agreement with human judgments. (B) Order-swap directional consistency; the starred annotation gives evidence-first atomic constraint-status stability.
Model
Publisher
Params
Quant.
Context (tok.)
GPT-OSS 20B
OpenAI
20B
MXFP4
16,384
Gemma 4 12B QAT
Google
12B
Q4_0
16,384
Llama 3.1 8B Instruct
Meta
8B
Q4_K_M
16,384
Granite 4 H Tiny
IBM
7B
Q4_K_M
16,384
Mistral 7B Instruct v0.3
Mistral AI
7B
Q4_K_M
16,384
Qwen3.5 9B
Qwen
9B
Q4_K_M
16,384 / 22,528 †
Appendix
Table 12: Models evaluated in the benchmark. All models were run locally with the listed 4-bit quantized variants.
Model
Official source model card
Selected local variant
License or terms
GPT-OSS 20B
openai/gpt-oss-20b
openai/gpt-oss-20b@mxfp4
Apache-2.0
Gemma 4 12B QAT
Google Gemma 4
google/gemma-4-12b-qat@q4_0
Apache-2.0
Llama 3.1 8B Instruct
meta-llama/Llama-3.1-8B-Instruct
llama-3.1-8b-instruct@q4_k_m
Llama 3.1 Community License
Granite 4 H Tiny
ibm-granite/granite-4.0-h-tiny
ibm/granite-4-h-tiny@q4_k_m
Apache-2.0
Mistral 7B Instruct v0.3
mistralai/Mistral-7B-Instruct-v0.3
mistralai/mistral-7b-instruct-v0.3@q4_k_m
Apache-2.0
Qwen3.5 9B
Qwen/Qwen3.5-9B
qwen/qwen3.5-9b@q4_k_m
Apache-2.0
Appendix
Table 13: Source-model and license documentation. “Selected variant” is the exact local identifier recorded in the run index.