Large language models (LLMs) can generate accurate responses, but these are void of diversity. We attempt to address this for the task of creative short story generation. Drawing on established writing conventions and known LLM limitations, we target variation in genre, tone, style, and named entities. To promote diversity across these dimensions, we introduce DivLM, an LLM post-training framework consisting of two phases. First, we perform continued pre-training on a creative writing corpus and restore instruction-following capabilities using weight residuals. We then apply reinforcement learning with a custom, composite reward function that jointly maximizes diversity across the targeted narrative dimensions while maintaining response quality. Our empirical results on two LLM families show that DivLM increases diversity metrics by more than 9% on average compared to alternative approaches, while preserving instruction following, overall response quality, and similarity to human outputs.
Figures & tables
Figure 1: Diversity absence (mode collapse) for Llama-3.1-8B-Instruct across generations for the same prompt.
Figure 2: Overview of DivLM , our proposed framework for improving diversity in short story generation, consisting of the CPT (section 4.1 ) and RL (section 4.2 ) phases.
Llama-3.1-8B
Gemma-2-9B
Metric
INS
CPT
DQO
DDPO
DLM
INS
CPT
DLM
Quality
Coherence
.997
.999
.999
.886
.997
1
.998
.995
Prompt adherence
.997
.995
.999
.834
.981
.999
.998
.990
No meta-comments
.968
.865
.952
.399
.998
.983
.956
.998
Completeness
.726
.111
.983
.017
.985
.633
.128
.979
Avg
.922
.742
.983
.534
.990
.904
.770
.990
Table 1: Comparison of the original Instruct (INS) LLM variants with CPT-tuned and DivLM (shortened to DLM ) models across quality and diversity metrics ranging in [0,1] . For Llama-3.1-8B , we include DQO and DDPO. To allow the calculation of meaningful averages, we report the no meta-commentary rate. Greater ( ↑ ) values indicate better performance.
Figure 3: Quality (left) and diversity (right) evaluation results for increasing proportions of the instruction residual ( Δθ ) using Llama-3.1-8B .
Llama-3.1-8B
Metric
INS
GRPO
DLM cos
DLM
Quality
Coherence
.997
.999
.999
.997
Prompt adherence
.997
.996
.987
.981
No meta-commentary
.968
.996
.999
.998
Completeness
.726
.986
.850
.985
Avg
.922
.994
.959
.990
Table 2: Ablation study based on Llama-3.1-8B . We compare the following model variants: Instruct (INS), GRPO-only, DLM cos , and DivLM (shortened to DLM ) across quality and diversity metrics ranging in [0,1] . Greater ( ↑ ) values indicate better performance.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1: Examples of semantically similar prompts identified during deduplication.
Figure S2: Genre, tone, and style diversity curves for different reward formulations on Llama. The score is the sum of the genre, tone, and style rewards, resulting in a range of [0,3] . (a) Training scores. (b) Validation scores evaluated every 500 steps.
Formulation
Score
Continuous
1.634
Proximity
1.712
Proximity + Penalty
1.917
Appendix
Table S1: Reward performance of different diversity reward formulations on the test set. The score ( ∈[0,3] ) is the sum of the genre, tone, and style rewards.
Variant
Avg. unique entities per group
Avg. repeated entities per group
INS
28
4.7
Original
62
2.8
Exponential
21
2.1
Clipped
26
1.5
Penalty
49
2.4
Appendix
Table S2: Results on the test set for different named entity diversity formulations with G=10 on the Llama-3.1-8B model.
Metric
Human-human
Judge-human
Quality metrics
Prompt adherence
92.30%
90.40%
Meta-commentary
97.60%
95.70%
Completeness
100.00%
94.10%
Coherence
97.00%
92.70%
Average
96.73%
93.23%
Appendix
Table S3: Human-human and judge-human agreement rates for quality and diversity metrics.
Figure S3: Validation results for different instruction residual scales s on Llama-3.1-8B . Diversity metrics consist of genre, tone, and style coverage values, as well as the named entity diversity reward Rne . All metrics are reported on a [0,1] scale.
Metric
GRPO
PPO
Average quality score
.894
.983
Average Genre/Tone/Style coverage
.454
.447
Average diversity reward
.360
.380
Cosine distance
.118
.135
Appendix
Table S4: Performance comparison between the GRPO and PPO variants of DQO.
Figure S4: Genre, tone, and style rewards obtained for different instruction residual scales s on Llama-3.1-8B . Rewards are reported on a [0,1] scale.
Figure S5: Training curves for Llama-3.1-8B under different values of s , showing the total reward R∈[−1,2] across training steps.
Figure S6: Training curves for Llama-3.1-8B under different values of h . (a) Percentage change in the entity diversity reward over training. (b) Average number of repeated entities per group over training.
Metric
INS
h=1
h=2
h=3
h=4
h=5
Ratio ↓
.107
.117
.066
.068
.078
.053
μ(Rne)↑
.581
.686
.838
.873
.813
.902
Appendix
Table S5: Entity repetition ratio (repeated entities per group over total entities per group) and mean Rne for INS and DivLM across values of h . Final evaluation uses h=5 for consistency.
Model
T=0.8
T=0.9
T=1.0
T=1.1
T=1.2
INS
.116
.118
.125
.131
.138
DivLM
.215
.225
.226
.231
.236
Appendix
Table S6: Average cosine distance on the test set for INS and DivLM under different values of T .
Figure S7: Mean diversity reward ( Rdiv ) on the test set for DivLM and INS across decoding temperatures T on Llama-3.1-8B .
Figure S8: DivLM diversity evaluation results across decoding temperatures T on Llama-3.1-8B .
Model
G=5
G=10
G=20
G=50
G=100
INS
.123
.125
.126
.126
.126
DivLM
.227
.226
.226
.226
.226
Appendix
Table S7: Average cosine distance on the test set for INS and DivLM under different values of G .
Figure S9: Mean diversity reward ( Rdiv ) on the test set for DivLM and INS across group sizes G on Llama-3.1-8B .
Figure S10: DivLM diversity evaluation results across group sizes G on Llama-3.1-8B .
Setting
INS
DivLM
Decay = 50
481.2
366.3
Decay = 150
481.2
410.0
Linear
481.2
483.6
Appendix
Table S8: Mean output length in words for the INS and DivLM models using different length penalties.
While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.g., story generation) as well as those that require it implicitly, e.g., reinforcement learning (RL). We instead propose CreativeInstruct, a scalable instruction-tuning method that teaches LLMs to balance creative, base-model-like generations with the quality of post-trained models, by learning to inject special [StartCreativity] spans that bias generation toward creativity. Furthermore, we introduce a structural diversity metric based on graph edit distance, which captures narrative level variation missed by purely lexical and semantic metrics. On narrative generation, CreativeInstruct matches or exceeds the diversity of both multi-model baselines and distilled variants of their outputs, without sacrificing quality or requiring multiple models at inference time. These results are mirrored in our human evaluation, where we find that annotators rate CreativeInstruct generations as more creative than the post-trained LLMs' generations in 70.3% of cases. We also show the benefits of creative models as a substrate for RL: GRPO applied to a CreativeInstruct checkpoint improves by ~4% on AMC and ~5% points on MATH over the same training applied to the post-trained checkpoint.
Ananya Sahu, Mohit Bansal, Elias Stengel-Eskin
Columbia · UNC Chapel Hill · University of Texas at Austin
Recent advances in large language models (LLMs) have enabled the generation of high-quality prose, yet whether these models are capable of generating diverse or creative artifacts remains a contested question. In this work, we investigate the diversity of LLM-generated stories through the framework of narrative similarity. Using a contrastive framework and a dataset of human-written stories and prompts from r/WritingPrompts, we collect narrative similarity judgments across 10 representative LLMs, utilizing both human evaluations and three different automatic annotation methods. Our findings reveal a clear trend: LLM-generated narratives are consistently more similar to each other than human-written stories are. We demonstrate that frontier models in particular converge on a "mean" generic narrative that approximates individual human stories but lacks the collective diversity of human authors. Finally, we show that common mitigation strategies, including negative prompting and temperature scaling, fail to meaningfully address this homogeneity.
LLM-generated stories are a popular use case, but they show very low variability. We sample 20,000 total stories from four current models using five prompts. We find that 11 words occur in 88.3% of generated stories, with little difference between models. These words include names (Elias, Mara, Elara), settings (lighthouses), and professions (clockmaker, librarian). These tokens do not often occur in published literature nor pre-training data, but they are found in preference data that is likely to have been used by all current models. Surprisingly, these "lighthouse" stories are infrequent when compared with the average post-training story, much of which contains references to copyrighted characters or adult content. This result demonstrates the potentially disproportionate impact of small datasets combined with powerful alignment algorithms.
Sil Hamilton, David Mimno
Department of Information Science Cornell University