Large language models (LLMs) can generate accurate responses, but these are void of diversity. We attempt to address this for the task of creative short story generation. Drawing on established writing conventions and known LLM limitations, we target variation in genre, tone, style, and named entities. To promote diversity across these dimensions, we introduce DivLM, an LLM post-training framework consisting of two phases. First, we perform continued pre-training on a creative writing corpus and restore instruction-following capabilities using weight residuals. We then apply reinforcement learning with a custom, composite reward function that jointly maximizes diversity across the targeted narrative dimensions while maintaining response quality. Our empirical results on two LLM families show that DivLM increases diversity metrics by more than 9% on average compared to alternative approaches, while preserving instruction following, overall response quality, and similarity to human outputs.
Figures & tables
Figure 1: Diversity absence (mode collapse) for Llama-3.1-8B-Instruct across generations for the same prompt.
Figure 2: Overview of DivLM , our proposed framework for improving diversity in short story generation, consisting of the CPT (section 4.1 ) and RL (section 4.2 ) phases.
Llama-3.1-8B
Gemma-2-9B
Metric
INS
CPT
DQO
DDPO
DLM
INS
CPT
DLM
Quality
Coherence
.997
.999
.999
.886
.997
1
.998
.995
Prompt adherence
.997
.995
.999
.834
.981
.999
.998
.990
No meta-comments
.968
.865
.952
.399
.998
.983
.956
.998
Completeness
.726
.111
.983
.017
.985
.633
.128
.979
Avg
.922
.742
.983
.534
.990
.904
.770
.990
Table 1: Comparison of the original Instruct (INS) LLM variants with CPT-tuned and DivLM (shortened to DLM ) models across quality and diversity metrics ranging in [0,1] . For Llama-3.1-8B , we include DQO and DDPO. To allow the calculation of meaningful averages, we report the no meta-commentary rate. Greater ( ↑ ) values indicate better performance.
Figure 3: Quality (left) and diversity (right) evaluation results for increasing proportions of the instruction residual ( Δθ ) using Llama-3.1-8B .
Llama-3.1-8B
Metric
INS
GRPO
DLM cos
DLM
Quality
Coherence
.997
.999
.999
.997
Prompt adherence
.997
.996
.987
.981
No meta-commentary
.968
.996
.999
.998
Completeness
.726
.986
.850
.985
Avg
.922
.994
.959
.990
Table 2: Ablation study based on Llama-3.1-8B . We compare the following model variants: Instruct (INS), GRPO-only, DLM cos , and DivLM (shortened to DLM ) across quality and diversity metrics ranging in [0,1] . Greater ( ↑ ) values indicate better performance.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1: Examples of semantically similar prompts identified during deduplication.
Figure S2: Genre, tone, and style diversity curves for different reward formulations on Llama. The score is the sum of the genre, tone, and style rewards, resulting in a range of [0,3] . (a) Training scores. (b) Validation scores evaluated every 500 steps.
Formulation
Score
Continuous
1.634
Proximity
1.712
Proximity + Penalty
1.917
Appendix
Table S1: Reward performance of different diversity reward formulations on the test set. The score ( ∈[0,3] ) is the sum of the genre, tone, and style rewards.
Variant
Avg. unique entities per group
Avg. repeated entities per group
INS
28
4.7
Original
62
2.8
Exponential
21
2.1
Clipped
26
1.5
Penalty
49
2.4
Appendix
Table S2: Results on the test set for different named entity diversity formulations with G=10 on the Llama-3.1-8B model.
Metric
Human-human
Judge-human
Quality metrics
Prompt adherence
92.30%
90.40%
Meta-commentary
97.60%
95.70%
Completeness
100.00%
94.10%
Coherence
97.00%
92.70%
Average
96.73%
93.23%
Appendix
Table S3: Human-human and judge-human agreement rates for quality and diversity metrics.
Figure S3: Validation results for different instruction residual scales s on Llama-3.1-8B . Diversity metrics consist of genre, tone, and style coverage values, as well as the named entity diversity reward Rne . All metrics are reported on a [0,1] scale.
Metric
GRPO
PPO
Average quality score
.894
.983
Average Genre/Tone/Style coverage
.454
.447
Average diversity reward
.360
.380
Cosine distance
.118
.135
Appendix
Table S4: Performance comparison between the GRPO and PPO variants of DQO.
Figure S4: Genre, tone, and style rewards obtained for different instruction residual scales s on Llama-3.1-8B . Rewards are reported on a [0,1] scale.
Figure S5: Training curves for Llama-3.1-8B under different values of s , showing the total reward R∈[−1,2] across training steps.
Figure S6: Training curves for Llama-3.1-8B under different values of h . (a) Percentage change in the entity diversity reward over training. (b) Average number of repeated entities per group over training.
Metric
INS
h=1
h=2
h=3
h=4
h=5
Ratio ↓
.107
.117
.066
.068
.078
.053
μ(Rne)↑
.581
.686
.838
.873
.813
.902
Appendix
Table S5: Entity repetition ratio (repeated entities per group over total entities per group) and mean Rne for INS and DivLM across values of h . Final evaluation uses h=5 for consistency.
Model
T=0.8
T=0.9
T=1.0
T=1.1
T=1.2
INS
.116
.118
.125
.131
.138
DivLM
.215
.225
.226
.231
.236
Appendix
Table S6: Average cosine distance on the test set for INS and DivLM under different values of T .
Figure S7: Mean diversity reward ( Rdiv ) on the test set for DivLM and INS across decoding temperatures T on Llama-3.1-8B .
Figure S8: DivLM diversity evaluation results across decoding temperatures T on Llama-3.1-8B .
Model
G=5
G=10
G=20
G=50
G=100
INS
.123
.125
.126
.126
.126
DivLM
.227
.226
.226
.226
.226
Appendix
Table S7: Average cosine distance on the test set for INS and DivLM under different values of G .
Figure S9: Mean diversity reward ( Rdiv ) on the test set for DivLM and INS across group sizes G on Llama-3.1-8B .
Figure S10: DivLM diversity evaluation results across group sizes G on Llama-3.1-8B .
Setting
INS
DivLM
Decay = 50
481.2
366.3
Decay = 150
481.2
410.0
Linear
481.2
483.6
Appendix
Table S8: Mean output length in words for the INS and DivLM models using different length penalties.