Latent Dirichlet Allocation (LDA) and models derived from it remain widely used topic models. LDA observes each document as a bag-of-words and models each topic by a categorical distribution over the vocabulary, so that it uses neither the internal structure of the document nor the similarity in meaning between words. Earlier work has responded to this limitation in two ways: many models have introduced embeddings, and some have assigned topics to sentences rather than to words. Their combination, a topic model that observes sentence embeddings, remains little explored. We propose vMF Sentence LDA (vSLDA), which keeps the admixture structure of LDA, observes each sentence as its L2-normalized embedding and models each topic by a von Mises-Fisher (vMF) distribution, which matches the cosine geometry of sentence embeddings. Its per-topic parameter count and per-iteration cost are linear in the embedding dimension, versus quadratic for the full-covariance Gaussian distribution in the existing model over sentence embeddings. We evaluate vSLDA where the limitation is expected to matter most, among topics that share much of their vocabulary: the topics that subdivide the one subject of a collection, and the narrow topics that result when a corpus is divided into a large number of topics. On two corpora, vSLDA attains the best mean rank against eight baselines when the fine categories within each coarse category are classified from the document-topic distributions. On the whole corpus, its advantage appears or widens as the number of topics grows. Weighting the word frequencies of each sentence by its topic posterior yields expected topic-word counts of the same form as those of LDA, so that the standard topic coherence and diversity measures apply to models that assign topics to sentences. On their product, topic quality, vSLDA leads in most within-category conditions of both corpora.
Figures & tables
Figure 1: Graphical model of vMF Sentence LDA (vSLDA). The plates repeat over the Nd sentences of document d , the D documents and the K topics. α is the Dirichlet parameter, θd the document-topic distribution of document d , zdi the topic assignment of its sentence i , wdi the normalized sentence embedding of that sentence, and (μk,κk) the mean direction and concentration parameter of the vMF distribution of topic k .
Observation distribution
Observation unit
Categorical over the vocabulary
Gaussian over embeddings, O(M2)
vMF over embeddings, O(M)
Word
LDA ( Blei et al., 2003 )
GLDA ( Das et al., 2015 )
vLDA ( Li et al., 2016 )
Sentence
SentLDA ( Balikas et al., 2016 )
GSLDA ( Cha & Lee, 2024 )
vSLDA (proposed)
Document
–
–
SAM ( Reisinger et al., 2010 )
Table 1: Compared models by observation unit (rows) and observation distribution (columns). Each cell names the model of Section 2 that we take as the representative of that combination. Other models discussed there fall in the same cells. vSLDA is the proposed model, and “–” marks a combination for which no model was discussed. The O(⋅) entries in the column headers give the cost of evaluating one observation’s likelihood under one topic as a function of the embedding dimension M (Section 3.4 ).
Figure 2: Topic-count sweep on the whole corpus (MiniLM). Rows: 20 Newsgroups (top) and NYT (bottom); columns: the accuracy (%) of fine-category classification by logistic regression on the document-topic distributions, CV topic coherence and topic quality, the product of CV and topic diversity (the diversity panels are in Appendix H ). The horizontal axis is the number of topics K . Each point is the mean over 5 runs ( n=5 ), shaded bands are ±1 standard deviation, and the eight models with runs across the whole grid are compared. Solid lines are the sentence-level models (SentLDA, GSLDA, vSLDA) and dashed lines the word- and document-level models.
(a)
X11/Unix windowing
UI events/widgets
( κ=267.3 )
( κ=140.2 )
( H=0.17 )
( H=0.91 )
xterm
button
xdm
expose
xmu
ctrl
Table 3: Groups of vSLDA topics matched to one SentLDA topic ( K=20 , MiniLM, run 0; every topic of both runs is listed in Appendix Q ). In each of the two runs, every vSLDA topic is matched to the SentLDA topic that receives the largest expected co-assignment of its sentences (not necessarily a majority of them; Appendix Q ); each panel shows the largest group of vSLDA topics matched to one SentLDA topic whose dominant fine category holds at least 50 % of its posterior mass. (a) 20 Newsgroups/ computer : two topics matched to the SentLDA topic with the representative words cursor , icon , event , ctrl , expose ; (b) NYT/ arts : three topics matched to the SentLDA topic with the representative words orchestra , philharmonic , carnegie , symphony , gilbert . Each column gives the fitted concentration parameter κ , the normalized fine-category entropy H of the topic’s posterior sentence mass and the top 5 representative words; within a panel the topics are ordered by ascending H . The italic label above each topic is a descriptive label assigned by the authors from its representative words and is not an output of the model.
Table A.1: Coarse categories (experimental units) of 20 Newsgroups (20news) and NYT, and their fine categories. For 20 Newsgroups, braces abbreviate a shared prefix (e.g., rec.{autos, motorcycles} denotes rec.autos and rec.motorcycles).
Dataset
Coarse cat.
# Fine cat.
# Docs
# Sentences
Sent./doc
Tokens/sent
20news
computer
5
4,727
39,881
8.44 ± 27.61
15.88 ± 11.35
ride
2
1,875
12,039
6.42 ± 9.78
15.84 ± 10.87
sports
2
1,868
14,300
7.66 ± 14.41
15.85 ± 11.36
science
4
3,807
37,181
9.77 ± 23.94
17.44 ± 11.89
religion
3
2,344
27,963
11.93 ± 18.46
17.76 ± 12.46
politics
3
2,532
33,089
13.07 ± 27.68
17.91 ± 13.52
Appendix
Table A.2: Statistics of the experimental units (coarse categories) of 20 Newsgroups (20news) and NYT. “# Docs” combines the training and test sets, “Sent./doc” is the mean number of sentences per document ( ± standard deviation), and “Tokens/sent” is the mean number of tokens per sentence ( ± standard deviation), counted over the evaluation vocabulary.
20 Newsgroups
NYT
Unit
∣F∣
JSD
cos
Unit
∣F∣
JSD
cos
Computer
5
0.164
0.781
Arts
4
0.224
0.760
Ride
2
0.129
0.838
Business
4
0.213
0.693
Sports
2
0.169
0.821
Politics
9
0.230
0.602
Science
4
0.203
0.512
Sports
7
0.231
0.695
Religion
3
0.105
0.935
Appendix
Table B.1: Distance between the fine categories inside each experimental unit. ∣F∣ is the number of fine categories of the unit. JSD is the mean Jensen–Shannon divergence (base 2) between their unigram distributions and cos the mean cosine similarity between their centroid directions, both averaged over the pairs of the unit; a larger JSD and a smaller cos indicate fine categories that lie further apart. JSD and cos are averaged over the 200 equal-sized draws described in the text; their standard deviation over the draws is at most 0.001 and 0.007. The Different units row gives the same pair averages over fine-category pairs drawn from different units. The sentence encoder is MiniLM.
Item
Setting
Common to all models
Runs and seeds
5 independent training runs of the topic model per condition; run i trains with seed 42+i , set jointly for the Python, NumPy and PyTorch generators; classifier random state fixed across runs
Topic counts
K∈{10,20,30} (main experiments); K∈{10,20,30,50,100,200,300} on the whole corpus and up to K=100 within the coarse categories (Section 5.3 ; GLDA only within the coarse categories at K≤30 , and GSLDA within the coarse categories only at K≤30 )
Sentence encoders
all-MiniLM-L6-v2 ( M=384 ; main text); all-mpnet-base-v2 and bge-base-en-v1.5 ( M=768 ; Appendix K ); outputs L2-normalized for vSLDA and GSLDA (except the refits of Appendix N ), unnormalized for ConTM; no instruction prefix for BGE
Word embeddings
word2vec-google-news-300 ( M=300 ) for GLDA, vLDA and ETM, whose vocabularies are the word types with a pretrained vector; L2-normalized for vLDA
Preprocessing
sentence segmentation with pysbd ; rule-based sentence filter (sentences of fewer than 4 or more than 120 word tokens and fragments without sentence structure dropped; URLs, e-mail addresses and non-ASCII characters removed); tokenization, lowercasing, WordNet lemmatization, numerals mapped to <NUM> , single-character tokens dropped; no stopword or part-of-speech filtering; sentences truncated at the encoder’s default maximum length (256/384/512 word pieces)
Appendix
Table D.1: Hyperparameters, preprocessing settings and random seeds of the experiments (Section 4.4 ). K is the number of topics, M the embedding dimension and i∈{0,…,4} the run index.
Model
Encoder ( M )
Per sweep (s)
Sweeps
Training (s)
vSLDA
MiniLM (384)
0.038 ± 0.000
200
7.74 ± 0.03
vSLDA
MPNet (768)
0.055 ± 0.000
200
11.0 ± 0.1
GSLDA
MiniLM (384)
60.2 ± 11.2
10
622.7 ± 82.9
GSLDA
MPNet (768)
288.0 ± 0.8
10
2,880 ± 8
Appendix
Table E.1: Wall-clock training time of vSLDA and GSLDA on the computer unit of 20 Newsgroups ( K=20 ; mean ± standard deviation over 3 runs). Per sweep is the time of one Gibbs sweep over the training sentences, averaged over all recorded sweeps and including the per-iteration work around the sweeps. Training is the time of the training length of the main experiments, 10 outer iterations (200 sweeps) for vSLDA and 10 Gibbs sweeps for GSLDA. The setting and what each column includes are described in the text.
Figure E.1: Sampler convergence against wall-clock time (20 Newsgroups/ computer , K=20 , MiniLM). Each panel shows the average log-likelihood of the training sentences against the cumulative training time in seconds, on a logarithmic scale, for vSLDA (left) and GSLDA (right), with one line per run (3 runs). The vMF likelihood is a density on the unit hypersphere and the GSLDA likelihood a Student- t posterior predictive density on RM , so the two are not comparable in value. Each panel therefore carries its own vertical axis and only the horizontal axis is shared. The dotted rule marks the training length of the main experiments, 10 outer iterations for vSLDA and 10 Gibbs sweeps for GSLDA. Both samplers reach their plateau within that length.
Figure G.1: Sample efficiency per category on 20 Newsgroups (MiniLM). Rows are the coarse categories computer , ride and sports (the other three follow on the next page) and columns the topic counts K=10 , 20 and 30 . The vertical axis is the accuracy (%) of fine-category classification by logistic regression on the document-topic distributions, and the horizontal axis is the fraction (%) of training documents whose labels were used to train the classifier. The three panels of a row share the vertical range. Each point is the mean over 5 runs and, below 100%, 5 labeled subsets drawn per run ( n=25 ; n=5 at 100%), shaded bands are ±1 standard deviation over the same values (Section 4.3 ), and the nine models of the main experiments are compared. Solid lines are the sentence-level models (SentLDA, GSLDA, vSLDA) and dashed lines the word- and document-level models.
Figure G.2: Sample efficiency per category on 20 Newsgroups (MiniLM), continued. Rows are the coarse categories science , religion and politics and columns the topic counts K=10 , 20 and 30 ; axes, points, bands and line styles as on the previous page.
Figure G.3: Sample efficiency per category on NYT (MiniLM). Rows are the four coarse categories and columns the topic counts K=10 , 20 and 30 ; axes, points, bands and line styles as in Figure G.1 .
Figure H.1: Topic diversity in the topic-count sweep on the whole corpus (MiniLM). Left: 20 Newsgroups; right: NYT. The vertical axis is topic diversity, the horizontal axis the number of topics K ; accuracy, CV coherence and topic quality on the same sweep are in Figure 2 . Each point is the mean over 5 runs ( n=5 ), shaded bands are ±1 standard deviation, and the eight models with runs across the whole grid are compared. Solid lines are the sentence-level models (SentLDA, GSLDA, vSLDA) and dashed lines the word- and document-level models.
Figure H.2: Topic-count sweep of classification accuracy within the coarse categories of 20 Newsgroups (MiniLM). Each panel is one coarse category. The vertical axis is the accuracy (%) of fine-category classification by logistic regression on the document-topic distributions, and the horizontal axis is the number of topics K (the per-category runs stop at K=100 ). Each point is the mean over 5 runs ( n=5 ), shaded bands are ±1 standard deviation, and the seven models with runs across the grid are compared. Solid lines are the sentence-level models (SentLDA, vSLDA) and dashed lines the word- and document-level models.
Figure H.3: Topic-count sweep of classification accuracy within the coarse categories of NYT (MiniLM). Each panel is one coarse category. The vertical axis is the accuracy (%) of fine-category classification by logistic regression on the document-topic distributions, and the horizontal axis is the number of topics K (the per-category runs stop at K=100 ). Each point is the mean over 5 runs ( n=5 ), shaded bands are ±1 standard deviation, and the seven models with runs across the grid are compared. Solid lines are the sentence-level models (SentLDA, vSLDA) and dashed lines the word- and document-level models.
Figure H.4: Topic-count sweep of topic quality within the coarse categories of 20 Newsgroups (MiniLM). Each panel is one coarse category. The vertical axis is topic quality, the product of CV coherence and topic diversity (as in Table F.4 ), and the horizontal axis is the number of topics K (the per-category runs stop at K=100 ). Each point is the mean over 5 runs ( n=5 ), shaded bands are ±1 standard deviation, and the seven models with runs across the grid are compared. Solid lines are the sentence-level models (SentLDA, vSLDA) and dashed lines the word- and document-level models.
Figure H.5: Topic-count sweep of topic quality within the coarse categories of NYT (MiniLM). Each panel is one coarse category. The vertical axis is topic quality, the product of CV coherence and topic diversity (as in Table F.4 ), and the horizontal axis is the number of topics K (the per-category runs stop at K=100 ). Each point is the mean over 5 runs ( n=5 ), shaded bands are ±1 standard deviation, and the seven models with runs across the grid are compared. Solid lines are the sentence-level models (SentLDA, vSLDA) and dashed lines the word- and document-level models.
Setting
20 Newsgroups
NYT
Acc. (%)
CV
Div.
Empty
Acc. (%)
CV
Div.
Empty
Fallback concentration κ0
1
73.50
0.456
0.996
0.0
86.35
0.553
0.995
0.0
10 (main)
73.51
0.457
0.996
0.0
86.27
0.554
0.995
0.0
100
73.51
0.457
0.996
0.0
86.33
0.554
0.995
0.0
1000
73.51
0.457
0.996
0.0
86.33
0.554
0.995
0.0
Appendix
Table L.1: Sensitivity of vSLDA to its own hyperparameters. Classification accuracy (%, logistic regression), CV topic coherence, topic diversity and number of empty topics out of K=20 as one hyperparameter at a time is varied from the setting of the main experiments ( κ0=10 , α0=50/K=2.5 , ζ=20 , B=8 , T=10 , T0=5 , a=1 ; the row marked “main” in each block), at K=20 with MiniLM embeddings. κ0 is the concentration assigned to a topic holding at most one sentence, α0 the initial Dirichlet parameter, ζ the number of Gibbs sweeps per E-step, B the number of stored sweeps, T the number of outer iterations, T0 the burn-in after which the statistics are averaged and α is held fixed ( T0=T : no averaging), and a the exponent of the step sizes γt=(t−T0)−a . Values are means over the coarse-category units of each dataset and 3 runs per condition; the main setting is read from the main runs with the same seeds and therefore differs slightly from Table F.1 . An empty topic is one to which no sentence is assigned at the last Gibbs sweep of training.
Figure M.1: Per-topic and per-document distributions of the structural diagnostics ( K=20 , MiniLM). Rows, from top: the normalized document entropy H(P(d∣k))/lnD of each topic (“topic spread over documents”), the rank-1 document fraction of each topic (“share as main topic”; the axis is linear below 0.01 and logarithmic above), and the normalized entropy H(θd)/lnK of each document (“document spread over topics”); columns: 20 Newsgroups (left) and NYT (right). The models are named on the horizontal axis. Each box pools the topics (or documents) of all coarse categories of the dataset and all 5 runs: boxes span the quartiles, whiskers extend to 1.5 interquartile ranges and points mark topics beyond the whiskers (not drawn in the bottom row). Dashed lines mark the cut-offs of Table M.1 , 0.95 for diffuse and 0.01 for dead topics, and the dotted line the equal-use value 1/K .
Model
Diffuse topics (%, ↓ )
Dead topics (%, ↓ )
H(P(d∣k))/lnD>
rank-1 document fraction <
0.85
0.90
0.95 (main)
0.99
0.001
0.01 (main)
0.02
0.05
LDA (Blei, 2003)
18.4
10.6
4.1
0.1
44.6
60.4
64.8
73.7
SentLDA (Balikas, 2016)
6.1
1.9
0.0
0.0
5.0
23.0
35.5
57.9
SAM (Reisinger, 2010)
99.4
82.7
14.8
5.6
7.6
17.9
27.7
54.7
GLDA (Das, 2015)
99.3
84.7
55.0
13.8
76.2
90.0
91.1
93.9
Appendix
Table M.2: Cut-off sensitivity of panels (a) and (b) of Table M.1 : the percentage of the K topics of a run that are diffuse (normalized document entropy H(P(d∣k))/lnD above the cut-off) and dead (rank-1 document fraction below the cut-off), computed per run and averaged with equal weight over the 150 runs of the main experiments (the 10 categories of the two datasets ×K∈{10,20,30}× 5 runs; MiniLM). The columns marked (main) are the cut-offs of Table M.1 . A lower value is preferable ( ↓ ) and bold indicates the lowest value within each column.
Figure N.1: Accuracy gap between vSLDA and GSLDA against the observations per dimension of GSLDA. Each point is one GSLDA training run (both datasets, every coarse category, K∈{10,20,30} , 5 runs, three sentence encoders; 450 points, 150 per encoder). The vertical axis is the classification accuracy of vSLDA minus that of GSLDA in the same condition (points), and the horizontal axis, on a logarithmic scale, is n/M , the median number of sentences assigned to a non-empty topic at the last Gibbs sweep divided by the embedding dimension M ( 384 for MiniLM, 768 for MPNet and BGE). The dotted vertical line marks n/M=1 , below which the sample covariance of a typical topic is singular.
Accuracy (%)
Topic assignments of GSLDA
Dataset
K
GSLDA
vSLDA
Δ
n/M
Largest (%)
Empty
20 Newsgroups
10
35.8±2.6
37.6±2.6
−1.8
26.4
16.9
0
20
41.6±1.0
48.6±1.0
−7.0
13.4
13.8
0
30
42.3±2.0
51.9±1.4
−9.6
3.3
13.7
0
NYT
10
53.0±3.2
66.6±3.2
−13.6
41.1
16.1
0
20
64.3±2.4
77.2±1.4
−13.0
19.9
13.6
0
Appendix
Table N.1: GSLDA on each corpus as a whole (MiniLM). Accuracy is that of fine-category classification by logistic regression on the document-topic distributions (%, mean ± standard deviation over 5 runs), and Δ the mean accuracy of GSLDA minus that of vSLDA . n/M is the median number of sentences assigned to a non-empty topic at the last Gibbs sweep of training divided by the embedding dimension M=384 , Largest the percentage of the sentences assigned to the largest topic and Empty the largest number of topics without a sentence in any run. The unit and the vSLDA runs are those of Figure 2 .
λ
20 Newsgroups
NYT
Acc. (%)
Empty
Div.
CV
Acc. (%)
Empty
Div.
CV
0.01
49.49
0.0
0.957
0.226
52.51
1.2
0.855
0.336
0.1 (main)
54.95
0.0
0.969
0.249
58.02
2.1
0.810
0.342
1
39.90
18.2
0.125
0.273
41.31
17.2
0.181
0.289
3
37.34
19.0
0.050
0.258
39.04
18.8
0.075
0.269
10
37.34
19.0
–
–
38.72
19.0
–
–
Appendix
Table N.2: Sensitivity of GSLDA to the scale matrix of its covariance prior. Classification accuracy (%, logistic regression), number of empty topics out of K=20 , topic diversity and CV of GSLDA as the prior scale λ , which sets the scale matrix of the inverse-Wishart prior on the topic covariances to λI , is varied at K=20 with MiniLM embeddings. λ=0.1 is the setting of the main experiments, λ=3 the value reported for Gaussian LDA, and λ=3M=1152 the default of the reference implementation. Values are means over the coarse-category units of each dataset and 3 runs per condition. An empty topic is one to which no sentence is assigned at the last Gibbs sweep of training; from λ=10 upwards 19 of the 20 topics are empty in every run, so the word-based measures cannot be computed (–). The last two rows give the same four quantities for vSLDA and SentLDA over the same runs, whose accuracy, diversity and CV therefore differ slightly from Tables F.1 and F.2 . A higher value is preferable in every column except Empty, and bold marks the highest accuracy of GSLDA on each dataset.
Model
Classification accuracy
Topic assignments
Δ vs. vSLDA
Ahead (of 30)
Empty (%)
Largest (%)
GSLDA, full
−22.4
0
4.0
51.0
GSLDA, diagonal
−0.8
9
30.2
16.2
GSLDA, isotropic
0.0
10
0.0
11.1
Unnormalized input
Full
−30.1
0
0.0
63.3
Appendix
Table N.3: Classification accuracy relative to vSLDA and topic-assignment diagnostics of the sentence-level Gaussian topic model under the three covariance structures (MiniLM). Full is the normal–inverse-Wishart model of Table F.1 ; diagonal and isotropic replace it by scaled-inverse- χ2 priors on the variances, leaving the prior on the mean, the prior width of an empty topic and every other setting unchanged. The rows under unnormalized input repeat the three structures on the pooled encoder output taken before the terminal normalization module of the released encoder. Δ is the classification accuracy of the structure (%, logistic regression) minus that of vSLDA , averaged over the 30 conditions of Table F.1 (5 runs each), on which vSLDA averages 77.5%; Ahead is the number of those conditions in which the structure’s mean accuracy exceeds that of vSLDA . Empty is the percentage of the K topics to which no sentence is assigned at the last Gibbs sweep of training, and Largest the percentage of the sentences assigned to the largest topic, both averaged over the same 150 runs. The word-based measures of the refits at K=20 are reported in the text.
Figure O.1: Expected log-density deficit of a topic against its number of observations. Each curve is the deficit b(n) of Eq. ( O.3 ) for one observation model, for vMF observations at M=384 and κ=250 (both axes logarithmic). The models are the full-covariance predictive of GSLDA at three prior scales λ , its diagonal and isotropic variants at λ=0.1 (all with τ0=0.1 ), and the vMF likelihood of vSLDA at the estimates of its M-step. Each point is the mean over 50 independent sets of n observations and 2,000 test draws, and two Monte Carlo standard errors are at most 5 % of each value. The dotted lines are M/(2n) and M(M+3)/(4n) , the large-sample deficits of correctly specified models with the numbers of parameters of the vMF and the full-covariance topic, and the dotted vertical line marks n=M . The vMF and isotropic curves follow M/(2n) and overlap, whereas under λ=0.1 and λ=0.01 the full-covariance deficit peaks near n=M , where both are below one nat.
Model
Scale matrix
Accuracy
Largest (%)
Empty
Released implementation, as released
diagonal, Eq. ( P.1 )
89.1±0.9
6.2±0.1
0.0±0.0
Released implementation, full forward substitution
full
56.1±1.2
80.2±0.1
0.0±0.0
GSLDA (this paper)
full
53.1±0.8
92.8±0.5
0.0±0.0
vSLDA
–
89.4±0.7
9.1±0.4
0.0±0.0
Appendix
Table P.1: Classification accuracy and topic-assignment diagnostics of the released SentenceLDA implementation, GSLDA and vSLDA on 20 Newsgroups/ sports at K=20 (all-mpnet-base-v2). The first two rows are the implementation released with Cha & Lee (2024) , run as released, where the posterior predictive has the diagonal scale matrix of Eq. ( P.1 ), and run with the forward substitution carried out in full and nothing else changed. The last two rows are the GSLDA of this paper, which evaluates the full quadratic form, and vSLDA , which has no scale matrix. Accuracy is the classification accuracy (%, logistic regression), Largest the percentage of the training sentences assigned to the largest topic and Empty the number of topics to which no sentence is assigned. Values are the mean ± standard deviation over 5 runs.
Dataset
K
ρ(κk,H)
ρ<0
ρ(κk,H∣lognk)
ρ<0
20news
10
−0.69 [ −0.95 , −0.21 ]
30/30
−0.60 [ −0.93 , 0.12 ]
28/30
20
−0.65 [ −0.91 , −0.27 ]
30/30
−0.63 [ −0.88 , −0.33 ]
30/30
30
−0.66 [ −0.88 , −0.32 ]
30/30
−0.65 [ −0.87 , −0.33 ]
30/30
NYT
10
−0.65 [ −0.90 , −0.16 ]
20/20
−0.57 [ −0.88 , 0.06 ]
18/20
20
−0.61 [ −0.86 , −0.16 ]
20/20
−0.59 [ −0.85 , −0.16 ]
20/20
30
−0.60 [ −0.82 , −0.24 ]
20/20
−0.58 [ −0.84 , −0.15 ]
20/20
Appendix
Table Q.2: Rank correlation, within a run, between the concentration of a topic and the breadth of its fine-category distribution. For every vSLDA run of the main experiments (MiniLM; 6 units of 20 Newsgroups and 4 of NYT, 5 runs each), the Spearman correlation across the K topics between the fitted κk and the normalized fine-category entropy H of the topic’s posterior sentence mass, and the same correlation with the ranks of lognk (the topic size) partialled out of both. Cells give the mean over the units and runs of the row with the range in brackets, and the number of runs in which the correlation is negative out of all runs of the row. Topics to which no sentence is assigned are excluded.
Topic
κk
nk
H
Fine category
Representative words
0
251.7
1254
0.87
os.ms-windows.misc (44%)
nt, msw, microsoft, window, apps
1
288.8
683
0.99
graphics (23%)
appreciate, thank, help, greatly, advance
2
138.7
1413
0.85
graphics (47%)
mahan, tgv, patrick, washer, printf
3
289.3
866
0.95
graphics (33%)
monitor, vga, resolution, hz, multisync
4
263.0
1385
0.43
graphics (80%)
graphic, visualization, ray, animation, plot
5
281.7
524
0.39
sys.mac.hardware (85%)
apple, mac, macintosh, appletalk, macweek
Appendix
Table Q.3: Every topic of vSLDA in the 20 Newsgroups/ computer run ( K=20 , MiniLM, run 0), with the fitted concentration parameter κk , its posterior sentence mass nk , its normalized fine-category entropy H , its dominant fine category with the share and its top 5 representative words; the columns are defined in the text of this appendix.
Topic
nk
H
Fine category
Representative words
0
1329
0.89
graphics (40%)
mail, subscribe, send, newsgroups, subscription
1
898
0.46
graphics (79%)
color, quantization, palette, quantize, lossless
2
2629
0.65
windows.x (63%)
cursor, icon, event, ctrl, expose
3
862
0.79
os.ms-windows.misc (45%)
irq, slip, xxxx, modem, dtr
4
1462
0.48
graphics (78%)
phigs, analysis, plot, conference, dimensional
5
874
0.55
graphics (74%)
format, hsi, jfif, pict, ppm
Appendix
Table Q.4: Every topic of SentLDA in the 20 Newsgroups/ computer run ( K=20 , MiniLM, run 0), with its posterior sentence mass nk , its normalized fine-category entropy H , its dominant fine category with the share and its top 5 representative words; the columns are defined in the text of this appendix.
Topic
κk
nk
H
Fine category
Representative words
0
258.7
365
0.53
movies (77%)
oscar, award, nominee, nomination, academy
1
164.6
802
0.90
music (39%)
cambridge, honoree, mandela, golf, roosevelt
2
262.1
949
0.31
movies (90%)
film, movie, cinema, cinemascope, ridley
3
239.1
851
0.40
television (86%)
episode, netflix, show, abc, network
4
266.8
550
0.35
music (88%)
clarinet, instrument, accordion, piano, trumpet
5
232.8
567
0.58
movies (72%)
killing, indonesia, trayvon, jfk, aftermath
Appendix
Table Q.5: Every topic of vSLDA in the NYT/ arts run ( K=20 , MiniLM, run 0), with the fitted concentration parameter κk , its posterior sentence mass nk , its normalized fine-category entropy H , its dominant fine category with the share and its top 5 representative words; the columns are defined in the text of this appendix.
Table Q.6: Every topic of SentLDA in the NYT/ arts run ( K=20 , MiniLM, run 0), with its posterior sentence mass nk , its normalized fine-category entropy H , its dominant fine category with the share and its top 5 representative words; the columns are defined in the text of this appendix.
Traditional topic modeling assigns a single topic to each document. In practice, however, many real-world documents, such as product reviews or open-ended survey responses, contain multiple distinct topics. This mismatch often leads to topic contamination, where unrelated themes are merged into a single topic, making it difficult to identify documents that truly focus on a specific subject. We address this issue by introducing segment-based topic allocation (SBTA), a reformulation of topic modeling that assigns topics not to entire documents, but to segments: short, coherent spans of text that each express a single theme. By modeling topical structure at the segment level, our approach yields cleaner and more interpretable topics and better supports analysis of multi-theme documents. To support systematic evaluation, we construct a SemEval-STM, a new dataset inspired by aspect-based sentiment analysis. Documents are first decomposed into topical segments using large language models (LLMs), followed by human refinement to ensure segment quality. We also propose a segment-level extension of the word intrusion task, enabling human evaluation of topical coherence at the granularity where topics are actually assigned. Across multiple models and evaluation metrics, we show that SBTA improves clustering quality and interpretability. Overall, this work provides a practical, scalable framework for fine-grained topic analysis in heterogeneous text corpora where documents naturally span multiple topics. URL: https://huggingface.co/datasets/LG-AI-Research/SemEval-STM
Hoonsang Yoon, Takyoung Kim, Wonkee Lee +3
1LG AI Research · University of Illinois Urbana-Champaign
Traditional topic models are effective at uncovering latent themes in large text collections. However, due to their reliance on bag-of-words representations, they struggle to capture semantically abstract features. While some neural variants use richer representations, they are similarly constrained by expressing topics as word lists, which limits their ability to articulate complex topics. We introduce Mechanistic Topic Models (MTMs), a class of topic models that operate on interpretable features learned by sparse autoencoders (SAEs). By defining topics over this semantically rich space, MTMs can reveal deeper conceptual themes with expressive feature descriptions. Moreover, uniquely among topic models, MTMs enable controllable text generation using topic steering vectors. To properly evaluate MTM topics against word list approaches, we propose \textit{topic judge}, an LLM-based pairwise comparison evaluation framework. Across eight datasets, MTMs match or exceed traditional and neural baselines on coherence metrics, are consistently preferred by topic judge, and enable effective LLM steering.
Carolina Zheng, Nicolas Beltran-Velez, Sweta Karlekar +5
Columbia University · Independent · Google Research
Topic models summarize large text corpora, but top-ranked words often provide only a limited representation of topic semantics. Sparse autoencoders (SAEs) offer a way to move beyond word-level descriptors by extracting interpretable features from dense representations, yet how feature interpretability relates to topic-inference quality remains unclear. We introduce \textbf{MonoTM}, an interpretable topic modeling framework that decouples these roles. Across three benchmark corpora, we show that document--topic mixture estimation and semantic interpretation favor different SAE configurations and feature subsets. MonoTM estimates mixtures from the full SAE bag-of-features representation and, with them fixed, learns topic descriptors over a separate vocabulary of corpus-grounded semantic features. This design preserves global topic structure while representing topics with semantic units more meaningful than individual words, making them more useful for downstream corpus analysis.