vMF Sentence LDA: A Spherical Topic Model over Sentence Embeddings
Organizations: Graduate School of Engineering, The University of Tokyo, Tokyo, Japan
Abstract
Latent Dirichlet Allocation (LDA) and models derived from it remain widely used topic models. LDA observes each document as a bag-of-words and models each topic by a categorical distribution over the vocabulary, so that it uses neither the internal structure of the document nor the similarity in meaning between words. Earlier work has responded to this limitation in two ways: many models have introduced embeddings, and some have assigned topics to sentences rather than to words. Their combination, a topic model that observes sentence embeddings, remains little explored. We propose vMF Sentence LDA (vSLDA), which keeps the admixture structure of LDA, observes each sentence as its L2-normalized embedding and models each topic by a von Mises-Fisher (vMF) distribution, which matches the cosine geometry of sentence embeddings. Its per-topic parameter count and per-iteration cost are linear in the embedding dimension, versus quadratic for the full-covariance Gaussian distribution in the existing model over sentence embeddings. We evaluate vSLDA where the limitation is expected to matter most, among topics that share much of their vocabulary: the topics that subdivide the one subject of a collection, and the narrow topics that result when a corpus is divided into a large number of topics. On two corpora, vSLDA attains the best mean rank against eight baselines when the fine categories within each coarse category are classified from the document-topic distributions. On the whole corpus, its advantage appears or widens as the number of topics grows. Weighting the word frequencies of each sentence by its topic posterior yields expected topic-word counts of the same form as those of LDA, so that the standard topic coherence and diversity measures apply to models that assign topics to sentences. On their product, topic quality, vSLDA leads in most within-category conditions of both corpora.
Figures & tables
| Observation distribution | |||
| Observation unit | Categorical over the vocabulary | Gaussian over embeddings, | vMF over embeddings, |
| Word | LDA ( Blei et al., 2003 ) | GLDA ( Das et al., 2015 ) | vLDA ( Li et al., 2016 ) |
| Sentence | SentLDA ( Balikas et al., 2016 ) | GSLDA ( Cha & Lee, 2024 ) | vSLDA (proposed) |
| Document | – | – | SAM ( Reisinger et al., 2010 ) |
| (a) | |
| X11/Unix windowing | UI events/widgets |
| ( ) | ( ) |
| ( ) | ( ) |
| xterm | button |
| xdm | expose |
| xmu | ctrl |
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Coarse category | Fine categories |
| 20news | computer | comp.{graphics, os.ms-windows.misc, sys.ibm.pc.hardware, sys.mac.hardware, windows.x} |
| ride | rec.{autos, motorcycles} | |
| sports | rec.sport.{baseball, hockey} | |
| science | sci.{crypt, electronics, med, space} | |
| religion | alt.atheism, soc.religion.christian, talk.religion.misc | |
| politics | talk.politics.{guns, mideast, misc} |
| Dataset | Coarse cat. | # Fine cat. | # Docs | # Sentences | Sent./doc | Tokens/sent |
| 20news | computer | 5 | 4,727 | 39,881 | 8.44 27.61 | 15.88 11.35 |
| ride | 2 | 1,875 | 12,039 | 6.42 9.78 | 15.84 10.87 | |
| sports | 2 | 1,868 | 14,300 | 7.66 14.41 | 15.85 11.36 | |
| science | 4 | 3,807 | 37,181 | 9.77 23.94 | 17.44 11.89 | |
| religion | 3 | 2,344 | 27,963 | 11.93 18.46 | 17.76 12.46 | |
| politics | 3 | 2,532 | 33,089 | 13.07 27.68 | 17.91 13.52 |
| 20 Newsgroups | NYT | ||||||
| Unit | JSD | Unit | JSD | ||||
| Computer | 5 | 0.164 | 0.781 | Arts | 4 | 0.224 | 0.760 |
| Ride | 2 | 0.129 | 0.838 | Business | 4 | 0.213 | 0.693 |
| Sports | 2 | 0.169 | 0.821 | Politics | 9 | 0.230 | 0.602 |
| Science | 4 | 0.203 | 0.512 | Sports | 7 | 0.231 | 0.695 |
| Religion | 3 | 0.105 | 0.935 | ||||
| Item | Setting |
| Common to all models | |
| Runs and seeds | 5 independent training runs of the topic model per condition; run trains with seed , set jointly for the Python, NumPy and PyTorch generators; classifier random state fixed across runs |
| Topic counts | (main experiments); on the whole corpus and up to within the coarse categories (Section 5.3 ; GLDA only within the coarse categories at , and GSLDA within the coarse categories only at ) |
| Sentence encoders | all-MiniLM-L6-v2 ( ; main text); all-mpnet-base-v2 and bge-base-en-v1.5 ( ; Appendix K ); outputs L2-normalized for vSLDA and GSLDA (except the refits of Appendix N ), unnormalized for ConTM; no instruction prefix for BGE |
| Word embeddings | word2vec-google-news-300 ( ) for GLDA, vLDA and ETM, whose vocabularies are the word types with a pretrained vector; L2-normalized for vLDA |
| Preprocessing | sentence segmentation with pysbd ; rule-based sentence filter (sentences of fewer than 4 or more than 120 word tokens and fragments without sentence structure dropped; URLs, e-mail addresses and non-ASCII characters removed); tokenization, lowercasing, WordNet lemmatization, numerals mapped to <NUM> , single-character tokens dropped; no stopword or part-of-speech filtering; sentences truncated at the encoder’s default maximum length (256/384/512 word pieces) |
| Model | Encoder ( ) | Per sweep (s) | Sweeps | Training (s) |
| vSLDA | MiniLM (384) | 0.038 0.000 | 200 | 7.74 0.03 |
| vSLDA | MPNet (768) | 0.055 0.000 | 200 | 11.0 0.1 |
| GSLDA | MiniLM (384) | 60.2 11.2 | 10 | 622.7 82.9 |
| GSLDA | MPNet (768) | 288.0 0.8 | 10 | 2,880 8 |
| Setting | 20 Newsgroups | NYT | ||||||
| Acc. (%) | Div. | Empty | Acc. (%) | Div. | Empty | |||
| Fallback concentration | ||||||||
| 73.50 | 0.456 | 0.996 | 0.0 | 86.35 | 0.553 | 0.995 | 0.0 | |
| (main) | 73.51 | 0.457 | 0.996 | 0.0 | 86.27 | 0.554 | 0.995 | 0.0 |
| 73.51 | 0.457 | 0.996 | 0.0 | 86.33 | 0.554 | 0.995 | 0.0 | |
| 73.51 | 0.457 | 0.996 | 0.0 | 86.33 | 0.554 | 0.995 | 0.0 | |
| Model | Diffuse topics (%, ) | Dead topics (%, ) | ||||||
| rank-1 document fraction | ||||||||
| (main) | (main) | |||||||
| LDA (Blei, 2003) | 18.4 | 10.6 | 4.1 | 0.1 | 44.6 | 60.4 | 64.8 | 73.7 |
| SentLDA (Balikas, 2016) | 6.1 | 1.9 | 0.0 | 0.0 | 5.0 | 23.0 | 35.5 | 57.9 |
| SAM (Reisinger, 2010) | 99.4 | 82.7 | 14.8 | 5.6 | 7.6 | 17.9 | 27.7 | 54.7 |
| GLDA (Das, 2015) | 99.3 | 84.7 | 55.0 | 13.8 | 76.2 | 90.0 | 91.1 | 93.9 |
| Accuracy (%) | Topic assignments of GSLDA | ||||||
| Dataset | GSLDA | vSLDA | Largest (%) | Empty | |||
| 20 Newsgroups | 10 | 26.4 | 16.9 | 0 | |||
| 20 | 13.4 | 13.8 | 0 | ||||
| 30 | 3.3 | 13.7 | 0 | ||||
| NYT | 10 | 41.1 | 16.1 | 0 | |||
| 20 | 19.9 | 13.6 | 0 | ||||
| 20 Newsgroups | NYT | |||||||
| Acc. (%) | Empty | Div. | Acc. (%) | Empty | Div. | |||
| 49.49 | 0.0 | 0.957 | 0.226 | 52.51 | 1.2 | 0.855 | 0.336 | |
| (main) | 54.95 | 0.0 | 0.969 | 0.249 | 58.02 | 2.1 | 0.810 | 0.342 |
| 39.90 | 18.2 | 0.125 | 0.273 | 41.31 | 17.2 | 0.181 | 0.289 | |
| 37.34 | 19.0 | 0.050 | 0.258 | 39.04 | 18.8 | 0.075 | 0.269 | |
| 37.34 | 19.0 | – | – | 38.72 | 19.0 | – | – | |
| Model | Classification accuracy | Topic assignments | ||
| vs. vSLDA | Ahead (of 30) | Empty (%) | Largest (%) | |
| GSLDA, full | 0 | 4.0 | 51.0 | |
| GSLDA, diagonal | 9 | 30.2 | 16.2 | |
| GSLDA, isotropic | 10 | 0.0 | 11.1 | |
| Unnormalized input | ||||
| Full | 0 | 0.0 | 63.3 | |
| Model | Scale matrix | Accuracy | Largest (%) | Empty |
| Released implementation, as released | diagonal, Eq. ( P.1 ) | |||
| Released implementation, full forward substitution | full | |||
| GSLDA (this paper) | full | |||
| vSLDA | – |
| Dataset | |||||
| 20news | 10 | [ , ] | 30/30 | [ , ] | 28/30 |
| 20 | [ , ] | 30/30 | [ , ] | 30/30 | |
| 30 | [ , ] | 30/30 | [ , ] | 30/30 | |
| NYT | 10 | [ , ] | 20/20 | [ , ] | 18/20 |
| 20 | [ , ] | 20/20 | [ , ] | 20/20 | |
| 30 | [ , ] | 20/20 | [ , ] | 20/20 |
| Topic | Fine category | Representative words | |||
| 0 | 251.7 | 1254 | 0.87 | os.ms-windows.misc (44%) | nt, msw, microsoft, window, apps |
| 1 | 288.8 | 683 | 0.99 | graphics (23%) | appreciate, thank, help, greatly, advance |
| 2 | 138.7 | 1413 | 0.85 | graphics (47%) | mahan, tgv, patrick, washer, printf |
| 3 | 289.3 | 866 | 0.95 | graphics (33%) | monitor, vga, resolution, hz, multisync |
| 4 | 263.0 | 1385 | 0.43 | graphics (80%) | graphic, visualization, ray, animation, plot |
| 5 | 281.7 | 524 | 0.39 | sys.mac.hardware (85%) | apple, mac, macintosh, appletalk, macweek |
| Topic | Fine category | Representative words | ||
| 0 | 1329 | 0.89 | graphics (40%) | mail, subscribe, send, newsgroups, subscription |
| 1 | 898 | 0.46 | graphics (79%) | color, quantization, palette, quantize, lossless |
| 2 | 2629 | 0.65 | windows.x (63%) | cursor, icon, event, ctrl, expose |
| 3 | 862 | 0.79 | os.ms-windows.misc (45%) | irq, slip, xxxx, modem, dtr |
| 4 | 1462 | 0.48 | graphics (78%) | phigs, analysis, plot, conference, dimensional |
| 5 | 874 | 0.55 | graphics (74%) | format, hsi, jfif, pict, ppm |
| Topic | Fine category | Representative words | |||
| 0 | 258.7 | 365 | 0.53 | movies (77%) | oscar, award, nominee, nomination, academy |
| 1 | 164.6 | 802 | 0.90 | music (39%) | cambridge, honoree, mandela, golf, roosevelt |
| 2 | 262.1 | 949 | 0.31 | movies (90%) | film, movie, cinema, cinemascope, ridley |
| 3 | 239.1 | 851 | 0.40 | television (86%) | episode, netflix, show, abc, network |
| 4 | 266.8 | 550 | 0.35 | music (88%) | clarinet, instrument, accordion, piano, trumpet |
| 5 | 232.8 | 567 | 0.58 | movies (72%) | killing, indonesia, trayvon, jfk, aftermath |
| Topic | Fine category | Representative words | ||
| 0 | 1175 | 0.57 | television (67%) | malick, episode, fx, clyde, pope |
| 1 | 971 | 0.83 | music (45%) | copyright, stream, netflix, beatles, bootleg |
| 2 | 143 | 0.51 | music (79%) | ferrando, mezzo, lukas, falstaff, soprano |
| 3 | 24 | 0.55 | movies (67%) | philippe, banter, exploitation, puckish, blanchett |
| 4 | 577 | 0.54 | music (77%) | euro, subsidy, gelb, bankruptcy, endowment |
| 5 | 3372 | 0.93 | music (32%) | you, lorne, me, didn, my |