Paper ID: 2312.17292

Effect of dimensionality change on the bias of word embeddings

Rohit Raj Rai, Amit Awekar

Word embedding methods (WEMs) are extensively used for representing text data. The dimensionality of these embeddings varies across various tasks and implementations. The effect of dimensionality change on the accuracy of the downstream task is a well-explored question. However, how the dimensionality change affects the bias of word embeddings needs to be investigated. Using the English Wikipedia corpus, we study this effect for two static (Word2Vec and fastText) and two context-sensitive (ElMo and BERT) WEMs. We have two observations. First, there is a significant variation in the bias of word embeddings with the dimensionality change. Second, there is no uniformity in how the dimensionality change affects the bias of word embeddings. These factors should be considered while selecting the dimensionality of word embeddings.

Submitted: Dec 28, 2023

Topics

Mixed Effect
Jina Embeddings
Absolute Stance Bias
Word Embeddings
Data Dimensionality
Text Data

Links

arXiv PDF