cs.CLSep 28, 2026

Word Similarity Datasets for Indian Languages: Annotation and Baseline Systems

Authors: Syed S. Akhtar, Arihant Gupta, Avijit Vajpayee, Arjit Srivastava, M. Shrivastava

Organizations: International Institute of Information Technology Hyderabad, Telangana, India

Abstract

With the advent of word representations, word similarity tasks are becoming increasing popular as an evaluation metric for the quality of the representations. In this paper, we present manually annotated monolingual word similarity datasets of six Indian languages - Urdu, Telugu, Marathi, Punjabi, Tamil and Gujarati. These languages are most spoken Indian languages worldwide after Hindi and Bengali. For the construction of these datasets, our approach relies on translation and re-annotation of word similarity datasets of English. We also present baseline scores for word representation models using state-of-the-art techniques for Urdu, Telugu and Marathi by evaluating them on newly created word similarity datasets.

Figures & tables

Explore similar work

CardsList
  1. BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources

    Apr 20, 2026Raghvendra Kumar, Devankar Raj, Sriparna SahaIndian LanguagesNatural Language Datasets

  2. DunbaaBERT: From Sacrifice to Semantics

    May 26, 2026Iffat Maab, Waleed Jamil, Raphael SchmittUrduMultilingual Language Models

  3. Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India

    Apr 21, 2026Kaushal Bhogale, Manas Dhir, Amritansh Walecha +11Automatic Speech Recognition EvaluationAutomatic Speech Recognition