cs.AIOct 6, 2026

OTel: Open Telco AI Datasets, Benchmarks, and Models

Authors: Farbod Tavakkoli, Gregory Diamos, Kenneth Church, David Kanter, Mark Austin, Imtiaz Karim, Mirza Masfiqur Rahman, Merouane Abdelkader Debbah, +10 more

Organizations: AT&T Chief Data Office · RelationalAI · Northeastern University · MLCommons · The University of Texas at Dallas · Purdue University · Khalifa University · University of Leeds · Yale University · Mantis NLP · GSMA · Essential AI

Abstract

We present Open Telco (OTel), an open telecom AI resource that releases derived telecom datasets for retrieval, reranking, instruction tuning, and safety/abstention, together with 30 full-parameter post-trained baselines spanning 10 embedding models, 3 rerankers, and 17 language models. The community has already engaged substantially with the resource: as of May 3, 2026, the released models have been downloaded over 16 million times and the project has received 157+ pieces of media coverage worldwide. Building on prior open telecom datasets and benchmarks, OTel provides documented telecom data sources, held-out evaluation partitions, trained embedding models, rerankers, context-grounded LLMs, and safety/abstention data in one unified resource. Each baseline starts from an open-weight model and is post-trained on OTel-derived data using an open training recipe, then evaluated on held-out OTel evaluation partitions. OTel post-training improves performance across all three model families: embedding retrieval reaches 93.1% NDCG@10, reranking reaches 0.947 MRR@10, and language-model correctness reaches 87.8%. We release OTel as a reproducible starting point and invite the community to expand the data, improve embedding and reranking models, and build stronger context-grounded telecom LLMs.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks

    Sep 21, 2026Bohao Wang, Chenwei Wu, Hang Zou +7Large Reasoning ModelsLarge Language Models(Llms

  2. TeleEmbedBench: A Multi-Corpus Embedding Benchmark for RAG in Telecommunications

    Apr 20, 2026Pranshav Gajjar, Vijay K ShahMobile NetworksEmbedder