cs.NIJul 20, 2026

Human Grounded Evaluation of Large Language Models for Optical Network Automation

Authors: Kiarash RezaeiOmran AyoubPaolo MontiCarlos Natalino

Organizations: Department of Electrical Engineering, Chalmers University of Technology, 412 96 Gothenburg, Sweden · University of Applied Sciences and Arts of Southern Switzerland, 6928 Lugano, Switzerland

Abstract

Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substantially across LLM families. We present HuGLEN, a stepwise evaluation pipeline that uses an LLM-as-a-judge together with a small set of expert ratings to enable scalable and reproducible comparison of candidate LLMs, and to rank them using a quality efficiency score (QES). We demonstrate HuGLEN for translating outputs from an explainable artificial intelligence (XAI) model for the optical network quality of transmission (QoT) estimation task into operator-friendly explanations. Our results show that a medium-sized LLM (12B parameters) achieves the highest QES, indicating the best trade-off between explanation quality and efficiency. Overall, HuGLEN reduces the human-labeling burden while supporting consistent model selection for operator-facing automation tasks.

Explore similar work

CardsList