cs.CLOct 7, 2026

Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale

Authors: Mikhail L. Arbuzov, Karan Dave, Evgeniya Dontsova, Yaodong Hu, Vincent Lao, Navita Jain, Sisong Bei, Dmitry Dimov

Organizations: Independent researcher

Abstract

Enterprise conversation analytics asks many questions of millions of interactions. Each question can require reconstructing what people mean and identifying which information matters, repeating costly interpretive work across the same transcripts. We propose a simple principle: clarify the text, then focus the reader. Statement normalization transforms dialogue into short, speaker-attributed statements with source references and semantic tags. The statements make meaning more explicit; the tags support selecting evidence for a particular question. Downstream models can use the full representation or a relevant subset, depending on what helps them make the decision. In an offer-suppression task on customer-service calls, normalization improves a supervised classifier without selection, while weaker prompted readers benefit from both normalization and selection. A small model can learn the normalization contract, while lightweight encoders handle tagging and downstream decisions. Sharing this preparation across questions supports an inference pipeline built entirely from small models, making analytics over millions of conversations substantially less expensive.

Figures & tables

Explore similar work

CardsList
  1. The Hidden States Cookbook: A Large-Scale Ablation Study for Noise-Robust Conversational Intent Classification in Industry

    Oct 4, 2026Bogdan Bogachov, Nikita Letov, Yaoyao Fiona ZhaoIntent ClassificationLanguage Model Robustness

  2. Beyond Supervised Clarification: Input Rewriting with LLMs for Dialogue Discourse Parsing

    Jul 2, 2026Yiming Liu, Ziyue Zhang, Zhichao Xu +4Selective PredictionLLM Prompting