cs.CLOct 5, 2026

Word-Level Text Unmixing via Evidence-Preserving Ownership Routing with Language Models

Authors: Jinglin He, Siyang Jiang, Lixing He, Guoliang Xing, Hongkai Chen

Organizations: The Chinese University of Hong Kong

Abstract

Text from multiple sources can become interleaved into a single sequence when attribution metadata is lost, such as overlapping speech transcripts, document reading flows, or concurrent agent streams. We formalize this challenge as Word-Level Text Unmixing: given an interleaved lexical stream and source count K, recover the original source sequences while preserving every word occurrence and its within-source order exactly. Directly generating separated texts with LLMs can omit, duplicate, or hallucinate words, violating this exact-reconstruction objective. We therefore propose Evidence-Preserving Ownership Routing (EPOR), which decouples source-ownership prediction from reconstruction. EPOR adapts a causal LLM to predict canonical ownership routes conditioned on the mixed stream and prior routing decisions. At inference, completion-safe constrained decoding is combined with deterministic indexed reconstruction, yielding structurally valid K-source partitions that preserve every observed occurrence exactly once. We also introduce UNMIXBENCH, covering controlled synthetic mixtures, timestamp-derived speech from AMI and ICSI, layout-derived document streams from ReadingBank, and simulated concurrent digital outputs. Across five evaluation tracks, a 4B EPOR model achieves the lowest mean minimum-permutation word error rate among finetuned baselines, reducing the five-track mean by 22.3% relative to compact source-array generation and remaining competitive with zero-shot frontier LLMs. These results show that when lexical evidence is fully observed, separating ownership inference from lexical regeneration provides a reliable alternative to direct generation.

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Natively Unlearnable Large Language Models

    Jun 11, 2026Gaurav R. Ghosal, Pratyush Maini, Aditi RaghunathanActivation SparsityMachine Unlearning

  2. TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection

    May 12, 2026Tom Sander, Hongyan Chang, Tomáš Souček +10Language Model FingerprintingLLM Watermarking

  3. Interleaved Speech Language Models Latently Work In Text

    Jun 21, 2026Talia Sternberg, Gallil Maimon, Yossi AdiLLM InterpretabilitySpeech Language Models