cs.CLAug 13, 2026

ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization

Authors: Lixing Li

Organizations: Cornell University Ithaca, NY 14853, USA

Abstract

Adaptive latent tokenization maps a fine-grained input to a shorter sequence of continuous representations associated with input-dependent spans. We introduce ReconSpan, which divides text into chunks that a backward decoder can reconstruct from a single contextual prefix code and retains one such code as the latent token for each chunk. The reconstruction criterion is applied when chunks are formed, allowing one trained autoencoder to produce average chunk lengths from 6.5 to 12.2. At matched average length, reconstruction-guided boundaries preserve more text than random boundaries. Readers of the resulting latent sequence recover topic information reliably but struggle to extract exact details.

Explore similar work

CardsList
  1. Effective Context in Transformers: An Analysis of Fragmentation and Tokenization

    May 13, 2026Amirmehdi Jafari Fesharaki, Mohammadamin Rami, Aslan TchamkertenTokenizationTransformer Architectures