cs.CLJul 4, 2026

Separating Representation from Reconstruction Enables Scalable Text Encoders

Authors: Megi DervishiMathurin VideauYann LeCun

Organizations: FAIR · 3New York University

Abstract

While decoders have rapidly scaled, encoders have remained largely unchanged since BERT. We revisit this disparity by frozen backbone evaluation via probing. Under this lens, the representations of BERT encoders become increasingly unexploitable\textit{unexploitable} by frozen probes, despite improved perplexity. The misalignment originates in BERT's flat design, which couples representation learning to the token reconstruction loss. We propose CrossBERT\textbf{CrossBERT}, a two-part architecture that separates the learning of high-quality encoded representations from the rigid grounding of token reconstruction. This design further enables high masking ratios (50%\ge 50\%) and gradient collection over all tokens via a Complementary Masking Strategy\textit{Complementary Masking Strategy}, respectively increasing throughput by 1.51.5 to 2×2\times and sample efficiency by 2×2\times. Overall, CrossBERT demonstrates monotonic scaling and superior performance on MTEB(eng, v2) and frozen GLUE benchmarks.

Explore similar work

CardsList
  1. Polish ModernBERT: The Long and Short of Polish Language Understanding

    Sep 1, 2026Michał Perełkiewicz, Sławomir Dadas, Rafał Poświata +1ModernbertLongbench