cs.LGOct 5, 2026

What Matters for Latent Reasoning with Flow Matching

Authors: Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

Organizations: Samsung AI Cambridge · Technical University of Iasi · Queen Mary University of London

Abstract

Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reasoning trajectories, explainable, so that a decoded chain of thought (CoT) reflects reasoning the answer actually follows, refinable with more inference compute, and efficient, costing less than an explicit CoT at comparable accuracy. Current methods rarely meet these requirements: they learn shortcuts from the question, distill the explicit CoT into their weights, or imitate it one token at a time. We focus on flow matching in a learned latent space, the family we argue is best placed to meet them, and identify the training choices that make it work. The result is Flow-based Latent Reasoning (FLaRe), a simple recipe covering what the latent space encodes and how to shape it, where to train the flow, how to read out the answer, and a final stage of training on the model's own verified thoughts. A probe for each requirement shows that FLaRe improves on prior latent methods in all five. It also compares favorably with them on arithmetic benchmarks, while reaching 97% of the accuracy of explicit CoT at a quarter of its latency.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Latent Thought Flow: Efficient Latent Reasoning in Large Language Models

    Jun 15, 2026Xiandong Zou, Jing Huang, Jianshu Li +1Efficient Latent ReasoningLatent Thoughts

  2. Latent Reasoning with Normalizing Flows

    Jun 4, 2026Guancheng Tu, Xiangjun Fu, Suhao Yu +5Efficient Latent ReasoningChain-of-Thought Reasoning

  3. LaTER: Efficient Test-Time Reasoning via Latent Exploration and Explicit Verification

    May 8, 2026Xuan Li, Yining Wang, Yuchen Liu +7Efficient Latent ReasoningChain-of-Thought Reasoning