cs.AIOct 6, 2026

PERSIST: Who-What-When Memory Across Sessions for Full-Duplex Spoken Dialogue

Authors: Achira Lin, Siyuan Hou, Wenyi Yu, Xinnian Zhao, Haoyu Niu, Wang Geng, Longshuai Xiao, Shihai Xiao, +2 more

Abstract

Modern voice assistants may be shared by multiple users and should be able to answer questions about earlier conversations such as "When did I originally plan to leave?" or adapt their behavior to individual users based on past interactions. This requires more than retrieving a topically similar passage: the assistant must identify the current speaker, recover the relevant past state, and distinguish it from later revisions. We present PERSIST, a persistent memory system for multi-session, multi-speaker spoken dialogue that explicitly models Who, What, and When. PERSIST structures cross-session histories into readable event records and retrieves them with a 3W joint scoring mechanism that combines semantic content, acoustic speaker identity, and temporal state. For real-time full-duplex interaction, PERSIST further reuses intermediate representations from the dialogue backbone, avoiding query-audio re-encoding and reducing retrieval latency from 578.42 ms to 7.03 ms. We also introduce SpokenTrace, a diagnostic benchmark that factorizes evaluation along memory tasks and speaker-query types, exposing failures in recall, speaker attribution, and temporal-state tracking. On SpokenTrace, PERSIST achieves 85.08% end-to-end task accuracy and improves all-support EM@3 from 49.01% with BGE-large to 82.10%.

Explore similar work

CardsList
  1. SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

    Sep 22, 2026Haobo Zheng, Tan Tang, Yan Chen +2Conversational MemoryLLM Memory

  2. VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models

    Sep 26, 2026Yang Xiao, Vidhyasaharan Sethu, Eun-Jung Holden +1