cs.CLOct 8, 2026

EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams

Authors: Heeseung Kim

Organizations: Department of AI, University of Seoul

Abstract

Wearable augmented reality (AR) assistants are moving toward continuous real-world interaction, where they perceive the user's activity through first-person video and audio and provide timely spoken guidance without being explicitly asked. While proactive video assistants, spoken dialog systems, and egocentric task understanding have each advanced rapidly, existing systems do not address the joint problem of deciding when to speak and what to say from continuous first-person streams. We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants. From HoloAssist video recordings of real human instructors, we construct clean audio streams through source separation and speech resynthesis, and convert each video session into a format where the model must decide at each moment whether to remain silent or provide spoken guidance. We fine-tune an omni-modal LLM with our data, and further improve its proactive intervention behavior with direct preference optimization. Experiments across closed and open-source models show that existing systems rarely produce well-timed, meaningful proactive interventions, while EgoVoice yields clear improvements in intervention timing, content relevance, and human preference over the zero-shot backbone.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

    Jul 13, 2026Gong Sitong, Tianyu Yan, Caixin Kang +6Egocentric Video UnderstandingProactive Assistance

  2. Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction

    Sep 21, 2026Lujia Bao, Qian Chen, Luyao Cheng +15Tool-Augmented Language Model AgentsSpeech Foundation Models