eess.ASJun 19, 2026

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users

Authors: Allan HenrySolange RossatoChristian GraffSylvain HuetJose-Ernesto Gomez-Balderas

Organizations: GIPSA-COPERNIC, GETALP, LPNC · 1LIG, Univ. Grenoble Alpes, Grenoble, France · 2LPNC, Univ. Grenoble Alpes, Grenoble, France · 3GIPSA-lab, Univ. Grenoble Alpes, Grenoble, France · GETALP · LPNC · GIPSA-COPERNIC

Abstract

Voice control offers an intuitive alternative to manual drone piloting, yet most existing systems rely on rigid command vocabularies that fail to handle the spontaneous, disfluent speech of naive users. This paper addresses this gap by proposing an End-to-End Spoken Language Understanding architecture for real-time human-drone interaction in French. Our model combines a frozen Self-Supervised Learning acoustic encoder with a lightweight LSTM-based classification head, augmented by a cross-modal knowledge distillation objective that aligns acoustic representations with semantic embeddings from a text teacher, without requiring transcription at inference time. We evaluate our approach on VoiceStick, a novel French corpus of spontaneous speech collected during real teleoperation sessions with 29 nonexpert dyads. On simple voice commands, our best configuration achieves 93% accuracy at 7 ms inference latency, outperforming cascade baselines (79%, 202 ms) with a 29x speedup. On the full spontaneous speech test set, our architecture reaches 82% accuracy, with crossmodal distillation consistently improving robustness across all configurations. These results demonstrate that End-to-End architectures are not only feasible but preferable for spontaneous voice-guided UAV teleoperation, combining semantic robustness, low latency, and calibrated confidence.

Explore similar work

CardsList
  1. OAA: Three Phases of Vocal Guidance in Human-Drone Teleoperation

    Aug 11, 2026Allan Henry, Christian Graff, Solange Rossato +2Neural GuidanceTeleoperation