FloorSAV: Elucidating Spatial Audio-Visual Context with 2D Floormap for AV-LLMs
Organizations: Korea Advanced Institute of Science and Technology
Abstract
While 3D spatial reasoning in dynamic egocentric environments is crucial for embodied intelligence, audio-visual large language models (AV-LLMs) lack explicit mechanisms to process and internalize global geometry directly from raw sensory streams. Existing approaches either require costly fine-tuning or underutilize the model's cross-modal reasoning capacities. In this paper, we propose FloorSAV, a novel framework that explicitly grounds spatial audio-visual context by rendering a dynamic 2D floormap. By integrating 3D point clouds, camera trajectories, spatial audio cues, and semantically grounded object landmarks, we inject this floormap into the AV-LLM as a synchronized stream with an egocentric video. AV-LLMs utilize their multi-modal capabilities to jointly reason over visual, auditory, and geometric cues in a single inference with floormap interpretation guidance. We further introduce SAVED-Bench (Spatial Audio-Visual Egocentric Benchmark with Dynamic Agents), constructing essential tasks of spatial capability in real-world scenarios: dynamic relativity, regional, and path reasoning QAs. FloorSAV improves AV-LLMs' spatial reasoning on various tasks from both SAVED-Bench and SAVVY-Bench. Studies with ground-truth floormaps demonstrate the substantial potential of FloorSAV with accurate spatial information.
Figures & tables
| Task | Baseline | FloorSAV | Partial GT Map † | Full GT Map ‡ |
|---|---|---|---|---|
| Dynamic Relativity | ||||
| Ego-to-Exo Direction | 56.0 | 65.0 {}_{{\color[rgb]{0,0.75,0.16}\text{+9.0}}} | 76.5 {}_{{\color[rgb]{0,0.75,0.16}\text{+20.5}}} | 72.0 {}_{{\color[rgb]{0,0.75,0.16}\text{+16.0}}} |
| Ego-to-Exo Distance (AbsMRA) | 34.3 | 34.6 {}_{{\color[rgb]{0,0.75,0.16}\text{+0.3}}} | 47.9 {}_{{\color[rgb]{0,0.75,0.16}\text{+13.6}}} | 45.9 {}_{{\color[rgb]{0,0.75,0.16}\text{+11.6}}} |
| Exo-to-Ego Direction | 32.0 | 50.5 {}_{{\color[rgb]{0,0.75,0.16}\text{+18.5}}} | 50.5 {}_{{\color[rgb]{0,0.75,0.16}\text{+18.5}}} | 53.5 {}_{{\color[rgb]{0,0.75,0.16}\text{+21.5}}} |
| Exo-to-Ego Distance (AbsMRA) | 46.4 | 44.7 {}_{{\color[rgb]{1,0,0}\text{--1.7}}} | 57.5 {}_{{\color[rgb]{0,0.75,0.16}\text{+11.1}}} | 55.0 {}_{{\color[rgb]{0,0.75,0.16}\text{+8.6}}} |
| 4-task average | 42.2 | 48.7 {}_{{\color[rgb]{0,0.75,0.16}\text{+6.5}}} | 58.1 {}_{{\color[rgb]{0,0.75,0.16}\text{+15.9}}} | 56.6 {}_{{\color[rgb]{0,0.75,0.16}\text{+14.4}}} |
| Qwen3-Omni-30B | Gemini-2.5-Flash | Gemini-2.5-Pro | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Task | Base | FloorSAV | Base | FloorSAV | Base | SAVVY | FloorSAV | Partial † | Full ‡ |
| Ego Direction | 71.7 | 72.8 | 74.2 | 71.3 | 75.2 | 84.7 | 75.8 | 70.2 | 68.3 |
| Ego Distance | 59.4 | 61.8 | 49.7 | 55.4 | 59.6 | 62.9 | 55.4 | 59.0 | 61.0 |
| Exo Direction | 29.4 | 32.9 | 29.8 | 44.5 | 31.7 | 44.0 | 52.8 | 52.3 | 69.9 |
| Exo Distance | 31.8 | 31.0 | 29.0 | 32.2 | 37.0 | 40.2 | 34.9 | 37.9 | 52.2 |
| Avg. | 48.1 | 49.6 | 45.7 | 50.9 | 50.9 | 58.0 | 54.7 | 54.9 | 62.8 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Egocentric | Exocentric | ||||
|---|---|---|---|---|---|
| # of Frames | Direction | Distance | Direction | Distance | Overall |
| 32 | 73.9 | 63.9 | 31.6 | 31.0 | 50.1 |
| 64 | 70.6 | 62.8 | 31.9 | 28.9 | 48.6 |
| 128 | 72.8 | 61.8 | 32.9 | 31.0 | 49.6 |
| Egocentric | Exocentric | ||||
| Method | Direction | Distance | Direction | Distance | Overall |
| Gemini-2.5-Flash ( Comanici et al., 2025 ) | 74.2 | 49.7 | 29.8 | 29.0 | 45.7 |
| Gemini-2.5-Flash- FloorSAV | 71.3 | 55.4 | 44.5 | 32.2 | 50.9 |
| (–) Semantic Grounding | 63.3 | 58.0 | 34.4 | 30.8 | 46.6 |
| Gemini-2.5-Pro ( Comanici et al., 2025 ) | 75.2 | 59.6 | 31.7 | 37.0 | 50.9 |
| Gemini-2.5-Pro- FloorSAV | 75.8 | 55.4 | 52.8 | 34.9 | 54.7 |
| Task | Gemini-3.6-Flash | Gemini-3.6-Flash- FloorSAV | ( ) Semantic Grounding |
|---|---|---|---|
| Dynamic Relativity | |||
| Ego-to-Exo Direction | 56.0 | 65.0 | 59.5 {}_{{\color[rgb]{1,0,0}\text{--5.5}}} |
| Ego-to-Exo Distance (AbsMRA) | 34.3 | 34.6 | 32.1 {}_{{\color[rgb]{1,0,0}\text{--2.5}}} |
| Exo-to-Ego Direction | 32.0 | 50.5 | 47.5 {}_{{\color[rgb]{1,0,0}\text{--3.0}}} |
| Exo-to-Ego Distance (AbsMRA) | 46.4 | 44.7 | 41.1 {}_{{\color[rgb]{1,0,0}\text{--3.6}}} |
| 4-task average | 42.2 | 48.7 | 45.0 {}_{{\color[rgb]{1,0,0}\text{--3.7}}} |