STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction
Organizations: Department of Computer Science, The University of Texas at Austin
Abstract
Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation. However, it remains unclear whether VLMs can accurately understand complex social navigation scenes (e.g., inferring the spatial-temporal relations among agents and human intentions), which is essential for safe and socially compliant robot navigation. While some recent works have explored the use of VLMs in social robot navigation, no existing work systematically evaluates their ability to meet these necessary conditions. In this paper, we introduce the Social Navigation Scene Understanding Benchmark (SocialNav-SUB), a Visual Question Answering (VQA) dataset and benchmark designed to evaluate VLMs for scene understanding in real-world social robot navigation scenarios. SocialNav-SUB provides a unified framework for evaluating VLMs against human and rule-based baselines across VQA tasks requiring spatial, spatiotemporal, and social reasoning in social robot navigation. Through experiments with state-of-the-art VLMs, we find that while the best-performing VLM achieves an encouraging probability of agreeing with human answers, it still underperforms simpler rule-based approach and human consensus baselines, indicating critical gaps in social scene understanding of current VLMs. Our benchmark sets the stage for further research on foundation models for social robot navigation, offering a framework to explore how VLMs can be tailored to meet real-world social robot navigation needs. An overview of this paper along with the code and data can be found at https://larg.github.io/stars.
Figures & tables
| Method | 1% | 5% | 10% | 25% | 50% | 100% | |
|---|---|---|---|---|---|---|---|
| SEAN-T Competence | MLP | ||||||
| Autoencoder | |||||||
| Random Forest | |||||||
| STARS-He (Ours) | |||||||
| STARS-Ho (Ours) | |||||||
| SEAN-T Surprise | MLP |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | # Scenarios | Sample Freq. | Window Size | Window Duration | Max Nodes | Max Edges |
|---|---|---|---|---|---|---|
| SEAN-T | 2969 | 5 Hz | 40 | 8s | 16 | 240 |
| SocialNav-SUB | 3052 | 8 Hz | 20 | 2.5s | 31 | 930 |
| Model | Dataset | Robot features | Human features |
|---|---|---|---|
| STARS-He | SEAN-T | 592D map and goal features | 8D learned embedding + 8D pooled edge features |
| STARS-He | SNS | 8D learned embedding + 8D pooled edge features | 8D learned embedding + 8D pooled edge features |
| STARS-Ho | SEAN-T | 592D map and goal features + type embedding | 592D zero vector + type embedding |
| STARS-Ho | SNS | 592D zero vector + type embedding | 592D zero vector + type embedding |
| Component | Homogeneous | Heterogeneous |
|---|---|---|
| Node feature projection | shared | per dataset and node type |
| Temporal edge encoder | shared ( ) | per dataset ( , ) |
| Message-passing layers | shared | shared |
| Edge-type embedding | not used | shared |
| Latent heads | shared | shared |
| Edge decoder | shared ( ) | per dataset |
| Model / Component | Stage | Trainable Params. | Frozen Params. | Notes |
|---|---|---|---|---|
| STARS-He Encoder/Decoder | Pretraining | 6,420,810 | N/A | Heterogeneous graph autoencoder |
| STARS-Ho Encoder/Decoder | Pretraining | 2,936,848 | N/A | Homogeneous graph autoencoder |
| Linear Prediction Head | Downstream adaptation | 1,285 (SNS), 771 (SEAN) | 6,420,810 | Used with frozen STARS encoder |
| Model / Pipeline | Dataset / Task | Pretraining Time | Downstream Training Time | Inference Time / Scenario | Peak GPU Memory |
|---|---|---|---|---|---|
| STARS-He + Linear Head | SEAN-T + SNS | 2.5 hrs / run | 1 sec | 12 ms | 13 GB |
| STARS-Ho + Linear Head | SEAN-T + SNS | 3 hrs / run | 12 sec | 30 ms | 17 GB |
| Task | Method | 1% | 5% | 10% | 25% | 50% | 100% |
|---|---|---|---|---|---|---|---|
| Competence | STARS-He (Ours) | ||||||
| STARS-Ho (Ours) | |||||||
| Surprise | STARS-He (Ours) | ||||||
| STARS-Ho (Ours) | |||||||
| Intention | STARS-He (Ours) | ||||||
| STARS-Ho (Ours) |
| Task | Model | 1% | 5% | 10% | 25% | 50% | 100% |
|---|---|---|---|---|---|---|---|
| Competence | STARS-He | ||||||
| Surprise | STARS-He | ||||||
| Intention | STARS-He |
| Model | Pretraining | 1% | 5% | 10% | 25% | 50% | 100% |
|---|---|---|---|---|---|---|---|
| STARS-He | SEAN-T only | ||||||
| STARS-He | SEAN-T+SNS | ||||||
| STARS-Ho | SEAN-T only | ||||||
| STARS-Ho | SEAN-T+SNS |
| Model | Pretraining | Task | 1% | 5% | 10% | 25% | 50% | 100% |
|---|---|---|---|---|---|---|---|---|
| STARS-He | SNS only | Competence | ||||||
| STARS-He | SNS only | Surprise | ||||||
| STARS-He | SNS only | Intention | ||||||
| STARS-He | SEAN-T+SNS | Competence | ||||||
| STARS-He | SEAN-T+SNS | Surprise | ||||||
| STARS-He | SEAN-T+SNS | Intention |
| Task | Method | 1% | 5% | 10% | 25% | 50% | 100% |
|---|---|---|---|---|---|---|---|
| Avoid | Hetero Self-Sup | ||||||
| Hetero Sup | |||||||
| Follow | Hetero Self-Sup | ||||||
| Hetero Sup | |||||||
| Not Consider | Hetero Self-Sup | ||||||
| Hetero Sup |
| Architecture | 1% | 5% | 10% | 25% | 50% | 100% | ||
|---|---|---|---|---|---|---|---|---|
| Random encoder: | Competence | Heterogeneous | ||||||
| Surprise | Heterogeneous | |||||||
| Intention | Heterogeneous | |||||||
| Avoid | Heterogeneous | |||||||
| Follow | Heterogeneous | |||||||
| Not Consider | Heterogeneous |
| Architecture | 1% | 5% | 10% | 25% | 50% | 100% | ||
|---|---|---|---|---|---|---|---|---|
| 128D robot-z only | Competence | Heterogeneous | ||||||
| Surprise | Heterogeneous | |||||||
| Intention | Heterogeneous | |||||||
| Avoid | Heterogeneous | |||||||
| Follow | Heterogeneous | |||||||
| Not Consider | Heterogeneous |
| Architecture | 1% | 5% | 10% | 25% | 50% | 100% | ||
|---|---|---|---|---|---|---|---|---|
| 128D byst only | Competence | Heterogeneous | ||||||
| Surprise | Heterogeneous | |||||||
| Intention | Heterogeneous | |||||||
| Avoid | Heterogeneous | |||||||
| Follow | Heterogeneous | |||||||
| Not Consider | Heterogeneous |
| Architecture | 1% | 5% | 10% | 25% | 50% | 100% | |
|---|---|---|---|---|---|---|---|
| Competence | Heterogeneous | ||||||
| Surprise | Heterogeneous | ||||||
| Intention | Heterogeneous |
| Architecture | 1% | 5% | 10% | 25% | 50% | 100% | ||
|---|---|---|---|---|---|---|---|---|
| WRB: | Competence | Heterogeneous | ||||||
| Surprise | Heterogeneous | |||||||
| Intention | Heterogeneous | |||||||
| Avoid | Heterogeneous | |||||||
| Follow | Heterogeneous | |||||||
| Not Consider | Heterogeneous |
| Model | Avoid | Follow | Overtake | Yield | NC | PA |
|---|---|---|---|---|---|---|
| Gemini | ||||||
| 4o | ||||||
| o4-mini |