Fine-Tuning VLM for Enhancing AI's Spatial Intelligence: Understanding 3D and 2D Rotations
Organizations: Department of Architecture Texas A&M University College Station, USA
Abstract
Spatial intelligence is a fundamental skill in multiple domains, such as Science, Technology, Engineering, and Mathematics (STEM), Medicine, Architecture, and Construction. Recent studies indicate that Vision-Language Models (VLMs) still face limitations in spatial reasoning, which inhibits artificial intelligence (AI) from performing practical spatial tasks. Using multiple object-rotation datasets developed for training and evaluation, our experiments demonstrated promising improvements in both 2D and 3D rotation detection. Fine-tuned Google DeepMind-built Gemma-4 mixture-of-experts (MoE) models significantly outperformed fine-tuned Gemma-4 generalist models in predicting rotations defined by both their axes and angles. Fine-tuning also substantially improved angle estimation for 2D representation without requiring an explicit coordinate system. Furthermore, identifiable objects did not improve angle-detection accuracy; instead, objects with prominent linear features showed improved performance.
Figures & tables
| Model | Baseline | Fine-tuned | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Axis Acc. | Axis Acc. | |||||||||
| Generalist | 34.0% | 0.0% | 2.1% | 2.1% | 4.3% | 31.9% | 0.0% | 8.5% | 17.0% | 31.9% |
| MoE | 53.0% | 6.3% | 11.8% | 12.5% | 18.8% | 98.0% | 18.8% | 58.8% | 87.5% | 93.8% |
| Object | Baseline | Fine-tuned Model | ||||
|---|---|---|---|---|---|---|
| Identifiable (Exp 2.1) | 2.58% | 4.50% | 7.30% | 62.36% | 84.40% | 87.86% |
| Arbitrary (Exp 2.2) | 2.49% | 3.98% | 7.32% | 82.15% | 85.73% | 85.92% |