cs.AIOct 3, 2026

Fine-Tuning VLM for Enhancing AI's Spatial Intelligence: Understanding 3D and 2D Rotations

Authors: Uttamasha Monjoree, Wei Yan

Organizations: Department of Architecture Texas A&M University College Station, USA

Abstract

Spatial intelligence is a fundamental skill in multiple domains, such as Science, Technology, Engineering, and Mathematics (STEM), Medicine, Architecture, and Construction. Recent studies indicate that Vision-Language Models (VLMs) still face limitations in spatial reasoning, which inhibits artificial intelligence (AI) from performing practical spatial tasks. Using multiple object-rotation datasets developed for training and evaluation, our experiments demonstrated promising improvements in both 2D and 3D rotation detection. Fine-tuned Google DeepMind-built Gemma-4 mixture-of-experts (MoE) models significantly outperformed fine-tuned Gemma-4 generalist models in predicting rotations defined by both their axes and angles. Fine-tuning also substantially improved angle estimation for 2D representation without requiring an explicit coordinate system. Furthermore, identifiable objects did not improve angle-detection accuracy; instead, objects with prominent linear features showed improved performance.

Figures & tables

Explore similar work

CardsList
  1. Rotated, but How Far? Diagnosing and Improving Object-Rotation Reasoning in VLMs

    Oct 5, 2026Zhaochen Wang, Yujun Cai, Huangbo Zou +4VLM EvaluationObject Pose Estimation

  2. How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study

    Apr 16, 2026Zhen Yang, Ping Jian, Zhongbin Guo +5Visual Spatial ReasoningLarge Vision-Language Models

  3. Agentic AI with Structured CoT for Enhancing AI's Spatial Intelligence: Visualization and Reasoning of Rotation

    Oct 3, 2026Uttamasha Monjoree, Wei YanSpatial Reasoning Benchmarks3D Spatial Reasoning