cs.CVOct 1, 2026

Task-Adaptive Grounded 3D-Programmers Using 2D VLMs

Authors: Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva, Jan-Nico Zaech, Luc Van Gool, Danda Pani Paudel

Organizations: TextQL · INSAIT, Sofia University “St. Kliment Ohridski”

Abstract

Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate Framing (CCF) and Task-Adaptive Feedback (TAF). CCF serves as a unified visual representation that anchors both inputs and outputs to a shared Euclidean coordinate system, solving common challenges in 3D grounding such as axis ambiguity, inconsistent metric scale, and floating references. Complementary to this structured framing of the 3D inputs, TAF closes the reasoning loop with task-adaptive dynamic feedback that enables 2D VLMs to perform varied open-vocabulary tasks within their native visual context. Building on this foundation, we introduce 3D-Prog, a 3D understanding, reasoning, and generation framework that jointly employs the capabilities of CCF and TAF together with powerful VLMs. Without requiring any retraining, 3D-Prog performs open-vocabulary 3D understanding, manipulation, and generation across both object-level and scene-level tasks. Our experiments show that the joint use of CCF and TAF transforms 2D VLMs into geometry-aware 3D programmers, achieving consistent, interpretable, and high-quality results across diverse 3D tasks.

Explore similar work

CardsList
  1. 3D-Aware VLMs with Implicit and Explicit Geometries

    Jul 23, 2026Wenhao Li, Xueying Jiang, Quanhao Qian +43D Spatial Reasoning3D Object Detection

  2. VLM3: Vision Language Models Are Native 3D Learners

    May 28, 2026Zhipeng Cai, Zhuang Liu, Yunyang Xiong +33D Scene Understanding3D Generation

  3. PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding

    Dec 24, 2025Seongmin Jung, Seongho Choi, Gunwoo Jeon +23D Visual GroundingRecent Vision-Language Models