cs.CVSep 6, 2026

3DHarnessBench: Probing Agentic 3D-to-Code Capabilities of Frontier Vision-Language Models

Authors: Ling Liu, Bingchen Gong, Amal Dev Parakkat, Maks Ovsjanikov

Organizations: Institut Polytechnique de Paris · École Polytechnique · Télécom Paris

Abstract

We introduce 3DHarnessBench, a benchmark that evaluates the agentic ability of frontier vision-language models (VLMs) to recover 3D geometry as Blender Python code from a variety of inputs. Unlike previous frameworks that prompt the VLMs with a fixed input (e.g., a single rendering or a text description), 3DHarnessBench evaluates four separate harness settings that progressively enable active agentic exploration, facilitated by recent Blender MCP functionality. Our hierarchy from Single-view, Multi-view, Active Visual (arbitrary viewpoint access), and Full 3D Interaction (complete access to the target object through Blender function calls) probes the models' abilities in both visual perception and active inference, tool calling, and self-correction. We observe that the ability of all frontier models to recover 3D geometry improves significantly with richer function call access, although the improvements are strongly model-dependent, revealing highly uneven agentic 3D-to-code capabilities. We will release the benchmark, code, outputs, and agent trajectories for reproducible 3D evaluation.

Explore similar work

CardsList
  1. SceneActBench: Can Agents Act on the 3D Scenes They See?

    Jul 24, 2026Yifei Zhao, Xiangxin Zhou, Wenhao Yang +11AI Agent BenchmarksTool-Using Vision-Language Agents

  2. 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code

    May 31, 2026Yipeng Gao, Lei Shu, Genzhi Ye +5VLM Evaluation3D Asset Generation

  3. 3D Primitives are a Spatial Language for VLMs

    May 12, 2026Junze Liu, Kun Qian, Florian Dubost +8Spatial Reasoning Benchmarks3D Spatial Reasoning