CGGT: Curve-Grounded Geometry Transformer for 3D Parametric Curve Reconstruction
Authors: Zhirui Gao, Renjiao Yi, Yunfan Ye, Ruizhen Hu, Chenyang Zhu, Wei Chen, Kai Xu
Organizations: National University of Defense Technology, China · Hunan University, China · Shenzhen University, Shenzhen, China · Institute of AI for Industries, Chinese Academy of Sciences, China
Recovering editable 3D parametric curves from 2D images is a fundamental challenge in computer graphics, bridging pixel-based perception and vector-based CAD modeling. Existing NeRF- and 3DGS-based methods often rely on dense calibrated views, precomputed 2D edge maps, and costly per-scene optimization, limiting their applicability to casually captured real-world inputs. We propose CGGT, a Curve-Grounded Geometry Transformer that directly grounds 3D-consistent 2D curve instances in the image space from sparse, unposed multi-view images. CGGT combines a geometry-aware transformer encoder for multi-view feature learning with a curve-aware masked-attention decoder for cross-view instance association. In a single forward pass, it predicts camera parameters, dense depth maps, and instance-level 2D curve masks, which are then lifted into 3D and refined through a fast parametric optimization stage to recover compact, editable 3D curve primitives. To support structured curve learning, we introduce Wireframe-100K, a large-scale dataset comprising 100,000 CAD models with diverse topologies, realistic multi-view renderings, and accurate parametric curve annotations. Extensive experiments show that our framework achieves substantial improvements in both reconstruction accuracy and efficiency, particularly under challenging sparse-view settings and in separating persistent 3D structural edges from view-dependent image edges caused by silhouettes, textures, and appearance variations. Despite being trained solely on synthetic data, CGGT generalizes well to real-world images, demonstrating its potential for practical CAD-style wireframe reconstruction from unconstrained visual inputs.
Parametric CAD reconstruction requires recovering both precise geometry and editable modeling operations from visual observations, making it challenging under limited and ambiguous views. Existing methods mainly rely on 2D appearance cues and lack strong multi-view geometric priors. In this work, we present VGGT-CAD, a geometry-aware framework for parametric CAD reconstruction from single- and multi-view observations. We transfer pretrained 3D geometric priors into CAD reconstruction by encoding camera parameters as condition tokens and jointly modeling them with image tokens. To handle varying numbers of viewpoints, we introduce a variable-view cross-view context aggregation module that adaptively fuses multi-view features. We further develop a training-free geometry-aware view selection strategy to select complementary and reliable frames during inference. The resulting representation is decoded into CAD command sequences using a non-autoregressive decoder. We also develop VideoCAD, a large-scale multi-view video benchmark derived from existing CAD data through multi-view re-rendering. Extensive experiments demonstrate the effectiveness of VGGT-CAD for visual CAD reconstruction under different observation configurations.
Chunan Yu, Tianrun Chen, Fu Shen +3
School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, China · KOKONI3D, Moxin (Huzhou) Technology Co., Ltd., China · Nanjing Institute of Agricultural Mechanization, Ministry of Agriculture and Rural Affairs, Nanjing 210014, China
Reconstructing editable parametric CAD models from a single-view image is of great practical value for modern manufacturing, yet remains challenging due to incomplete geometric observations and complex inter-part relationships. To address it, we propose CADForge, an agentic framework that progressively converts a single image into CadQuery programs. CADForge decomposes an object into CAD-meaningful components and performs explicit geometric reasoning for each component, a process that first identifies CAD-relevant constraints and then translates them into precise modeling parameters through mathematical code. The inferred parameters then drive component-wise synthesis of executable CadQuery programs, with a review agent evaluating the resulting geometry and providing targeted feedback for iterative refinement. To further improve robustness and efficiency, CADForge incorporates a failure-guided toolkit construction mechanism to distill accumulated experience into tools, and maintains a compact parametric CAD memory for retrieving modeling context on demand. Experiments on diverse single- and multi-part objects show that CADForge consistently outperforms existing baselines in reconstruction fidelity and perceptual quality, demonstrating an effective approach to accurate single-view CAD reconstruction.
Keyang Lu, Zhifei Yang, Tianao Dong +3
Peking University · Shenzhen Loop Area Institute · Beijing Normal University
Reconstructing 3D geometry from 2D engineering line drawings is an inherently ambiguous problem: while visible strokes determine the object's projected structure, they do not specify the depth of each stroke. Rather than treating this problem as sketch-based asset generation, where models often infer unobserved structure, we study projection-faithful wireframe reconstruction: lifting a user-provided drawing into 3D according to its visible strokes. We formulate this task as conditional depth estimation over line drawings, predicting a depth value for each drawn pixel to produce a 3D wireframe. To model the ambiguities of orthographic projection, we implement a Latent Diffusion Model with spatial conditioning on the input sketch and optional partial-depth conditioning for iterative reconstruction. We train and evaluate our models on over three million synthetic image-depth pairs derived from CAD wireframes, including a newly curated corpus of roughly 90,000 shapes. Across varying shape complexities, our framework achieves robust reconstruction performance; scaling from 256 to 512 resolution with a retrained latent space roughly halves reconstruction error, reaching a 3.9% best-of-five (7.0% average) normalized depth error. These results demonstrate the potential of projection-faithful depth estimation as a user-controlled approach for iterative 3D wireframe creation in engineering design.
Elton Cao, Hod Lipson
Creative Machines Lab, Columbia University New York, NY