Beyond Pixels: A Vector-to-Graph Framework for Reliable Schematic Auditing
Organizations: Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ), Shenzhen, China · Guangdong Power Grid Co., Ltd., Yangjiang Power Supply Bureau, Yangjiang, China · Shanghai University, Shanghai, China · Carleton University, Ottawa, Canada
Abstract
Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual understanding, yet they suffer from a critical limitation: structural blindness. Even state-of-the-art models fail to capture topology and symbolic logic in engineering schematics, as their pixel-driven paradigm discards the explicit vector-defined relations needed for reasoning. To overcome this, we propose a Vector-to-Graph (V2G) pipeline that converts CAD diagrams into property graphs where nodes represent components and edges encode connectivity, making structural dependencies explicit and machine-auditable. On a diagnostic benchmark of electrical compliance checks, V2G yields large accuracy gains across all error categories, while leading MLLMs remain near chance level. These results highlight the systemic inadequacy of pixel-based methods and demonstrate that structure-aware representations provide a reliable path toward practical deployment of multimodal AI in engineering domains. To facilitate further research, we release our benchmark and implementation at https://github.com/gm-embodied/V2G-Audit.
Figures & tables
| Summary (avg., ) | Conn. (%) | Ground. (%) | Wiring (%) | Overall (%) |
| Baseline avg. | 6% | 17% | 13% | 12% |
| +V2G avg. | 67% | 44% | 33% | 47% |
| ( +V2G Base) | +61% | +27% | +20% | +35% |