cs.CVSep 21, 2026

What Survives on Real Drawings: Active Sampling, Connectome Wiring, and Matched Baselines in Architectural Document Vision

Authors: Dmitry Kuklev

Organizations: Independent Researcher

Abstract

A connectome-constrained model of the fly visual system, optimized for motion and then frozen, can be driven over architectural drawings by prescribed motion and used as a texture representation. We compare it with information-matched baselines that see the same 721 photoreceptor samples. On clean synthetic data the frozen model transfers but loses to task training: 0.857 area-weighted accuracy in one-shot hatch matching versus 0.959 for a 5,888-parameter CNN, and 0.619 IoU in wall segmentation versus 0.905 for a matched network. Under scan noise and thickened strokes, the trained networks lose up to 0.188 accuracy while the frozen pipeline loses 0.030. On fourteen production sheets, opened once, a 1,876-parameter fly model reaches 0.505 average precision versus 0.415 for a network two hundred times larger. A preregistered held-out split confirms the clean-data ordering: 0.835 for the circuit, 0.894 for receptors only, and 0.971-0.980 for trained CNNs. Rewiring the connectome while preserving degrees or type pairs and transmitter signs costs 0.271-0.356 accuracy across three seeds, so the exact wiring is load-bearing. Yet the intact circuit does not beat its moving retina, and T4/T5 silencing leaves both tasks intact. Longer observations reverse the circuit-receptor ordering once the stimulus spans a period, but not through T4/T5. Thus active sampling and exact structure matter, while clean-data practical performance remains dominated by task-trained networks and the useful transfer margin is largely retinal.

Explore similar work

Aug 1, 2026cs.CV

CrossProjection: Geometric Grounding Beyond Viewpoint Change in Architectural Drawings

Architectural drawings violate the usual assumption behind multi-view reasoning: plans and sections are cuts, while elevations are facade projections, so corresponding components change appearance in ways camera motion cannot explain. We introduce CrossProjection, an anchor-grounded diagnostic of whether vision-language models preserve component identity and externalize geometry across heterogeneous architectural views. It evaluates Matching, Registration, and Geometric Grounding through categorical judgments, candidate selection, and free point, line, and region localization. Across 23 real drawing sets and 1,954 categorical conditions per model, GPT-5.5 scores 82.4%, Qwen3-VL-32B-Instruct 62.2%, and GLM-4.5V 57.2%. A matched 200-target study crosses natural and vector-text-suppressed drawings with closed-candidate and free-geometry outputs. Candidate-supported performance is often higher, but free localization remains fragile: on natural drawings, point/region PCK@.05 is 54-76% for GPT, 8-10% for Qwen, and 14-36% for GLM; line endpoint PCK@.05 is 22%, 4%, and 0%. A coordinate grid recovers some GPT point/region precision but not lines. Three architecture-trained participants reach 87.3-93.3% categorical accuracy and 76-92% GT-region hit, supporting task feasibility rather than a population-level human ceiling. Because the categorical families do not form a same-item Matching-Registration contrast and interface controls alter multiple burdens, we avoid mechanistic claims. The supported conclusion is narrower: closed-choice or marked-element success does not entail reliable explicit geometric grounding. For drawing-guided CAD/BIM systems, categorical correctness should not be treated as evidence of candidate-free spatial reliability. Reusable on-sheet anchors, fixed-denominator scoring, and hash-locked artifacts establish an audit trail for this gap.
Kaho Li, Pengyu Zeng, Yuqin Dai +3
Jul 21, 2026cs.CV

Benchmarking Deep Learning Approaches for AEC Engineering Drawing Layout Detection and Information Extraction

Information Extraction (IE) from Architecture, Engineering, and Construction (AEC) drawings remains hindered by manual inefficiency, while Layout Detection, a vital 'middleware' organizing graphical and textual hierarchies, is underexplored. General document layout models, optimized for text-centric content, lack validation on engineering drawings. This study constructs a custom AEC-specific layouts dataset and benchmarks five deep learning architectures. RF-DETR achieves state-of-the-art performance with an mAP50mAP_{50} of 0.949, while the Vision-Language Model Qwen3-VL attains a leading F1-score of 0.911. Conversely, models pre-trained on general document datasets suffer from "domain interference", causing performance degradation. This establishes a robust technical foundation for automated IE in AEC.
Tianyang Huang, Alessio Lombardi, Ahmed Elnagar +7
May 29, 2026cs.CV

MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding

Multimodal Large Language Models (MLLMs) have demonstrated significant achievements in general visual question answering (VQA) tasks. However, they remain brittle on mechanical engineering drawings, where high annotation density and weak domain knowledge, compounded by unreliable spatial relation reasoning under strict projection rules and geometric constraints, make decisive cues easy to miss and frequently lead to wrong answers. To bridge this gap, we introduce the first comprehensive mechanical drawing understanding dataset, MechVQA, created through a semi-automated construction and quality-control pipeline. MechVQA contains 3.3k high-density pictures with 21K question-answer pairs, spanning 10 different fine-grained tasks across three capability levels: Recognition, Reasoning, and Judging, providing a testbed to evaluate and improve MLLM understanding on real-world mechanical drawings. On top of MechVQA, we then develop the MechVL model through a multi-stage training paradigm, building a strong domain-specialized baseline. Extensive experimental results demonstrate that MechVL outperforms the strongest closed-source baseline by 7.57 percentage points on the MechVQA total score, significantly enhancing mechanical drawing understanding ability and providing a reusable foundation for deploying MLLMs in mechanical design and inspection scenarios.
Qian Kou, Xiaofeng Shi, Yulin Li +4