cs.AISep 28, 2026

K-OPSD: Verifiable On-Policy Self-Distillation for Post-Training Vision-Language Models on AEC Drawings

Authors: Yunfei Bai, Enrico Chionna, Akash Amol, Kawaljit Singh KC, Joern Tinnemeyer

Organizations: Amazon Web Services

Abstract

Interpreting architecture, engineering, and construction (AEC) drawings is hard for general Multimodal Large Language Models (MLLMs) and vision-language models (VLMs). We introduce K-OPSD, a VLM post-training methodology for improving AEC drawing understanding. Building on On-Policy Self-Distillation (OPSD) with verifiable supervision, we construct a teacher from the model's own best-of-N generations, certified by a process-level verifier, and rescue failed prompts by resampling under a hint that exposes the verified answer. We then perform an on-policy model update by training on verified completions with a cross-entropy inner-loss, outperforming the bounded token-wise generalized Jensen-Shannon divergence (JSD) used by on-policy distillation. Using K-OPSD, we fine-tune Qwen3-VL models on the AECV-Bench dataset. The resulting models attain the top average judge score (0.819) and combined accuracy (0.738), achieving competitive results against open-source baseline models. The recipe transfers to the out-of-domain ArchCAD dataset, where the 8B model gains most. We present the verifier suite and the continual learning and self-improving pipeline, our results provide preliminary evidence that verifier-guided self-distillation is a promising route toward more reliable machine reading of architecture drawings.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding

    May 29, 2026Qian Kou, Xiaofeng Shi, Yulin Li +4Visual ReasoningMultimodal Large Language Models

  2. Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation

    Jul 25, 2026Shuai Wang, Daoan Zhang, Zhe Tang +2Recent Vision-Language ModelsUnsupervised On-Policy Self-Distillation

  3. OPD-V: Visual On-Policy Self-Distillation with Modality Balance

    Aug 5, 2026Aniri, Jinhe Bi, Peng Liao +5Multimodal Large Language ModelsRecent Vision-Language Models