K-OPSD: Verifiable On-Policy Self-Distillation for Post-Training Vision-Language Models on AEC Drawings
Organizations: Amazon Web Services
Abstract
Interpreting architecture, engineering, and construction (AEC) drawings is hard for general Multimodal Large Language Models (MLLMs) and vision-language models (VLMs). We introduce K-OPSD, a VLM post-training methodology for improving AEC drawing understanding. Building on On-Policy Self-Distillation (OPSD) with verifiable supervision, we construct a teacher from the model's own best-of-N generations, certified by a process-level verifier, and rescue failed prompts by resampling under a hint that exposes the verified answer. We then perform an on-policy model update by training on verified completions with a cross-entropy inner-loss, outperforming the bounded token-wise generalized Jensen-Shannon divergence (JSD) used by on-policy distillation. Using K-OPSD, we fine-tune Qwen3-VL models on the AECV-Bench dataset. The resulting models attain the top average judge score (0.819) and combined accuracy (0.738), achieving competitive results against open-source baseline models. The recipe transfers to the out-of-domain ArchCAD dataset, where the 8B model gains most. We present the verifier suite and the continual learning and self-improving pipeline, our results provide preliminary evidence that verifier-guided self-distillation is a promising route toward more reliable machine reading of architecture drawings.
Figures & tables
| AECV-Bench | ArchCAD | ||||
| Avg Judge Score | Combined Accuracy | Avg Judge Score | Combined Accuracy | ||
| Open-source model | Qwen3-VL-235B | 0.698 | 0.667 | 0.538 | 0.268 |
| Pixtral-Large-2502 | 0.687 | 0.643 | 0.514 | 0.293 | |
| Kimi-K2.5 | 0.765 | 0.714 | 0.619 | 0.415 | |
| Gemma-3-27B | 0.731 | 0.714 | 0.543 | 0.268 | |
| Llama4-Maverick-17B | 0.795 | 0.738 | 0.478 | 0.171 | |
| Qwen3-VL-2B | Qwen3-VL-4B | Qwen3-VL-8B | ||||
| Avg Judge Score | Combined Accuracy | Avg Judge Score | Combined Accuracy | Avg Judge Score | Combined Accuracy | |
| AECV-Bench | ||||||
| Base model | 0.417 | 0.619 | 0.771 | 0.738 | 0.738 | 0.643 |
| OPD | 0.665 | 0.595 | 0.780 | 0.714 | 0.792 | 0.738 |
| OPSD | 0.588 | 0.476 | 0.713 | 0.690 | 0.760 | 0.690 |
| VS-OPSD | 0.676 | 0.667 | 0.730 | 0.667 | 0.736 | 0.690 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Teacher model | Verification Gate | Inner-loop Loss |
| OPD | External (Claude Fable 5) | None | generalized JSD |
| OPSD | On-policy self-teacher | None | generalized JSD |
| VS-OPSD | On-policy self-teacher | process-level verifier | generalized JSD |
| K-OPSD (Ours) | On-policy self-teacher | process-level verifier | Cross-entropy |