DCM-SAM: Defect-Conditioned Mixture of LoRA Experts for NPU-Deployed AM Defect Segmentation
Organizations: Department of Industrial and Management Systems Engineering West Virginia University, Morgantown, WV · School of Mathematical and Data Sciences West Virginia University, Morgantown, WV · Lane Department of Computer Science & Electrical Engineering West Virginia University, Morgantown, WV
Abstract
Metal additive manufacturing parts are inspected by X-ray computed tomography, where labelled data is scarce, the pores and inclusions that matter span a few pixels, and inspection must happen at the machine. We present DCM-SAM, a defect-conditioned adaptive mixture of LoRA experts: one frozen Segment Anything backbone carries a separate Conv-LoRA expert bank and mask decoder per defect class, each trained in its own pass, without prompts, on synthetic slices alone, updating only 4.4% of the parameters. On benchmarks that XCT-SAM reports, DCM-SAM improves on every baseline for both classes from a ViT-B backbone against their ViT-H, and reaches 64.2% pore IoU on real NIST scans having seen no real images during training. Deployment then exposes what adaptation work rarely measures: on a Qualcomm Hexagon NPU, ViT-H and ViT-L compile yet cannot allocate at 1024x1024 image resolution, since activations rather than weights exceed the device ceiling, and quantizing weights does not help. ViT-B alone runs, but the adapted encoder then fails to allocate where the stock one succeeds, until a numerically identical rewrite of the attention lets the complete DCM-SAM run in FP16 at 1024x1024, with no operator falling back to the CPU, masks within 0.01% of pixels of the FP32 reference. Code: https://github.com/MushfiqShovon/DCM-SAM.
Figures & tables
| Role | Synth. pore | Synth. incl. | NIST pore (OOD) | |
|---|---|---|---|---|
| corner | 32.8 | 49.6 | 64.4 | |
| pore sweep | 29.4 | 47.7 | 65.1 | |
| composed optimum | 35.0 | 46.3 | 64.3 | |
| (selected) | pore sweep | 36.1 | 48.9 | 64.2 |
| pore sweep | 34.6 | 44.9 | 65.8 | |
| incl. sweep | 36.5 | 47.5 | 65.3 |
| Method | Pores (synthetic) | Inclusions (synthetic) | Pores (NIST, OOD) |
|---|---|---|---|
| SAM Kirillov et al. (2023) (zero-shot) | 14.3 | 29.5 | 49.4 |
| UNet++ Zhou et al. (2020) | 14.0 | 10.2 | 13.4 |
| MedSAM Ma et al. (2024) | 16.3 | 32.5 | 33.8 |
| SAM-Med2D Cheng et al. (2023) | 18.7 | 36.0 | 57.8 |
| Conv-LoRA-SAM Zhong et al. (2024) | 27.7 | 24.1 | 54.6 |
| XCT-SAM Hasan et al. (2026) | 32.5 | 36.4 | 58.5 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| ViT-H, | ViT-H, (raw / rescaled) | ViT-B , | |
|---|---|---|---|
| Pore IoU, image-averaged | 37.3 | 31.7 / 30.8 | 36.1 |
| Pore IoU, defect-only | 35.8 | 27.9 / 27.6 | 32.7 |
| Inclusion IoU, image-averaged | 52.5 | 51.8 / 46.6 | 48.9 |
| Inclusion IoU, defect-only | 15.5 | 9.3 / 20.7 | 11.8 |
| Real NIST pore IoU | – | – | 64.2 |
| Runs on the IQ-9075 with adapters | no | – | yes, verified |
| Pore sweep ( ) | Incl. sweep ( ) | |||
| Pore | Incl. | Incl. | Pore | |
| 1 | 35.5 | 53.8 | 54.3 | 34.5 |
| 2 | 36.3 | 53.2 | 53.3 | 36.6 |
| 4 | 36.2 | 54.3 | 54.4 | 32.6 |
| 8 | 36.1 | 53.5 | 53.5 | 36.1 |
| Direct runs: / ; (selected) / | ||||
| Component | ViT-H | Share | ViT-B | Share |
|---|---|---|---|---|
| Frozen SAM backbone (encoder, prompt encoder) | 637.0 M | frozen | 89.7 M | frozen |
| Conv-LoRA | 327,680 | 0.051% | 73,728 | 0.079% |
| Convolutional experts | 9,728 | 0.002% | 3,648 | 0.004% |
| MoE gating ( ) | 1,024 | 0.001% | 384 | 0.001% |
| Mask decoder (incl. IoU head) | 4.06 M | 0.633% | 4.06 M | 4.33% |
| Trainable per head | 4.40 M | 0.69% | 4.13 M | 4.41% |
| Graph | Conv-LoRA | Attention | On device | Cosine |
| Stock ViT-B | no | SDPA | runs | 0.99994 |
| Stock ViT-B | no | eager | runs | 0.99994 |
| DCM-SAM pore head, as trained | yes | eager | fails, needs 4.07 GB | n/a |
| DCM-SAM pore head, rewritten | yes | SDPA | runs | 0.99988 |
| DCM-SAM inclusion head, rewritten | yes | SDPA | runs | 0.99995 |
| Static one-hot routing vs. dynamic top- (PyTorch, eight slices): of M mask pixels differ | ||||
| Artifact | Experts | NPU ops | Inference (ms) | Peak mem. (MB) | Load (ms) |
|---|---|---|---|---|---|
| Pore encoder | 1012 | 1980.8 | 11.8 | 825 | |
| Pore decoder | n/a | 205 | 3.8 | 9.6 | 160 |
| Inclusion encoder | 1444 | 8134.3 | 13.9 | 868 | |
| Inclusion decoder | n/a | 205 | 3.8 | 9.6 | 167 |
| Split | IoU (%) | Dice (%) | Tol-F1 (%) | Precision (%) | Recall (%) |
|---|---|---|---|---|---|
| NIST test2 | 50.0 | 66.5 | 85.1 | 74.2 | 100.0 |
| NIST test3 | 64.2 | 78.0 | 94.1 | 99.2 | 90.0 |
| NIST test4 | 54.5 | 70.4 | 89.0 | 80.6 | 99.9 |
| NIST test5 | 66.6 | 79.9 | 96.2 | 98.3 | 94.2 |
| NIST test6 ∗ | 85.4 | 92.0 | 96.8 | 93.8 | 100.0 |
| Mean | 64.2 | 77.4 | 92.2 | 89.2 | 96.8 |