Automated insertion of board-to-board (BTB) connectors in 3C manufacturing requires both high visual accuracy and strong deployment robustness. This problem remains challenging because multi-variant connectors exhibit significant morphological and appearance variations, making stable cross-variant generalization difficult, while the mismatch between training and deployment under fixed-view inspection settings induces background spurious correlation and degrades real-world performance. To address these issues, this paper proposes MCFR, a Mask-Guided Coarse-to-Fine Regression framework for multi-variant BTB connector assembly. By introducing an object-aware mask prior and explicit photometric refinement, the proposed method suppresses background interference and improves alignment accuracy and robustness in practical deployment. Experiments on a self-constructed multi-variant dataset, a BTB batch insertion testbed, and a real smartphone assembly task show that MCFR consistently outperforms representative baselines and achieves an average real-world insertion success rate of 99.25%. These results demonstrate the effectiveness and practical potential of MCFR for automated assembly of multi-variant BTB connectors.
Figures & tables
Fig. 1: Overview of the proposed method. (a) Initial grasping introduces pose misalignment, and fixed-viewpoint images are acquired. (b) The images are fed into the proposed mask-guided coarse-to-fine regression framework for relative pose estimation. (c) The estimated pose is used for insertion correction, improving insertion success in real-world open scenarios.
Fig. 2: Three-stage architecture of MCFR. Stage I constructs an object-aware probability prior from SAM responses to suppress background interference. Stage II performs mask-guided coarse regression to estimate an initial relative pose. Stage III refines the estimate through image-domain photometric alignment, yielding the final relative pose.
Fig. 3: Experimental platform and representative test objects. (a) Real assembly setup, including the recognition camera, vacuum gripper, BTB mating testbed, detection camera, real smartphone, and component tray. (b) Representative BTB connectors and smartphone camera-module components used in the experiments.
Model
Validation Set
Test Set
Trans. MAE (mm)
Rot. MAE ( ∘ )
Trans. Max. (mm)
Rot. Max. ( ∘ )
Trans. MAE (mm)
Rot. MAE ( ∘ )
Trans. Max. (mm)
Rot. Max. ( ∘ )
YOLO+LDA
0.095
0.117
0.283
0.331
0.094
0.142
0.280
0.428
ResNet50+LSTM
0.103
0.316
0.339
1.027
1.062
1.335
3.690
4.712
MCFR w/o OAMG
0.021
0.047
0.076
0.213
0.793
0.192
3.049
1.138
MCFR w/o RPR
0.021
0.025
0.159
0.245
0.052
0.257
0.533
1.295
MCFR (Ours)
0.013
0.015
0.064
0.498
0.027
0.126
0.196
0.833
TABLE I: Quantitative comparison of different methods on the validation and test sets.
Fig. 4: Comparison of mean errors on the validation and test sets. The figure reports the performance changes under the fixed-background test setting and the stability advantage of the proposed method.
Fig. 5: Occlusion sensitivity visualization. (a) Input image. (b) Occlusion sensitivity map. (c) Overlay result. The top row corresponds to the model without the object-aware mask, while the bottom row corresponds to the model with the mask prior, illustrating the difference in attention regions.
Fig. 6: Complete execution pipeline in the real smartphone assembly task, including recognition, grasping, detection, alignment, insertion, and homing, demonstrating the closed-loop deployment of the proposed method in a practical task.
Fig. 7: Real-world insertion success rates of three methods on 8 BTB connector variants. The horizontal-axis labels encode the pin pitch and pin count of each connector variant, reflecting the difference in assembly difficulty across specifications.
High-precision connector insertion remains challenging for robotic systems due to tight mechanical tolerances, partial observability during contact, and multimodal uncertainty arising from occlusion and contact ambiguity. Successful insertion requires closed-loop contact guidance that continuously integrates global alignment cues with local contact feedback to produce stable corrective actions under interaction. In this work, we present ContactDP (Contact-Guided Diffusion Policy for Tight Insertion Tasks), a multimodal diffusion-policy framework for contact-rich insertion. ContactDP jointly integrates wrist RGB observations, fingertip tactile sensing, and wrist-mounted force-torque measurements to infer contact state and generate temporally consistent corrective motions during insertion. To ensure stable execution under contact, the learned policy operates together with a hybrid position-force controller that provides compliant low-level interaction. We evaluate our approach on a suite of industrial-grade connector insertion tasks with varying connector geometries, grasp conditions, and initial misalignment. Across all tasks, ContactDP significantly outperforms vision-only diffusion policies for performance, reliability and generalization.
Chengyi Xing, Shaoxiong Yao, Diego Romeres +1
Mitsubishi Electric Research Labs, Cambridge, MA, US
Robotic assembly in high-mixture settings requires adaptable systems that can handle diverse parts, yet current approaches typically rely on policies specialized to each insertion task. Although this can reach high success rates, it makes the process of deploying systems for new problems tedious and time consuming. We present a framework for generalizable insertion using world models that combine robot proprioceptive information with raw visual observations captured by a wrist-mounted camera. Our model-based approach trains a single world model on up to 90 insertion tasks with geometrically diverse parts, achieving 56% zero-shot success on unseen objects with unknown geometry compared to just 7% with a model-free baseline. Importantly, performance improves as more objects are included in the training dataset, demonstrating strong scalability. Lastly, finetuning the generalist model on held-out objects significantly enhances data-efficiency compared to training from scratch and, in some cases, achieves better asymptotic performance. To our knowledge, this is the first system capable of assembling unseen objects in an entirely data-driven manner, and thus represents a significant step toward scalable, generalizable robotic assembly systems.
Nicklas Hansen, Iretiayo Akinola, Yijie Guo +7
NVIDIA · University of California San Diego · Work completed during an internship at NVIDIA +1
High-precision insertion remains a fundamental challenge in robotic manipulation due to the strict alignment requirements and contact-rich interactions involved. Although peg-in-hole tasks are widely used for evaluation, existing bench- marks often rely on fixed task configurations, limiting their ability to assess robustness and generalization across different insertion scenarios. This paper introduces a reconfigurable peg-in-hole benchmark designed to evaluate task generalization in high-precision insertion. The benchmark consists of a set of fully 3D-printable modular components, including multiple peg geometries, tolerance levels, and configurable base structures that can be combined to generate a large variety of insertion and assembly tasks. By varying object layouts, orientations, and task structures while maintaining controlled physical conditions, the benchmark enables systematic evaluation of adaptation to unseen scenarios. To support reproducibility, we additionally provide a scenario generation tool capable of producing standardized task configurations and machine-readable task descriptions. The scenario generation tool and the STL files of the benchmark pieces are available through the project repository: https://github.com/aistairc/peg-in-bench.
Yosel Delgado, José G. Buenaventura-Carreón, Floris Erich +4
National Institute of Advanced Industrial Science and Technology (AIST), Japan