Organizations: Materials Informatics Initiative, RD Technology and Digital Transformation Center, JSR Corporation, 3-103-9 Tonomachi, Kawasaki-ku, Kawasaki, Kanagawa, Japan, 210-0821. · Division of Information Science, Graduate School of Science and Technology, Nara Institute of Science and Technology, 8916-5 Takayama-cho, Ikoma, Nara, Japan, 630-0192.
Robotic manipulation of labware is difficult when transparent or reflective objects must be identified and localized. Coded planar fiducials are a practical retrofit: easy to print, they leave the marked face flat and graspable. Yet a single planar tag is least reliable in near-frontal views, where perspective cues fade. Non-planar geometries restore those cues but intrude on the flat face that a parallel-jaw gripper must contact. Our idea is to tilt multiple tags within one compact footprint, so that each tag is seen at a non-frontal angle even when the marker faces the camera. We propose the quARtet marker, a 3D-printable fiducial embodying this idea: all detected corners of its four tilted AprilTags enter one Perspective-n-Point solve, and a shared configuration defines the fabricated geometry and the detector model. Because tilting consumes flat area, its three layouts trade pose-estimation consistency against graspability. In robot-referenced, same-setup fixed-camera experiments, all three layouts reduced the mean frontal orientation error from 2.18 degree for a single planar tag to 0.24-0.47 degree and the root-mean-square position error from 1.50 to 0.17-0.20 mm. A robot-mounted-camera pose-hold test confirmed this separation under closed-loop visual feedback. In swing-down trials under identical conditions, the two layouts with flat contact strips retained the object with about 2 mm of in-grasp slip, whereas the layout without flat strips slipped by roughly 100 mm. For the tested conditions, the results support a rule: the layout without flat strips when pose-estimation consistency dominates, a layout with flat strips when the marked face must remain graspable.
Figures & tables
Figure 1 : Target context and the four practical demands on a laboratory fiducial. (a) A robot operates among weakly textured, transparent, or specular labware within the hand-camera field of view. (b) The hand-camera view, annotated for illustration: the camera must distinguish the pipette to be picked (green box) from two visually similar transparent vials (gray boxes). (c)–(f) The marker must provide identification, reliable pose estimation during a near-frontal approach, access for grasping, and low-cost printability.
Figure 2 : Overview of the quARtet system and its shared-parameter workflow, shown for the pitch layout (quARtet-P). (a) Design and print: a single configuration file ( setting.json ) is read by an add-in for the CAD software Fusion 360. The add-in constructs the marker as an editable model, and the model is then fabricated by FDM printing. (b) Detection: the 16 tag corners observed in the camera image are paired with the corresponding 3D corner coordinates reconstructed from the same parameters, and a single solvePnP call over these 16 2D–3D correspondences yields the marker pose, visualized by the orientation axes overlaid at the parameter-derived marker center.
Figure 3 : The three implemented quARtet layouts, one per row: (a)–(c) quARtet-D (diagonal), (d)–(f) quARtet-P (pitch), (g)–(i) quARtet-rP (reversed-pitch). Each layout is shown as a 3D overview rendering (left), orthographic top and side views (center), and a photograph of the FDM-printed marker (right). The top views show the arrangement of the four tags on the shared square footprint, and the side views show their out-of-plane tilt ( 18\tcdegree in the present designs). The layouts share the four-tag concept and differ in how the tags are tilted and, consequently, in the graspability of the marker face. Geometric parameters are summarized in Table 1 .
Parameter name
Description
AR{0,1,2,3}_file_path
File path for the input image of AprilTag 0–3 (one parameter per tag).
tag_family
Family name of the AprilTags.
tilt_mode
Marker layout type: “diagonal”, “pitch”, or “pitch_r” (reversed-pitch).
whole_size_mm
Overall dimension (mm) of the square defined by the outer tag corners.
tilt_degree
Rotational inclination (degrees) of each tag relative to the base plane.
tag_size_mm
Edge length (mm) of an individual AprilTag.
Table 1 : Configurable parameters of the shared-parameter generation-and-detection framework. A single setting.json containing these parameters is consumed by both the model-generation and detection pipelines.
Figure 4 : Fixed-camera setup for the pose-estimation evaluation. The robot provided repeatable changes in orientation and position while the industrial camera remained stationary.
Figure 5 : Closed-loop pose-hold setup and protocol. (a) The camera is mounted on the end-effector and observes a quARtet marker fixed on the table. (b) Close-up of the 3D-printed fixture that holds the camera on the end-effector. (c) One cycle of the closed loop, drawn as a side view of the camera and the fixed marker with the current camera image as an inset. The initial near-frontal view of the marker is locked as the target (dashed camera). The camera is then orbited about the marker to the displaced start, 12.8\tcdegree off the marker normal with a 6\tcdegree rotation about the viewing axis, so that the marker stays in view (0). At every control step, the camera-to-marker pose is estimated from the current image (1) and compared with the locked target view (2), which gives the orientation hold error and the end-effector pose that restores the target view. The end-effector is moved there (3), and the rotation between consecutive end-effector poses is the per-step reorientation. Steps 1–3 repeat for 120 s.
Figure 6 : Mean orientation error per commanded cell over the broad tilt grid (up to 10 consecutive frames per cell, Supplementary Section S3.1), with per-cell values annotated. Errors are relative to the robot-commanded reference after the constant alignment of Section 4.2.1 . Panels are ordered quARtet-D, quARtet-P, quARtet-rP, and Single, with one color scale shared by all. Frames within a cell are repeated images of one held pose, not independent re-mountings. Cell-wise summaries are given in Supplementary Section S2.3.
Figure 7 : Mean Euclidean position error per commanded cell over the translation grid (up to 10 consecutive frames per cell, Supplementary Section S3.1), with per-cell values annotated. Panels are ordered quARtet-D, quARtet-P, quARtet-rP, and Single, with one color scale shared by all. Errors are residuals to the robot-commanded relative translation after the same-dataset constant alignment of Section 4.2.1 . Cell-wise summaries are given in Supplementary Section S2.3.
Figure 8 : Representative closed-loop pose-hold traces, one run per marker type, in the order quARtet-D, quARtet-P, quARtet-rP, and Single. Over the common ∼117 s window, the plotted medians are 0.07 – 0.08\tcdegree for the three quARtet layouts and 0.47\tcdegree for Single in panel (a), and 0.04 – 0.08\tcdegree and 0.24\tcdegree , respectively, in panel (b). Dotted horizontal lines mark the pooled quARtet (blue) and Single (red) medians.
Fixed camera
Closed-loop pose hold
Marker
Frontal orientation error [ \tcdegree ]
Position RMSE [mm]
Orientation hold error [ \tcdegree ]
Per-step end-effector reorientation [ \tcdegree ]
quARtet-D
0.24±0.03
0.17
0.08±0.01∗∗
0.05±0.01∗
quARtet-P
0.27±0.03
0.20
0.09±0.02∗∗
0.07±0.02∗
quARtet-rP
0.47±0.07
0.19
0.11±0.01∗∗
0.10±0.01∗
Single
2.18±0.80
1.50
0.68±0.06
0.31±0.10
Table 2 : Summary of the RQ1 results, with the Single control last. Fixed camera: orientation is the mean ± SD over the 10 consecutive frames at the frontal cell of the broad grid, and position is the RMSE over the translation grid, both under the operational-reference and constant-alignment conventions of Section 4.2.1 . Closed-loop pose hold: mean ± SD over five runs per marker type of the within-run mean orientation hold error and per-step end-effector reorientation, evaluated over the common ∼117 s window of Fig. 8 . The fixed-camera entries summarize technical repeats within one setup and carry no inferential marks. In the closed-loop columns, asterisks mark quARtet entries that differ significantly from Single (two-sided Welch t -tests on the five run means, Holm-corrected within each metric, ∗ : adjusted p<0.05 , ∗∗ : adjusted p<0.01 ). The smallest value in each column is bold.
Figure 9 : Swing-down test object and marker contact surfaces. (a) The 40×40×200 mm object grasped near its marker end by the two-finger gripper. The flat fingers close directly on the marker faces, and the silver spheres on the gripper and object are the retroreflective spheres of the motion-capture system. (b)–(e) Close-ups of the markers affixed to the grasped end, in the order quARtet-D, quARtet-P, quARtet-rP, and Single (control).
Figure 10 : One-way swing-down motion (quARtet-P shown). The red arcs indicate the commanded rotation of the gripper about the swing axis: raising to +25\tcdegree , then swinging down to −25\tcdegree .
Figure 11 : Swing-down outcome and representative time series. (a),(b) Snapshots at the post-swing hold for quARtet-D and quARtet-P. The dashed line extends the gripper’s finger axis, along which the object would lie if the grasp had remained rigid: the quARtet-P object stays on it, whereas the quARtet-D object has pivoted away and slid outward (quARtet-rP and Single behaved like quARtet-P, Table 3 ). (c)–(f) Gripper (solid green) and object (dashed red) rotation about the swing axis relative to the pre-swing hold, for the trial closest to the median position slip of each marker type, with the downward swing plotted as a positive rotation. The dashed style keeps both curves visible where they coincide. Time is aligned at swing onset, each window spans the two bracketing static holds, and all panels share one scale. All five trials per marker type are shown in Supplementary Section S7.
Marker
Position slip [mm]
Orientation slip [ \tcdegree ]
quARtet-D
103.53±14.97∗∗
24.85±3.69∗∗
quARtet-P
1.71±0.12∗∗
0.31±0.02
quARtet-rP
2.02±0.14∗∗
0.32±0.05
Single
1.02±0.03
0.30±0.03
Table 3 : In-grasp slip after the one-way swing-down (mean ± SD, N=5 independently re-grasped trials per marker type). Smaller is better, the best value in each column is bold, and Single is last as the control. Asterisks mark quARtet entries that differ significantly from Single (two-sided Welch t -tests, Holm-corrected within each metric, ∗∗ : adjusted p<0.01 ). All starred entries have adjusted p<0.001 (Supplementary Section S7).
Figure S1 : Simulation-based rationale for selecting a moderate tag tilt. (a) Heatmap of the mean rotation error Eˉ as a function of roll and pitch, computed from the 16 one-corner pixel-perturbation cases at each pose. (b) Scatter plot of the same mean rotation error against the geodesic angular distance from the frontal reference orientation.
Figure S2 : Per-image orientation estimates on the narrow grid ( −10\tcdegree to +10\tcdegree in 5\tcdegree increments). Black plus signs mark commanded orientations, and dot color distinguishes the commanded cells without encoding error. Coordinates follow the operational-reference and constant-alignment conventions of main-text Section 4.2.1 and Section S3.
Figure S3 : Per-image orientation estimates on the broad grid ( −30\tcdegree to +30\tcdegree in 15\tcdegree increments). Black plus signs mark commanded orientations. Conventions as in Fig. S2 .
Figure S4 : Per-image position estimates over the commanded translation grid. Black plus signs mark commanded positions, and dot color distinguishes the commanded cells without encoding error. Coordinates follow the operational-reference and constant-alignment conventions of main-text Section 4.2.1 and Section S3.
Figure S5 : Difference maps of the per-cell mean orientation error on the broad grid (quARtet layout minus Single, so blue means the layout’s error is smaller). The map is a descriptive summary of the cell-mean differences. No frame-level test is applied, because the frames of a cell are consecutive technical repeats of one held pose. The large improvement is concentrated at the frontal cell, while at oblique cells the pitch-based layouts pay penalties of up to about +1.3\tcdegree in cell-mean error and quARtet-D of up to about +0.4\tcdegree . Differences are computed from unrounded cell means, so end digits can differ by 0.01 from the rounded values of Table S1 .
Figure S6 : Difference maps of the per-cell mean position error (quARtet layout minus Single, so blue means the layout’s error is smaller). The map is a descriptive summary of the cell-mean differences. No frame-level test is applied, because the frames of a cell are consecutive technical repeats of one held pose. All three layouts have smaller errors over almost the entire grid, the single exception being the cell at (−20,0) (at most +0.11 mm). Differences are computed from unrounded cell means.
(θx,θy) [ \tcdegree ]
quARtet-D
quARtet-P
quARtet-rP
Single
(−30,−30)
0.42±0.02
1.63±0.88
1.09±0.21
0.31±0.03
(−30,−15)
0.33±0.03
0.56±0.21
0.34±0.02
0.31±0.04
(−30,+0)
0.27±0.01
0.60±0.03
0.39±0.03
0.31±0.05
(−30,+15)
0.22±0.02
0.27±0.28
0.58±0.01
0.18±0.04
(−30,+30)
0.24±0.01
0.63±0.03
0.67±0.26
0.32±0.01
(−15,−30)
0.60±0.02
0.97±0.02
1.32±0.02
0.45±0.02
Table S1 : Cell-wise orientation residual on the broad grid: mean ± SD over the valid consecutive frames of each commanded cell, in degrees, in the common result order with the Single control last. The smallest mean in each row is bold (ties resolved on unrounded means). Frames within a cell are technical repeats of one held pose at one mounting, so these are descriptive within-acquisition summaries, and no inferential asterisks are attached. The companion difference maps (Figs. S5 and S6 ) display the same cell-mean differences graphically.
(Δx,Δz) [mm]
quARtet-D
quARtet-P
quARtet-rP
Single
(−20,−40)
0.25±0.01
0.06±0.01
0.10±0.01
1.87±0.26
(−20,−20)
0.29±0.01
0.11±0.02
0.03±0.01
1.02±0.14
(−20,+0)
0.24±0.02
0.26±0.02
0.21±0.03
0.15±0.04
(−20,+20)
0.23±0.03
0.16±0.02
0.10±0.02
0.61±0.20
(−20,+40)
0.26±0.03
0.17±0.01
0.16±0.02
0.37±0.20
(−10,−40)
0.10±0.01
0.17±0.01
0.21±0.01
0.80±0.18
Table S2 : Cell-wise Euclidean position residual on the translation grid: mean ± SD over the valid consecutive frames of each commanded cell, in mm, in the common result order with the Single control last. The smallest mean in each row is bold (ties resolved on unrounded means). Frames within a cell are technical repeats of one held pose at one mounting, so these are descriptive within-acquisition summaries, and no inferential asterisks are attached.
quARtet-D
quARtet-P
quARtet-rP
Tag
Corner
x
y
z
x
y
z
x
y
z
0
c1
-17.50
-17.50
0.00
-17.50
-17.50
0.00
-1.50
-1.50
0.00
c2
-1.89
-17.89
3.50
-1.50
-17.50
0.00
-17.50
-1.50
0.00
c3
-2.28
-2.28
6.99
-1.50
-2.28
4.94
-17.50
-16.72
4.94
c4
-17.89
-1.89
3.50
-17.50
-2.28
4.94
-1.50
-16.72
4.94
1
c1
17.50
-17.50
0.00
17.50
-17.50
0.00
1.50
-1.50
0.00
Table S3 : Detector-side 3D corner coordinates (mm) of the three implemented layouts for the configuration used in this study ( whole_size_mm=35 , tag_size_mm=16 , tilt_degree=18 ), generated by the released detector code from the shared configuration (with the detector’s tilt argument set to −tilt_degree , see the text). Coordinates are in the marker frame of Section 3.3 (origin at the footprint center on the base plane, z normal to it and pointing toward the tag side). Tag t∈{0,1,2,3} is the AprilTag with ID t ; corners c1 – c4 follow the AprilTag corner convention of the local tag square before tilting. These 16 points per layout are the object points of the joint PnP solve.
Entry
Δx [m]
Δy [m]
Δz [m]
Δrx [ \tcdegree ]
Δry [ \tcdegree ]
Δrz [ \tcdegree ]
1
0.00
0.00
0.00
0.0
0.0
0.0
2
0.03
0.00
0.00
0.0
10.0
0.0
3
-0.03
0.00
0.00
0.0
-10.0
0.0
4
0.00
0.03
0.00
10.0
0.0
0.0
5
0.00
-0.03
0.00
-10.0
0.0
0.0
6
0.00
0.00
0.03
0.0
0.0
10.0
Table S4: Robot pose displacements used for hand–eye estimation in the robot-mounted-camera configuration (Section S4). The robot was moved from a near-frontal initial pose to the displaced poses listed in the table. Δx , Δy , and Δz are translational displacements from the initial pose in meters, and Δrx , Δry , and Δrz are increments added to the components of the initial pose’s rotation vector in degrees. At each pose, the robot kinematics and the marker observation were recorded and used for hand–eye transform estimation. Entry 1 is the undisplaced initial pose.
Table S5: Estimated hand–eye transforms Tcamee for the robot-mounted-camera configuration, expressed as 4×4 homogeneous matrices. Translations are in meters. A separate estimate was performed for each marker type on the day of the pose-hold runs, because the marker observations enter the optimization objective described in Section S4.
Figure S7 : Orientation hold error of all five closed-loop runs per marker type over the common ∼117 s window, in the order quARtet-D, quARtet-P, quARtet-rP, and Single, with shared axis limits. The bold trace is the representative run shown in the main text.
Figure S8 : Per-step end-effector reorientation of all five closed-loop runs per marker type over the common ∼117 s window, in the order quARtet-D, quARtet-P, quARtet-rP, and Single, with shared axis limits. The bold trace is the representative run shown in the main text.
Marker
Run
Hold error mean [ \tcdegree ]
Hold error median [ \tcdegree ]
Reorientation mean [ \tcdegree ]
Reorientation median [ \tcdegree ]
quARtet-D
1
0.088
0.069
0.054
0.043
quARtet-D
2 †
0.079
0.067
0.050
0.041
quARtet-D
3
0.079
0.065
0.064
0.054
quARtet-D
4
0.062
0.049
0.042
0.034
quARtet-D
5
0.072
0.059
0.054
0.044
mean ± SD
0.076±0.010
0.053±0.008
Table S6: Per-run summary of the closed-loop pose-hold experiment (five runs per marker type, common ∼117 s window): within-run mean and median of the orientation hold error and of the per-step end-effector reorientation, in degrees. The run marked † is the representative run plotted in the main text (mean hold error closest to the median of the five runs). Group rows give the mean ± SD of the five run means, the values reported in the closed-loop columns of main-text Table 2.
Figure S9 : Idealized planar-grasp model and effective graspable regions on the 35 mm marker footprint. The angle labeled θ in panel (a) is the contact angle α of Section S6.1. quARtet-D exposes no effective region under this definition, whereas the pitch-based layouts retain strips. Single is the fully planar reference.
Marker
w=8.75 mm
w=17.5 mm
w=35 mm
quARtet-D
0
0
0
quARtet-P
0.12
0.16
0.22
quARtet-rP
0.18
0.28
0.45
Single
1
1
1
Table S7 : Normalized mean area-ratio for the idealized model.
Marker
Amplitude [ \tcdegree ]
Duration [s]
Peak speed [ \tcdegree /s]
quARtet-D
49.99±0.01
0.65±0.01
182±2
quARtet-P
49.99±0.01
0.63±0.01
181±1
quARtet-rP
49.96±0.00
0.64±0.02
181±3
Single
50.01±0.00
0.65±0.02
181±3
All 20 swings
49.99±0.02
0.64±0.02
181±2
Table S8 : Repeatability of the realized swing-down motion, measured by motion capture (mean ± SD over the N=5 swings of each marker type, and pooled over all 20 swings). Amplitude is the change in median gripper tilt between the pre- and post-swing holds. Duration is the time to traverse 5 – 95% of that amplitude. Peak speed is the sustained peak of the swing-axis tilt rate ( 5 -frame median filter at 120 Hz). Brief unfiltered frame-to-frame peaks reach ∼230\tcdegree /s for at most 17 ms. The near-identical values across trials and marker types show that every marker type experienced the same commanded motion and inertial load. The contact conditions at the fingers are layout-specific by design, and the grip force was not measured in situ.
Figure S10 : Gripper tilt profiles of all 20 swings (4 marker types × 5 trials), aligned at swing onset with the swing shown positive. The trajectories coincide in shape across trials and marker types, complementing the scalar metrics of Table S8 .
Figure S11 : All five swing-down time series per marker type (gripper solid green, object dashed red, with the dashed style keeping both visible where they coincide). Time is aligned at swing onset and each window spans the two bracketing holds. The five gripper curves of each panel coincide within 3\tcdegree at any instant and separate visibly only during the fast transient, so the trials repeat the same motion. The object curves separate persistently only for quARtet-D.
Figure S12 : Per-trial in-grasp orientation slip traces: for each trial, the geodesic angle between the current gripper-relative object rotation and its pre-swing-hold mean. The mean of each curve over the central 0.5 s of the post-swing hold is that trial’s orientation slip, which enters the per-marker-type mean of the main-text swing-down table and is plotted in Fig. S13 . Insets magnify the 0 – 2\tcdegree range for the three retaining marker types. All five quARtet-D trials show the same failure sequence, a transient during the swing followed by a large permanent offset ( 19 – 28\tcdegree ), whereas quARtet-P, quARtet-rP, and Single show a brief transient of at most ∼1.5\tcdegree (elastic deformation of the grasp and tracking noise during the fast motion) that returns to ∼0.3\tcdegree .
Figure S13 : Per-trial in-grasp position and orientation slip ( N=5 independently re-grasped trials per marker type, logarithmic axes), in the order quARtet-D, quARtet-P, quARtet-rP, and Single. Red bars mark per-marker-type means.
In many autonomous applications requiring real-time localization, active marker-based systems are preferred due to their low latency and ease of deployment compared to computationally demanding feature-based methods. Event~\mbox{cameras} offer high temporal resolution and minimal delay and are commonly used with active LED markers for robust real-time localization. Existing methods typically rely on Perspective-n-Point (PnP) solvers for pose estimation. However, structured marker layouts can be challenging to deploy in space-constrained scenarios, while partial self-motion information (e.g., gravity direction and altitude) is readily available from onboard sensors. We derive a robust and accurate minimal solver that estimates camera pose from only two LED markers by incorporating known tilt angle and camera height measured by an onboard sensor, such as an IMU or an altimeter. The proposed formulation uniquely determines the camera pose through both a closed-form and a linear least-squares solution. We further analyze degenerate configurations and characterize the conditions under which height information does not contribute to rotation estimation. For evaluation, we developed an event-based active marker system to collect real-world data with ground truth from a motion capture system. Experiments on both synthetic and real data demonstrate improved accuracy over the state-of-the-art P2P solver and competitive performance relative to P3P.
Runze Yuan, Alexander Kappler, Jun Zhang +5
Institute of Visual Computing, Graz University of Technology · MIS laboratory, University of Picardie Jules Verne · ICB laboratory, University of Burgundy Europe
Accurate 6-DoF object pose estimation is fundamental to robotic manipulation, yet vision-based methods often fail under occlusion, poor lighting, and reflective or transparent surfaces. We present YOTO, a tactile-only pose estimation system that recovers the full 6-DoF object pose from a single pair of simultaneous contacts, without requiring contact history. YOTO represents each tactile contact as a local 3D point cloud and localizes it on the object surface through a coarse-to-fine network. The two localized contacts, together with the calibrated sensor poses, are then fed to a closed-form normal-aware SVD solver that recovers the full 6-DoF object pose in one step. To reduce real-data requirements, the localization network is pretrained on virtual tactile patches sampled from the object model and fine-tuned with a small number of real contacts. We further show that YOTO can operate on object models reconstructed from consumer-grade mobile scans, and quantify the gap relative to CAD-based models. Experiments on four geometrically diverse objects demonstrate accurate tactile contact localization and pose estimation, outperforming vision-based and geometric baselines, especially when visual perception is unreliable. Code, trained models, and the real GelSight dataset will be released upon publication.
Pengfei Ye, Yuxiang Ma, Haonan Chen +5
MIT CSAIL · Harvard University · University of Cambridge
Flexible endoscopic continuum manipulators offer high dexterity and access to complex anatomy, but nonlinear hysteresis limits feedforward control accuracy. Closed-loop control can compensate for these errors but requires accurate six-degree-of-freedom (6D) end-effector pose feedback. We present MarUco, a markerless 6D pose estimation framework for closed-loop control using only stereo vision during operation. A photorealistic pseudo-rigid-body simulation pipeline generates large-scale annotated training data without manual labeling. A multifeature fusion network integrates masks, keypoints, heatmaps, and bounding boxes from both stereo views to estimate an initial pose, followed by a learned single-pass render-and-compare module that refines the pose without iterative optimization. Kinematics-free hand-eye calibration estimates camera-to-robot-base extrinsics, and self-supervised adaptation uses unlabeled real stereo pairs to mitigate sim-to-real pose error. Across 1,000 real samples, MarUco achieves translation and rotation errors of 0.78 ± 0.50 mm and 3.07 ± 1.25°, respectively, with a total processing time of 52.4 ms per stereo pair. In closed-loop experiments over eight reference paths, MarUco achieves a mean terminal translation error of 1.8 mm, an 88% reduction relative to uncompensated open-loop control. To the best of our knowledge, this is the first markerless position-based visual servoing framework for continuum manipulators.
Junhyun Park, Chunggil An, Myeongbo Park +3
Department of Robotics and Mechatronics Engineering, DGIST, Daegu 42988, Republic of Korea · AI Research Lab, DEEPNOID, Seoul, Republic of Korea