Object reconstruction and inspection tasks play a crucial role in various robotics applications. Identifying paths that reveal the most unknown areas of the object is paramount in this context, as it directly affects reconstruction efficiency. Current methods often use sampling based path planning techniques, evaluating views along the path to enhance reconstruction performance. However, these methods are computationally expensive as they require evaluating several candidate views on the path. To this end, we propose a computationally efficient solution that relies on calculating a focus point in the most informative region and having the robot maintain this point in the camera field of view along the path. In this way, object reconstruction related information is incorporated into the whole body control of a mobile manipulator employing a visibility constraint without the need for an additional path planner. We conducted comprehensive and realistic simulations using a large dataset of 114 diverse objects of varying sizes from 57 categories to compare our method with a sampling based planning strategy and a strategy that does not employ informative paths using Bayesian data analysis. Furthermore, to demonstrate the applicability and generality of the proposed approach, we conducted real world experiments with an 8 DoF omnidirectional mobile manipulator and a legged manipulator. Our results suggest that, compared to a sampling based strategy, there is no statistically significant difference in object reconstruction entropy, and there is a 52.3% probability that they are practically equivalent in terms of coverage. In contrast, our method is 6.2 to 19.36 times faster in terms of computation time and reduces the total time the robot spends between views by 13.76% to 27.9%, depending on the camera FoV and model resolution.
Figures & tables
Fig. 1 : First, an NBV is calculated. Then, the robot moves to the NBV while focusing on a Focus Point and avoiding obstacles in the workspace, enabling the robot to reveal more unknown areas safely.
Fig. 2 : Illustration of the general framework, in which the dashed box highlights our contribution. Given the current partial object model and a set of candidate views, an NBV algorithm (RSV) [ 1 ] selects the next view to visit. Our method then uses this target NBV and the current partial model to select an informative focus point and employ a whole-body control strategy to reach the target while keeping the focus point within the camera FoV with a visibility constraint [ 11 ] , avoiding obstacles with vector field inequalities (VFIs) [ 12 ] , and mitigating local minima at obstacle boundaries with a circulation constraint [ 13 ] .
Fig. 3 : A 2D representation of the general concept adopted by sampling-based IPP strategies.
Fig. 4 : The search space for candidate view generation is defined by 40 positions evenly distributed around a cylinder with a radius of 3 m. Each position has five orientations, resulting in a search space comprising 200 views.
Fig. 5 : Illustration of the stages for focus point calculation. Steps 1–5 are repeated until the robot reaches the NBV.
Fig. 6 : Top view of the geometrical relationship between the focus point and the angle of the camera FoV. The restricted zone ( ΩR ) and the visibility zone ( ΩV ) are represented with gray and white, respectively, whereas blue dashed lines are the left and right FoV planes. Furthermore, pf is the focus point and pc is the origin of the camera FoV.
Fig. 7 : The illustration of the position and direction objectives.
Fig. 8 : Box and cylinder obstacle boundaries are represented as compositions of infinite primitives.
Fig. 9 : Illustration of the obstacles objects and their boxes and cylinders.
Fig. 10 : Average object coverage, entropy, and total time for consecutive NBV steps across 114 runs in simulation. Our method ( focus ) and the sampling method both demonstrate similar object coverage and entropy performance, consistently outperforming the no path method. Additionally, our method requires significantly less time than the sampling method.
Parameter
Value
Description
r
0.03 m
model resolution
dmaxsensor
4.5 m
maximum distance of depth sensor
FoV
( 74∘ , 60∘ )
horizontal and vertical angles of the camera FoV
Expanded FoV
( 90∘ , 90∘ )
horizontal and vertical angles of the FoV used for the focus point calculation
df
2.5 m
distance between the focus point and the camera center
dth
0.75 m
the threshold distance for the visibility constraint
TABLE I : Parameters used in our implementation
Focus
Sampling
No path
μ(σ)
μ(σ)
μ(σ)
Coverage
0.833 ( 0.093 )
0.84 (0.094)
0.794 (0.103)
AUC
0.684 ( 0.083 )
0.693 (0.086)
0.613 (0.103)
Entropy ( ×103 )
1796.25 (128.95)
1805.75 ( 123.30 )
1989.85 (155.32)
Tcomp (s)
11.7 ( 3.24 )
203.61 (74.03)
—
Texec (s)
588.76 (124.43)
629.81 (136.74)
555.37 ( 81.65 )
TABLE II : The mean ( μ ) and standard deviation ( σ ) of object coverage, AUC, entropy, and total time ( Ttotal=Tcomp+Texec ) along with computation time ( Tcomp ) and execution time ( Texec ) for the focus , sampling and no path methods at the 10 th NBV.
Fig. 11 : The difference of means from the simulation data between the three methods, focus ( μfocus ), sampling ( μsampling ), and no path ( μNoPath ), for the object coverage ( a–c ), entropy ( d–f ) and total time, Ttotal ( g–i ). Our method ( focus ) and the sampling method show similar performance regarding object coverage and entropy, while both significantly outperform the no path method. Additionally, our method significantly reduces computation time compared to the sampling method.
FoV=(74∘,60∘)
FoV=(60∘,48∘)
FoV=(45∘,36∘)
r=0.12\penaltym
r=0.03\penaltym
r=0.03\penaltym
Focus
Sampling
No path
Focus
Sampling
No path
Focus
Sampling
No path
Coverage
0.826 (0.093)
0.828 ( 0.092 )
0.794 (0.103)
0.806 ( 0.1 )
0.814 ( 0.1 )
0.76 (0.113)
0.78 (0.107)
0.78 ( 0.106 )
0.701 (0.127)
AUC
0.679 (0.09)
0.683 ( 0.085 )
0.61 (0.103)
0.657 (0.09)
0.665 ( 0.089 )
0.561 (0.108)
0.62 ( 0.099 )
0.616 (0.1)
0.486 (0.117)
Entropy ( ×103 )
1825.21 (121.096)
1848.14 ( 117.952 )
1989.85 (155.32)
1880.64 ( 133.466 )
1911.60 (146.925)
2109.81 (168.758)
2024.99 ( 133.559 )
2080.97 (143.440)
2252.76 (161.614)
Tcomp (s)
3.27 ( 1.02 )
63.31 (28.77)
—
11.81 ( 2.70 )
136.59 (53.16)
—
12.57 ( 2.65 )
77.98 (26.08)
—
TABLE III : Mean (sd) object reconstruction performance in terms of coverage, AUC, entropy, and total time ( Ttotal=Tcomp+Texec ) along with computation time ( Tcomp ) and execution time ( Texec ) for varying camera FoV and model resolution.
Only Focus
Only Visibility
Both
μ(σ)
μ(σ)
μ(σ)
Coverage
0.794 (0.103)
0.798 (0.101)
0.833 ( 0.093 )
AUC
0.613 (0.103)
0.622 (0.097)
0.684 ( 0.083 )
Entropy ( ×103 )
1989.85 (155.32)
2051.85 (148.2)
1796.25 ( 128.95 )
Ttotal (s)
555.37 ( 81.65 )
551.46 (93.75)
600.46 (126.72)
TABLE IV : The mean ( μ ) and standard deviation ( σ ) of object coverage, AUC, entropy, and total time ( Ttotal ) are reported for three cases at the 10 th NBV: using only the focus point estimation strategy ( Only Focus), using only the visibility constraint ( Only Visibility), and using both strategies ( Both ).
Fig. 12 : Illustration of collision avoidance for the robot base and end-effector. The base and end-effector paths are shown in blue and red, respectively.
Fig. 13 : The first experimental setup consists of the object to be reconstructed, enclosed within a cylinder (2 m in diameter), along with four circular obstacles and an 8-DoF holonomic mobile manipulator. The robot base is enclosed by a cylinder with a diameter of 0.5 m, whereas obstacles O1,…,O3 have a diameter of 0.3 m and O4 has a diameter of 0.24 m.
Fig. 14 : Covered volume, entropy, and total time for consecutive NBV steps in the real-world experiment with the holonomic wheeled mobile manipulator for a single run. Our method ( focus ) and the sampling method both demonstrate similar covered volume and entropy performance, outperforming the no path method except for final NBVs where all three methods converge to similar values. Additionally, our method requires significantly less time than the sampling method.
Without Obstacles
With Obstacles
μ(σ)
μ(σ)
Coverage
0.844 (0.078)
0.862 ( 0.069 )
AUC
0.699 (0.075)
0.71 ( 0.068 )
Entropy ( ×103 )
1770.51 (71.141)
1719.60 ( 67.974 )
Ttotal (s)
600.93 ( 139.01 )
942.57 (193.24)
TABLE V : The mean ( μ ) and standard deviation ( σ ) of object coverage, AUC, entropy, and total time ( Ttotal ) are reported at the 10 th NBV for 20 objects using our method in environments with and without obstacles.
Fig. 15 : Snapshots of the reconstruction process for a barrier and a chair using a legged manipulator with a camera mounted on its end-effector.
Fig. 16 : Covered volume and entropy for consecutive NBV steps in the real-world experiment with the legged manipulator when reconstructing a chair and a barrier.
Fig. 17 : Planning an informative path with an RRT* sampling-based strategy.
Active 3D reconstruction of moving objects requires selecting informative viewpoints while accounting for object motion uncertainty during the decision-to-execution delay. Existing methods address only parts of this problem: next-best-view (NBV) planners for object reconstruction typically optimize surface coverage but assume static objects, while motion-aware active perception for moving targets accounts for target motion but prioritizes tracking or visibility over reconstruction coverage. This work presents a motion-uncertainty-aware NBV framework for reconstructing an unknown rigid object undergoing planar motion, using noisy planar position measurements of the object and depth observations from a mobile robot. The key idea is to evaluate each candidate viewpoint by its expected observation quality over plausible future object states induced by motion and measurement uncertainty, rather than at a single predicted object pose. To obtain this predictive belief, a fixed-lag Gaussian Process smoother estimates and predicts the object state from noisy position measurements. The resulting belief is used to generate candidate viewpoints around the predicted object location, filter them by reachability, and estimate their expected coverage-driven scores. Simulation and real-world experiments demonstrate improved reconstruction completeness over non-predictive NBV and prediction-only tracking methods, bridging coverage-driven active reconstruction and prediction-driven tracking.
Karen Li, Mattia Mantovani, Robert J. Wood +2
Harvard University, Cambridge, MA 02134, USA. · University of Modena and Reggio Emilia, 42122 Reggio Emilia, Italy.
Object-centric view planning is a core component of active geometric 3D reconstruction in robotics, yet existing evaluations often conflate object complexity, planning difficulty, budget assumptions, and physical reachability constraints. As a result, conclusions drawn from idealized view-planning evaluations may not reliably predict performance under realistic reconstruction settings. We introduce ObjView-Bench, an evaluation framework for rethinking difficulty and deployment in object-centric view planning. First, we disentangle three quantities underlying view-planning evaluation: omnidirectional self-occlusion as an object-side attribute, observation saturation difficulty, and protocol-dependent planning difficulty defined through a set-cover formulation. This separation supports controlled dataset construction, analysis of slow-saturation objects, and a case study showing that planning difficulty-aware sampling can improve learned view planners. Second, we design deployment-oriented evaluation protocols that reveal how budget regimes and reachable-view constraints alter method behavior. Across classical, learned, and hybrid planners, ObjView-Bench shows that difficulty, budget, and reachability constraints substantially change method rankings and failure modes.
Fast dexterous grasping while a mobile base remains in motion requires coordinated whole-body control and rapid adaptation to physical contact. We propose FastGrasp, a two-stage learning framework that integrates grasp guidance, whole-body control, and tactile feedback. First, a pretrained conditional variational autoencoder generates diverse grasp candidates from object point clouds. Candidates are filtered by approach direction and supporting-surface constraints, then ranked using an envelopment-based criterion combining grasp width and depth coverage measures. Second, a reinforcement learning policy jointly controls the mobile base, arm, and dexterous hand, using the selected grasp as guidance and tactile observations for online grasp adjustment. The policy is trained with domain randomization and deployed with command filtering and tactile-triggered grasp tightening. Simulation experiments show higher grasp success rates than the evaluated baselines under full and partial point-cloud observations, while real-world experiments demonstrate sim-to-real transfer across diverse object geometries.