Active visual exploration of tabletop objects often requires reorienting an unknown resting object onto a different stable support face to expose occluded surfaces. To identify such placements without exhaustive physical search, we learn a probabilistic placement prior from a single-view point cloud. Stable placement prediction is inherently multimodal, and conventional 6-DoF regression introduces further ambiguity by modeling translation and in-plane yaw. We therefore propose FlipToSee, a probabilistic framework that removes this representational ambiguity by parameterizing placements as unit support normals on S2 while modeling their multimodal conditional distribution via a von Mises--Fisher mixture density network. To decouple mode diversity from physical robustness, FlipToSee deterministically extracts a compact candidate set from the mixture components and applies robustness-aware reranking using an auxiliary head trained with candidate-aligned supervision. In simulation, FlipToSee achieves 98.4% first-proposal success on in-distribution objects, 95.3% on out-of-distribution shapes, and 90.0% under zero-shot transfer to household YCB objects. We further demonstrate the learned placement prior on a physical robot by integrating it with grasp and motion planning for exploratory regrasping.
Task-oriented grasping (TOG) requires robots to grasp functional parts of objects (e.g., the handle of a mug for pouring), yet these affordance regions are frequently occluded in cluttered scenes. Active perception via next-best-view (NBV) planning can resolve such occlusions by moving the camera for more informative observations. However, existing NBV methods typically optimize viewpoints for grasping the target object as a whole without distinguishing which part is task-relevant. A naive adaptation, fully scanning the target object before predicting the affordance, wastes most of the viewpoint budget on task-irrelevant surfaces (e.g., the mug body for pouring). To address this, we propose ATAP, an Affordance-Targeted Active Perception framework that shifts viewpoint planning from exhaustive target scanning to targeted affordance verification. ATAP hypothesizes the occluded target geometry via a generative shape prior and predicts the affordance distribution over the imagined complete surface. In cluttered scenes, severe occlusion can make the location of the hidden affordance ambiguous, leaving multiple locations plausible given the partial observation. ATAP therefore introduces an uncertainty-aware viewpoint planner that jointly optimizes expected entropy reduction over these competing hypotheses and expected affordance verification gain from real observations. This process iterates until the affordance is sufficiently verified for grasp execution. Experiments in simulation and real-world cluttered scenes show that ATAP substantially improves the functional grasp success rate over fixed-view TOG baselines, and outperforms reconstruction-based active perception with over 57% fewer NBV steps.
A geometrically valid placement can still be difficult to execute because the selected grasp changes the required end-effector pose, collision geometry, and transport motion. Placement is formulated as a pre-execution ranking problem in which supplied grasp-placement candidates are scored before planning. The model combines a typed target-conditioned point cloud with three pose descriptors and hierarchical heads for planning success and execution success conditioned on planning. On a 30-object, 1,235-scene dataset with scene-group-held-out splits, three-seed top-1 success on covered test groups reaches 85.63 +/- 1.08% for joint selection and 79.84 +/- 0.16% for fixed-target ranking. For the designated frozen seed-42 checkpoint, top-1 success improves from 72.84% to 85.78% over full-pool cuMotion for joint ranking and from 59.65% to 79.67% for fixed-target ranking. Frozen transfer to xArm7/MoveIt requires no xArm-specific retraining. Across 27 locked cases, 13 complete end to end (48.15%). Of the 16 cases that pass Top-5 preflight and begin execution, 13 succeed (81.25%). Candidate-level deployment-feasibility prediction reaches 81.25% recall, 85.20% specificity, and 83.23% balanced accuracy.
Tianyuan Liu, Rutherford Agbeshi Patamia, Benjamin Champion +2
Humans can reproduce the viewpoint specified by a target image through active head and body motion, yet spatial intelligence in foundation models has largely been studied as passive understanding of pre-collected observations. We introduce Target Viewpoint Reproduction (TVR) -- an active task where an agent adjusts its viewpoint in a 3D environment until its observation matches a given target image -- and TVRBench, an indoor-simulation benchmark spanning scene scale and target-view visual richness. TVR is far from solved: on the evaluation split, the strongest open-source and closed-source models reach only 7.8% and 12.0% success. Fine-grained analysis identifies two consistent bottlenecks: off-the-shelf models struggle with multi-turn visual history, and performance drops sharply when viewpoint reproduction requires body translation rather than in-place rotation, exposing a gap in mapping spatial discrepancies to embodied movement. To study reducing this gap, we build a unified TVR post-training framework covering expert-trajectory SFT, rationale-supervised CoT-SFT, offline Single-turn GRPO, and on-policy Multi-turn GRPO from live simulator rollouts. Visual-action SFT supplies the main gain, raising a 9B open-source model to 50.8% success; Multi-turn GRPO provides targeted multi-room refinement and reaches 51.4% overall, while CoT supervision and Single-turn GRPO degrade closed-loop performance. These results establish TVRBench as a testbed for measuring and training foundation models that actively perceive and act in 3D environments. Our code, data, and models are available at https://github.com/aim-uofa/TVRBench.