Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This fragmentation limits the learning of transferable object-affordance semantics across visual and geometric spaces. We propose Token Router for Tasks, a multitask training paradigm for MLLM-based systems that routes contextual hidden states to task-specific branches without requiring the language head to generate predefined markers. Routed states are supervised directly by branch-specific objectives, enabling dense prediction losses to shape shared MLLM representations. We instantiate this paradigm as UniAfford, a unified framework for generalizable 2D-3D affordance perception, together with UniAfford-Data, a dataset integrating pixel-level 2D annotations, point-level 3D annotations, and language instructions under a shared object-affordance taxonomy, supporting heterogeneous supervision through semantic-level 2D-3D pairing. UniAfford adopts an MLLM as a shared semantic hub and a modality-aware token router to produce image- and point-cloud-affordance queries. These queries respectively condition a SAM-style pixel decoder and a SONATA-based point decoder, enabling flexible 2D, 3D, and joint affordance inference from image-only, point-cloud-only, or paired multimodal inputs. Experiments demonstrate strong zero-shot generalization across 2D and 3D affordance benchmarks without target-specific fine-tuning, alongside state-of-the-art branch-wise performance under modality-isolated protocols. Ablations validate token routing, joint 2D-3D supervision, and decoder coupling, while language-head diagnostics show that routed latent states carry meaningful object-affordance semantics. Project page: https://4dvlab.github.io/UniAfford
Figures & tables
Figure 1: Overview of UniAfford-Data and UniAfford. Left: Pixel-level and point-level annotations organized under a shared object–affordance taxonomy with semantic pairing across instances. Right: UniAfford achieves strong performance across OOD zero-shot transfer and modality-isolated branch-wise evaluations in both 2D and 3D affordance grounding.
Dataset
Modality
#Objects
#Affordances
#2D Samples
#3D Samples
Annotation
UMD ( Myers et al., 2015 )
2D
17
7
10k
×
2D
HANDAL ( Guo et al., 2023 )
2D
17
1
308k
×
2D
AGD20K ( Luo et al., 2022 )
2D
50
36
26k
×
2D
RAGNet ( Wu et al., 2025a )
Text, 2D
180
–
273k
×
2D
ReasonAff ( Liu et al., 2025b )
Text, 2D
48
30
2.5k
×
Text, 2D
3D-AffordanceNet ( Jia et al., 2021 )
3D
23
17
×
23k
3D
Table 1: Comparison with Existing affordance datasets. UniAfford-Data provides both 2D and 3D annotations under a unified object–affordance indexing scheme.
Figure 2: Overview of UniAfford for unified 2D–3D affordance grounding.
Table 2: OOD zero-shot transfer results on unseen 2D and 3D affordance benchmarks. UniAfford is trained on UniAfford-Data and evaluated on target benchmarks without target-specific training or fine-tuning. Best zero-shot results are in bold , and second-best zero-shot results are underlined .
Table 3: Branch-wise performance under modality-isolated protocols. (a) 2D branch results on ReasonAff. (b) 3D branch results under 3D-only training protocols. Best results are in bold , and second-best results are underlined ; for (b), rankings are computed within each training block.
Type
Variant
2D Metrics
3D Metrics
gIoU ↑
cIoU ↑
AUC ↑
mIoU ↑
SIM ↑
MAE ↓
Full
Full model
68.79
58.36
84.43
34.56
0.583
0.105
Routing
Fixed-anchor routing
62.28
55.55
84.24
17.46
0.535
0.112
Joint learning
2D-only training
41.41
35.74
–
–
–
–
3D-only training
–
–
82.28
30.07
0.535
0.109
Coupling
Prompt-style 2D coupling
36.97
23.88
74.47
22.34
0.589
0.100
Table 4: Ablation studies on routing, joint 2D–3D learning, and decoder coupling. All variants are trained on the same subset of UniAfford-Data and evaluated on the same held-out subset for controlled comparison. Bold numbers mark the best value for each metric; coupling variants should be interpreted mainly by the branch they modify.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Stored content
Instruction.csv
Language instructions, object and affordance labels, and optional img_id / pc_id bindings
Image/
RGB images and pixel-level affordance masks grouped by affordance label
PointCloud/
CSV point clouds with x,y,z coordinates followed by point-wise affordance labels
train.json , val.json , test.json
Split-specific instruction and modality-sample identifiers indexed by object and affordance
Appendix
Table 5: Storage layout of UniAfford-Data. Semantic labels provide a shared index, while sample identifiers link instructions to the corresponding observations and annotations.
Figure 3: Representative annotations in UniAfford-Data. Top: RGB images with pixel-level affordance masks. Bottom: Point clouds with point-wise affordance annotations. Examples are organized under a shared object–affordance taxonomy and do not imply instance-level spatial correspondence.
Protocol
Image Resolution
Points
MLLM LR
2D Decoder LR
3D Decoder LR
Router LR
Mixed-training OOD zero-shot
1024×1024
2048
1e-5
5e-6
5e-4
1e-3
2D modality-isolated
1024×1024
–
1e-5
1e-5
–
1e-3
3D modality-isolated
–
2048
1e-5
–
1e-4
1e-3
Appendix
Table 6: Protocol-specific training configurations. The MLLM learning rate applies to its LoRA parameters.
Figure 4: Language-head readouts of routed representations. Each panel summarizes decoded-token frequencies for the indicated training and evaluation setting. N denotes the number of distinct decoded-token categories recorded in the corresponding diagnostic run.
Method
2D Metrics
3D Metrics
gIoU ↑
cIoU ↑
AUC ↑
mIoU ↑
SIM ↑
MAE ↓
AffordanceNet ( Wu et al., 2025a )
31.44
21.37
–
–
–
–
GREAT ( Yang et al., 2024 )
–
–
71.02
13.14
0.415
0.154
UniAfford
68.79
58.36
84.43
34.56
0.583
0.105
Appendix
Table 7: Performance on the shared UniAfford-Data ablation subset. Specialized baselines use their respective modalities, whereas the full UniAfford model uses heterogeneous 2D–3D supervision.
Table 10: Single-GPU memory footprint of UniAfford.
Figure 5: Qualitative 2D zero-shot comparisons on AGD20K. Columns show the input image, ground-truth annotation, UniAfford, the RAGNet baseline, and Affordance-R1. Red overlays visualize annotated or predicted affordance regions.
Figure 6: Qualitative 3D zero-shot comparisons on GEAL*. Columns show the input point cloud, ground-truth annotation, UniAfford, the GREAT baseline, and IAGNet. Red colors represent point-wise affordance scores, with blue highlighting the highest-confidence regions.
Open-vocabulary 3D affordance detection requires localizing interaction regions on point clouds given novel affordance descriptions. Recent methods extend multimodal large language models (MLLMs) with special output tokens that are decoded into segmentation masks. However, these tokens are produced through autoregressive generation, which models sequential dependencies rather than spatial neighborhood relations, leaving them semantically rich but spatially impoverished for 3D localization. We propose Voxel-enhanced Affordance detection (VoxAfford), which bypasses this bottleneck by injecting multi-scale geometric features from a frozen pre-trained 3D VQVAE encoder into the output tokens after generation. Each output token uses its affordance semantics as a query to retrieve relevant geometric patterns from its paired voxel scale via cross-attention, with a learned compatibility gate controlling the injection strength. The enhanced tokens are then aggregated into a spatially-aware affordance prompt through semantic-conditioned attention and propagated alongside per-point features to generate the final mask. Experiments on open-vocabulary affordance detection tasks show that VoxAfford achieves state-of-the-art performance with approximately an 8% improvement in mIoU, and real robot experiments confirm zero-shot transfer to novel objects.
Haowen Sun, Shaolong Zhang, Mingyang Li +6
National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University, Xi’an, 710049
Task-driven 3D affordance grounding aims to localize the functional region in a cluttered 3D scene that enables an action specified by a natural-language instruction. Existing methods either predict 3D masks directly or construct them by selecting and fusing intermediate 2D/3D regions. However, they remain vulnerable to two intertwined failure modes: the predicted or selected regions may miss the target interaction area or have unsuitable granularity, while language grounding may confuse visually similar alternatives under relational instructions. To this end, we introduce ThinkAfford, which decouples high-recall affordance proposal generation from instruction-grounded reasoning. Specifically, the Affordance Proposal Generation module first uses learnable affordance prompts and multi-level visual features to predict interaction-conditioned heatmaps, extracting a variable number of fine-grained proposals without parsed object or part names as segmentation prompts. Visual-Prompted Affordance Reasoning then reasons over labeled proposal overlays using the full instruction, returning identifiers in a structured "think-then-answer" response. Moreover, Group Relative Policy Optimization uses proposal-level rewards from lifted 3D overlap to align VPAR selection with final 3D grounding. On the SceneFun3D validation split, ThinkAfford achieves 10.69% AP50 and 25.46% AP25 under the official evaluator, outperforming comparable 3D open-vocabulary and vision-language-model-based 2D-to-3D baselines. Module-level diagnostics further show that APG attains 77.5% recall at 25% intersection-over-union, while GRPO-trained VPAR achieves 72.1% selection accuracy on APG-covered queries, compared with 63.4% under supervised fine-tuning.
Xinrui Lin, Sha Zhang, Shumin Wang +3
University of Science and Technology of China, Hefei, China · The Chinese University of Hong Kong, Hong Kong SAR, China
3D affordance grounding aims to understand how diverse objects can be manipulated, making it a cornerstone of embodied interaction. However, prior works struggle to generalize to out-of-distribution, open-world scenarios, leaving a critical gap between limited dataset performance and real-world application needs. Inspired by the saying: \textit{\textbf{``What I can not create, I do not understand''}}, we find generative models can generate semantically valid HOI images, which indicates inherent encoding of affordance concepts. Building on this insight, we propose DAG, the first innovative diffusion-based 3D affordance grounding framework that extracts general affordance knowledge from text-to-image diffusion models for 3D affordance prediction. Specifically, we extract the affordance priors from a diffusion model to encode HOI priors, and design an affordance block with a multi-source affordance decoder for dense 3D affordance prediction. Extensive experiments show that DAG consistently outperforms state-of-the-art methods and exhibits strong open-world generalization, even in the challenging one-shot setting. The code of our method is released on \textcolor{blue}{\textit{https://github.com/hq-King/DAG}}.
Hanqing Wang, Zhenhao Zhang, Kaiyang Ji +12
The Hong Kong University of Science and Technology, Guangzhou, China · ShanghaiTech University, Shanghai, China · Zhejiang University, Hangzhou, China +3