Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This fragmentation limits the learning of transferable object-affordance semantics across visual and geometric spaces. We propose Token Router for Tasks, a multitask training paradigm for MLLM-based systems that routes contextual hidden states to task-specific branches without requiring the language head to generate predefined markers. Routed states are supervised directly by branch-specific objectives, enabling dense prediction losses to shape shared MLLM representations. We instantiate this paradigm as UniAfford, a unified framework for generalizable 2D-3D affordance perception, together with UniAfford-Data, a dataset integrating pixel-level 2D annotations, point-level 3D annotations, and language instructions under a shared object-affordance taxonomy, supporting heterogeneous supervision through semantic-level 2D-3D pairing. UniAfford adopts an MLLM as a shared semantic hub and a modality-aware token router to produce image- and point-cloud-affordance queries. These queries respectively condition a SAM-style pixel decoder and a SONATA-based point decoder, enabling flexible 2D, 3D, and joint affordance inference from image-only, point-cloud-only, or paired multimodal inputs. Experiments demonstrate strong zero-shot generalization across 2D and 3D affordance benchmarks without target-specific fine-tuning, alongside state-of-the-art branch-wise performance under modality-isolated protocols. Ablations validate token routing, joint 2D-3D supervision, and decoder coupling, while language-head diagnostics show that routed latent states carry meaningful object-affordance semantics. Project page: https://4dvlab.github.io/UniAfford
Figures & tables
Figure 1: Overview of UniAfford-Data and UniAfford. Left: Pixel-level and point-level annotations organized under a shared object–affordance taxonomy with semantic pairing across instances. Right: UniAfford achieves strong performance across OOD zero-shot transfer and modality-isolated branch-wise evaluations in both 2D and 3D affordance grounding.
Dataset
Modality
#Objects
#Affordances
#2D Samples
#3D Samples
Annotation
UMD ( Myers et al., 2015 )
2D
17
7
10k
×
2D
HANDAL ( Guo et al., 2023 )
2D
17
1
308k
×
2D
AGD20K ( Luo et al., 2022 )
2D
50
36
26k
×
2D
RAGNet ( Wu et al., 2025a )
Text, 2D
180
–
273k
×
2D
ReasonAff ( Liu et al., 2025b )
Text, 2D
48
30
2.5k
×
Text, 2D
3D-AffordanceNet ( Jia et al., 2021 )
3D
23
17
×
23k
3D
Table 1: Comparison with Existing affordance datasets. UniAfford-Data provides both 2D and 3D annotations under a unified object–affordance indexing scheme.
Figure 2: Overview of UniAfford for unified 2D–3D affordance grounding.
Table 2: OOD zero-shot transfer results on unseen 2D and 3D affordance benchmarks. UniAfford is trained on UniAfford-Data and evaluated on target benchmarks without target-specific training or fine-tuning. Best zero-shot results are in bold , and second-best zero-shot results are underlined .
Table 3: Branch-wise performance under modality-isolated protocols. (a) 2D branch results on ReasonAff. (b) 3D branch results under 3D-only training protocols. Best results are in bold , and second-best results are underlined ; for (b), rankings are computed within each training block.
Type
Variant
2D Metrics
3D Metrics
gIoU ↑
cIoU ↑
AUC ↑
mIoU ↑
SIM ↑
MAE ↓
Full
Full model
68.79
58.36
84.43
34.56
0.583
0.105
Routing
Fixed-anchor routing
62.28
55.55
84.24
17.46
0.535
0.112
Joint learning
2D-only training
41.41
35.74
–
–
–
–
3D-only training
–
–
82.28
30.07
0.535
0.109
Coupling
Prompt-style 2D coupling
36.97
23.88
74.47
22.34
0.589
0.100
Table 4: Ablation studies on routing, joint 2D–3D learning, and decoder coupling. All variants are trained on the same subset of UniAfford-Data and evaluated on the same held-out subset for controlled comparison. Bold numbers mark the best value for each metric; coupling variants should be interpreted mainly by the branch they modify.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Stored content
Instruction.csv
Language instructions, object and affordance labels, and optional img_id / pc_id bindings
Image/
RGB images and pixel-level affordance masks grouped by affordance label
PointCloud/
CSV point clouds with x,y,z coordinates followed by point-wise affordance labels
train.json , val.json , test.json
Split-specific instruction and modality-sample identifiers indexed by object and affordance
Appendix
Table 5: Storage layout of UniAfford-Data. Semantic labels provide a shared index, while sample identifiers link instructions to the corresponding observations and annotations.
Figure 3: Representative annotations in UniAfford-Data. Top: RGB images with pixel-level affordance masks. Bottom: Point clouds with point-wise affordance annotations. Examples are organized under a shared object–affordance taxonomy and do not imply instance-level spatial correspondence.
Protocol
Image Resolution
Points
MLLM LR
2D Decoder LR
3D Decoder LR
Router LR
Mixed-training OOD zero-shot
1024×1024
2048
1e-5
5e-6
5e-4
1e-3
2D modality-isolated
1024×1024
–
1e-5
1e-5
–
1e-3
3D modality-isolated
–
2048
1e-5
–
1e-4
1e-3
Appendix
Table 6: Protocol-specific training configurations. The MLLM learning rate applies to its LoRA parameters.
Figure 4: Language-head readouts of routed representations. Each panel summarizes decoded-token frequencies for the indicated training and evaluation setting. N denotes the number of distinct decoded-token categories recorded in the corresponding diagnostic run.
Method
2D Metrics
3D Metrics
gIoU ↑
cIoU ↑
AUC ↑
mIoU ↑
SIM ↑
MAE ↓
AffordanceNet ( Wu et al., 2025a )
31.44
21.37
–
–
–
–
GREAT ( Yang et al., 2024 )
–
–
71.02
13.14
0.415
0.154
UniAfford
68.79
58.36
84.43
34.56
0.583
0.105
Appendix
Table 7: Performance on the shared UniAfford-Data ablation subset. Specialized baselines use their respective modalities, whereas the full UniAfford model uses heterogeneous 2D–3D supervision.
Table 10: Single-GPU memory footprint of UniAfford.
Figure 5: Qualitative 2D zero-shot comparisons on AGD20K. Columns show the input image, ground-truth annotation, UniAfford, the RAGNet baseline, and Affordance-R1. Red overlays visualize annotated or predicted affordance regions.
Figure 6: Qualitative 3D zero-shot comparisons on GEAL*. Columns show the input point cloud, ground-truth annotation, UniAfford, the GREAT baseline, and IAGNet. Red colors represent point-wise affordance scores, with blue highlighting the highest-confidence regions.
May 2, 2026·Haowen Sun, Shaolong Zhang, Mingyang Li +6Voxel3D Generation
National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University, Xi’an, 710049
The Hong Kong University of Science and Technology, Guangzhou, China · ShanghaiTech University, Shanghai, China · Zhejiang University, Hangzhou, China +3