Point-Focused Attention Meets Context-Scan State Space: Robust Biological Visual Perception for Point Cloud Representation
Organizations: College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics · Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education · School of Mathematical Sciences, Beijing University of Posts and Telecommunications
Abstract
Synergistically capturing intricate local structures and global contextual dependencies has become a critical challenge in point cloud representation learning. To address this, we introduce PointLearner, a point cloud representation learning network that closely aligns with biological vision which employs an active, foveation-inspired processing strategy, thus enabling local geometric modeling and long-range dependency interactions simultaneously. Specifically, we first design a point-focused attention, which simulates foveal vision at the visual focus through a competitive normalized attention mechanism between local neighbors and spatially downsampled features. The spatially downsampled features are extracted by a pooling method based on learnable inducing points, which can flexibly adapt to the non-uniform distribution of point clouds as the number of inducing points is controlled and they interact directly with point clouds. Second, we propose a context-scan state space that mimics eye's saccade inference, which infers the overall semantic structure and spatial content in the scene through a scan path guided by the Hilbert curve for the bidirectional S6. With this focus-then-context biomimetic design, PointLearner demonstrates remarkable robustness and achieves state-of-the-art performance across multiple point cloud tasks.
Figures & tables
| Networks | Operator | Params | Latency | Memory | mIoU |
| HydraMamba ( Qu et al., 2025 ) | SSM | 63.14M | 54ms | 5.9G | 73.6 |
| PTv3 ( Wu et al., 2024a ) | Attention | 46.17M | 49ms | 6.3G | 73.4 |
| Swin3D ( Yang et al., 2025 ) | Attention | 71.15M | 365ms | 10.7G | 72.5 |
| PointLearner | Hybrid | 52.78M | 63ms | 6.5G | 74.3 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Configurations | ModelNet40 | ScanObjectNN | ShapeNet | S3DIS |
| Training epochs | 500 | 500 | 600 | 500 |
| Optimizer & Scheduler | Adamw & CosLR | AdamW & CosLR | Adamw & CosLR | Adamw & CosLR |
| Weight decay | 0.01 | 0.01 | 0.01 | 0.01 |
| Learning rate | 8e-4 | 4e-4 | 1e-3 | 1e-3 |
| Warmup epochs | 10 | 20 | 10 | 10 |
| Batch size | 24 | 24 | 24 | 12 |
| Serialization | Params | FLOPs | Throughput | OA |
| None | 7.36M | 0.610G | 219FPS | 91.34 |
| Hilbert | 7.36M | 0.610G | 163FPS | 94.17 |
| Z-Order | 7.36M | 0.610G | 209FPS | 93.06 |
| Hilbert & Trans-Hilbert | 7.36M | 0.610G | 133FPS | 93.78 |
| Hilbert & Z-Order | 7.36M | 0.610G | 155FPS | 93.52 |
| Learnable Serialization | 8.04M | 0.723G | 168FPS | 92.78 |