SemRD-V2X: Closure-Guided Communication with Bounded Inference for Cooperative Perception
Organizations: Shanghai Jiao Tong University, Shanghai, China
Abstract
Vehicle-to-Everything (V2X) cooperative perception improves 3-D detection by sharing intermediate features, but dense remote features may repeat context that the ego agent can infer locally. Most communication-efficient designs optimize masks or codes empirically, leaving a more basic question open: which remote evidence is indispensable given the receiver's own observation? We introduce a closure-fidelity perspective on ego conditioned remote perception. Under a finite deductive abstraction and explicit conditions, its rate--distortion function decomposes over an irredundant core, and the exact zero-distortion rate becomes . This analysis suggests a concrete design principle: transmit compact evidence and recover derivable context with bounded receiver-side inference. Guided by this principle, SemRD-V2X is an operational neural proxy that combines exact-budget BEV support selection, pointwise channel compression, and masked shared-weight reconstruction before standard fusion. Experiments on simulated V2XSet and real-world DAIR-V2X validate the resulting design. In a controlled five-run V2XSet comparison against a locally reproduced V2X-ViT-v1 baseline on one Tesla V100, SemRD-V2X reduces the analytical feature payload by while improving AP@0.5/AP@0.7 by 4.13/8.57 points, with 3.81% additional mean compute latency. These results position closure fidelity as both an analytical lens and an actionable design principle for communication-efficient cooperative perception.
Figures & tables
| V2XSet | DAIR-V2X | |||||
|---|---|---|---|---|---|---|
| Perfect | Noisy | Mean(P,N) | ||||
| Method | AP@0.5 | AP@0.7 | AP@0.5 | AP@0.7 | AP@0.5 | AP@0.7 |
| Late Fusion | 0.727 | 0.620 | 0.549 | 0.307 | 0.520 | 0.422 |
| Early Fusion | 0.819 | 0.710 | 0.720 | 0.384 | 0.525 | 0.366 |
| F-Cooper | 0.840 | 0.680 | 0.715 | 0.469 | 0.550 | 0.415 |
| OPV2V | 0.807 | 0.664 | 0.709 | 0.487 | 0.548 | 0.420 |
| Method | AP@0.5 | AP@0.7 | MiB/frame | Measured compute (mean SD) | Analytical comm. (150 Mbit/s) | Illustrative total (mean SD) |
|---|---|---|---|---|---|---|
| V2X-ViT-v1 | 0.8033 | 0.5606 | 13.6416 | 144.348 0.465 | 762.895 | 907.243 0.465 |
| SemRD-V2X (Full) | 0.8445 | 0.6463 | 0.5131 | 149.845 3.589 | 28.697 | 178.542 3.589 |
| Variant | MiB/frame | AP@0.5 | AP@0.7 | PCF@0.5 | SA@K | ||
|---|---|---|---|---|---|---|---|
| Full | 0.5131 | 0.8445 | 0.6465 | 0.9357 | 0.4917 | 18.1582 | 7.2013 |
| Random selector | 0.5131 | 0.6528 | 0.5032 | 0.9178 | 0.2998 | 8.9167 | 11.1486 |
| w/o rate regularization | 0.5131 | 0.8148 | 0.5974 | 0.9301 | 0.4650 | 17.4119 | 9.5521 |
| w/o reconstruction supervision | 0.5131 | 0.8272 | 0.5951 | 0.9115 | 0.4710 | 21.9847 | 9.3176 |
| w/o codec | 4.0935 | 0.7976 | 0.5996 | 0.9254 | 0.4714 | 0.0000 | 11.5870 |
| Setting | Payload | AP@0.5 | AP@0.7 | Miss. S- | |
| Decoder depth (Noisy protocol, ) | |||||
| 0.2 | 0.5131 | 0.8344 | 0.6451 | 0.002193 | |
| 0.2 | 0.5131 | 0.8382 | 0.6440 | 0.002035 | |
| 0.2 | 0.5131 | 0.8445 | 0.6465 | 0.001979 | |
| 0.2 | 0.5131 | 0.8399 | 0.6399 | 0.002356 | |
| 0.2 | 0.5131 | 0.8187 | 0.6213 | 0.003886 | |