Collaborative perception improves autonomous perception by sharing intermediate Bird's-Eye-View (BEV) features across connected agents, but dense feature exchange is difficult to deploy under strict Vehicle-to-Everything (V2X) bandwidth limits. Existing efficient methods typically either compress the full feature map uniformly, spending bits on low-value background, or sparsify communication, risking the loss of useful context. We propose a coverage-refinement design for byte-constrained cooperative perception: each agent transmits a highly compressed coarse layer over the full BEV map and allocates the remaining budget to selected high-resolution patches. A Task-Aware Benefit Selector ranks cells by estimated downstream utility, enabling deterministic budgeted refinement and zero-retraining adaptation to changing bandwidth. The receiver reconstructs a dense BEV tensor compatible with standard fusion modules. Experiments on DAIR-V2X and OPV2V show strong accuracy-payload trade-offs at kilobyte-scale budgets. On DAIR-V2X, our method reaches 0.60 AP@0.7 at only 1.87 KB per non-ego agent, compared with 0.52 at 4.61 KB for uniform SimVQ compression. Controlled diagnostics further show that the gain arises from coverage-refinement allocation rather than quantization alone. Code will be published.
Figures & tables
Figure 1 : Overview of our Dual-Resolution Budget-Aware Transmission Framework. CAV i extracts BEV features Fi ( Sec. 3.1 ) and computes a pooled task-confidence prior ψh,w ( Sec. 3.2 ). Guided by ψh,w , the selector predicts a benefit map Bi to output a soft mask Misoft for budgeted training via dual ascent ( Secs. 3.3 and 3.4 ), or a deterministic Top- Nfine mask Mi for strict inference constraints. The ego vehicle decodes the highly compressed payload (VQ indices and spatial headers) to reconstruct Mi and the dense feature map F^ for fusion ( Sec. 3.6 ). The dashed arrow ( dˉh,w ) denotes the regularized offline supervision target ( Sec. 3.5 ).
Figure 2 : Spatial hierarchy and bitstream composition. The selector sends selected cells as fine patches, while a globally downsampled BEV map provides the dense coarse base layer. Both branches share the VQ-VAE codebook; the bitstream stores spatial headers, fine indices, and coarse indices.
Test
Variant
AP@0.5 ↑
AP@0.7 ↑
KB ↓
(A) Frozen ranking
Random
0.680
0.573
1.878
Prior only ( ψ )
0.713
0.589
1.878
Learned benefit ( β )
0.736
0.604
1.878
(B) No VQ
Raw full
0.720
0.580
6291.46
Raw C
0.651
0.546
393.22
Raw C+F (25% fine cells)
0.715
0.583
1966.12
Table 1 : Controlled diagnostics on DAIR-V2X. (A) Frozen-checkpoint selector isolation under the same 2 KB cap. (B) No-VQ diagnostic separating coverage-refinement allocation from quantization. Payload is per non-ego agent; C/F denotes coarse/fine.
Method
DAIR-V2X [ 46 ]
OPV2V [ 36 ]
AP@0.5 ↑
AP@0.7 ↑
KB ↓
AP@0.5 ↑
AP@0.7 ↑
KB ↓
Dense & Sparse Baselines (High Bandwidth)
CoBEVT (Original)
0.72
0.58
6291.46
0.95
0.88
8650
Where2comm [ 4 ]
0.72
0.57
2415.22
N/A (MB-scale)
ERMVP [ 49 ]
0.70
0.57
1230.00
N/A (MB-scale)
EffiComm [ 42 ]
0.72
0.57
785.00
N/A (MB-scale)
Table 2 : Accuracy-payload comparison under ideal pose. Payload (KB) is the average decodable message size per non-ego agent. Threshold-based methods report the closest mean payload from a sweep. OPV2V is denser; MB-scale baselines are omitted from the kilobyte-regime comparison and shown as N/A.
Fusion
Communication
AP@0.5 ↑
AP@0.7 ↑
KB ↓
MaxFusion
SimVQ
0.686
0.457
4.61
MaxFusion
DR
0.682
0.525
1.878
Where2comm
SimVQ
0.700
0.532
4.61
Where2comm
DR
0.721
0.575
1.878
Table 3 : Fusion generalization on DAIR-V2X. DR improves high-IoU accuracy with MaxFusion and Where2comm-style fusion while reducing payload to the kilobyte regime. Payload is per non-ego agent.
Method
AP@0.5 ↑
AP@0.7 ↑
KB ↓
SECOND
0.712
0.499
6291.46
SECOND + SimVQ
0.681
0.404
4.61
SECOND + Ours
0.707
0.518
1.878
Table 4 : Backbone generalization with SECOND on DAIR-V2X. DR improves over uniform SimVQ with the SECOND backbone at lower payload. Payload is per non-ego agent.
Figure 3 : Integrated Scalability Analysis. Visual logic 3(a) and mask expansion 3(b) illustrate the learned refinement policy used for fixed and dynamic budget adaptation.
Budget
AP@0.5 ↑
AP@0.7 ↑
R@0.5 ↑
Coarse
0.645
0.554
0.670
1.0 KB
0.706
0.588
0.743
1.5 KB
0.730
0.601
0.773
2.0 KB
0.736
0.604
0.779
3.0 KB
0.738
0.605
0.782
Table 5 : Budget adaptation. (A) Fixed-budget scalability on DAIR-V2X using a single trained model. (B) Dynamic per-frame budget scheduling on OPV2V, where each agent receives a budget from {0,1,2} KB and 0 KB means no transmission.
Figure 4 : Robustness Analysis. (a) Pose Error : High-bandwidth baselines ( Full ) degrade rapidly under noise, while threshold-constrained versions ( Adjusted ) suffer from low performance ceilings. (b) Time Delays : CoBEVT-DR and InfoCom show comparable stability, while CoBEVT-DR maintains a stronger clean-pose operating point.
Param.
Set
Base ↓
AP@0.5 ↑
AP@0.7 ↑
Sbase
8
∼ 0.08
0.726
0.594
2
∼ 1.16
0.730
0.593
4
∼ 0.29
0.736
0.604
Ccell
8
–
0.735
0.600
32
–
0.718
0.589
16
–
0.736
0.604
Table 6 : Structural ablation at 2.0 KB. Effect of base-layer stride Sbase and cell granularity Ccell .
Table 8 : No-coarse diagnostic at 1.0 KB. Removing the coarse base layer degrades performance, especially when missing regions are zero-filled.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value / Setting
Training Setup
Hardware
1 × RTX 4090
Epochs
30 / 40
Batch size
4 / 2
Optimizer
AdamW
LR (detector)
1×10−3 / 2×10−3
Appendix
Table 9 : Methodological Configuration. Condensed hyperparameters for training, architecture, and the differentiable selector (Train ∣ Inference). Unless stated otherwise, dataset-dependent values are reported as DAIR-V2X / OPV2V.
Method
Factor
AP@0.5
AP@0.7
Full Res
1×
0.71
0.52
Global Downsample
2×
0.65
0.39
Appendix
Table 10 : Justification for Fine Layer. Preliminary analysis on DAIR-V2X shows that global downsampling harms precision.
Figure 5 : Comparison of codebook dynamics at Epoch 11. (a) and (c) show that standard VQ collapses, with only a few vectors being utilized or updated. (b) and (d) demonstrate that SimVQ maintains a healthy, distributed codebook usage and consistent updates across the entire grid.
Method
(VAE)
(SimVQ)
(FSQ v1)
(FSQ v2)
Input Channels
256
256
256
256
Latent Channels
4
64
32
16
Compression Ratio
64 ×
-
-
-
Codebook Size ( K )
-
64
-
-
Num. Codebook
-
-
16
16
Level
-
-
4
4
Appendix
Table 11 : Model Configuration Comparison. A detailed breakdown of architectural parameters for AE, SimVQ and FSQ.
Method
AP@0.5 ↑
AP@0.7 ↑
KB ↓
FSQ v1
0.71
0.58
49.15
FSQ v2
0.73
0.57
24.58
CoBEVT-VAE
0.61
0.47
98.3
CoBEVT-SimVQ
0.71
0.52
4.61
Appendix
Table 12 : Justification for Learned Codebooks. Comparison of Learned VQ against FSQ and the Autoencoder compression shows that Learned VQ achieves the best balance between accuracy and bandwidth.
K
Latent Channels
AP@0.7 ↑
Inference time (ms)
Params (M)
64
64
0.523
19.6
0.596
64
128
0.508
21.6
1.79
64
256
0.500
24.0
5.97
Appendix
Table 13 : Impact of Latent Dimension. Reducing latent channels improves accuracy and efficiency.
Model
Stages
Payload (KB) ↓
AP@0.5 ↑
AP@0.7 ↑
Codebook Size 64
VQ
1
4.61
0.71
0.52
RVQ
1
4.61
0.71
0.51
RVQ
3
13.82
0.70
0.54
Codebook Size 128
VQ
1
5.38
0.71
0.51
Appendix
Table 14 : VQ vs. Residual VQ. While RVQ benefits from additional stages, standard VQ achieves better accuracy-per-bit compared to the first RVQ stage.
Strategy
AP@0.5 ↑
AP@0.7 ↑
Staged (Fine-Tuned)
0.70
0.49
Joint Training (Frozen Codebook)
0.69
0.49
End-to-End (Joint from Scratch)
0.71
0.52
Appendix
Table 15 : Impact of Training Strategy. Joint optimization from scratch outperforms both staged fine-tuning and frozen codebook strategies.
Figure 6 : Selection robustness under different priors. Top: binary cell selection masks; bottom: corresponding prior heatmaps.
Budget
Architecture
AP@0.5
AP@0.7
Params (M)
GFLOPs
1.0 KB
Dilated CNN
0.686
0.555
0.39
9.54
1.0 KB
Pyramid Attn.
0.706
0.588
1.27
28.84
2.0 KB
Dilated CNN
0.725
0.580
0.39
9.54
2.0 KB
Pyramid Attn.
0.736
0.604
1.27
28.84
Appendix
Table 16 : Selector architecture trade-off. Pyramid Attention improves ranking under strict budgets compared to a local Dilated CNN (both with task-aware supervision Ψ ).
Component
Count
Size (Bits)
Size (KB)
1. Coarse Base Layer
1 (Full 12×32 Grid)
2,304
0.288
2. Fine Patches (VQ Indices)
8 patches
12,288
1.536
3. Per-Patch Headers
8 headers
384
0.048
Total Payload
14,976
1.872
Appendix
Table 17 : Bitstream composition under a 2.0 KB budget (DAIR-V2X). Under a 2.0 KB cap, the system transmits the coarse floor and the top-8 highest-utility refinement patches. The 48-bit header includes the spatial coordinate required to reconstruct Mi at the receiver.
Figure 7 : Scaling with number of agents under fixed budgets (OPV2V). AP@0.5 versus non-ego agent count N at a 2.0 KB/agent cap for CoBEVT-DR. Dense CoBEVT uses ∼ 6.3 MB/agent ( ∼ 31.4 MB total network payload at N=5 ), whereas CoBEVT-DR uses ∼ 2.0 KB/agent ( ∼ 10 KB total at N=5 ).
Method
Mode
Compute
Trips
Est. E2E
CoBEVT (Dense)
Broadcast
17 ms
1
27 ms
CoBEVT + VQ
Broadcast
21 ms
1
31 ms
CoSDH [ 33 ]
Interactive
35 ms
2
55 ms
InfoCom [ 30 ]
Broadcast
37 ms
1
47 ms
Ours (Dual-Res)
Broadcast
40 ms
1
50 ms
Appendix
Table 18 : Inference and End-to-End (E2E) Latency Analysis. We compare compute runtime alongside a theoretical E2E latency model ( TE2E=Tcompute+Ntrips×10 ms). Interactive methods require multiple network trips, resulting in higher E2E latency than one-shot broadcast methods.
Figure 8 : Per-agent Top- K normalized cumulative marginal detection gain ( K≤20 ). Ranking patches by our predicted benefit score β successfully concentrates the majority of attainable detection gain within a strict low- K budget cap regime, vastly outperforming random allocation.
Figure 9 : Qualitative comparison. Green: ground truth. Red: detections. (a) InfoCom relies on pure sparsification, discarding unselected regions and missing several vehicles in the upper-right cluster. (b) CoBEVT-DR at a strict 2 KB budget recovers the cluster; the dense coarse floor provides a continuous semantic anchor even for unselected objects.
Method
Dynamic IoU ↑
KB ↓
CoBEVT
0.47
524.28
CoBEVT-SimVQ
0.43
0.72
CoBEVT-DR
0.47
0.40
CoBEVT-DR
0.46
0.20
CoBEVT-DR
0.37
0.09
Appendix
Table 19 : Task transfer to dynamic-object BEV segmentation. We evaluate the same coverage-refinement communication design on dynamic-object segmentation using the original CoBEVT segmentation setup. Payload is reported per non-ego agent.
Method
Car AP@0.3/0.5 ↑
Ped. AP@0.3/0.5 ↑
Truck AP@0.3/0.5 ↑
mAP@0.3 ↑
mAP@0.5 ↑
No Fusion [ 32 ]
38.7/35.9
25.5/13.1
20.2/14.5
28.2
21.2
Early Fusion [ 32 ]
51.1/47.6
31.6/16.0
32.5/23.6
38.4
29.1
F-Cooper [ 32 ]
57.3/54.2
30.0/14.1
27.0/21.2
38.1
29.8
AttFuse [ 32 ]
62.6/59.4
32.2/15.5
32.6/26.6
42.5
33.8
V2X-ViT [ 32 ]
62.7/60.3
36.7/18.6
35.1/28.3
44.8
35.8
CooPre [ 50 ]
71.5/70.2
46.9/28.0
61.9/58.3
60.1
52.2
Appendix
Table 20 : Multi-class detection performance on the V2X-Real-VC test set. Car, Ped., and Truck report AP@0.3/AP@0.5. Published results are taken from the respective papers and may use different detector architectures and evaluation implementations.
Method
Vehicle
Bicyclist
Pedestrian
mAP30
AP30
AP50
AP70
AP30
AP50
AP30
AP50
No Fusion
0.89
0.84
0.73
0.40
0.30
0.41
0.24
0.57
Late Fusion
0.88
0.86
0.81
0.43
0.38
0.45
0.27
0.59
F-Cooper
0.93
0.82
0.68
0.44
0.29
0.56
0.33
0.64
V2X-ViT
0.93
0.91
0.84
0.50
0.36
0.41
0.12
0.61
CoopDet3D
0.93
0.90
0.81
0.48
0.41
0.53
0.31
0.65
Appendix
Table 21 : Multi-class detection performance on V2XVerse. Our DR variants operate under strict per-agent communication budgets. Values are AP. The results are taken from CoDriving [ 11 ] .
Vehicle-to-Everything (V2X) cooperative perception improves 3-D detection by sharing intermediate features, but dense remote features may repeat context that the ego agent can infer locally. Most communication-efficient designs optimize masks or codes empirically, leaving a more basic question open: which remote evidence is indispensable given the receiver's own observation? We introduce a closure-fidelity perspective on ego conditioned remote perception. Under a finite deductive abstraction and explicit conditions, its rate--distortion function decomposes over an irredundant core, and the exact zero-distortion rate becomes PAH(πA). This analysis suggests a concrete design principle: transmit compact evidence and recover derivable context with bounded receiver-side inference. Guided by this principle, SemRD-V2X is an operational neural proxy that combines exact-budget BEV support selection, pointwise channel compression, and masked shared-weight reconstruction before standard fusion. Experiments on simulated V2XSet and real-world DAIR-V2X validate the resulting design. In a controlled five-run V2XSet comparison against a locally reproduced V2X-ViT-v1 baseline on one Tesla V100, SemRD-V2X reduces the analytical feature payload by 26.6× while improving AP@0.5/AP@0.7 by 4.13/8.57 points, with 3.81% additional mean compute latency. These results position closure fidelity as both an analytical lens and an actionable design principle for communication-efficient cooperative perception.
Collaborative-perception enables multi-robot systems to enhance situational awareness by sharing perceptual information. Existing collaborative-perception systems face an inherent trade-off between communication bandwidth requirements and perception accuracy, where methods that exchange more information achieve better perception results at the cost of increased communication overhead. However, real-world communication networks impose bandwidth constraints that require minimizing communication overhead without sacrificing perception performance. To address this challenge, we propose HydraCollab, an adaptive collaborative-perception framework that (i) selectively transmits the most informative sensor features and (ii) dynamically employs collaboration strategies (intermediate or late) based on spatial confidence maps. Extensive evaluations on the V2X-R, V2X-Radar and UAV3D-mini datasets demonstrate that HydraCollab achieves the best overall trade-off between accuracy and communication cost among existing collaborative-perception methods. Relative to SOTA Where2comm, HydraCollab uses only 41% of the bandwidth on V2X-R and 26% on V2X-Radar while improving performance by 0.78% and 0.75% respectively. Our code and models are available at https://github.com/AICPS/HydraCollab.
Luke Chen, Cheng-Ju Wu, David R. Martin +3
Department of Electrical Engineering and Computer Science, University of California, Irvine, USA.
Cooperative perception through Vehicle-to-Everything (V2X) communication offers significant potential for enhancing vehicle perception by mitigating occlusions and expanding the field of view. However, past research has predominantly focused on improving accuracy metrics without addressing the crucial system-level considerations of efficiency, latency, and real-world deployability. Noticeably, most existing systems rely on full-precision models, which incur high computational and transmission costs, making them impractical for real-time operation in resource-constrained environments. In this paper, we introduce \textbf{QuantV2X}, the first fully quantized multi-agent system designed specifically for efficient and scalable deployment of multi-modal, multi-agent V2X cooperative perception. QuantV2X introduces a unified end-to-end quantization strategy across both neural network models and transmitted message representations that simultaneously reduces computational load and transmission bandwidth. Remarkably, despite operating under low-bit constraints, QuantV2X achieves accuracy comparable to full-precision systems. More importantly, when evaluated under deployment-oriented metrics, QuantV2X reduces system-level latency by 3.2× and achieves a +9.5 improvement in mAP30 over full-precision baselines. Furthermore, QuantV2X scales more effectively, enabling larger and more capable models to fit within strict memory budgets. These results highlight the viability of a fully quantized multi-agent intermediate fusion system for real-world deployment. The system will be publicly released to promote research in this field: https://github.com/ucla-mobility/QuantV2X.