Collaborative perception improves autonomous perception by sharing intermediate Bird's-Eye-View (BEV) features across connected agents, but dense feature exchange is difficult to deploy under strict Vehicle-to-Everything (V2X) bandwidth limits. Existing efficient methods typically either compress the full feature map uniformly, spending bits on low-value background, or sparsify communication, risking the loss of useful context. We propose a coverage-refinement design for byte-constrained cooperative perception: each agent transmits a highly compressed coarse layer over the full BEV map and allocates the remaining budget to selected high-resolution patches. A Task-Aware Benefit Selector ranks cells by estimated downstream utility, enabling deterministic budgeted refinement and zero-retraining adaptation to changing bandwidth. The receiver reconstructs a dense BEV tensor compatible with standard fusion modules. Experiments on DAIR-V2X and OPV2V show strong accuracy-payload trade-offs at kilobyte-scale budgets. On DAIR-V2X, our method reaches 0.60 AP@0.7 at only 1.87 KB per non-ego agent, compared with 0.52 at 4.61 KB for uniform SimVQ compression. Controlled diagnostics further show that the gain arises from coverage-refinement allocation rather than quantization alone. Code will be published.
Figures & tables
Figure 1 : Overview of our Dual-Resolution Budget-Aware Transmission Framework. CAV i extracts BEV features Fi ( Sec. 3.1 ) and computes a pooled task-confidence prior ψh,w ( Sec. 3.2 ). Guided by ψh,w , the selector predicts a benefit map Bi to output a soft mask Misoft for budgeted training via dual ascent ( Secs. 3.3 and 3.4 ), or a deterministic Top- Nfine mask Mi for strict inference constraints. The ego vehicle decodes the highly compressed payload (VQ indices and spatial headers) to reconstruct Mi and the dense feature map F^ for fusion ( Sec. 3.6 ). The dashed arrow ( dˉh,w ) denotes the regularized offline supervision target ( Sec. 3.5 ).
Figure 2 : Spatial hierarchy and bitstream composition. The selector sends selected cells as fine patches, while a globally downsampled BEV map provides the dense coarse base layer. Both branches share the VQ-VAE codebook; the bitstream stores spatial headers, fine indices, and coarse indices.
Test
Variant
AP@0.5 ↑
AP@0.7 ↑
KB ↓
(A) Frozen ranking
Random
0.680
0.573
1.878
Prior only ( ψ )
0.713
0.589
1.878
Learned benefit ( β )
0.736
0.604
1.878
(B) No VQ
Raw full
0.720
0.580
6291.46
Raw C
0.651
0.546
393.22
Raw C+F (25% fine cells)
0.715
0.583
1966.12
Table 1 : Controlled diagnostics on DAIR-V2X. (A) Frozen-checkpoint selector isolation under the same 2 KB cap. (B) No-VQ diagnostic separating coverage-refinement allocation from quantization. Payload is per non-ego agent; C/F denotes coarse/fine.
Method
DAIR-V2X [ 46 ]
OPV2V [ 36 ]
AP@0.5 ↑
AP@0.7 ↑
KB ↓
AP@0.5 ↑
AP@0.7 ↑
KB ↓
Dense & Sparse Baselines (High Bandwidth)
CoBEVT (Original)
0.72
0.58
6291.46
0.95
0.88
8650
Where2comm [ 4 ]
0.72
0.57
2415.22
N/A (MB-scale)
ERMVP [ 49 ]
0.70
0.57
1230.00
N/A (MB-scale)
EffiComm [ 42 ]
0.72
0.57
785.00
N/A (MB-scale)
Table 2 : Accuracy-payload comparison under ideal pose. Payload (KB) is the average decodable message size per non-ego agent. Threshold-based methods report the closest mean payload from a sweep. OPV2V is denser; MB-scale baselines are omitted from the kilobyte-regime comparison and shown as N/A.
Fusion
Communication
AP@0.5 ↑
AP@0.7 ↑
KB ↓
MaxFusion
SimVQ
0.686
0.457
4.61
MaxFusion
DR
0.682
0.525
1.878
Where2comm
SimVQ
0.700
0.532
4.61
Where2comm
DR
0.721
0.575
1.878
Table 3 : Fusion generalization on DAIR-V2X. DR improves high-IoU accuracy with MaxFusion and Where2comm-style fusion while reducing payload to the kilobyte regime. Payload is per non-ego agent.
Method
AP@0.5 ↑
AP@0.7 ↑
KB ↓
SECOND
0.712
0.499
6291.46
SECOND + SimVQ
0.681
0.404
4.61
SECOND + Ours
0.707
0.518
1.878
Table 4 : Backbone generalization with SECOND on DAIR-V2X. DR improves over uniform SimVQ with the SECOND backbone at lower payload. Payload is per non-ego agent.
Figure 3 : Integrated Scalability Analysis. Visual logic 3(a) and mask expansion 3(b) illustrate the learned refinement policy used for fixed and dynamic budget adaptation.
Budget
AP@0.5 ↑
AP@0.7 ↑
R@0.5 ↑
Coarse
0.645
0.554
0.670
1.0 KB
0.706
0.588
0.743
1.5 KB
0.730
0.601
0.773
2.0 KB
0.736
0.604
0.779
3.0 KB
0.738
0.605
0.782
Table 5 : Budget adaptation. (A) Fixed-budget scalability on DAIR-V2X using a single trained model. (B) Dynamic per-frame budget scheduling on OPV2V, where each agent receives a budget from {0,1,2} KB and 0 KB means no transmission.
Figure 4 : Robustness Analysis. (a) Pose Error : High-bandwidth baselines ( Full ) degrade rapidly under noise, while threshold-constrained versions ( Adjusted ) suffer from low performance ceilings. (b) Time Delays : CoBEVT-DR and InfoCom show comparable stability, while CoBEVT-DR maintains a stronger clean-pose operating point.
Param.
Set
Base ↓
AP@0.5 ↑
AP@0.7 ↑
Sbase
8
∼ 0.08
0.726
0.594
2
∼ 1.16
0.730
0.593
4
∼ 0.29
0.736
0.604
Ccell
8
–
0.735
0.600
32
–
0.718
0.589
16
–
0.736
0.604
Table 6 : Structural ablation at 2.0 KB. Effect of base-layer stride Sbase and cell granularity Ccell .
Table 8 : No-coarse diagnostic at 1.0 KB. Removing the coarse base layer degrades performance, especially when missing regions are zero-filled.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value / Setting
Training Setup
Hardware
1 × RTX 4090
Epochs
30 / 40
Batch size
4 / 2
Optimizer
AdamW
LR (detector)
1×10−3 / 2×10−3
Appendix
Table 9 : Methodological Configuration. Condensed hyperparameters for training, architecture, and the differentiable selector (Train ∣ Inference). Unless stated otherwise, dataset-dependent values are reported as DAIR-V2X / OPV2V.
Method
Factor
AP@0.5
AP@0.7
Full Res
1×
0.71
0.52
Global Downsample
2×
0.65
0.39
Appendix
Table 10 : Justification for Fine Layer. Preliminary analysis on DAIR-V2X shows that global downsampling harms precision.
Figure 5 : Comparison of codebook dynamics at Epoch 11. (a) and (c) show that standard VQ collapses, with only a few vectors being utilized or updated. (b) and (d) demonstrate that SimVQ maintains a healthy, distributed codebook usage and consistent updates across the entire grid.
Method
(VAE)
(SimVQ)
(FSQ v1)
(FSQ v2)
Input Channels
256
256
256
256
Latent Channels
4
64
32
16
Compression Ratio
64 ×
-
-
-
Codebook Size ( K )
-
64
-
-
Num. Codebook
-
-
16
16
Level
-
-
4
4
Appendix
Table 11 : Model Configuration Comparison. A detailed breakdown of architectural parameters for AE, SimVQ and FSQ.
Method
AP@0.5 ↑
AP@0.7 ↑
KB ↓
FSQ v1
0.71
0.58
49.15
FSQ v2
0.73
0.57
24.58
CoBEVT-VAE
0.61
0.47
98.3
CoBEVT-SimVQ
0.71
0.52
4.61
Appendix
Table 12 : Justification for Learned Codebooks. Comparison of Learned VQ against FSQ and the Autoencoder compression shows that Learned VQ achieves the best balance between accuracy and bandwidth.
K
Latent Channels
AP@0.7 ↑
Inference time (ms)
Params (M)
64
64
0.523
19.6
0.596
64
128
0.508
21.6
1.79
64
256
0.500
24.0
5.97
Appendix
Table 13 : Impact of Latent Dimension. Reducing latent channels improves accuracy and efficiency.
Model
Stages
Payload (KB) ↓
AP@0.5 ↑
AP@0.7 ↑
Codebook Size 64
VQ
1
4.61
0.71
0.52
RVQ
1
4.61
0.71
0.51
RVQ
3
13.82
0.70
0.54
Codebook Size 128
VQ
1
5.38
0.71
0.51
Appendix
Table 14 : VQ vs. Residual VQ. While RVQ benefits from additional stages, standard VQ achieves better accuracy-per-bit compared to the first RVQ stage.
Strategy
AP@0.5 ↑
AP@0.7 ↑
Staged (Fine-Tuned)
0.70
0.49
Joint Training (Frozen Codebook)
0.69
0.49
End-to-End (Joint from Scratch)
0.71
0.52
Appendix
Table 15 : Impact of Training Strategy. Joint optimization from scratch outperforms both staged fine-tuning and frozen codebook strategies.
Figure 6 : Selection robustness under different priors. Top: binary cell selection masks; bottom: corresponding prior heatmaps.
Budget
Architecture
AP@0.5
AP@0.7
Params (M)
GFLOPs
1.0 KB
Dilated CNN
0.686
0.555
0.39
9.54
1.0 KB
Pyramid Attn.
0.706
0.588
1.27
28.84
2.0 KB
Dilated CNN
0.725
0.580
0.39
9.54
2.0 KB
Pyramid Attn.
0.736
0.604
1.27
28.84
Appendix
Table 16 : Selector architecture trade-off. Pyramid Attention improves ranking under strict budgets compared to a local Dilated CNN (both with task-aware supervision Ψ ).
Component
Count
Size (Bits)
Size (KB)
1. Coarse Base Layer
1 (Full 12×32 Grid)
2,304
0.288
2. Fine Patches (VQ Indices)
8 patches
12,288
1.536
3. Per-Patch Headers
8 headers
384
0.048
Total Payload
14,976
1.872
Appendix
Table 17 : Bitstream composition under a 2.0 KB budget (DAIR-V2X). Under a 2.0 KB cap, the system transmits the coarse floor and the top-8 highest-utility refinement patches. The 48-bit header includes the spatial coordinate required to reconstruct Mi at the receiver.
Figure 7 : Scaling with number of agents under fixed budgets (OPV2V). AP@0.5 versus non-ego agent count N at a 2.0 KB/agent cap for CoBEVT-DR. Dense CoBEVT uses ∼ 6.3 MB/agent ( ∼ 31.4 MB total network payload at N=5 ), whereas CoBEVT-DR uses ∼ 2.0 KB/agent ( ∼ 10 KB total at N=5 ).
Method
Mode
Compute
Trips
Est. E2E
CoBEVT (Dense)
Broadcast
17 ms
1
27 ms
CoBEVT + VQ
Broadcast
21 ms
1
31 ms
CoSDH [ 33 ]
Interactive
35 ms
2
55 ms
InfoCom [ 30 ]
Broadcast
37 ms
1
47 ms
Ours (Dual-Res)
Broadcast
40 ms
1
50 ms
Appendix
Table 18 : Inference and End-to-End (E2E) Latency Analysis. We compare compute runtime alongside a theoretical E2E latency model ( TE2E=Tcompute+Ntrips×10 ms). Interactive methods require multiple network trips, resulting in higher E2E latency than one-shot broadcast methods.
Figure 8 : Per-agent Top- K normalized cumulative marginal detection gain ( K≤20 ). Ranking patches by our predicted benefit score β successfully concentrates the majority of attainable detection gain within a strict low- K budget cap regime, vastly outperforming random allocation.
Figure 9 : Qualitative comparison. Green: ground truth. Red: detections. (a) InfoCom relies on pure sparsification, discarding unselected regions and missing several vehicles in the upper-right cluster. (b) CoBEVT-DR at a strict 2 KB budget recovers the cluster; the dense coarse floor provides a continuous semantic anchor even for unselected objects.
Method
Dynamic IoU ↑
KB ↓
CoBEVT
0.47
524.28
CoBEVT-SimVQ
0.43
0.72
CoBEVT-DR
0.47
0.40
CoBEVT-DR
0.46
0.20
CoBEVT-DR
0.37
0.09
Appendix
Table 19 : Task transfer to dynamic-object BEV segmentation. We evaluate the same coverage-refinement communication design on dynamic-object segmentation using the original CoBEVT segmentation setup. Payload is reported per non-ego agent.
Method
Car AP@0.3/0.5 ↑
Ped. AP@0.3/0.5 ↑
Truck AP@0.3/0.5 ↑
mAP@0.3 ↑
mAP@0.5 ↑
No Fusion [ 32 ]
38.7/35.9
25.5/13.1
20.2/14.5
28.2
21.2
Early Fusion [ 32 ]
51.1/47.6
31.6/16.0
32.5/23.6
38.4
29.1
F-Cooper [ 32 ]
57.3/54.2
30.0/14.1
27.0/21.2
38.1
29.8
AttFuse [ 32 ]
62.6/59.4
32.2/15.5
32.6/26.6
42.5
33.8
V2X-ViT [ 32 ]
62.7/60.3
36.7/18.6
35.1/28.3
44.8
35.8
CooPre [ 50 ]
71.5/70.2
46.9/28.0
61.9/58.3
60.1
52.2
Appendix
Table 20 : Multi-class detection performance on the V2X-Real-VC test set. Car, Ped., and Truck report AP@0.3/AP@0.5. Published results are taken from the respective papers and may use different detector architectures and evaluation implementations.
Method
Vehicle
Bicyclist
Pedestrian
mAP30
AP30
AP50
AP70
AP30
AP50
AP30
AP50
No Fusion
0.89
0.84
0.73
0.40
0.30
0.41
0.24
0.57
Late Fusion
0.88
0.86
0.81
0.43
0.38
0.45
0.27
0.59
F-Cooper
0.93
0.82
0.68
0.44
0.29
0.56
0.33
0.64
V2X-ViT
0.93
0.91
0.84
0.50
0.36
0.41
0.12
0.61
CoopDet3D
0.93
0.90
0.81
0.48
0.41
0.53
0.31
0.65
Appendix
Table 21 : Multi-class detection performance on V2XVerse. Our DR variants operate under strict per-agent communication budgets. Values are AP. The results are taken from CoDriving [ 11 ] .