Organizations: Hong Kong JC STEM Lab of Smart City and the Department of Computer Science, City University of Hong Kong, Hong Kong · School of Cyber Science and Technology, Beihang University, Beijing, China
Unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs) often lose satellite positioning in urban canyons, indoor facilities, and jammed or spoofed environments, making vision-based matching with geo-tagged databases important for absolute positioning. However, limited onboard computation and energy often require localization to be offloaded over wireless links with time-varying throughput. We present a network-adaptive task-oriented communication framework that jointly determines when to offload, which views and semantic rate to transmit, and which client to serve. The framework combines scalable orthogonality-regularized variational information bottleneck (O-VIB) encoding, value-of-information (VOI)-guided request control, and VOI-weighted Lyapunov scheduling. O-VIB supports importance-ordered latent prefixes and different view subsets, while edge assistance is requested only when its predicted reduction in localization risk exceeds the communication and service cost. Under a matched per-route traffic budget, VOI-guided control reduces mean and p95 route errors by 24.8% and 31.0% over budgeted periodic offloading on CARLA multi-view UAV data. In real-world indoor UAV and UGV experiments, the pipeline reduces mean position error by 28.0% and 14.4% over uncompressed all-view CLIP retrieval while cutting descriptor traffic by 98.6% and 98.2%, respectively. Under high congestion, value-aware shaping reduces edge-side p95 latency for the top-10% high-value requests by 76.2%, from 137.7 ms to 32.8 ms.
Figures & tables
Fig. 1: Motivating scenario and edge-assisted multi-view localization pipeline. Panel (a) illustrates localization drift and collision risk when GNSS is unavailable or unreliable. Panel (b) shows the proposed pipeline: the mobile client observes multiple views, selects a task-oriented semantic representation, transmits it over the client-edge link, and receives a corrected pose after edge decoding and database-assisted localization.
Feature
[ 16 ]
[ 10 ]
[ 30 , 31 ]
[ 6 , 7 , 5 ]
[ 40 ]
[ 9 ]
[ 24 , 47 , 25 ]
[ 1 ]
Proposed work
Task-oriented compression
✓
✓
△
×
△
×
×
✓
✓
Variable-rate representation
△
✓
✓
×
△
×
×
×
✓
Request control
×
×
△
×
✓
△
×
×
✓
View selection
×
×
△
×
×
×
×
×
✓
Coarse-to-fine retrieval
×
×
×
△
×
△
×
×
✓
Uncertainty feedback
×
×
×
×
✓
△
×
×
✓
TABLE I: Feature coverage across representative works.
Fig. 2: System architecture of network-adaptive edge-assisted multi-view localization. The client maintains a local fast loop, requests edge localization only when the predicted value is high, and receives correction feedback after shared wireless and edge-scheduling stages.
Fig. 3: Scalable O-VIB encoder for view-rate adaptive semantic localization. View masks select the available camera subset, while nested latent prefixes provide low-, medium-, and high-rate semantic payloads for one deployed encoder.
Fig. 6: Self-collected multi-scene multi-view UAV data rendered in the CARLA simulator. The synchronized five-view RGB observations and logged simulator poses support retrieval, feature extraction, and model training.
Setting
Data path
Scenes
Frames/routes
Views
Ground truth
Main purpose
Single-scene split
Conference-version Town05
1 scene
4,994; 3,495/749/750
5
Simulator pose
Scalable compression and ablations
Raw multi-scene data
CARLA UAV
5 scenes
6,684 frames
5
Simulator pose
Raw samples and retrieval database
Sampled benchmark
CARLA subset
5 scenes
4,593 frames
5
Simulator pose
Coarse-to-fine retrieval and plots
Control workload
Town02 route replay
Held-out route
593 slots / 7 policies
Variable
Simulator pose
gψ request and mode control
Retrieval queries
Offline query set
5 scenes
400 queries
5
Simulator pose
Scene coarse + tile-pruned fine search
Indoor UAV
Qualisys capture
4.1 m × 4.0 m
76 ref. / 58 query
5 sequential
6-DoF mocap
Real-scene compact localization
TABLE II: Datasets and evaluation configurations used in the study.
Method
Models
Storage
Rates
Latency @ 8 KB/s
Fixed-rate O-VIB-32
1
2.8 MB
1
50.4 ms
Fixed-rate O-VIB-128
1
2.9 MB
1
100.6 ms
Fixed-rate bank {32,128}
2
5.7 MB
2
50.4–100.6 ms
Scalable O-VIB k=35
1
13.9 MB
128
52.0 ms
Scalable O-VIB k=16
1
13.9 MB
128
42.0 ms
TABLE III: Deployment overhead of fixed-rate and scalable O-VIB configurations under the Jetson Orin NX latency profile.
Fig. 9: Reproduction of the conference-version single-scene O-VIB experiments. Panel (a) shows ARD-controlled active dimensions and error. Panels (b) and (c) report the fixed- k=16 sweep, and panel (d) reports Jetson-profile compute cost.
Fig. 11: Calibration and evaluation of VOI-Control. Panel (a) shows the train-only context/action median residual. F, FS, H4, and A5 denote front, front-plus-sides, four-horizontal, and all-five views. C and R denote k=16 and k=32 . Panel (b) shows tuning-seed budget search, and panel (c) shows held-out policies with the oracle action.
High congestion: N=30 4 KB/s weak link
Scheduler
E2E p95
Edge p95
Stale
Cost
MaxWeight
198 ms
142.4 ms
0.022 m
1.87
Base-DPP
193 ms
137.7 ms
0.022 m
1.82
VOI-Lyap.
88.5 ms
32.8 ms
0.012 m
1.12
TABLE IV: High-congestion scheduler comparison for top-10% high-value requests.
Fig. 15: Real-world UAV/UGV testbed and edge-assisted localization pipeline. The platforms transmit visual inputs over Wi-Fi, while Qualisys provides pose ground truth.
Plat.
Representation
KB
Mean
p90
R@0.5
(m)
(m)
UAV
CLIP
10.000
0.407
0.758
86.2%
UAV
O-VIB-8
0.051
0.286
0.458
89.7%
UAV
O-VIB-32
0.145
0.293
0.447
91.4%
UAV
O-VIB-128
0.520
0.274
0.459
93.1%
UGV
CLIP
8.000
0.346
0.396
98.6%
TABLE V: Held-out real-scene localization for representative O-VIB prefixes.
Real-time, drift-free UAV geo-localization is essential for autonomous missions in GNSS-denied environments. The pioneering system, PiLoT, achieves high precision via Neural Pixel-to-3D Registration, aligning UAV video streams with a single rendered reference view from 3D meshes. However, its reliance on heavy 3D meshes incurs massive storage overheads, complex map acquisition, and significant computational rendering costs, severely hindering deployment on embedded platforms. To address these bottlenecks, we propose PiLoT v2, a lightweight yet robust evolution that shifts the paradigm to direct pixel-to-orthogonal map registration for free-view UAV geo-localization. By leveraging True Digital Orthophoto Maps (TDOMs) and Digital Surface Models (DSMs) as the reference substrate, PiLoT v2 replaces GPU-intensive 3D rendering with a highly efficient, CPU-friendly map cropping operation. To bridge the severe geometric discrepancy between these 2.5D orthogonal crops and free-view oblique UAV imagery, we train a cross-view feature registration network using a novel, large-scale geometrically annotated dataset. Furthermore, we integrate onboard sensor prior--specifically gravity direction and single-point laser rang--directly into the pose optimization manifold to enhance robustness against cross-view visual degradation. Experimental results demonstrate that PiLoT v2 achieves performance comparable to, or even exceeding, its Pixel-to-3D predecessor, while offering drastically lower storage and computational costs.
This paper investigates a multi-Unmanned Aerial Vehicle (UAV) joint base station-assisted Internet of Vehicles (IoV) task offloading system in dense urban environments. To minimize system delay and energy consumption under strict coupling constraints, the complex non-convex optimization problem is decoupled into a hierarchical execution framework. First, a sequential distributed optimization algorithm based on Second-Order Cone Programming (SOCP) is proposed to optimize the 3D flight trajectory of each UAV, ensuring adaptive network coverage. Second, a novel hybrid resource scheduling paradigm synergizing Deep Reinforcement Learning (DRL) and Large Language Models (LLMs) is developed. Within this framework, the DRL agent dictates the initial resource allocation, while the LLM acts as a semantic macro-scheduler to rectify long-tail allocation imbalances for failed and surplus tasks. Crucially, a reward decoupling mechanism is introduced to isolate DRL training from external LLM interventions, thereby ensuring policy convergence. Finally, the task offloading ratios are precisely determined via Linear Programming (LP) within an alternating optimization loop. Simulation results demonstrate that the proposed method significantly outperforms traditional multi-agent reinforcement learning baselines in terms of task success rate and system efficiency.
Maoxin Ji, Qiong Wu, Pingyi Fan +4
School of Internet of Things Engineering, Jiangnan University, Wuxi 214122, China · School of Information Engineering, Jiangxi Provincial Key Laboratory of Advanced Signal Processing and Intelligent Communications, Nanchang University, Nanchang 330031, China · Department of Electronic Engineering, State Key laboratory of Space Network and Communications, Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing 100084, China +4
Robust and efficient cooperative exploration with multiple unmanned ground vehicles (UGVs) in unknown, GPSdenied, and bandwidth-limited environments without prior maps remains challenging, as localization drift degrades map consistency and induces redundant coverage. This paper presents a fully distributed exploration framework that couples descriptoraided inter-UGV loop closure with loop-aware hierarchical planning while enabling autonomous localization and exploration. We develop a lightweight LiDAR global descriptor with range-image prealignment to enable robust cross-UGV place recognition under large yaw and lateral variations, and use verified loop closures to maintain globally consistent trajectories and a sparse topological representation. We further introduce an uncertainty-aware crossUGV loop-closure selection module that scores candidate loop closures under pose uncertainty and retains high-utility loop closures as planning anchors for global task allocation and local route refinement. Simulations and real-UGV experiments show that the loop-closure module achieves AR@1/AR@1% of 89.9%/95.5%, distributed optimization reduces absolute trajectory error, the system substantially reduces two-way communication volume, and the overall framework reduces exploration time and travel distance by 15% and 14%, respectively, compared with an mTSP baseline.
Zhiwei Li, Haiou Liu, Xijun Zhao +3
School of Mechanical Engineering, Beijing Institute of Technology, Beijing 100081, China · China North Artificial Intelligence & Innovation Research Institute, Collective Intelligence & Collaboration Laboratory (CIC), Beijing 100081, China