Uncrewed aerial vehicle base stations (UAV-BSs) are expected to cover traffic demand that shifts across space and time, yet most repositioning schemes either re-solve an optimization problem per slot or learn reactive policies without an explicit demand model. We cast demand-driven fleet repositioning as latent-space decision-time planning and propose DSWM, a decomposed spatio-temporal world model: an agentic controller that perceives the demand field through a rolling observation window, retains operational context in a latent recurrent state, reasons about candidate motions by imagined rollouts under an uncertainty penalty, and coordinates the fleet through replanned first actions. DSWM learns a recurrent state-space model shaped by an exponential-moving-average (EMA) based latent predictive objective with variance regularization. It attaches a differentiable service simulator that replays the association, probabilistic line-of-sight channel, and Shannon rate chain inside latent rollouts. Planning uses a cross-entropy method whose imagined demand is anchored on the current observation window with mixing coefficient ρ=0.95. On a unified pipeline over three real datasets (Milan CDR (call detail record), Shanghai Telecom, YJMob100K) and 14 methods including five reproduced IEEE baselines, DSWM attains weekday served ratios of 0.889, 0.908, and 0.898, ranking first among non-ablated configurations on every dataset. On Milan it improves over the strongest non-learning baseline (Greedy, 0.780) by 0.109, a margin that comes from decision-time use of observations rather than prediction accuracy.
Figures & tables
Fig. 1: Scenario and motivation. A fleet of K UAV-BSs at altitude hfly covers a gridded demand map xt that drifts over the day (commute bimodal pattern). Because the demand map is strongly structured in space and time, the bottleneck for coverage is how the current observation window is exploited at decision time, not raw forecasting accuracy. DSWM therefore plans in a learned latent space with the imagination anchored on the current observation.
Symbol
Meaning
t , T
slot index; slots per day ( T=144 )
Tslot
slot duration (10 min)
k , K
UAV-BS index; fleet size ( K=16 )
g , G
cell index; number of cells ( G=400 )
c
cell side length (235/470/500 m)
xt∈R+G
normalized demand map
TABLE I: Main Notation
Fig. 2: DSWM framework. The encoder and RSSM maintain a latent state from the observation window. The CEM planner imagines H -slot rollouts, anchors the imagined demand on the current observation with mixing coefficient ρ=0.95 , scores each rollout with the decomposed service head minus an uncertainty penalty, and executes the first action. Training (dashed) is offline over replayed sequences.
Fig. 3: Decomposed demand-service physics head. An E -head ensemble predicts the next demand map in symlog space; a differentiable service simulator replays association → probabilistic LoS ( 1 ) → Shannon rate ( 3 ) → capped per-cell fulfillment → aggregate SR ( 17 ). Ensemble disagreement defines the uncertainty ut .
Fig. 4: Observation-anchored CEM planning. Candidate action sequences are imagined for H=15 slots through the RSSM prior; at each imagined step the demand estimate is anchored on the current observation with mixing coefficient ρ=0.95 ( 20 ); rollouts are scored by J in ( 19 ); elites refit the sampling distribution and the first action is executed.
Parameter
Value
Basis
Fleet size K
16
city-wide; {2,8} scaling
Altitude hfly / speed Vmax
100 m / 25 m/s
low-altitude [ 5 ]
Move limit Dmax / slot
1500 m / 10 min
per-slot motion limit
Carrier fc / bandwidth B
2 GHz / 20 MHz
sub-6GHz access
Power / noise N0
0.5 W / 10−20.4 W/Hz
small-cell
Rate demand r0 / threshold Upeak
0.5 Mbps / 0.8
per-user
TABLE II: Scenario Parameters and Justification
Method
Milan
Shanghai
YJMob
DSWM (ours)
0.889/0.903
0.908/0.862
0.898/0.891
TD-MPC
0.769/0.805
0.875/0.839
0.818/0.831
PPO
0.723/0.791
0.810/0.748
0.879/0.870
GA-MATR
0.770/0.854
0.745/0.708
0.820/0.790
Greedy
0.780/0.792
0.746/0.741
0.762/0.760
INS-WOA
0.744/0.752
0.816/0.788
0.753/0.756
TABLE IV: Cross-Dataset Results: Weekday / Holiday Served Ratio (mean; 5 seeds for Milan, 3 seeds for Shanghai and YJMob; columns: Milan (2013, CDR), Shanghai (2014, sessions), YJMob (2023, 4.5 km)). Per-seed std where it affects adjacent rankings: DSWM ± 0.006 (Shanghai), ± 0.002 (YJMob); GA-MATR ± 0.058 (Shanghai), ± 0.039 (YJMob).
Fig. 14: World-model prediction visualization on Milan. From left to right: full-city context, ground-truth demand map, predicted demand map, and per-cell error. Predictions track the commute-driven hotspot migration; errors concentrate at cell boundaries during peak transitions.
Fig. 15: Shanghai dataset visualization. From left to right: full-region demand, ground truth, weekday profile, and weekend (OOD) profile. The weekday–weekend shift is a spatial redistribution of sparse sessions (85.3% zero cells), not a uniform intensity change.
Fig. 16: YJMob100K dataset visualization. From left to right: full-region demand, ground truth, weekday profile, and pseudo-weekend (OOD) profile. Demand semantics are distinct-user presence counts (a 5% sampled mobility proxy), as declared in Section V-A ; weekend labels are inferred from the autocorrelation-recovered week structure.
Using Unmanned Aerial Vehicle (UAV) for urban sensing has emerged as a powerful paradigm to monitor the status of the city, e.g., air quality and noise levels, through agile aerial crowdsourcing. Despite this potential, existing UAV-based sensing approaches overlook environmental disturbances like wind that drastically impact drone velocity and energy efficiency. Consequently, directly applying existing methods to this joint delivery and sensing paradigm in dynamic environments faces two severe challenges: (1) scalability bottlenecks as fleet sizes expand; and (2) multi-timescale decision heterogeneity between macro task dispatching and micro velocity control. To tackle these, we formalize the problem as SensUAV and propose a Two TimeScale Reinforcement Learning framework (TSRL). Specifically, TSRL separates decision-making into two cooperative layers. At the macro level, a task-embedding sensing dispatcher handles scalability by separately encoding distinct task features and sequentially evaluating UAV suitability before task selection. At the micro level, a wind-aware velocity controller learns fine-grained velocity scheduling to adapt to dynamic environmental variations. Extensive experiments on real-world datasets demonstrate that TSRL significantly outperforms baselines, achieving average system profit improvements of 20.1% in Hangzhou and 46.6% in Shanghai.
Xin Ouyang, Songxin Lei, Xusen Guo +3
The Hong Kong University of Science and Technology (Guangzhou) Guangzhou, China · Beijing Institute of Technology Beijing, China
UAVs have emerged as highly flexible platforms for data sensing in Wireless Sensor Networks (WSNs). Path planning for UAVs in such tasks plays a key role to assure remote sensing effectiveness and friendly energy consumption. However, existing approaches show two key limitations: i) they are primarily hand-crafted with certain design biases that harm adaptation on unseen tasks. ii) they predominantly assume idealized spatial complexities of actual environments through simplified simulation, causing them to underperform during real-world deployment. In this paper, we propose a novel learning-assisted planning framework, termed Landscape-Aware Meta Differential Evolution (LAMDE), to tackle the mentioned limitations. The major contributions come from the following aspects. We first re-formulate such UAV path planning problem to embrace challenging constraints. To efficiently navigate this highly constrained space, we propose a bi-level learning to optimize approach, where the meta-level is a trainable algorithm configuration policy that meta-learns an adaptable planning strategy for low-level planning algorithm. To address the potential training data scarcity and distribution shift in real-world environments, we introduce a landscape-aware automatic augmentation scheme that enriches training data. At the low-level, a Differential Evolution algorithm is deployed for solving the path planning tasks. To enhance the solving flexibility, we further design a variable-length encoding strategy that dynamically prunes redundant hover points and optimizes continuous flight parameters concurrently within a unified search space. Based on all proposed designs, we meta-train LAMDE and compare it with representative baselines. Comprehensive experiments demonstrate that LAMDE achieves state-of-the-art performance on the tested complex UAV path planning tasks in WSN data collection scenarios.
Sijie Ma, Zeyuan Ma, Weijia Cao +4
South China University of Technology · South China Normal University · Information Research Institute, CAS +2
Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate trajectories using predicted futures, but typically rely on only one or a few point predictions, which is inadequate for large-scale outdoor environments with substantial future-state uncertainty. To address this limitation, we propose the Uncertainty-Aware Navigation World Model (UA-NWM), an efficient latent world model for aerial image-goal navigation, which formulates trajectory scoring as conditional out-of-distribution detection. UA-NWM represents plausible futures with an uncertainty subspace and decomposes the prediction--goal discrepancy into uncertainty-explainable and unexplainable components. Only the unexplainable residual is used for scoring, enabling robust selection without multiple future samples. Extensive experiments demonstrate that UA-NWM consistently outperforms existing navigation world models while maintaining low inference latency. Real-world UAV experiments further validate its practical applicability. Project page: https://duryi.github.io/UA-NWM-Project-Page
Deyi Zhu, Haoyu Fan, Yinan Zhu +4
Tsinghua Shenzhen International Graduate School, Tsinghua University