Memory in the Sky: Low-Altitude Question Answering with Multi-Agent Memory Aggregation
Authors: Chengyang Li, Yujie Wan, Shuai Wang, Kejiang Ye, Weijie Yuan, Boyu Zhou, Yik-Chung Wu, Chengzhong Xu, +1 more
Organizations: The University of Hong Kong, Hong Kong SAR, China · Southern University of Science and Technology, Shenzhen, China · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China · University of Macau, Macau SAR, China · Istanbul Medipol University, Istanbul, Turkey
This paper studies low-altitude question answering (LAQA), in which distributed unmanned aerial vehicle (UAV) memories are aggregated at a ground server to answer questions about observations over a long horizon. Unlike conventional resource allocation based on sensing, communication, control, or computation metrics, LAQA requires an explicit measure of memory value. We propose a generative adversarial exam (GAE) that uses forward simulation to evaluate memory retrieval and exam scores to quantify memory quality. This enables the downstream QA value of candidate memories to be measured and optimized without accessing the internal mechanisms of the black-box captioning, retrieval, and reasoning pipeline. Building on this metric, we develop a memory-centric (MemCen) framework that jointly selects UAVs and allocates transmit power to maximize memory quality under communication constraints. In the noise-limited regime, we derive a QoM-aware capped water-filling law that explicitly connects task utility with physical-layer power allocation. We further develop penalty successive optimization (PSO) and learning to memorize (L2M) solvers. MemCen achieves QA accuracies of 92.4% and 84.0% in CARLA Town04 and Town05 under static and dynamic communication conditions, respectively. In real-world experiments, MemCen achieves 88.5% QA accuracy on the panoramic multi-agent system (PMAS) benchmark. Finally, UAV-to-robot-dog demonstrations further validate the practical utility of the acquired memories for environmental understanding and navigation.
Figures & tables
Fig. 1 : System architecture of LAQA with memory collection and memory-empowered query.
Fig. 2 : Architecture of GAE for QoM computation.
Fig. 3 : L2M for real-time decision making.
Fig. 5 : Captioning and QA in the 5 -UAV scenario.
UAV ID
SemCom
Qwen3-8B ( 6.43 s per QA)
Qwen3-14B ( 12.17 s per QA)
Qwen3-8B
[ 15 ]
GAE Score
GAE QA Accuracy
GAE Score
GAE QA Accuracy
Downstream QA Accuracy
UAV 01
0.8913
1.05
79%
0.60
88%
39%
UAV 02
0.8664
2.75
45%
1.65
67%
48%
UAV 03
0.8746
3.25
35%
1.20
76%
59%
UAV 04
0.9009
3.65
27%
2.25
55%
70%
UAV 05
0.9394
1.40
72%
0.20
96%
38%
TABLE I : Evaluation of GAE. GAE scores are mean unanswered-question counts for five-question exams. Purple denotes the best memory selected by the method. Bold denotes the highest value.
Fig. 6 : Visualization of ten UAV configurations, image frames, and the associated captions.
Fig. 7 : Convergence analysis of PSO.
Method
↑ Sum GAE
UAVs
↑ Sum rate (Mbps)
↓ Time (s)
MemCen+PSO
24.62±9.00
3.50±1.41
45.03±16.57
4.60±0.40
MemCen+L2M
23.38±12.99
3.10±1.69
42.44±21.91
0.04±0.03
MemCen+Relax-Round
20.75±7.12
3.20±1.20
35.26±14.19
9.16±0.43
MemCen+DQN
22.20±8.83
2.93±1.14
47.81±18.67
0.017±0.0003
TABLE II : Comparison of solvers under the same MemCen formulation.
Fig. 8 : Accuracy–rate tradeoff and power-budget sensitivity.
Metric
QA Accuracy ↑
Sum GAE Score ↑
UAVs ↑
Sum Rate ↑
ComCen
[ 12 ]
65.6%
10.2
4.1
74.70 Mbps
SenCen
[ 10 ]
88.6%
19.9
8.1
25.48 Mbps
FairCen
[ 29 ]
46.0%
2.5
1.0
21.46 Mbps
Greedy
[ 30 ]
89.8%
17.3
4.4
44.56 Mbps
Remember
[ 3 ]
41.0%
–
–
–
MemCen
( Ours )
92.4%
20.7
7.4
24.31 Mbps
TABLE III: Quantitative results of QA tasks.
Method
↑ Sum GAE Score
↑ UAVs
↑ Sum Rate
↑ Min. Rate
ComCen
[ 12 ]
12
4
63.9 Mbps
0 Mbps
SenCen
[ 10 ]
16
7
21.9 Mbps
0 Mbps
FairCen
[ 29 ]
0
0
17.2 Mbps
1.72 Mbps
Greedy
[ 30 ]
16
3
33.5 Mbps
0 Mbps
MemCen
(Ours)
21 (+ 31.25% )
6
20.1 Mbps
0 Mbps
TABLE IV : Evaluation of the 10 -UAV Case
Fig. 9 : GAE scores, channel conditions, and UAV selections of different schemes.
Fig. 10 : Comparison between MRC and IRC.
Fig. 11 : Dynamic Town05 setting. (a)–(b) show the trajectories and geometry-induced channel blockage. (c)–(f) show task switch on the same map and routes. Hatched and solid bars denote fixed-wing and multirotor profiles, respectively.
Method
↑ QA accuracy
↑ Sum QoM
Selected UAVs
↑ Sum rate (Mbps)
Remember
33.0±14.9%
0.000±0.000
0.00±0.00
0.00±0.00
FairCen
33.0±14.9%
0.000±0.000
0.00±0.00
31.06±0.29
ComCen
59.0±16.5%
1.505±0.544
4.00±0.00
369.10±0.29
SenCen
71.0±15.2%
2.435±0.509
7.00±0.00
289.84±0.22
SemCom
79.0±16.5%
2.770±0.540
7.80±0.41
230.18±25.07
MemCen
84.0±10.5%
2.775±0.454
6.70±0.73
285.91±0.29
TABLE V : Town05 results over 20 random trials.
FW frames per UAV
MR frames per UAV
Mean frame size FW/MR (MB)
↑ Sum rate (Mbps)
Selected UAVs
↑ Sum QoM
↑ QA accuracy
1,000
1,000
0.666/0.286
184.87±0.03
9.00±0.00
3.390±0.409
93.0±9.8%
1,500
1,000
0.667/0.286
244.66±4.14
8.45±0.83
3.185±0.627
89.0±12.1%
2,000
1,000
0.666/0.286
285.91±0.29
6.70±0.73
2.775±0.454
84.0±10.5%
TABLE VI : Sensitivity of MemCen to image loads in Town05. FW and MR denote fixed-wing and multirotor UAVs.
Fig. 13 : Case-level comparison of spatially grounded QA and the visual evidence retrieved by MemCen.
Fig. 14 : QoM and GAE matrices for each initial memory, and a case study with M0=M1 .
Method
↑ QA accuracy
↑ YES/NO
↑ Where
↑ Sum QoM
Avg. uploads
↑ Sum rate
Remember
45.25±4.32%
50.5%
40.0%
0.000
0.00
–
SenCen
76.50±8.67%
83.5%
69.5%
0.640
1.90
36.53 Mbps
ComCen
61.50±12.05%
71.0%
52.0%
0.320
1.15
36.68 Mbps
FairCen
62.00±12.19%
72.5%
51.5%
0.310
1.30
36.63 Mbps
MemCen
88.50±4.50%
98.0%
79.0%
0.990
1.90
25.65 Mbps
TABLE VII : Quantitative results over 20 random trials on the PMAS data.
Fig. 15 : QoM stability and downstream QA versus the pilot sampling ratio. Error bars show the standard deviation of random sampling. Dashed lines show the full-data reference.
VLM
LLM
Embedding
QA accuracy
Mean QoM
Comparison of different VLMs
(tested on the 20-question benchmark)
4B
30B
Mxbai
80.00%
0.500
8B
30B
Mxbai
90.00%
0.567
32B
30B
Mxbai
100.00%
0.450
Comparison of different LLMs and embeddings
TABLE VIII : Model sensitivity on the real-world data.
Fig. 16 : Experimental setup of the real-world UAV-to-ground-robot collaboration task.
Fig. 17 : Single-UAV views. Their positions are marked by purple boxes in Fig. 16 (c).
Fig. 18 : GAE accuracy and score in real-world experiments.
Fig. 19 : Robot-dog QA and navigation using the UAV memory. Dog positions are marked by yellow boxes in Fig. 16 (c).
This paper considers multi-agent embodied question answering (MA-EQA), which aims to query robot teams on what they have seen over a long horizon. In contrast to existing edge resource management methods that emphasize sensing, communication, or computation performance metrics, MA-EQA emphasizes the memory qualities. To cope with this paradigm shift, we propose a quality of memory (QoM) model based on generative adversarial exam (GAE), which leverages forward simulation to assess memory retrieval and uses the resulting exam scores to compute QoM values. Then we propose memory centric power allocation (MCPA), which maximizes the QoM function under communication resource constraints. Through asymptotic analysis, it is found that the transmit powers are proportional to the GAE error probability, thus prioritizing towards high-QoM robots. Extensive experiments demonstrate that MCPA achieves significant improvements over extensive benchmarks in terms of diverse metrics in various scenarios.
Chengyang Li, Shuai Wang, Kejiang Ye +5
Department of Electrical and Computer Engineering, The University of Hong Kong, Hong Kong · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China · Southern University of Science and Technology, Shenzhen, China +2
Long-term memory is essential for LLM agents to reason coherently across extended interactions, personalize responses, and reuse past experience. However, existing memory-augmented methods typically treat memory as a fixed resource: text-space approaches concatenate retrieved memories into the context window, causing substantial token overhead and sensitivity to noisy evidence, while latent-space approaches reduce textual cost but still rely on rigid retrieval or fixed-capacity memory interfaces. This creates a mismatch between query-dependent memory utility and fixed memory allocation. We propose ElasticMem, a memory-augmented LLM framework that learns to use memory as an elastic latent resource. ElasticMem builds an offline latent memory bank with retrieval keys and content caches, retrieves memories adaptively from the reasoner's hidden state, assigns each retrieved memory a variable latent budget through a learned policy, and injects selected latent states as soft memory tokens for generation. The full memory-use process is optimized with downstream task rewards through group-relative policy optimization. We evaluate ElasticMem on MemorySuite, covering memory-intensive QA and embodied agent control. Across Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct backbones, ElasticMem improves weighted average QA accuracy by 26.2% and 24.6%, and improves ALFWorld success rate by 66.3% and 27.2%, respectively, over the strongest baselines, while achieving the lowest ALFWorld token cost. Ablations and qualitative analyses further show that adaptive retrieval and elastic budget allocation help ElasticMem prioritize useful evidence and transferable plans beyond rigid cosine similarity. Our code for ElasticMem will be released at https://github.com/ulab-uiuc/ElasticMem.
Tao Feng, Chongrui Ye, Tianyang Luo +5
University of Illinois Urbana-Champaign · Nanyang Technological University
Memory data are ubiquitous in Large Language Model (LLM)-based agents (e.g., OpenClaw and Manus). A few recent works have attempted to exploit agents'memory for improving their performance on the question-answering (QA) task, but they lack a principled mechanism for effectively modeling how memory data evolves over time and retrieving memory data effectively, leading to poor performance in memory utilization. To fill this gap, we present H-Mem, a novel memory mechanism via a hybrid structure that can not only effectively model the evolution of agent memory over a long period of time, but also provide an efficient memory retrieval approach. Particularly, H-Mem builds a temporal and semantic tree structure that allows the short-term memory data to evolve progressively into long-term memory data, where the latter provides summarized information about the former, while simultaneously constructing a knowledge graph to capture the relationships between entities in memory. Moreover, it offers an effective memory retrieval approach by exploiting the hybrid structure of the tree and graph structures. Extensive experiments on three agent memory benchmarks show that H-Mem achieves state-of-the-art performance on the QA task.
Jiawei Yu, Yixiang Fang, Xilin Liu +1
The Chinese University of Hong Kong, Shenzhen · Huawei Cloud Computing Technologies CO., LTD.