cs.ROOct 7, 2026

AeroEval: Staged Program and Execution Validation for AI-Generated Drone Missions

Authors: Kautuk Astu, Naina Rabha, Yogesh Simmhan

Organizations: Department of Computational and Data Sciences, Indian Institute of Science, Bangalore 560012, India

Abstract

Large Language Models (LLMs) can generate drone programs from natural-language mission descriptions, but syntactically valid programs may still violate user intent, environmental constraints, and mission-level behavior. This problem is pronounced in cyber-physical applications, where correctness depends on the interaction among generated code, mobile sensing, environmental geometry, event-driven analytics, and physical execution. Existing drone code-generation systems primarily use prompt guardrails or simulator outcomes and provide limited failure localization. We present AeroEval, an agent-assisted middleware for staged validation of AI-generated drone missions. AeroEval combines deterministic program analysis with context-grounded LLM agents. It first validates program syntax, platform API usage, and mission intent, and then evaluates the realized behavior using execution trajectories, mission requirements, and environmental context. Each stage returns structured failure information for iterative regeneration. In our evaluation using 20 navigation tasks and five analytical mission types over AirSim and Gazebo simulators, AeroEval improves navigation success from 55% to 95%. In a stagewise ablation study, our Code and Trajectory Validators by themselves achieve mean run-level success rates of 44% and 56%, respectively, while the full AeroEval pipeline achieves 88%; the stages detect complementary failures in program structure, API usage, mission intent, obstacle avoidance, altitude, coverage, and event-driven transitions and the guided regeneration corrects for them. Across the main analytics missions, AeroEval increases aggregate run-level success from 34% for one-shot AeroGen to 88% within the regeneration budget. These results demonstrate the benefit of combining program-level and execution-grounded agentic validation for AI-generated drone applications in the evaluated environment.

Figures & tables

Explore similar work

CardsList
  1. AeroCopilotBench: Safety-Gated Evaluation of LLM Agents on Aircraft Emergency Procedures in an Executable Cockpit

    Aug 17, 2026Yuchen Yuan, Zhenghuang Wu, Yuangan Li +2LLM Safety BenchmarksAI Agent Safety Benchmarks

  2. Say the Mission, Execute the Swarm: Agent-Enhanced LLM Reasoning in the Web-of-Drones

    May 5, 2026Andrea Iannoli, Lorenzo Gigli, Luca Sciullo +2Aerial RoboticsMulti-Robot Systems