cs.CLOct 7, 2026

AI4Fire: Evaluating Large Language Models on Wildfire Tasks

Authors: Yue Zhao, Xiyang Hu, Zuobin Xiong, Zhangyu Wang, Ruolin Li

Organizations: University of Southern California · Arizona State University · University of Nevada, Las Vegas · University of Maine

Abstract

Large language models (LLMs) are entering wildfire management, where overstated evaluations can cost property and lives. How do they perform on wildfire tasks, with and without grounding? Bare means a model receives the task input alone. Grounded means it also receives one task-specific addition: for smoke detection, a smoke-free reference frame from the same camera. AI4Fire runs six core models bare and grounded on five wildfire tasks, zero-shot; a sweep adds 29 more. Our literature search on fire tasks found 138 works; none combines this roster, task coverage, and paired bare and grounded runs. We report three findings. (1) Grounding helped most where the addition carried the answer: a read-only SQL tool lifted every core model's database accuracy from at most 16 to at least 88 percent. (2) Simple rules were hard to beat: no core model outperformed repeating today's staffing count, and two open-weight models mostly copied the median of similar earlier fire-days, a worse forecast. (3) Public releases carry hazards: a fire-danger column separates the holdout perfectly, and 67 aerial fire frames carry smoldering or fire-free labels read from a clipped thermal maximum. We release prompts, responses, scores, code, and the survey record.

Figures & tables

Appendix figures & tables42 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

    Sep 7, 2026Pengfei Li, Naufal Suryanto, Sicheng Zhang +2LLM Safety BenchmarksLLM Safety Evaluation

  2. Does Your Wildfire Prediction Model Actually Work, or Just Score Well?

    May 14, 2026Yangshuang Xu, Yuyang Dai, Liling Chang +2Geospatial Foundation ModelsWildfire Forecasting

  3. Evaluating the Generalizability of Foundation Models for Extreme Environmental Events: Case Study of California Wildfire PM2.5

    Jul 8, 2026Yongcan Huang, Li Jiang, Ze Yu LiuWildfire ForecastingOOD Generalization