AI4Fire: Evaluating Large Language Models on Wildfire Tasks
Organizations: University of Southern California · Arizona State University · University of Nevada, Las Vegas · University of Maine
Abstract
Large language models (LLMs) are entering wildfire management, where overstated evaluations can cost property and lives. How do they perform on wildfire tasks, with and without grounding? Bare means a model receives the task input alone. Grounded means it also receives one task-specific addition: for smoke detection, a smoke-free reference frame from the same camera. AI4Fire runs six core models bare and grounded on five wildfire tasks, zero-shot; a sweep adds 29 more. Our literature search on fire tasks found 138 works; none combines this roster, task coverage, and paired bare and grounded runs. We report three findings. (1) Grounding helped most where the addition carried the answer: a read-only SQL tool lifted every core model's database accuracy from at most 16 to at least 88 percent. (2) Simple rules were hard to beat: no core model outperformed repeating today's staffing count, and two open-weight models mostly copied the median of similar earlier fire-days, a worse forecast. (3) Public releases carry hazards: a fire-danger column separates the holdout perfectly, and 67 aerial fire frames carry smoldering or fire-free labels read from a clipped thermal maximum. We release prompts, responses, scores, code, and the survey record.
Figures & tables
| Coverage | Compares | ||||
| Work | Survey | Cat. | Models | B/G | Ref. |
| Fire AI reports with no model evaluation | |||||
| OGC D-123 ( 2024 ) | |||||
| OGC 25-012 ( 2026 ) | |||||
| Fire benchmarks that evaluate no LLM or agent | |||||
| WILDFIRE-FM ( 2026 ) | 1 | ||||
| Grounded arm | ||||
|---|---|---|---|---|
| Task | Addition | Ans. | Data source | |
| Detection and Perception | ||||
| Smoke | Reference frame | FIgLib ( 2022 ) | 196 | |
| Aerial QA | Thermal block | WildFireVQA ( 2026b ) | 408 | |
| Forecasting and Prediction | ||||
| Fire danger | Climatology | Mesogeos ( 2023 ) | 386 | |
| Evaluation | Score range, core six | Paired contrast | ||||
|---|---|---|---|---|---|---|
| Metric | Items (clusters) | Comparator | Bare | Grounded | span | Excl. 0 |
| Personnel allocation, ICS-209-PLUS. Grounding: up to six retrieved analogues (rule v1). | ||||||
| Norm. error | 300 (245 incidents) | Persistence 0.146 | 0.157–0.193 | 0.170–0.248 | to | 3 |
| Smoke detection, FIgLib, first design. Grounding: one earlier clear frame from the same camera. | ||||||
| Smoke recall | 196 (17 fires) | Frame diff. 0.920 † | 0.509–0.643 | 0.554–0.741 | to | 4 ∗ |
| Aerial question answering, WildFireVQA. Grounding: a seven-number thermal summary block. | ||||||
| Entry (task) | Evidence |
|---|---|
| Model failure patterns | |
| Median copying (allocation) | Open-weight forecasts copy the displayed analogue median on 284 and 259 of 300 items |
| Guessed date format (tool use) | 30 of the open-weight models’ 35 failures filter dates in a guessed format |
| Properties of the benchmark or a source | |
| Unit gloss (aerial, ours) | Our prompt calls the percent fields shares of pixels; 7 of the 9 misses the block added fit a fraction reading |
| Clipped maximum (aerial) | 15 of 34 misses, on frames clipped at 187.4 ∘ C and labeled smoldering or fire-free |
Appendix figures & tables42 assets
Supplementary material from the paper’s appendix.
Appendix
| Grouping | Value | Works |
| Paper kind | ||
| Kind | system-with-eval | 75 |
| Kind | evaluation | 21 |
| Kind | benchmark | 14 |
| Kind | survey | 14 |
| Kind | position | 8 |
| Work | What It Does | What It Leaves Undone |
|---|---|---|
| Hyun et al. (2025) | Benchmarks four multi-agent frameworks using GPT-4o in a procedurally generated wildfire environment. | Evaluates one task family (operations in simulation), tests one base model, and provides no literature survey. |
| Zhang et al. (2026) | Evaluates 21 baseline MLLMs plus DisasterVL on nine UAV disaster-response tasks, including wildfire propagation. | Focuses on general disasters using UAV imagery only; contains no survey of fire literature. |
| Chen et al. (2025a) | Evaluates about 14 LLM configurations on wildfire personnel and cost allocation, with and without geospatial grounding. | Covers only two related decision tasks. |
| Tricomi (2024) | Formulates domain adaptation, National Incident Management System and Incident Command System (NIMS/ICS) operational lifecycle mapping, and governance frameworks for generative AI and agent workflows in wildland fire management. | No model evaluation across tasks. |
| Open Geospatial Consortium (2026) | Assesses over 200 Canadian wildfire datasets for generative AI readiness and outlines prototype architectures for LLMs and agents. | No model evaluation across tasks. |
| Xu et al. (2026) | Compares a wildfire foundation model with ten Earth foundation models across six task forms (occupancy, spread, burned area, analog retrieval, smoke PM 2.5 , and extreme heat) under a fixed evaluation contract. | No LLMs or agents. |
| Task or Source | Category | Domain | Verdict | Access | License |
|---|---|---|---|---|---|
| Evacuation guidance dialog, BEACON | Communication and Alerts | wildfire | not feasible now | none found; Watch Duty feed closed to automated access | paper CC BY 4.0; no data license |
| Crisis translation and urgency classification | Communication and Alerts | both | needs work | public download for three of four corpora | CC BY 4.0 on the one fire-bearing multilingual corpus |
| Fire database and tool use, FPA-FOD and public APIs | Data Retrieval and Tool Use | both | needs work | public download or free API for all five sources | usable without additional permissions or fees |
| Fire radiative power forecasting, FIRMS | Forecasting and Prediction | wildfire | needs work | paper table unreleased; FIRMS behind a free key | NASA open data; no license on the paper table |
| Wildfire risk rasters, FireScope-Bench | Forecasting and Prediction | wildfire | needs work | public download, 127 GB | CC BY 4.0 plus upstream provider terms |
| Risk sub-criteria scoring against expert ranks | Forecasting and Prediction | wildfire | needs work | public repository download, no registration | article CC BY 4.0; repository states none |
| Task or Source | Category | Domain | Verdict | Access | License |
|---|---|---|---|---|---|
| Foundation-model transfer tasks, WILDFIRE-FM | Forecasting and Prediction | wildfire | not feasible now | anonymous repository serves code, metadata, and paper-result files; the Hugging Face link returned HTTP 401 to anonymous requests on 2026-09-18; raw inputs and large feature arrays are excluded | paper CC0; code MIT; source data under provider terms |
| Fire code compliance, RegCheck audit set | Knowledge Question Answering | structure fire | needs work | audit set none found; code text public | audit set none stated; Approved Document B under OGL v3.0 |
| Fire science and safety question answering | Knowledge Question Answering | both | needs work | public download for all four source corpora | OGL v3.0 and CC BY 4.0 |
| Barriers to fire spread, WFDSS incident text | Document Understanding | wildfire | not feasible now | restricted for sensitivity | article CC BY-NC-ND 4.0; no data license |
| Incident-command transcript task state | Document Understanding | structure fire | needs work | public repository download, no registration | none stated on the repository |
| Bushfire damage assessment reports | Document Understanding | wildfire | not feasible now | no dataset released | code CC BY-NC-ND 4.0, which blocks derivatives |
| Task or Source | Category | Domain | Verdict | Access | License |
|---|---|---|---|---|---|
| Operations by incident-command phase, OGC 24-071 | Decision Support and Operations | wildfire | not feasible now | report public; no evaluation items in it | OGC document license; NWCG source text public domain |
| Risk and insurance retrieval, OGC 25-012 | Decision Support and Operations | wildfire | not feasible now | report public; no evaluation items in it | OGC document license; no license given for listed sources |
| Suppression and rescue, CREW-Wildfire | Simulation-Coupled Agents | wildfire | ready | public download; environment and scoring released | code Apache-2.0; paper CC BY-NC-ND 4.0 |
| Policy adaptation, MOASEI and free-range-zoo | Simulation-Coupled Agents | wildfire | needs work | public download; frozen evaluation tag | MIT on the environment; AGPL-3.0 on the starter |
| Zero-shot active fire classification, Roboflow set | Detection and Perception | wildfire | not feasible now | Roboflow download with a free account | CC BY 4.0 |
| Smoke classification and localization, SmokeBench | Detection and Perception | wildfire | needs work | none found for the benchmark; FIgLib images public | preprint CC BY 4.0; benchmark none |
| Task or Source | Category | Domain | Verdict | Access | License |
|---|---|---|---|---|---|
| Satellite risk reasoning, Wildfire-Smoke-RS | Detection and Perception | wildfire | not feasible now | none found; the repository holds only a README | none stated |
| UAV disaster reasoning, DisasterBench | Detection and Perception | both | needs work | public download of the test split | Apache-2.0 |
| Smoke sequences, FIgLib | Detection and Perception | wildfire | ready | public download, no registration | FIgLib page names no license and requires credit; HPWREN data-use page lists CC BY-NC-ND 4.0 |
| Fire-scene severity items, DetectiumFire | Detection and Perception | both | not feasible now | Kaggle download with a free account | two authoritative sources disagree, both non-commercial |
| Structure-fire video questions, Fire360 | Detection and Perception | structure fire | needs work | public Box download, no login | paper: MIT with research-only restrictions; released license file: standard MIT text |
| Post-fire change captions, RSCC | Detection and Perception | wildfire | not feasible now | public download, not gated | paper: CC BY 4.0; repository binds the xBD-derived subset to the xBD terms and research use |
| Source | What reproduces | What it does not establish |
|---|---|---|
| Mesogeos Track A | A released static column separates the holdout at average precision 1.000; a rule over two window columns reaches 0.814 on the last day | That the published baselines read those columns. All three released configurations name their inputs and exclude them |
| Fire360 | The released 100 questions carry no reference to a video, a frame, or a timestamp, and the answer key puts the correct choice first on 25 of 50 | That the released items are the scored ones. The release calls them sample prompts, and no mapping ties them to the reported evaluation |
| WildFireVQA | Four question types are reproduced exactly by closed-form rules on the supplied thermal block, on 6,097 of 6,097 frames each. On 67 frames filed as fire, the maximum is clipped at 187.4 ∘ C, and every temperature-derived label reads smoldering or no fire | That the block is illicit, that a model missing these labels ignored the block, or where the clipping arises. The source sets out to study thermal retrieval; what is missing is a text-only arm beside the pooled score |
| This benchmark | A few-hundred-token output cap truncated a reasoning model’s answer on 40 of 196 items, against 8 for another family | That the affected runs measure model skill. The cap selected which items each family was scored on |
| Allocation nMAE | Fire danger AUPRC | Fire-call rate | Tool use acc. | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Cap | bare | grd. | bare | grd. | bare | grd. | bare | tool |
| Nova Micro | 1,536 | 0.166 | 0.289 | 0.339 | 0.332 | 0.000 | 0.000 | 0.019 | 0.821 |
| Llama 3.3 70B | 1,536 | 0.927 | 0.248 | 0.413 | 0.458 | 0.446 | 0.440 | 0.090 | 0.288 |
| Llama 3.1 70B | 1,536 | 0.205 | 0.257 | 0.400 | 0.427 | 0.145 | 0.202 | 0.006 | 0.000 |
| Mistral Small 2402 | 1,536 | 0.491 | 1.958 | 0.401 | 0.407 | 0.858 | 0.671 | 0.058 | 0.038 |
| Devstral 2 123B | 1,536 | 0.179 | 0.237 | 0.522 | 0.522 | 0.681 | 0.674 | 0.096 | 0.872 |
| Model | Condition | Mean | ECE | Brier | Reliability | Resolution | Departures | Omitted |
| Reference points | ||||||||
| Calendar-month prior | – | 0.333 | 0.037 | 0.182 | 0.003 | 0.046 | – | – |
| Proprietary models | ||||||||
| claude-opus-4.8 | bare | 0.187 | 0.154 | 0.206 | 0.034 | 0.051 | 50 | 0 |
| claude-opus-4.8 | grounded | 0.220 | 0.119 | 0.200 | 0.016 | 0.040 | 50 | 2 |
| claude-opus-5 | bare | 0.117 | 0.222 | 0.230 | 0.069 | 0.061 | 11 | 46 |
| Model | Condition | Effective | ECE | Brier | Rel. | Res. | AUPRC | Brier minus prior |
| prior | and 95% CI | |||||||
| As stated | ||||||||
| Calendar-month prior | – | 0.331 | 0.037 | 0.182 | 0.003 | 0.046 | 0.555 | – |
| claude-opus-4.8 | bare | 0.149 | 0.154 | 0.206 | 0.034 | 0.051 | 0.613 | 0.024 [ , 0.049] |
| claude-opus-4.8 | grounded | 0.184 | 0.119 | 0.200 | 0.016 | 0.040 | 0.581 | 0.018 [ , 0.043] |
| claude-opus-5 | bare | 0.083 | 0.222 | 0.230 | 0.069 | 0.061 | 0.688 | 0.048 [0.016, 0.079] |
| ECE | ||||||||
|---|---|---|---|---|---|---|---|---|
| Model | Arm | Mean | First | Rerun | Brier | Brier minus prior | Calls | AUPRC |
| Calendar prior | – | 0.333 | 0.037 | – | 0.182 | – | – | 0.555 |
| claude-opus-4.8 | bare | 0.187 0.305 | 0.154 | 0.068 | 0.206 0.179 | [ , 0.014] | 0.215 0.311 | 0.613 0.607 |
| claude-opus-4.8 | grounded | 0.220 0.355 | 0.119 | 0.076 | 0.200 0.189 | +0.006 [ , 0.024] | 0.249 0.399 | 0.581 0.581 |
| claude-opus-5 | bare | 0.117 0.267 | 0.222 | 0.087 | 0.230 0.168 | [ , 0.006] | 0.065 0.174 | 0.688 0.701 |
| claude-opus-5 | grounded | 0.116 0.289 | 0.224 | 0.066 | 0.232 0.173 | [ , 0.010] | 0.109 0.241 | 0.682 0.663 |
| Fit | MAE | Norm. | Beats | Stable | Moving | False | Missed | Dir. |
|---|---|---|---|---|---|---|---|---|
| base | days | days | move | move | right | |||
| Reference points | ||||||||
| Boosted regressor, next-day ratio | 12.97 | 0.144 | 0.48 | 0.035 | 0.317 | 0.03 | 0.93 | 0.72 |
| Same, plus the displayed analogue median | 13.03 | 0.145 | 0.48 | 0.035 | 0.319 | 0.03 | 0.94 | 0.71 |
| Boosted regressor, next-day count | 13.68 | 0.201 | 0.42 | 0.126 | 0.319 | 0.14 | 0.81 | 0.66 |
| Ridge regression, next-day count | 19.05 | 0.416 | 0.28 | 0.412 | 0.423 | 0.63 | 0.47 | 0.57 |
| Model | Prompt | AUPRC | F1 on fire | Call rate | Omitted | Difference in AUPRC | Spearman |
| Proprietary models | |||||||
| claude-opus-4.8 | p0 | 0.613 | 0.561 | 0.215 | 0 | ||
| claude-opus-4.8 | p1 | 0.605 | 0.234 | 0.060 | 0 | [ , 0.035] | 0.93 |
| claude-opus-4.8 | p2 | 0.631 | 0.059 | 0.010 | 0 | 0.018 [ , 0.057] | 0.91 |
| claude-opus-5 | p0 | 0.688 | 0.269 | 0.065 | 46 | ||
| claude-opus-5 | p1 | 0.731 | 0.308 | 0.065 | 0 | 0.043 [0.009, 0.080] | 0.92 |
| Model | Bare | v1 | v2 | v2 bare | v2 v1 | False move | Copies |
| Proprietary models | |||||||
| claude-opus-4.8 | 0.157 | 0.170 | 0.166 | 0.010 | 0.14 / 0.29 / 0.22 | 123 / 145 | |
| [ , 0.023] | [ , 0.014] | ||||||
| claude-opus-5 | 0.165 | 0.173 | 0.166 | 0.000 | 0.25 / 0.27 / 0.20 | 82 / 93 | |
| [ , 0.011] | [ , 0.006] | ||||||
| gemini-3.1-pro | 0.179 | 0.177 | 0.166 | 0.23 / 0.29 / 0.20 | 154 / 152 | ||
| Model | Condition | 0 to 10 min | 10 to 25 min | 25 min or more |
|---|---|---|---|---|
| Proprietary models | ||||
| claude-opus-4.8 | bare | 0.43 | 0.57 | 0.62 |
| claude-opus-4.8 | grounded | 0.46 | 0.66 | 0.68 |
| claude-opus-5 | bare | 0.50 | 0.66 | 0.62 |
| claude-opus-5 | grounded | 0.50 | 0.72 | 0.65 |
| gemini-3.1-pro | bare | 0.25 | 0.64 | 0.65 |
| Model | Condition | CL | CMR | DS | FP | LD | PD |
|---|---|---|---|---|---|---|---|
| Per-question majority | 0.708 | 0.438 | 0.573 | 0.604 | 0.542 | 0.771 | |
| Proprietary models | |||||||
| claude-opus-4.8 | bare | 0.625 | 0.562 | 0.625 | 0.625 | 0.521 | 0.781 |
| claude-opus-4.8 | grounded | 0.569 | 0.708 | 0.604 | 0.604 | 0.542 | 0.771 |
| claude-opus-5 | bare | 0.694 | 0.625 | 0.594 | 0.646 | 0.562 | 0.729 |
| claude-opus-5 | grounded | 0.681 | 0.750 | 0.625 | 0.646 | 0.542 | 0.688 |
| Model | Diff. | By frame | By question type | Two-way |
|---|---|---|---|---|
| Grounded bare | ||||
| claude-opus-4.8 | [ , ] | [ , ] | [ , ] | |
| claude-opus-5 | [ , ] | [ , ] | [ , ] | |
| gemini-3.1-pro | [ , ] | [ , ] | [ , ] | |
| gpt-6-astra | [ , ] | [ , ] | [ , ] | |
| Qwen3-VL | [ , ] | [ , ] | [ , ] | |
| Work | Category | Domain | Models confirmed | Bare or grounded | Headline number |
|---|---|---|---|---|---|
| RegCheck-Hybrid ( Chen and Zhou, 2025 ) | Knowledge QA | Structure | Rule-LLM hybrid, RAG LLM baseline | Both | 89% F1 on building fire code compliance |
| Triple-Phase control ( Sun et al., 2026 ) | Knowledge QA | Structure | Commercial RAG LLMs, pure LLMs | Both | Hallucination on numerical thresholds eliminated |
| WFDSS barrier analysis ( Epstein and Seielstad, 2025 ) | Documents | Wildfire | GPT-4 (1) | Bare | 13 barrier types over 24,254 entries; roads 42% |
| Task-state monitoring ( Grünert et al., 2026 ) | Documents | Fire service | gpt-5.2, gpt-5.4 (2) | Bare | Strong incrementally; unit-assignment timing weakest |
| DisasTeller ( Chen et al., 2026 ) | Documents | Multi-hazard | GPT-4o, Gemma3 (2) | Grounded | Damage reports on earthquake, flood, and bushfire |
| Fire protection specifications ( Kim and Shin, 2025 ) | Documents | Structure | Gemini, Claude, ChatGPT (3) | Bare | Adequate structures; expert review still required |
| Work | Category | Domain | Models confirmed | Bare or grounded | Headline number |
|---|---|---|---|---|---|
| DisasterBench ( Zhang et al., 2026 ) | Detection | Multi-hazard | 21 baseline MLLMs plus DisasterVL | Bare | 2B DisasterVL matched GPT-4o at 72.60% |
| WildfireGPT assessment ( Ramesh et al., 2025 ) | Forecasting | Wildfire | WildfireGPT (GPT-4), TabNet (2) | Grounded | MAE 14.849 and 0 against MAE 0.055 |
| FireScope ( Markov et al., 2026 ) | Forecasting | Wildfire | Qwen2.5-VL-7B, GPT-5 (3) | Grounded | Out-of-distribution transfer from US to Europe |
| Retrieval over Response ( Cheng et al., 2026 ) | Forecasting | Wildfire | Not listed | Both | direct against guided |
| IC-EO ( Lahouel et al., 2026 ) | Geospatial | Wildfire | IC-EO agent, GPT-4o, LLaVA (3) | Both | 50% against 0% for GPT-4o and LLaVA |
| Geospatial-aware agents ( Chen et al., 2025a ) | Operations | Wildfire | About 14 LLM configurations | Both | Text-only over-allocated personnel 2–3 times |
| Category | Definition | Example Re-Checked Fire Works |
|---|---|---|
| knowledge-qa | Fire science, engineering, codes, or safety question answering. | ( Chen and Zhou, 2025 ; Sun et al., 2026 ) |
| document-understanding | Incident reports, after-action reviews, records, and extraction. | ( Epstein and Seielstad, 2025 ; Grünert et al., 2026 ; Kim and Shin, 2025 ; Tao et al., 2026 ; Chen et al., 2026 ) |
| detection-perception | Fire or smoke recognition from imagery or video by VLMs or MLLMs. | ( Seidel et al., 2025 ; Qi et al., 2026 ; Kohli, 2025 ; Habibpour et al., 2026b ; Ayanzadeh et al., 2026 ; Zhang et al., 2026 ; Al-Mohannadi et al., 2026 ) |
| forecasting-prediction | Spread, risk, danger, occurrence, or behavior prediction by LLMs or agents. | ( Ramesh et al., 2025 ; Cheng et al., 2026 ; Markov et al., 2026 ) |
| geospatial-analysis | GIS reasoning, spatial computation, map or remote-sensing interpretation. | ( Lahouel et al., 2026 ) |
| decision-support-operations | Resource allocation, dispatch, evacuation, and incident command. | ( Chen et al., 2025a ; Open Geospatial Consortium, 2026 ) |
| Category | What the records show | What they do not settle | |
|---|---|---|---|
| Detection and perception | 7 | Binary fire recognition scores high; early-stage smoke and quantitative estimates do not | The datasets fix camera set, prevalence, and smoke stage; no study shares a model set with another |
| Document understanding | 5 | Models sort incident records, command transcripts, and investigation cases at working scale | None of the five reports an error rate validated against human agreement |
| Operations and allocation | 3 | Geospatial grounding improved personnel and cost allocation | One study against baselines its authors chose; filed staffing is a proxy for decision quality |
| Forecasting and prediction | 3 | The one direct comparison put an LLM agent below tabular deep learning on fire radiative power | A single comparison on a GPT-4-era model; no ranking against other categories follows |
| Simulation-coupled agents | 2 | Coordination held on small suppression tasks and broke down on large ones | One base model; the second record is a separate result rather than a replication |
| Knowledge question answering | 2 | Rule-first and retrieval systems removed hallucination on numerical thresholds | Both are structure-fire code tasks, each against baselines its authors chose |
| Survey pattern (Section 2 ) | Design decision | Result (Section 4 ) |
|---|---|---|
| Grounding helped in each re-checked study that tried it | Same models bare and grounded on every task, with an information-only comparator per source | Where an addition carried the answer and reached the model, every resolved difference was a gain; additions lacking it moved scores in both directions, and surveyed systems are not replicated (Section 4 ) |
| Recognition high; early smoke and quantities weak | Smoke frames scored by minutes since the plume | Smoke accuracy is 0.25 to 0.57 in the first ten minutes (Section 4.1 ). On aerial question answering the thermal block moved pooled accuracy by at most 0.015, and no core model’s paired interval against the 0.627 majority comparator lies wholly above zero (Section 4.1 ) |
| One direct forecasting comparison favored a tabular model | Trivial rules and published reference points beside every model | Four of twelve prompted configurations beat a temperature rule, and trained tabular classifiers beat all twelve (Section 4.1 ) |
| Evaluations hard to reconstruct; two records test models behind the frontier | Access, reference answers, and leakage audited; every response stored; frontier models of the run date | Three hazards in public releases and one in our harness (Section 3.5 ); 11 of 34 candidate sources not feasible |
| Model | Condition | MAE | Norm. | Beats | Stable | Moving | False | Missed | Dir. |
|---|---|---|---|---|---|---|---|---|---|
| base | days | days | move | move | right | ||||
| Reference points | |||||||||
| Persistence baseline | 13.76 | 0.146 | 0.027 | 0.337 | 0.00 | 1.00 | |||
| Analogue-only rule | 19.88 | 0.249 | 0.23 | ||||||
| Trained ratio regressor | 12.97 | 0.144 | 0.48 | 0.035 | 0.317 | 0.03 | 0.93 | 0.72 | |
| Proprietary models | |||||||||
| Model | Condition | Items | Accuracy | Recall on smoke | FPR | Sequences detected |
|---|---|---|---|---|---|---|
| Reference points | ||||||
| Constant: always smoke | 196 | 0.571 | 1.000 | 1.000 | 28 of 28 | |
| Constant: always clear | 196 | 0.429 | 0.000 | 0.000 | 0 of 28 | |
| Frame difference, held out | 196 | 0.770 | 0.920 | 0.429 | 28 of 28 | |
| Proprietary models | ||||||
| claude-opus-4.8 | bare | 196 | 0.745 | 0.554 | 0.000 | 23 of 28 |
| Model | Condition | AUPRC | F1 on fire | Call rate | Omitted |
| Reference points | |||||
| Calendar-month prior, training years | 0.555 | ||||
| Last-day maximum 2 m temperature | 0.654 | ||||
| Published baselines, full holdout | 0.853 to 0.858 | ||||
| Trained boosted classifier, prompt numbers | 0.851 | 0.783 | 0.316 | 0 | |
| Trained logistic regression, prompt numbers | 0.768 | 0.697 | 0.285 | 0 | |
| Model | Arm | Accuracy | Single lookup | Filtered aggregate | Multi-step | Abstained | Calls | Tool bare |
| Reference point | ||||||||
| Best constant per family | 0.141 | 0.077 | 0.077 | 0.231 | ||||
| Proprietary models | ||||||||
| claude-opus-4.8 | bare | 0.103 | 0.000 | 0.000 | 0.246 | 99 | 0.00 | |
| claude-opus-4.8 | tool | 1.000 | 1.000 | 1.000 | 1.000 | 0 | 1.13 | 0.897 [0.756, 1.000] |
| claude-opus-5 | bare | 0.160 | 0.000 | 0.038 | 0.338 | 35 | 0.00 | |
| Model | Original | Phrasing | De-schem. | Phrasing orig. | De-schem. orig. |
|---|---|---|---|---|---|
| Proprietary models | |||||
| claude-opus-4.8 | 1.000 | 1.000 | 1.000 | +0.000 | +0.000 |
| claude-opus-5 | 1.000 | 1.000 | 1.000 | +0.000 | +0.000 |
| gemini-3.1-pro | 1.000 | 1.000 | 1.000 | +0.000 | +0.000 |
| gpt-6-astra | 0.994 | 0.994 | 0.994 | +0.000 | +0.000 |
| Open-weight models | |||||
| Allocation nMAE | Smoke recall | Fire danger AUPRC | Tool use acc. | Aerial acc. | ||||||
| Model | bare | grd. | bare | grd. | bare | grd. | bare | tool | bare | grd. |
| Core set (Section 3.3 ) | ||||||||||
| claude-opus-4.8 | 0.157 | 0.170 | 0.554 | 0.616 | 0.613 | 0.581 | 0.103 | 1.000 | 0.642 | 0.642 |
| claude-opus-5 | 0.165 | 0.173 | 0.607 | 0.643 | 0.688 | 0.682 | 0.160 | 1.000 | 0.650 | 0.657 |
| gemini-3.1-pro | 0.179 | 0.177 | 0.545 | 0.732 | 0.697 | 0.682 | 0.160 | 1.000 | 0.632 | 0.647 |
| gpt-6-astra | 0.161 | 0.220 | 0.643 | 0.741 | 0.620 | 0.590 | 0.000 | 0.994 | 0.623 | 0.635 |
| Work | Surv. | Tasks | LLMs | Fam. | B/G | Non-LLM | Scale |
| Fire AI reports with no model evaluation | |||||||
| OGC D-123 ( Tricomi, 2024 ) | – | – | – | – | – | – | NIMS/ICS phases |
| OGC 25-012 ( Open Geospatial Consortium, 2026 ) | – | – | – | – | – | – | 200 datasets |
| Fire benchmarks and datasets that evaluate no LLM or agent | |||||||
| WILDFIRE-FM ( Xu et al., 2026 ) | – | 1 | – | – | – | ✓ | 6 task forms, 11 FMs |
| Land8Fire ( Tran et al., 2025 ) | ✓ | 1 | – | – | – | ✓ | 4 nets, 55 GB |
| Clear frames | Smoke frames | ||||||
|---|---|---|---|---|---|---|---|
| Gap (min) | 10 | 20 | 30 | 45 | 55 | 65 | 79 |
| (a) Changed verdicts by ladder position | |||||||
| claude-opus-4.8 | 0/0 | 0/0 | 0/0 | 1/0 | 2/0 | 3/0 | 1/0 |
| claude-opus-5 | 1/0 | 1/0 | 1/1 | 0/0 | 1/0 | 2/0 | 1/0 |
| gemini-3.1-pro | 1/0 | 0/0 | 1/0 | 7/0 | 4/0 | 4/0 | 6/0 |
| gpt-6-astra | 1/0 | 0/1 | 2/0 | 3/0 | 3/0 | 3/0 | 2/0 |
| Recall on 84 smoke frames | 84 clear | Both labels | |||||
|---|---|---|---|---|---|---|---|
| Model | Bare | Grounded | Change | Gained/lost | FP / | Bal. acc. change | First design |
| claude-opus-4.8 | 0.536 | 0.583 | +0.048 ▲ [0.004, 0.117] | 4/0 | 0/0 | +0.024 [0.002, 0.058] | +0.071 |
| claude-opus-5 | 0.607 | 0.667 | +0.060 ▲ [0.017, 0.112] | 5/0 | 5/0 | 0.000 [ , 0.048] | +0.036 |
| gemini-3.1-pro | 0.595 | 0.726 | +0.131 ▲ [0.063, 0.220] | 11/0 | 1/0 | +0.060 [0.020, 0.104] | +0.179 |
| gpt-6-astra | 0.655 | 0.774 | +0.119 ▲ [0.046, 0.227] | 10/0 | 2/2 | +0.060 [0.000, 0.134] | +0.107 |
| Qwen3-VL | 0.571 | 0.607 | +0.036 [ , 0.072] | 3/0 | 1/0 | +0.012 [ , 0.024] | +0.060 |
| Model | Bare acc. | Grounded acc. | Grd. Bare | Gained / lost | Closed-form 48 | Other 360 |
| Reference points | ||||||
| Per-question majority | 0.627 | 0.627 | 0.688 | 0.619 | ||
| Majority + closed-form | 0.664 | 1.000 | 0.619 | |||
| Closed-form rule | 1.000 | |||||
| Proprietary models | ||||||
| claude-opus-4.8 | 0.642 | 0.642 | +0.000 | 17 / 17 | 0.667 0.896 | 0.639 0.608 |
| 48 closed-form items | 347 remaining items | ||||
| Model | Gained / lost | Bare grounded | Gained / lost | ||
| Core set | |||||
| claude-opus-4.8 | 11 / 0 | 0.001 | 0.640 0.608 | 6 / 17 | 0.035 |
| claude-opus-5 | 13 / 0 | 0.001 | 0.654 0.622 | 4 / 15 | 0.019 |
| gemini-3.1-pro | 13 / 0 | 0.001 | 0.628 0.608 | 9 / 16 | 0.230 |
| gpt-6-astra | 7 / 0 | 0.016 | 0.617 0.608 | 7 / 10 | 0.629 |
| Label source | Applicability | ||||||
| All | Formula, telemetry, or detector | Same, without closed form | Model- verified | ||||
| Model | Arm | 408 (34) | 132 (11) | 84 (7) | 276 (23) | 284 (33) | 124 (25) |
| Held-out majority | 0.627 | 0.614 | 0.571 | 0.634 | 0.637 | 0.605 | |
| Proprietary models | |||||||
| claude-opus-4.8 | bare | 0.642 | 0.500 | 0.405 | 0.710 ▲ | 0.651 | 0.621 |
| claude-opus-4.8 | grounded | 0.642 | 0.583 | 0.405 | 0.670 | 0.676 | 0.565 |
| Model | Alloc. | Smoke | Aerial | Danger | Tool |
|---|---|---|---|---|---|
| claude-opus-4.8 | 0/0 | 0/0 | 0/0 | 0/0 | 0 |
| claude-opus-5 | 0/0 | 0/0 | 0/0 | 0/0 | 0 |
| gemini-3.1-pro | 431/355 | n/r/296 | 266/197 | 484/572 | 1,158 |
| gpt-6-astra | 33/31 | 0/0 | 30/21 | 55/45 | 21 |