GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks
Authors: Gabriel Diaz-Ireland, Diego Prieto-Herráez, Mario García Peces, Javier Velázquez, Benjamin Zaitchik, Devika Jain
Organizations: Universidad Católica de Ávila (UCAV) Ávila, Spain · Independent Researcher Madrid, Spain · Department of Earth and Planetary Sciences, Johns Hopkins University Baltimore, MD, USA · Center for Geographic Analysis, Harvard University Cambridge, MA, USA
Before tool-using LLM agents are deployed in environmental and geospatial workflows, teams need evidence that an agent reliably selects the right operations against real APIs. We introduce GeoNatureAgent (GNA), a framework for pre-production evaluation of tool-using agents: a fixed sixteen-tool geospatial interface published as a Model Context Protocol (MCP) server, so the agent under test is the only variable, scored against an identical tool layer, task suite, and deterministic scorer. Its flagship instance is a 103-task benchmark (a 93-task main suite across 18 categories plus a ten-task comparison expansion) evaluated against an open, self-hostable geospatial API serving three environmental indicators across Spain and Portugal. We evaluate nine LLMs under three temperature-1.0 seeds, reporting capability and per-case cost as orthogonal axes. (1) Claude Sonnet 4 achieves the highest capability (61.7% +/- 0.7% on all 103 tasks; 60.8% on the main suite), followed closely by DeepSeek V3.2 (57.9%), while no other model exceeds 53%; (2) the cost-accuracy Pareto frontier is mostly open-weight, with DeepSeek V3.2 offering 93% of Claude's capability at 11.3x lower list-price cost; (3) under strict all-checks scoring the best model sits 24-36 points below the 85-97% reported on general-purpose GIS benchmarks, whereas per-check partial credit for the top four models (86-90%) is comparable, so much of that gap reflects scoring strictness rather than task difficulty alone. The MCP server, evaluation harness, benchmark, and API are publicly available; swapping the tool executors and task suite instantiates an equivalent benchmark for any geospatial domain.
Figures & tables
GeoBenchX
GeoNatureAgent Benchmark
Tasks
202
103
Categories
5
18
Tools
24 primitives
12 principal ops
Tool abstraction
Low-level
High-level
Data source
Local files
Cloud API (COG + JSON)
Domain
General GIS
Environmental
Table 1 . Comparison of structured tool-calling geospatial benchmarks. Accuracies are not same-protocol: GeoBenchX uses an LLM judge, GNA strict all-checks scoring (best per-check partial credit: 89.6% ).
Figure 1 . GNA system architecture. The LLM agent interacts with a production API serving COG data; the eval harness logs all tool calls and scores responses. NL denotes natural language.
Tool
Purpose
lookup_province
Resolve province → boundary geometry
lookup_municipality
Resolve municipality (+ province hint)
analyze_area
Zonal statistics for indicator in AOI
analyze_multi_layer
Multi-indicator analysis in single call
compare_areas
Compare two areas on the same indicator
find_top_n
Rank provinces by indicator value
Table 2 . Principal agent tools. The first eight are domain-specific; the last four are adapted from GeoBenchX ( Krechetova and Kochedykov, 2025 ) . The agent additionally exposes four auxiliary tools not listed here (layer discovery and erosion statistics).
Category
Tasks
Tests
Tool selection
21
Correct tool choice
Cross-indicator
8
CO 2 × erosion × land cover synthesis
Deep dive
6
Multi-tool environmental profiling
Interpretation
7
Reasoning over analysis results
Error handling
6
Non-existent entities, invalid indicators
Habitat analysis
7
BigEarthNet V2 land cover (Portugal)
Table 3 . GeoNatureAgent Benchmark main-suite (v5) task categories ( 93 of the 103 tasks, 18 categories).
Check
Pass condition
expected_tools
Every expected tool was called (recall = 1.0). Extra tools do not cause failure. a
expected_actions
Every expected UI action was generated (recall = 1.0).
must_contain
Each required keyword appears in the answer (case-insensitive substring).
must_not_contain
Forbidden keyword does not appear in the answer.
numeric_accuracy
For each ground-truth entry, the label is found in the answer and the nearest percentage value is within tolerance. b
chart_generated
At least one chart URL was produced (only checked when generate_chart is an expected tool).
Table 4 . Scoring checks. Binary capability pass requires every applicable check except max_cost_usd to pass.
Model
Provider
Params
Access
DeepSeek V3.2
DeepSeek AI
671B MoE
Vertex MaaS
GLM-5
Zhipu AI
—
Vertex MaaS
Gemini 2.5 Pro
Google
—
Vertex native
Claude Sonnet 4
Anthropic
—
Anthropic API
GPT-4o
OpenAI
—
OpenRouter
Qwen3-235B
Alibaba
235B MoE
Vertex MaaS
Table 5 . Models evaluated. Six via Vertex AI (five MaaS plus Gemini native), two via OpenRouter, Claude via the Anthropic Messages API.
#
Model
All (103)
Main (93)
Cost
$/case
1
Claude Sonnet 4 †
61.7±0.7
60.8±0.8
\11.84$
0.127
2
DeepSeek V3.2
57.9±3.9
56.3±3.1
\1.05$
0.011
3
GLM-5
52.4±2.6
50.2±2.2
\3.58$
0.038
4
Gemini 2.5 Pro
49.2±3.9
48.0±3.3
\4.87$
0.052
5
Qwen3-235B
44.3±5.0
41.2±4.3
\0.89$
0.010
6
GPT-4o
43.7±2.6
41.6±2.7
\6.49$
0.070
Table 6 . Capability leaderboard (accuracy in %, combined rank): nine models on the combined 103-task suite and the 93-task v5 main suite (Section 6.5 ). Per-seed mean ± one sd, decoupled from cost (Section 5 ); cost is the mean per-seed main-suite run total, $/case the per-call mean.
Figure 2 . Cost-accuracy trade-off (bubble size ∝ total tokens). The Pareto frontier (dashed) runs Scout → Qwen3-235B → DeepSeek V3.2 → Claude Sonnet 4; three of the four are open-weight.
Model
Accuracy
gCO 2 /case
Claude Sonnet 4
60.8%
6.44
DeepSeek V3.2
56.3%
6.42
GLM-5
50.2%
5.61
Gemini 2.5 Pro
48.0%
4.92
GPT-4o
41.6%
3.56
Qwen3-235B
41.2%
4.99
Table 7 . Estimated energy and carbon per case (tokens ×0.40 Wh/1k; 400 gCO 2 /kWh). Order-of-magnitude estimate; footprint is token-driven.
Failure mode
% of failures
Tool missing (expected tool not called)
68.0
Keyword missing (answer lacks required term)
19.3
Wrong data (incorrect value)
9.7
Rounds exceeded
1.3
Forbidden keyword
1.3
Chart missing
0.3
Table 8 . Failure-mode distribution over all 1,429 failed runs (nine models).
Model
Cap.
Raw
Comp. cap/raw
103-task comb.
Qwen3-235B
73%
37%
78/22%
44.3%
GLM-5
73%
37%
83/33%
52.4%
DeepSeek V3.2
73%
27%
83/17%
57.9%
Claude Sonnet 4
73%
20%
83/0%
61.7%
GPT-4o
63%
50%
72/50%
43.7%
Gemini 2.5 Pro
60%
60%
67/67%
49.2%
Table 9 . v6 expansion (10 comparison-heavy tasks, 3 seeds) and the combined 103-task suite. Cap. is capability scoring (budget/round gates excused; the v5 leaderboard excuses cost only); raw is strict pass. 103-task comb. is capability accuracy over the full v5+v6 suite (per-seed mean), the inclusive headline metric for this paper.
Figure 3 . v6 comparison tasks. Left: pooled comparison accuracy vs. value gap, raw and capability scoring—flat in the gap, and the 45 pp control is no easier, so difficulty is not near-equal discrimination. Right: per-model comparison accuracy, capability vs. raw; the bar gap is round/budget exhaustion while composing repeated analyze_area calls.
Figure 4 . Accuracy by category across the nine models.
Environmental scientists spend disproportionate effort on data wrangling rather than analysis. New AI agents can be a helpful tool, but no benchmark exists to evaluate AI agents that automate environmental geospatial workflows through structured tool calling against real APIs. We introduce the GeoNatureAgent Benchmark, the first benchmark for environmental analysis agents that operate via structured tool calls to a production-style geospatial API. The benchmark comprises 93 tasks across 18 categories. Tasks are evaluated against an open, self-hostable geospatial API that serves three environmental indicators across Spain and Portugal via sixteen tools. We evaluate nine frontier and open-weight LLMs, reporting capability and per-case cost as orthogonal axes. Results manifest that (1) Claude Sonnet 4 achieves the highest capability at 60.8% +/- 0.8%, followed closely by DeepSeek V3.2 at 56.3% +/- 3.1%, while no other model exceeds 51%; (2) the cost-accuracy Pareto frontier is occupied mostly by open-weight models, with DeepSeek V3.2 offering 93% of Claude's capability at 11.6x lower cost; and (3) structured tool calling against a real API provides a more discriminative measure of real-world agent capability, with mean accuracies 25-35 percentage points below those reported on general-purpose GIS benchmarks.
Gabriel Diaz-Ireland, Diego Prieto-Herráez, Mario García Peces +2
Universidad Católica de Ávila (UCAV) Ávila, Spain · Johns Hopkins University Baltimore, MD, USA · Independent Researcher Madrid, Spain +1
In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo-related tasks. We leverage a careful selection of twelve publicly available datasets from diverse geo-related tasks and domains, and evaluate a set of LLMs on geo-spatial and temporal understanding using our benchmark. Our results show that reasoning and size have a strong impact on overall performance. GeoBenchLLM is publicly available at https://github.com/Rfr2003/GeoBenchLLM.
Rodrigo Ferreira Rodrigues, Karim Radouane, Jose G Moreno +1
Geographic Information System (GIS) professionals rely on multi-step spatial analysis workflows to support decision-making in urban planning, disaster response, and environmental monitoring. The process is tedious, time-consuming, and error-prone. While recent large language model (LLM) agents equipped with external tools have the potential to automate geospatial analysis, their ability to perform realistic GIS workflows remains largely unexplored. Existing GIS agent benchmarking datasets are mostly drawn from textbooks, tutorials, or LLM-generated seeds and remain limited in size and trajectory depth. More importantly, none provides ground truth outputs. They therefore rely on surrogate signals such as code similarity, trajectory matching, or LLM and VLM judges, which can conflate workflow resemblance with task correctness. To address this gap, we introduce GISAgentBench, a benchmark of 349 multi-step GIS tasks curated from GIS Stack Exchange and instantiated on real public data across six selected geographic areas of interest. Each task ships with an executable reference trajectory and an exact ground truth output file, enabling strict, deterministic, tolerance-aware output matching beyond LLM judging. Evaluations of six LLM models reveal that realistic GIS workflows remain challenging: the best agent completes only 32.7% of tasks under strict tolerance-aware scoring, although most models produce outputs that are close to the ground truth.
Abhinav Pothuri, Zhe Jiang, Zelin Xu +1
Department of Computer & Information Science & Engineering, University of Florida · Department of Geography, University of Florida