GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks
Authors: Gabriel Diaz-Ireland, Diego Prieto-Herráez, Mario García Peces, Javier Velázquez, Benjamin Zaitchik, Devika Jain
Organizations: Universidad Católica de Ávila (UCAV) Ávila, Spain · Independent Researcher Madrid, Spain · Department of Earth and Planetary Sciences, Johns Hopkins University Baltimore, MD, USA · Center for Geographic Analysis, Harvard University Cambridge, MA, USA
Before tool-using LLM agents are deployed in environmental and geospatial workflows, teams need evidence that an agent reliably selects the right operations against real APIs. We introduce GeoNatureAgent (GNA), a framework for pre-production evaluation of tool-using agents: a fixed sixteen-tool geospatial interface published as a Model Context Protocol (MCP) server, so the agent under test is the only variable, scored against an identical tool layer, task suite, and deterministic scorer. Its flagship instance is a 103-task benchmark (a 93-task main suite across 18 categories plus a ten-task comparison expansion) evaluated against an open, self-hostable geospatial API serving three environmental indicators across Spain and Portugal. We evaluate nine LLMs under three temperature-1.0 seeds, reporting capability and per-case cost as orthogonal axes. (1) Claude Sonnet 4 achieves the highest capability (61.7% +/- 0.7% on all 103 tasks; 60.8% on the main suite), followed closely by DeepSeek V3.2 (57.9%), while no other model exceeds 53%; (2) the cost-accuracy Pareto frontier is mostly open-weight, with DeepSeek V3.2 offering 93% of Claude's capability at 11.3x lower list-price cost; (3) under strict all-checks scoring the best model sits 24-36 points below the 85-97% reported on general-purpose GIS benchmarks, whereas per-check partial credit for the top four models (86-90%) is comparable, so much of that gap reflects scoring strictness rather than task difficulty alone. The MCP server, evaluation harness, benchmark, and API are publicly available; swapping the tool executors and task suite instantiates an equivalent benchmark for any geospatial domain.
Figures & tables
GeoBenchX
GeoNatureAgent Benchmark
Tasks
202
103
Categories
5
18
Tools
24 primitives
12 principal ops
Tool abstraction
Low-level
High-level
Data source
Local files
Cloud API (COG + JSON)
Domain
General GIS
Environmental
Table 1 . Comparison of structured tool-calling geospatial benchmarks. Accuracies are not same-protocol: GeoBenchX uses an LLM judge, GNA strict all-checks scoring (best per-check partial credit: 89.6% ).
Figure 1 . GNA system architecture. The LLM agent interacts with a production API serving COG data; the eval harness logs all tool calls and scores responses. NL denotes natural language.
Tool
Purpose
lookup_province
Resolve province → boundary geometry
lookup_municipality
Resolve municipality (+ province hint)
analyze_area
Zonal statistics for indicator in AOI
analyze_multi_layer
Multi-indicator analysis in single call
compare_areas
Compare two areas on the same indicator
find_top_n
Rank provinces by indicator value
Table 2 . Principal agent tools. The first eight are domain-specific; the last four are adapted from GeoBenchX ( Krechetova and Kochedykov, 2025 ) . The agent additionally exposes four auxiliary tools not listed here (layer discovery and erosion statistics).
Category
Tasks
Tests
Tool selection
21
Correct tool choice
Cross-indicator
8
CO 2 × erosion × land cover synthesis
Deep dive
6
Multi-tool environmental profiling
Interpretation
7
Reasoning over analysis results
Error handling
6
Non-existent entities, invalid indicators
Habitat analysis
7
BigEarthNet V2 land cover (Portugal)
Table 3 . GeoNatureAgent Benchmark main-suite (v5) task categories ( 93 of the 103 tasks, 18 categories).
Check
Pass condition
expected_tools
Every expected tool was called (recall = 1.0). Extra tools do not cause failure. a
expected_actions
Every expected UI action was generated (recall = 1.0).
must_contain
Each required keyword appears in the answer (case-insensitive substring).
must_not_contain
Forbidden keyword does not appear in the answer.
numeric_accuracy
For each ground-truth entry, the label is found in the answer and the nearest percentage value is within tolerance. b
chart_generated
At least one chart URL was produced (only checked when generate_chart is an expected tool).
Table 4 . Scoring checks. Binary capability pass requires every applicable check except max_cost_usd to pass.
Model
Provider
Params
Access
DeepSeek V3.2
DeepSeek AI
671B MoE
Vertex MaaS
GLM-5
Zhipu AI
—
Vertex MaaS
Gemini 2.5 Pro
Google
—
Vertex native
Claude Sonnet 4
Anthropic
—
Anthropic API
GPT-4o
OpenAI
—
OpenRouter
Qwen3-235B
Alibaba
235B MoE
Vertex MaaS
Table 5 . Models evaluated. Six via Vertex AI (five MaaS plus Gemini native), two via OpenRouter, Claude via the Anthropic Messages API.
#
Model
All (103)
Main (93)
Cost
$/case
1
Claude Sonnet 4 †
61.7±0.7
60.8±0.8
\11.84$
0.127
2
DeepSeek V3.2
57.9±3.9
56.3±3.1
\1.05$
0.011
3
GLM-5
52.4±2.6
50.2±2.2
\3.58$
0.038
4
Gemini 2.5 Pro
49.2±3.9
48.0±3.3
\4.87$
0.052
5
Qwen3-235B
44.3±5.0
41.2±4.3
\0.89$
0.010
6
GPT-4o
43.7±2.6
41.6±2.7
\6.49$
0.070
Table 6 . Capability leaderboard (accuracy in %, combined rank): nine models on the combined 103-task suite and the 93-task v5 main suite (Section 6.5 ). Per-seed mean ± one sd, decoupled from cost (Section 5 ); cost is the mean per-seed main-suite run total, $/case the per-call mean.
Figure 2 . Cost-accuracy trade-off (bubble size ∝ total tokens). The Pareto frontier (dashed) runs Scout → Qwen3-235B → DeepSeek V3.2 → Claude Sonnet 4; three of the four are open-weight.
Model
Accuracy
gCO 2 /case
Claude Sonnet 4
60.8%
6.44
DeepSeek V3.2
56.3%
6.42
GLM-5
50.2%
5.61
Gemini 2.5 Pro
48.0%
4.92
GPT-4o
41.6%
3.56
Qwen3-235B
41.2%
4.99
Table 7 . Estimated energy and carbon per case (tokens ×0.40 Wh/1k; 400 gCO 2 /kWh). Order-of-magnitude estimate; footprint is token-driven.
Failure mode
% of failures
Tool missing (expected tool not called)
68.0
Keyword missing (answer lacks required term)
19.3
Wrong data (incorrect value)
9.7
Rounds exceeded
1.3
Forbidden keyword
1.3
Chart missing
0.3
Table 8 . Failure-mode distribution over all 1,429 failed runs (nine models).
Model
Cap.
Raw
Comp. cap/raw
103-task comb.
Qwen3-235B
73%
37%
78/22%
44.3%
GLM-5
73%
37%
83/33%
52.4%
DeepSeek V3.2
73%
27%
83/17%
57.9%
Claude Sonnet 4
73%
20%
83/0%
61.7%
GPT-4o
63%
50%
72/50%
43.7%
Gemini 2.5 Pro
60%
60%
67/67%
49.2%
Table 9 . v6 expansion (10 comparison-heavy tasks, 3 seeds) and the combined 103-task suite. Cap. is capability scoring (budget/round gates excused; the v5 leaderboard excuses cost only); raw is strict pass. 103-task comb. is capability accuracy over the full v5+v6 suite (per-seed mean), the inclusive headline metric for this paper.
Figure 3 . v6 comparison tasks. Left: pooled comparison accuracy vs. value gap, raw and capability scoring—flat in the gap, and the 45 pp control is no easier, so difficulty is not near-equal discrimination. Right: per-model comparison accuracy, capability vs. raw; the bar gap is round/budget exhaustion while composing repeated analyze_area calls.
Figure 4 . Accuracy by category across the nine models.