WeaveAgent: A Two-Stage Tool-Routing Agent for Ultra-High-Resolution Remote Sensing Imagery
Organizations: Department of Electronics, National University of Defense Technology, Changsha, China
Abstract
Problem. Ultra-high-resolution (UHR) remote sensing with vague user intents has two bottlenecks: visual tokens are expensive, and tool calling must be format-reliable (pretrained models emit zero tool calls zero-shot). Method. WeaveAgent, a two-stage tool-routing agent, decouples routing from visual perception. Stage A is routing-first: emission is trained, not elicited. Stage B executes conditionally: intrinsic queries enter visual answering (full-scene thumbnail; a WeaveEarth-style evidence board as an optional fixed-budget, approx. 5k-token compression interface); extrinsic queries execute tool call on original full-resolution imagery, answering from tool observations in a second, observation-masked round. Training: alignment SFT, then GRPO under reward R_WA2. Results. Alignment SFT lifts extrinsic routing from 0% to 80.75% (323/400); GRPO suppresses 9 intrinsic mis-emissions while tool selection is unchanged. The trained 2B system does not beat the zero-shot 8B baseline overall (0.263 vs. 0.250), a diagnostic contribution. Oracle attribution separates two repair ingredients: loading the observation into context lifts extrinsic answer accuracy from 0.025 to 0.425 under marker-free cross-mode returns, and the two-turn SFT stage adds a further +9.3 points to 0.518 at a small routing cost. A +/- image ablation shows emission suppression is visually grounded, and a query-register matrix shows LLM-rewritten queries cost trained checkpoints 2-11 points. Scope. All training and evaluation use the 5,000 / 3,273 / 1,000-record VagueUHR corpus (600 intrinsic + 400 tool-requiring; the base seeds synthesis and is not used for optimization). Single-pass evidence construction runs at 7.31 s per image on an RTX 4090. Code, data, and evaluation protocols will be released.
Figures & tables
| Policy (on extrinsic tasks) | ||
|---|---|---|
| Correct routing (right tool call) | 0.35 | 2.25 |
| Wrong tool call | 0.05 | 0.25 |
| Silent skip (no call, no answer) | 0.06 | 0.0 |
| Always answer directly: correct / wrong | 1.11 / 0.11 | 1.25 / 0.25 |
| Split | Records | Tool-traj. records | Tool pool | Task shapes |
|---|---|---|---|---|
| Train (template) | 5,000 | — (all intrinsic) | — | 5 (training-visible) |
| Routing-train ( train_tools ) | 3,273 | 873 (870 rendered tool-call demos; 3 skipped: gt trajectory has no real tool step; 0 render failures) | 2 visible: ref-seg 579 / skysense-det 291 | 2,400 intrinsic 873 tool |
| Test (template) | 1,000 | 400 | 3: skysense 138 sm3det 62 (unseen) ref-seg 200 | 10 (5 unseen in training); 600 intrinsic 400 tool |
| Test (LLM-rewritten) | 992 kept 8 removed (gt–image conflicts) 1,000 | 593 intrinsic 399 extrinsic (8 removed) | same | same |
| Model / training stage | Intr. | Extr. | T-sel. | Calls | Mean |
| Qwen3-VL-8B zero-shot (E3a) | 1.000 (600/600) | 0.000 (0/400) | 0.600 ∗ | 0 | 0.2630 |
| GRPO from 2B base (E3b) | 1.000 | 0.000 | 0.600 ∗ | 0 | 0.2420 |
| GRPO from 2B base (E3b) | 1.000 | 0.000 | 0.600 ∗ | 0 | 0.2570 |
| Base 2B, native tools= channel (zero-shot) | 0.015 (9/600) | 0.12 (48/400) | 0.057 | 986 | — |
| Aligned SFT (E13, 2B) | 0.970 | 0.8075 (323/400) | 0.905 | 398 | 0.2430 |
| GRPO (E15) | 0.985 | 0.8075 (323/400) | 0.914 | 384 | 0.2500 |
| Model | Family | Inst | Tool | ArgN | ArgV | Summ |
|---|---|---|---|---|---|---|
| Aligned SFT (E13) | detection | 0.905 | 0.685 | 0.994 | 0.909 | 0.035 |
| segmentation | 0.995 | 1.000 | 0.997 | 0.990 | 0.005 | |
| ALL | 0.950 | 0.850 | 0.996 | 0.951 | 0.020 | |
| GRPO (E15) | detection | 0.880 | 0.705 | 0.994 | 0.932 | 0.045 |
| segmentation | 0.995 | 1.000 | 0.997 | 0.990 | 0.005 | |
| ALL | 0.938 | 0.861 | 0.996 | 0.963 | 0.025 |
| Metric | GRPO from base (E7a) | SFT cold start GRPO (E12) | Aligned SFT GRPO (E15) |
|---|---|---|---|
| Initial reward | 0.058 | 0.55 (first-6: 0.5455) | 1.18 |
| frac_reward_zero_std | 0.85–0.9 | 0.43–0.78 | 0–0.1 early, for most of training (max 0.47) |
| Completions mean length (tokens) | 59–69 | 11–27 (converges to direct answers; the arm collapses to 8.5–9.7) | 165–180 (stable) |
| Benchmark | Thumbnail 2048 (interface off) | Thumbnail board ( ) | Paired |
|---|---|---|---|
| LRS-VQA (1,500 strat.) | 33.20 | 32.20 | |
| MME-RealWorld-RS (1,500 strat.) | 39.33 | 35.93 | |
| XLRS-Bench-lite (1,500 strat.) | 41.47 | 38.60 |
| Variant | GCC | MSES | SEM | TPEB | Accuracy |
|---|---|---|---|---|---|
| w/o GCC | – | ✓ | ✓ | ✓ | 0.3040 |
| w/o MSES | ✓ | – | ✓ | ✓ | 0.3140 |
| w/o SEM | ✓ | ✓ | – | ✓ | 0.3080 |
| w/o TPEB | ✓ | ✓ | ✓ | – | 0.3000 |
| WeaveAgent (full) | ✓ | ✓ | ✓ | ✓ | 0.2940 |
| 2 | 4 | 6 | 8 | 10 | |
|---|---|---|---|---|---|
| Accuracy | 0.3080 | 0.3040 | 0.2940 ∗ | 0.3100 | 0.3020 |
| Method | Hardware / protocol | s / sample | Relative |
|---|---|---|---|
| Thumbnail baseline 2048 (official script) | 4090, , greedy | 2.84 | 0.39 |
| Zoom baseline (ours, E6, multi-round) | 4090, | 4.33 | 0.59 |
| Single-pass evidence construction (ours, E1a, greedy) | 4090, | 7.31 | 1 |
| WeaveAgent E15 single-pass (VagueUHR-test) | 4090, | 6.17 | 0.84 |
| WeaveEarth (paper-reported) | A100 | 7.59 | 1.04 |
| ZoomSearch (paper-reported, multi-round search) | — | 54.25 | 7.4 |
| Clear | Template-vague | LLM-rewritten | LLM per-type | |||||
|---|---|---|---|---|---|---|---|---|
| Checkpoint | overall | extr | overall | extr | overall | extr | det | seg |
| 8B zero-shot | 0.268 | 0.000 | 0.263 | 0.000 | 0.296 | 0.0025 | 0.231 | 0.405 |
| E15 (aligned) | 0.241 | 0.800 | 0.250 | 0.8075 | 0.219 | 0.667 | 0.085 | 0.005 |
| E17 (two-turn) | 0.476 | 0.7875 | 0.488 | 0.800 | 0.379 | 0.659 | 0.382 | 0.235 |
| Split | Overall acc. | Intr. routing | Extr. routing | T-sel. | Calls |
|---|---|---|---|---|---|
| Template (992 paired) | 0.2651 | 1.000 | 0.000 | 0.598 ∗ | 0 |
| LLM-rewritten (992) | 0.2964 | 0.998 (592/593) | 0.0025 (1/399) | 0.598 ∗ | 2 |
| Arm | Detection | Segmentation | Overall | No-call |
|---|---|---|---|---|
| Strict (marker-free) | 0.784 (29/37) | 0.339 (21/62) | 0.505 | 1/100 |
| Embedded (answer-bearing) | 0.974 (37/38) | 0.758 (47/62) | 0.840 | 0/100 |
| Protocol | Overall | Class. | Count. | Reas. | Det. | Seg. | Extr. | Intr. | T-sel. | Calls |
|---|---|---|---|---|---|---|---|---|---|---|
| Eval A (strict fallback) | 0.488 | 0.560 | 0.330 | 0.515 | 0.635 | 0.400 | 0.800 | 0.913 | 0.868 | 433 |
| Eval B (embedded) | 0.667 | 0.560 | 0.330 | 0.515 | 0.960 | 0.970 | 0.800 | 0.913 | 0.868 | 433 |
| E13 loaded strict returns (control) | — | — | — | — | 0.510 | 0.335 | extr. ans. | |||
| E15 loaded strict returns (control) | — | — | — | — | 0.510 | 0.340 | extr. ans. | |||
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| # | Difference | Training prompt | Evaluation prompt |
|---|---|---|---|
| 1 | Opening sentence | “You are WeaveAgent answering…” | “Answer the question about…” |
| 2 | Dual-image description block | none | thumbnail evidence-board description |
| 3 | Category line | none | Category: <task_type> |
| 4 | Query prefix | User query: | Question: |
| 5 | Catalog position | before the query | after the query |
| 6 | Decision-instruction position | last line, mentions both options | mid-prompt, followed by 3 more blocks |
| Prompt family | Thumbnail | TPEB board | SEM block |
|---|---|---|---|
| Routing-train prompts (E13/E15) | path (lazily decoded) | none | none (sem_text 0/3,273) |
| Main evaluation prompts | present | present (real injection) | present |
| E19 text-only arm (E15 ckpt.) | image removed; description text retained | image removed; description text retained | retained as text (byte-identical prompt) |
| E17 turn-1 / turn-2 | present (E13 demo images) | present (evaluation-identical image format) | none (training side); turn-2 adds the observation text |
| ID | Configuration | Result |
|---|---|---|
| E1a | Single-pass evidence pipeline, interface on (thumbnail board, ); three benchmarks, stratified 1,500 each, 8B, greedy; timing anchor | Tables 6 and 9 |
| E1b | Evidence interface off (2048 thumbnail), same samples | Table 6 |
| E3a | Zero-shot Qwen3-VL-8B routing evaluation | Table 3 |
| E3b | GRPO from the 2B base, / arms (merged E7a/E7b weights) | Table 3 |
| E4 / E4_full | Component ablations with the same-subset full-system control ( , greedy) | Table 7 |
| E5 | Evidence-budget scan | Table 8 |
| Case | Answer indicator | Tool indicator | Format indicator | Total |
|---|---|---|---|---|
| Intrinsic, direct answer correct | 1.0 | 1.0 | 0.5 | 1.35 |
| Intrinsic, direct answer wrong | 0 | 1.0 | 0.5 | 0.35 |
| Intrinsic, emits T_call | 0 | 0.0 | 0.0 | 0.00 |
| Extrinsic, correct T_call ( SFT behavior) | 0 | 1.0 | 0.5 | 0.35 |
| Extrinsic, wrong T_call | 0 | 0.0 | 0.5 | 0.05 |
| Extrinsic, no call wrong answer | 0 | 0.2 | 0.5 | 0.11 |
| Measurement | Numbers | Source |
|---|---|---|
| WeaveEarth paper-reported (LRS-VQA full set) | 26.68 33.38 ( ) | paper |
| Released script, 300-record stratified paired, both arms greedy | 29.67 32.33 ( , 24/32 discordant, McNemar exact ) | this work, Section 5.1 |
| Category task-type line asymmetry (released baseline lacks it, treatment arm has it) | paired control (E20): points, McNemar — no answer leakage; the earlier -point cross-protocol inference is refuted | this work (E20) |
| MCQ option-instruction order (self-audit run, indicative) | 25.47 35.93 ( -point swing; 34% of outputs degrade to prose) | quarantined audit |
| Task type | Recall@8 anchors | Recall@6 final board | |
|---|---|---|---|
| Classification | 0.285 | 0.240 | 200 |
| Counting | 0.0 | 0.0 | 2 |
| Detection | 0.323 | 0.258 | 62 |
| Reasoning | 0.296 | 0.204 | 196 |
| Segmentation | 0.285 | 0.240 | 200 |
| All spatial records | 0.291 | 0.230 | 660 |
| Case | Query | Stage-A decision and T_call | Turn-2 tool return | Output vs. GT |
|---|---|---|---|---|
| C1 (det.) | “What is the amount of boat in the image?” ( 4498 ) | extrinsic: T_call(skysense_detection, classes=[object], confidence_threshold=0.5) | “Detection finished; 3 objects localized.” | 3 vs. 3 ✓ |
| C2 (seg.) | “hey, what is the shape of the top-most tank?” ( 0177 ) | extrinsic: T_call(ref-seg, prompt= query ) | “Segmentation mask generated.” | circular vs. circular ✓ |
| C3 (intr.) | “hey, what is the activity of the top-most plane?” ( 983 ) | intrinsic: no call; answers from thumbnail (+ evidence board) | — (single turn) | parked vs. parked ✓ |
| C4 (fail) | “What is the amount of bridge in the image, roughly?” ( 11069 ) | extrinsic: T_call(skysense_detection, classes=[bridge], confidence_threshold=0.5) | “Detection finished; 10 objects localized.” | 3 vs. 10 |