Urban environments are shaped by design choices with long-term implications for health, safety, and quality of life, yet evaluating proposed interventions remains costly, time-consuming, and often impractical. Existing geospatial vision methods largely focus on monitoring urban indicators from aerial and street-view imagery, rather than proposing interventions and estimating their effects on such indicators. Moving beyond recognition, we introduce the problem of discovering interventions that improve target indicators for a given aerial or street-view image. We argue that a black-box indicator model, combined with a generative editing model, can serve as an implicit digital twin for testing intervention hypotheses. We present VIDA-Geo , a multi-agent system that explores this intervention space by coordinating segmentation, diffusion-based inpainting, and indicator scoring models to produce interventions that are both perceptually realistic and aligned with real-world policies. We evaluate our system on 8 indicators across aerial and street-view imagery, measuring changes in factors such as perceived safety and greenery. Our approach outperforms existing baselines in many cases, achieving up to 2X higher perceptual quality and policy alignment scores. Finally, our model provides users with multiple candidate interventions, supporting an expert city-planner-in-the-loop workflow.
Figures & tables
Figure 1: Given a geospatial image (street view or aerial), a user can ask VIDA-Geo to suggest interventions that improve a target indicator, such as liveliness. Our system proposes multiple realistic, policy-grounded interventions to support decision-making. The user can then choose the intervention most suitable for their use case.
Figure 2: Overview of the VIDA-Geo pipeline. Given an image and a target urban-improvement instruction, a group of agents identifies ROIs, segments, and edits with task-specific constraints to produce candidate intervention edits. Quality checks assess realism and policy preservation. Candidates are finally evaluated by a black-box indicator scorer.
Task
Perceptual Quality
Policy Alignment
Method
Output Avg. (%)
Delta Avg. (%)
FID-Proxy ( ↓ ) ± 1.00
Visual Quality (%) ± 0.70
Realism (%) ± 0.87
Policy Pres. (%) ± 1.25
LLM Judge Avg. (%) ± 0.76
Greenery
VIDA-Geo
35.6
6.3
44.8
72.2
67.2
64.1
67.8
LANCE
25.3
-4.1
57.3
37.5
28.6
16.9
27.7
DIFFusion
59.5
30.1
59.6
26.1
25.5
26.0
25.9
NB2.5 (ZS)
43.0
13.6
44.7
71.9
66.5
66.1
68.2
Road Risk
VIDA-Geo
82.1
4.0
47.5
72.8
67.4
57.8
66.0
Table 1: Task performance, perceptual quality, and policy alignment for edited images judged with GPT5.2 . Green and red tasks favor higher and lower task scores respectively. Delta Avg. is signed toward the task objective. FID-Proxy and judge-score columns report average 95% confidence-intervals . The best judge score for each task and evaluation is boldfaced.
Figure 3: Comparison of intervention outputs from VIDA-Geo vs. DIFFusion for street view (first 2 rows) and aerial inputs (last row). The interventions provided by VIDA-Geo are more practical and realizable in the real world, e.g., adding bollards and tactile pavements around the intersection in the second. Moreover, VIDA-Geo can provide text instructions along with the images, unlike DIFFusion, which is more useful to a user.
Figure 4: Using VIDA-Geo with multiple sequential edits. (a) It suggests multiple edits in different regions and increases the indicator score. For example, in the bottom row, it improves the building facade, then adds a lamppost, followed by roadside flowers, increasing the beauty score from 29.2% to 92.4%. (b) Quantitative progression over successive iterations.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Examples of failure cases. (a) shows policy violations that our system steers away from. (b) shows segmentation masks rejected by the QC sub-agent, with the top image segmenting "medium rise building" and the bottom "grass shrubs". Whereas subfigure (c) illustrates rejected edits, the top edit includes a low-res and unrealistic edit, and the bottom edit inpainting the wrong region. Fallbacks and fixes are explained in appendix A .
Figure 6: Qualitative examples from the Segmentation and Generation Agents outputs, indicating originally failed outputs, that were later corrected to pass the QC.
Evaluation
VIDA-Geo vs. DIFFusion
VIDA-Geo vs. LANCE
Q1 – Metric Preference
123/210 (58.6%)
189/210 (90.0%)
↪ Lively
28/35 (80.0%)
34/35 (97.1%)
↪ Beautiful
25/35 (71.4%)
34/35 (97.1%)
↪ Less Boring
25/35 (71.4%)
29/35 (82.9%)
↪ Less Depressing
18/35 (51.4%)
32/35 (91.4%)
↪ Safe
17/35 (48.6%)
34/35 (97.1%)
Appendix
Table 2: Graduate Student preference for VIDA-Geo over DIFFusion and LANCE. Bold rows report the overall result for each evaluation question. For Q1, we additionally report preference by targeted perceptual attribute. Each entry gives the number and percentage of judgments favoring VIDA-Geo .
Evaluation
VIDA-Geo vs. DIFFusion
VIDA-Geo vs. LANCE
Q1 – Metric Preference
92/150 (61.3%)
130/150 (86.7%)
↪ Lively
21/25 (84.0%)
22/25 (88.0%)
↪ Beautiful
18/25 (72.0%)
24/25 (96.0%)
↪ Less Boring
15/25 (60.0%)
19/25 (76.0%)
↪ Less Depressing
10/25 (40.0%)
23/25 (92.0%)
↪ Safe
15/25 (60.0%)
24/25 (96.0%)
Appendix
Table 3: Expert preference for VIDA-Geo over DIFFusion and LANCE. Bold rows report the overall result for each evaluation question. For Q1, we additionally report preference by targeted perceptual attribute. Each entry gives the number and percentage of judgments favoring VIDA-Geo .
Judge
Human Group
Q1 Metric
Q2 Realism
Q3 Planning
VIDA-Geo vs. DIFFusion
Qwen3-VL
Students
46.4%
93.1%
89.3%
Experts
50.0%
96.6%
92.9%
GPT-5.2
Students
48.3%
92.9%
96.6%
Experts
51.7%
96.4%
100.0%
VIDA-Geo vs. LANCE
Appendix
Table 4: Agreement between LLM judges and human majority preferences. For each image pair and evaluation question, agreement indicates whether the LLM judge selected the same preferred output as the corresponding human majority. Human ties are excluded.
Figure 7: Sequential editing behavior of VIDA-Geo across six Street View perception tasks. Each panel reports the mean goal-aligned perception score together with Qwen3-VL realism, policy-preservation, visual-quality, and average LLM-Judge scores over sequential edits of 10 inputs. For Boring and Depressing , the raw perception score is inverted ( 10−raw score ) so that higher values consistently indicate progress toward the target. Editing stops for an input when no acceptable candidate passes QC or when the target is reached; stopped inputs retain their last accepted result in subsequent cohort averages.
Figure 8: Counterfactual scene edits and corresponding predictor outputs. VIDA-Geo generates alternative versions of the same street-view image through two distinct editing directions. As vegetation is removed and the environment becomes less maintained, the black-box "beauty" predictor assigns progressively lower scores. The consistent and interpretable score changes suggest that the predictor captures meaningful scene attributes under counterfactual interventions, supporting its use as a component of a world model.
Metric
Successful
No Improving Edit
Failure per 100
Safety
96
4
4%
Lively
100
0
0%
Beautiful
98
2
2%
Wealthy
99
1
1%
Boring
99
1
1%
Depressing
98
2
2%
Appendix
Table 5: Edit success and failure counts by metric. Each metric contains 100 evaluated inputs.
Task
Perceptual Quality
Policy Alignment
Method
Output Edit Avg. (%)
Delta Avg. (%)
FID-Proxy ( ↓ ) ± 1.97
Visual Quality (%) ± 2.48
Realism (%) ± 2.73
Policy Pres. (%) ± 3.45
LLM Judge Avg. (%) ± 2.42
Greenery
VIDA-Geo
44.6
15.3
45.6
68.3
59.2
46.1
57.9
LANCE
36.5
7.1
57.5
37.3
27.4
14.6
26.4
DIFFusion
77.8
48.5
58.3
21.8
22.1
27.2
23.7
NB2.5 (ZS)
56.3
26.9
47.7
65.0
57.3
51.0
57.8
Road Risk
VIDA-Geo
70.6
15.5
49.6
69.7
61.5
45.8
59.0
Appendix
Table 6: Human-aligned quality, intervention strength, and distributional realism proxy for the Best edited satellite and Street View images across urban intervention tasks, evaluated by GPT-5.2 . Tasks shown in green correspond to indicators where higher scores are desirable, while tasks shown in red correspond to indicators where lower scores are desirable. Bold judge scores indicate that the best-performing method is statistically separated from all alternatives by non-overlapping 95% confidence intervals.
Task
Method
Visual Quality (%) ± 0.67
Realism (%) ± 0.86
Policy Pres. (%) ± 1.15
LLM Judge Avg. (%) ± 0.76
Greenery
VIDA-Geo
80.0
68.7
69.1
72.6
LANCE
41.5
37.3
23.8
34.2
DIFFusion
28.1
25.2
30.0
27.7
NB2.5 (ZS)
78.1
70.2
77.7
75.3
Road Risk
VIDA-Geo
80.5
71.5
65.6
72.6
LANCE
45.1
43.4
26.9
38.5
Appendix
Table 7: Perceptual quality and policy alignment for edited satellite and Street View images. Green tasks correspond to indicators where higher scores are desirable, while red tasks correspond to indicators where lower scores are desirable. Visual Quality, Realism, and Policy Preservation are evaluated using Qwen3-VL 32B , and LLM Judge Avg. is their mean. All columns report average 95% confidence-interval half-widths. Bold values indicate the best-performing method for each task and metric.
Task
Method
Perceptual Quality
Policy Alignment
LLM Judge
FID-Proxy ↓± 1.97
Visual Quality (%) ± 2.77
Realism (%) ± 3.01
Policy Pres. (%) ± 3.19
Avg. (%) ± 2.50
Greenery
VIDA-Geo
45.6
75.6
59.0
51.5
62.0
LANCE
57.5
40.3
36.4
22.4
33.0
DIFFusion
58.3
22.3
22.3
30.7
25.1
NB2.5 (ZS)
47.7
68.0
58.4
70.3
65.6
Road Risk
VIDA-Geo
49.6
77.7
66.1
54.2
66.0
Appendix
Table 8: Human-aligned quality and distributional realism proxy for the Best edited satellite and Street View images across urban intervention tasks, evaluated by Qwen3-VL 32B . Tasks shown in green correspond to indicators where higher scores are desirable, while tasks shown in red correspond to indicators where lower scores are desirable. FID-Proxy and judge-score columns report average 95% confidence-interval half-widths. Bold judge scores indicate that the best-performing method is statistically separated from all alternatives by non-overlapping 95% confidence intervals.
Indicator
Median Score
Score Range
Safety
13.0
3.0–49.0
Lively
11.0
2.0–43.0
Beautiful
15.0
2.0–47.0
Wealthy
14.0
2.3–48.0
Boring
83.0
54.0–98.0
Depressing
83.0
55.0–98.0
Appendix
Table 9: Summary statistics of indicator scores.
Task
Input
Seg. Mask
Gen. Prompt
Gen. Edit
Score
Greenery
Seg: LISAt
“Introduce a dense mixed-species tree canopy with distinct rounded crown shapes and natural shadows.”
Gen: FLUX
+ 7.1%
Seg: SAM3
“Convert the masked roadside area into a linear park with a dense tree canopy, walking paths, and seating areas.”
Gen: NanoBanana
+ 5.5%
Road Risk
Seg: SAM
“Transform the masked bridge into a pedestrian and cycle bridge with barriers separating it from vehicle lanes.”
Gen: NanoBanana
− 9.0%
Appendix
Table 11: Qualitative examples of counterfactual edits and indicator changes.
Figure 9: Qualitative comparison of VIDA-Geo interventions with real-world changes. From left to right: the original scene, the observed real-world intervention, and the intervention generated by VIDA-Geo . In the top example, VIDA-Geo similarly improves the pedestrian crossing, sidewalk, and theater marquee; in the bottom example, it introduces a bike lane similar to the one observed in the real-world intervention.
Configuration
Edit QC
Policy
Mask QC
Suggestor
FID-Proxy ↓
Vis. Qual. (%)
Realism (%)
Policy Pres. (%)
Judge Avg. (%)
Full Pipeline ( VIDA-Geo )
✓
✓
✓
✓
54.13
71.9
75.6
85.5
77.7
w/o Edit QC
✗
✓
✓
✓
57.05
65.0
68.8
67.5
67.1
w/o Policy Restrictor
✓
✗
✓
✓
56.09
68.0
64.6
62.5
65.0
w/o Mask QC
✓
✓
✗
✓
56.79
62.5
60.0
70.1
64.2
w/o Suggestor
✓
✓
✓
✗
55.87
60.0
65.5
80.0
68.5
Appendix
Table 12: Ablation study results. Each row disables one pipeline module (marked with ✗) while keeping all others active (✓). Visual Quality , Realism , and Policy Preservation are scored by an LLM judge; Judge Avg. is their mean. FID-Proxy measures domain typicality (lower is better).
Figure 10: Instructions given to the human evaluators. The instructions were intentionally kept simple to prevent users from biasing themselves to select our method.
Figure 11: User interface with questions shown to the human annotators. They are randomly shown the best-performing DIFFusion intervention and the best VIDA-Geo intervention, and asked three questions about the indicator, perceptual quality, and policy alignment.
School of Architecture, Tsinghua University, Beijing, China · Department of Geography, University College London, London, United Kingdom · School of Engineering, Cardiff University, Cardiff, United Kingdom