CrossTimeEdit: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation
Organizations: Sun Yat-sen University
Abstract
Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures while preserving persistent scene content. We construct VIGOR-his, a decade-spanning cross-view dataset containing 43,653 location-level quadruplets across 11 cities on three continents. Its automated pipeline performs spatial pairing, consistency screening, change classification, and the generation and validation of satellite-based change descriptions and local editing instructions. Based on VIGOR-his, we propose CrossTimeEdit, a model that reformulates historical street-view generation as editing, using recent street views to constrain viewpoint and unchanged appearance and temporal satellite differences as change evidence. Starting from FLUX.2 [Klein] 4B, we train CrossTimeEdit through supervised fine-tuning (SFT) followed by online reinforcement learning (RL). We design three street-view editing criteria, namely Instruction Alignment (IA), Background Preservation (BP), and Quality and Physical Plausibility (QP), as both RL reward dimensions and evaluation metrics. We optimize this multi-reward objective using Within Group Relative Policy Optimization for flow-matching models (Flow-GRPO) with Group reward-Decoupled Normalization Policy Optimization (GDPO), which normalizes each reward dimension before aggregation. CrossTimeEdit improves overall performance across the three editing criteria by 17.12% over the pretrained baseline and outperforms cross-view generation models in scene consistency, visual realism, and perceptual quality. The implementation code, dataset, and model weights are available at https://luhanwen67.github.io/CrossTimeEdit-release/.
Figures & tables
| Dataset | Cross-view | Temporal | Temporal | Cross-view | Text | Continents | Size |
| pairing | street views | satellite views | time alignment | modality | |||
| CVUSA | ✓ | ✗ | ✗ | ✗ | ✗ | 1 | 44,416 pairs |
| CVACT | ✓ | ✗ | ✗ | ✗ | ✗ | 1 | 128,334 pairs |
| VIGOR | ✓ | ✗ | ✗ | ✗ | ✗ | 1 | 105,214 panoramas |
| CVGlobal | ✓ | ✓ | ✗ | ✗ | ✗ | 6 | 134,233 pairs |
| CV-Cities | ✓ | ✗ | ✗ | ✗ | ✗ | 6 | 223,736 pairs |
| Method | IA ( ) | BP ( ) | QP ( ) | Overall ( ) |
| Baseline | 5.960 | 8.008 | 6.861 | 6.800 |
| w/o RL | 6.919 | 6.698 | 6.543 | 6.759 |
| w/o SFT | 6.085 | 8.112 | 6.886 | 6.893 |
| w/o GT | 7.938 | 7.356 | 7.508 | 7.656 |
| w/o DGN | 8.148 | 7.433 | 7.432 | 7.755 |
| Full | 7.950 | 8.264 | 7.631 | 7.964 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Phase 1 | Phase 2 |
| Backbone | FLUX.2 [Klein] 4B; bf16 weights | |
| Trainable modules | Attention-only LoRA; rank 32, alpha 32; text-stream projections included | |
| Frozen modules | Base transformer, MLPs, text encoder, and VAE | |
| Data | 8,797 change + 13,196 no-change; shuffled; one pass per epoch | |
| Epochs | 0–4 (5 epochs) | 5–9 (5 epochs) |
| Initialization | Pretrained backbone | Own epoch-4 LoRA checkpoint |
| Setting | Value |
| Initialization / reference | SFT epoch 9 / frozen SFT epoch 9 |
| Trainer / aggregation | Flow-GRPO; DGN, weighted aggregation, then batch normalization |
| Group sampling | 50 conditions per round; 8 candidates per condition; 2 updates per round |
| LoRA / precision | Rank 32, alpha 32; fp32 master weights; bf16 mixed precision |
| Optimizer | AdamW; ; weight decay ; ; |
| Gradient norm / RL seed | 1.0 / 42 |
| Model | Human win rate (%) | Human rank (11) | Editing Overall | Editing rank (8) | Cross-view Overall | Cross-view rank (4) |
| Qwen-Image-3.0-Pro | 80% | 1 | 8.042 | 3 | – | – |
| GPT-Image-2 | 72% | 2 | 8.581 | 1 | – | – |
| Nano Banana 2 Lite | 72% | 2 | 8.302 | 2 | – | – |
| CrossTimeEdit | 64% | 4 | 7.964 | 4 | 3.546 | 1 |
| FLUX.2 [Klein] 4B | 58% | 5 | 6.800 | 7 | – | – |
| LongCat-Image-Edit | 58% | 5 | 7.360 | 6 | – | – |