cs.CVAug 17, 2026

Remote-Sensing City Layout Extraction with MLLM

Authors: Zigan Zhou, Kai Li, Yupeng Deng

Abstract

Remote-sensing systems usually describe urban content with detection boxes, semantic masks, or vector boundaries. Such outputs locate classes and support image-plane scoring, yet they do not by themselves constitute an executable layout that retains object identities, typed relations, topology, and regeneration rules. Code-as-City instead casts urban-layout extraction from a single top-down image as constrained code generation with a multimodal large language model (MLLM). An image model first produces an aligned five-class semantic layout prior. Three ordered MLLM passes use the image and this prior to recover roads, land-cover regions and relations, and buildings. Deterministic normalization converts the accumulated records into a city graph and a restricted layout program. Executing the program creates a renderable 3D city layout and an orthographic semantic projection over shared geometry. The projection admits pixel-level comparison with remote-sensing masks, while named objects, relations, and editing operations remain available for synchronized regeneration of both views. Evaluated on the 100 scenes of CityLayout-100, the complete framework obtains 41.1% mean intersection-over-union and 48.3% global intersection-over-union. This result provides quantitative evidence that visual observations can be translated into inspectable, editable city code with coupled planar and 3D outputs.

Explore similar work

May 8, 2026cs.CL

Built Environment Reasoning from Remote Sensing Imagery Using Large Vision--Language Models

This work investigates the use of large language models (LLMs) for tasks in smart cities. The core idea is to leverage remote sensing imagery to characterize the built environment, including design suggestions, constructability assessment, landuse patterns, and risk identification. We examine remote sensing imagery at multiple spatial scales as inputs for multimodal language modeling and evaluate their effects on built-environment-related reasoning. In addition, we compare state-of-the-art LLMs, including InternVL and Qwen, in terms of accuracy and reliability when generating built environment recommendations. The results demonstrate the potential of integrating remote sensing imagery with large language models to assist smart cities and decision-making.
Jul 16, 2026cs.CL

CityLLM: A framework for natural-language querying of semantic 3D city models

Semantic 3D city models provide rich geometric and semantic information, but remain challenging for non-experts and interdisciplinary researchers to access and query due to their complex structures and specialized data formats. To address this issue, we present CityLLM, a framework for natural-language querying of semantic 3D city models alongside complementary urban datasets. The framework combines spatial and graph databases within an LLM-based workflow that supports iterative query refinement and cross-database chaining. We evaluate CityLLM on a CityJSON dataset of Rotterdam (853 LoD2 buildings) using GPT-OSS, Gemini 3.1, and GPT-5.4, along with selected variants, across multiple metrics: answer correctness, visualization correctness, query success, and retry attempts. A total of 54 natural-language queries are curated across four scenarios: spatial, graph, cross-database, and conversational. Results show strong overall performance, with answer correctness ranging from 85.2% to 100%, visualization correctness from 92.9% to 100%, a 100% query success rate, and fewer than three retries across all 54 queries. Overall, the findings suggest that CityLLM provides a lightweight and extensible approach for conversational access to semantic 3D city data.
Sep 26, 2026cs.CV

CityToolVQA: Tool-Augmented Visual Question Answering for 3D Spatial Cognition in Urban Low-Altitude Environments

CityToolVQA addresses the weak performance of Vision-Language Models (VLMs) on quantitative tasks in urban low-altitude visual question answering. We divide the seven tasks into qualitative and quantitative groups: qualitative questions are answered directly by the VLM, whereas quantitative questions are handled by an external visual-geometric toolchain that performs object grounding, segmentation, depth back-projection, and spatial computation. The toolchain can be attached to different VLMs in a zero-shot manner; a Depth-Assisted Prompt Inference (DAPI) fallback is triggered when the main-chain detection is invalid or unreliable, and CityToolVQA-SFT adapts the 8B backbone to tool-conditioned inputs. On the 73,324-question Open3D-VQA-v2 test set, CityToolVQA-SFT (Qwen3-VL-8B) reaches 67.6% overall accuracy, and attaching the toolchain to ten open-source VLMs improves quantitative-task accuracy by 12.7-36.9 percentage points. These results indicate that externalizing explicit 3D geometric computation effectively complements the limited ability of RGB-only VLMs to estimate metric distances and object sizes.