Lane and road markings provide critical guidance for vehicle navigation and multi-agent coordination, yet authoring them at scale remains a manual workflow that limits quantitative analysis and scenario testing. We introduce Controllable Road Marking Generation, which synthesizes a missing center-region marking layout from a drivable-area mask, optional outer-ring markings, and a textual description. Our benchmark uses deterministic, metadata-derived prompts and three output channels: lane dividers, road dividers, and pedestrian crossings. We develop a conditional bird's-eye-view (BEV) pipeline that combines (i) a text-conditioned latent rectified-flow DiT trained with a topology-aware auxiliary loss, (ii) Gaussian-blurred training targets that stabilize learning of thin, sparse markings, and (iii) Structured Gaussian Render (SGR), a training-free post-process that recovers crisp divider geometry by extracting polylines, fitting cubic Bézier curves, and re-rendering them as anisotropic super-Gaussian primitives. On 4,597 Argoverse~2 test tiles, our system achieves Buffered F1 of 80.8 and clDice of 50.2, compared with 38.8 and 24.6 for an adapted state-of-the-art mask-refinement baseline. On Waymo dataset, it yields 88.0 Buffered F1 and 66.2 clDice. Component ablations show complementary connectivity gains from topology-aware supervision and SGR. Text-editing experiments reveal that stronger guidance improves edit success but also increases changes to non-target structures. We see this framework as a step toward simulation-ready road-marking variation, automated map completion, and early-stage infrastructure design exploration.
Figures & tables
Method
Raster
Generative
Layout
Text
Thin geometry
RoadMark-cGAN [ 7 ]
✓
✓
×
×
△
LSR-DM [ 41 ]
✓
✓
×
×
△
MapDiffusion [ 32 ]
×
✓
△
×
△
LaneDiffusion [ 45 ]
×
✓
×
×
✓
PolyDiffuse [ 6 ]
×
✓
△
×
△
Ours
✓
✓
✓
✓
✓
Table 1: Capability positioning. ✓ direct support; △ partial or adjacent capability; × outside scope. Only capabilities are compared in this table.
Figure 1: Overview of the proposed controllable road marking generation pipeline.
Figure 2: Structured Gaussian Render (SGR). Starting from the raw soft raster d~ , SGR thresholds the divider channels, extracts lane- and road-divider polylines, fits cubic Bézier curves, and re-renders sampled curve primitives with anisotropic super-Gaussian splats. The pedestrian-crossing channel is passed through as a raster mask.
Method
Overall
Lane Div.
Road Div.
Ped. Cross.
F1 ↑
clDice ↑
F1 ↑
clDice ↑
F1 ↑
clDice ↑
F1 ↑
clDice ↑
LSR-DM † [ 41 ]
38.8
24.6
37.3
15.6
43.6
27.9
35.4
30.4
Ours w/ CNN deblur
62.7
46.3
50.2
18.5
64.9
37.3
73.0
83.1
Ours w/ SGR deblur
80.8
50.2
74.0
28.4
66.6
38.0
89.7
84.1
Table 2: Structural metrics on 4,597 AV2 test tiles, evaluated on the generated center 256×256 region. Per-channel scores use GT-positive tiles; Overall is the macro-average over the three channels. Values are percentages; bold/underline mark best/second-best.
Figure 3: Qualitative text control. Columns show ground truth and edited-prompt generations; green boxes mark the generated center region.
Control family
Nelig
Success ↑
False act. ↓
fa-MSE ↓
Δtgt↑
Add crosswalk
2,771
49.6
9.2
0.071
+1,772.2
Delete crosswalk
957
75.1
19.7
0.119
+4,070.4
Single → double yellow
1,750
31.0
7.4
0.041
+29.8
Table 3: Mask-parsed controllability for single-clause prompt edits on the AV2 test split. Each family is evaluated only on eligible non-trivial tiles ( Nelig ), with (s,p) and seed fixed between baseline and edited prompts. Success and False activation are percentages; fa-MSE is a τ -normalized MSE over non-target channel changes, and Δtgt is the signed target-channel change. Protocol details are in App. A and App. E .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Semantic case
Example prompt
Two-way crossing tile
“two-way, 2 lanes each direction, solid white dividers, single yellow center line, 4 crosswalks at intersection”
One-way non-crossing tile
“one-way, 1 lane, no double yellow, no crosswalks”
Double-yellow center line
“two-way, 2 lanes each direction, solid white dividers, double yellow center line, 1 crosswalk at intersection”
Lane-type context
“two-way, 2 lanes each direction, solid white dividers, double yellow center line, 2 crosswalks at intersection”
Appendix
Table 4: Representative deterministic prompt strings used for text conditioning.
Dataset
Topology loss
SGR
Buffered F1 ↑
clDice ↑
AV2
No
No
80.5
43.4
No
Yes
81.8
47.2
Yes
No
78.2
48.1
Yes
Yes
80.8
50.2
Waymo
No
No
88.6
59.1
No
Yes
87.3
61.9
Appendix
Table 5: Topology and SGR ablation. Values are percentages; bold marks the best result for each metric within each dataset.
Train data
Steps
Buf. F1 ↑
clDice ↑
0% †
–
3.9
2.4
10%
18,000
16.2
7.3
30%
54,000
20.2
9.3
50% †
90,000
56.0
39.7
70% †
126,000
69.3
47.4
100%
180,000
80.8
50.2
Appendix
Table 6: Data efficiency under matched-epoch training.
Method
Channel
Buf. F1 ↑
Buf. IoU ↑
clDice ↑
LSR-DM † [ 41 ]
Lane divider
37.3
18.7
15.6
Road divider
43.6
25.2
27.9
Pedestrian crossing
35.4
23.4
30.4
Ours w/ CNN deblur
Lane divider
50.2
45.9
18.5
Road divider
64.9
59.7
37.3
Pedestrian crossing
73.0
76.5
83.1
Appendix
Table 7: Per-channel structural metrics on GT-positive AV2 test tiles.
Method
Buf. F1 ↑
clDice ↑
r=1
r=3
r=5
LSR-DM † [ 41 ]
32.8
38.8
41.7
24.6
Ours w/ CNN deblur
59.8
62.7
68.2
46.3
Ours w/ SGR deblur
64.1
73.7
77.5
49.9
Appendix
Table 8: Buffered-F1 sensitivity to tolerance radius.
Asset
Version
License
Argoverse 2 Sensor / Map Dataset
v2.0
CC BY-NC-SA 4.0
CLIP text encoder
ViT-L/14 (OpenAI)
MIT
PyTorch
2.5.1
BSD-3-Clause
diffusers (HuggingFace)
0.30.x
Apache-2.0
xFormers
0.0.28
BSD-3-Clause
NumPy / SciPy / scikit-image
current
BSD-3-Clause
Appendix
Table 9: Existing assets, versions, and licenses.
Figure 4: Additional qualitative comparisons on Argoverse 2 BEV tiles.
Dataset
CFG
Success ↑
FA ↓
fa-MSE ↓
Δtgt↑
AV2
1
7.7
6.0
0.0170
+26.8
3
24.3
12.7
0.0406
+66.1
5
29.3
17.3
0.0707
+74.1
7
35.3
20.7
0.0812
+98.5
Waymo
1
12.1
4.4
0.0110
+45.9
3
28.7
7.4
0.0369
+71.9
Appendix
Table 10: Fixed-seed classifier-free guidance sweep on 300 eligible tiles per dataset, evaluated on raw outputs before SGR. Success and false activation (FA) are percentages. Δtgt denotes direction-normalized target change.
Dataset
Overall
Lane Div.
Road Div.
Ped. Cross.
F1 ↑
clDice ↑
F1 ↑
clDice ↑
F1 ↑
clDice ↑
F1 ↑
clDice ↑
nuScenes-SG †
30.4
8.1
40.7
4.6
33.1
8.7
17.4
11.0
AV2
80.8
50.2
74.0
28.4
66.6
38.0
89.7
84.1
Waymo
88.0
66.2
89.5
55.9
85.7
56.5
88.9
86.3
Appendix
Table 11: Cross-dataset structural metrics for Ours w/ SGR deblur, evaluated on the generated center 256×256 region. Results are reported on 261 nuScenes-SG held-out tiles, 4,597 AV2 test tiles, and 1,000 Waymo test tiles. Per-channel scores use GT-positive tiles; Overall is the macro-average over the three channels. Values are percentages.
Tsinghua University Baidu Beijing, China · University of Macau Macao, China · Institute of Information Engineering, Chinese Academy of Sciences Beijing, China +2