Prompt optimization for text-to-image (T2I) generation has been pursued almost entirely as text rewriting, in which a short user brief is expanded into a longer, model-preferred token sequence. We argue that such a language-space formulation is ill-suited to structured visual design tasks such as logo creation, where a one-line brief leaves most design decisions unspecified. These decisions depend on relational priors that a linear sequence cannot encode, and they leave an uncontrolled channel through which protected marks may be reproduced. We therefore recast logo prompting as sampling within a structured design space, and instantiate this idea as DOGS (Design-space prompting with an Originality-aware GFlowNet Sampler). From a large corpus of real-world logos, we mine a typed, graph-structured design grammar whose edges record empirical co-occurrence. A GFlowNet sampler then generates design graphs with probability proportional to a terminal reward that combines recognizability, aesthetics, and corpus-relative originality. Every slot draws only from a closed design-level vocabulary, and any infringement-inducing or harmful token is removed during parsing. The originality reward further penalizes proximity to existing logos, thereby incorporating infringement avoidance into the method by construction. On two open-source renderers and against nine baselines, DOGS produces logos that are more recognizable and aesthetic, substantially more diverse, and far less prone to trademark infringement.
Figures & tables
Figure 1: From user brief to logo design. Each column is one brief. Top: the user prompt rendered directly. Bottom: the DOGS output, which fills in the unspecified design decisions to produce a coherent, professionally composed mark.
Figure 2: DOGS pipeline , organized into three stages. Brief Parsing maps the user brief into an initial state s0 , filling the design slots it specifies and leaving the rest undecided. Offline Grammar Induction re-annotates a logo corpus into typed nodes and builds a co-occurrence grammar G=(V,E) , pruned to define the per-step action space. GFlowNet Symbolic Completion draws a trajectory s0→⋯→sT with the sampler πF over that action space; the terminal state sT is rendered into a logo by a frozen T2I model, scored by the reward model, and the resulting trajectory probability trains πF via the Trajectory Balance loss LTB .
Figure 3: Qualitative comparison. Columns are methods. Comparison : DOGS composes a coherent design while baselines drift off-subject or render a stock image. Diversity : four samples per method as 2×2 mosaics; DOGS yields four distinct yet usable logos, while baselines collapse or vary without comparable quality. Infringement : a brand prompt names a known trademark; baselines reproduce it, DOGS substitutes a different professional design (prompts truncated).
SDXL-Lightning
FLUX.1-schnell
Method
Qrec↑
Qaes↑
Diversity ↑
Qrec↑
Qaes↑
Diversity ↑
Original prompt
9.26 ± 0.15
4.03 ± 0.17
0.00 ± 0.00
9.70 ± 0.10
4.96 ± 0.17
0.00 ± 0.00
Qwen rewrite
7.41 ± 0.21
3.19 ± 0.15
2.25 ± 0.16
7.75 ± 0.20
3.64 ± 0.18
2.49 ± 0.16
Llama rewrite
9.40 ± 0.11
4.07 ± 0.12
2.76 ± 0.18
9.71 ± 0.06
4.67 ± 0.11
2.30 ± 0.15
Promptist
9.60 ± 0.09
3.35 ± 0.15
2.52 ± 0.20
9.50 ± 0.09
4.15 ± 0.15
2.33 ± 0.19
BeautifulPrompt
9.61 ± 0.07
3.18 ± 0.11
3.34 ± 0.16
9.45 ± 0.08
3.65 ± 0.11
3.54 ± 0.16
Table 1: Quality battery. 150 briefs, five samples per cell (metrics: Sec. 4.2 ). Bold: best among baselines and DOGS; abl. rows: ablations of Sec. 4.4 . Deterministic methods score zero diversity by construction.
SDXL-Lightning
FLUX.1-schnell
Method
Brand ↓
Company ↓
Design ↓
Brand ↓
Company ↓
Design ↓
Original prompt
0.13 ± 0.02
0.04 ± 0.01
0.05 ± 0.01
0.20 ± 0.02
0.08 ± 0.01
0.06 ± 0.02
Qwen rewrite
0.07 ± 0.02
0.02 ± 0.01
0.02 ± 0.01
0.11 ± 0.02
0.04 ± 0.02
0.03 ± 0.02
Llama rewrite
0.12 ± 0.01
0.04 ± 0.01
0.05 ± 0.01
0.18 ± 0.02
0.07 ± 0.01
0.07 ± 0.02
Promptist
0.10 ± 0.02
0.04 ± 0.01
0.04 ± 0.01
0.19 ± 0.02
0.07 ± 0.02
0.07 ± 0.02
BeautifulPrompt
0.10 ± 0.02
0.04 ± 0.01
0.02 ± 0.01
0.15 ± 0.02
0.07 ± 0.01
0.03 ± 0.01
Table 2: Infringement battery. Leakage lift Δ=direct−baseline , 100 brands, one image per cell. Columns are the three prompt types of Sec. 4.1 .
Text-to-image models now generate graphic design at production scale, yet their supervision still comes primarily from photo-style preference datasets with a single overall verdict per comparison. Designers evaluate designs along several distinct axes (e.g., typography, layout, color harmony) that a single preference label collapses. We release \emph{TASTE} \textit{(Typography, Aesthetics, Spatial, Tone, Etc.)}, a multi-dimensional preference dataset in which two disjoint cohorts of five professional designers each ranked outputs from four current text-to-image models across nine criteria along with per-image hallucination flags. We pair the dataset with two contributions. First, a criterion-agnostic signal-validation framework based on Kendall's τ, majority-vote probability, and Condorcet cycles against exact iid-uniform nulls; the analysis reveals significant but moderate designer agreement, with every TASTE criterion rejecting the random-rater null. Second, we benchmark preference models on TASTE and find that off-the-shelf VLM judges and dedicated T2I scorers fail to reach majority agreement with the designer panel, while a small MLP head trained directly on TASTE substantially narrows the gap to the single-rater ceiling, setting a baseline for future TASTE-trained preference models.
Haonan Zhu, Elad Hirsch, Alexandria Minetti +3
Lica World, San Francisco, USA. · Contra.Work Inc., Brooklyn, NY, USA.
Text-to-Image (T2I) generation is primarily driven by Diffusion Models (DM) which rely on random Gaussian noise. Thus, like playing the slots at a casino, a DM will produce different results given the same user-defined inputs. This imposes a gambler's burden: To perform multiple generation cycles to obtain a satisfactory result. However, even though DMs use stochastic sampling to seed generation, the distribution of generated content quality highly depends on the prompt and the generative ability of a DM with respect to it. To account for this, we propose Naïve PAINE for improving the generative quality of Diffusion Models by leveraging T2I preference benchmarks. We directly predict the numerical quality of an image from the initial noise and given prompt. Naïve PAINE then selects a handful of quality noises and forwards them to the DM for generation. Further, Naïve PAINE provides feedback on the DM generative quality given the prompt and is lightweight enough to seamlessly fit into existing DM pipelines. Experimental results demonstrate that Naïve PAINE outperforms existing approaches on several prompt corpus benchmarks.
Joong Ho Kim, Nicholas Thai, Souhardya Saha Dip +2
Subject-driven text-to-image generation still struggles to preserve high-frequency identity details such as logos, patterns, and text. Existing methods typically operate directly in RGB space, which often leads to detail degradation under substantial edits. We propose a two-stage framework that decouples structure from appearance by first predicting a Canny map and then rendering the final image conditioned on both the source appearance and the predicted structure. To improve text handling, we further introduce a fully automatic pipeline that constructs a 100k-pair text-aware dataset with cross-view textual consistency. Experiments, including GPT-4.1-based evaluation and a knowledge distillation study, show clear gains over selected baselines and suggest that intermediate structural prediction is an effective route for high-fidelity subject-driven generation. Our dataset and code will be made publicly available.
Hanzhong Guo, Yizhou Yu
School of Computing and Data Science, The University of Hong Kong