Aesthetic image cropping aims to identify the optimal crop of an image in terms of aesthetics and composition. While supervision based on annotated data is fundamental, the field has been hindered by a long-standing problem: existing datasets suffer from (1) human subjectivity and (2) rigid discreteness confined to fixed sampling grids. These flawed annotations not only limit the accuracy and generalization of trained models but also severely distort fair evaluation. To overcome this, we propose to model human cropping preference as a multi-peaked, continuous, and sharp field over the crop space. We introduce the Continuous Preference Field (CPF), which recovers a dense preference landscape from discrete annotations through (1) peak clustering, (2) off-lattice refinement, (3) negative shaping, and (4) field assembly. Based on this, we train CPIC, a VLM-based cropping model optimized via GRPO with the CPF reward, which overcomes template collapse, achieving state-of-the-art performance and exceptional out-of-domain generalization. Finally, to resolve the long-standing benchmark evaluation crisis, we introduce CPICD, a comprehensive recalibration of existing ground-truth boxes. By leveraging the CPF to correct grid-bound artifacts across mainstream benchmarks, CPICD establishes a rigorous and reliable foundation for future cropping research. Extensive experiments and user studies demonstrate the superiority of our CPF, CPIC, and CPICD. Code, model, and data are available at https://github.com/zzqingz/CPIC.
Figures & tables
Figure 1: Illustration of three key properties of our Continuous Preference Field (CPF) : (1) multi-peakness, (2) continuity, and (3) sharpness. (Zoom-in for best view)
Table 2
Figure 2: Overview of CPF construction and its applications. (1) Peak clustering identifies distinct preference modes from human-scored crops. (2) Off-lattice refinement adjusts each peak using preference, composition, and subject-preservation scores. (3) Negative shaping constructs five types of undesirable crops. (4) Field assembly combines peak-distance and negative gating into CPF. CPF serves as the reward for GRPO training of CPIC and guides annotation recalibration for CPICD.
Figure 3: Comparison of original benchmark annotations and CPICD recalibrated boxes. Recalibration improves unbalanced compositions with truncated subjects or overly narrow crops.
Table 3: Comparison on the four benchmarks. Best per column in bold , second underlined .
Figure 4: Qualitative results on in-domain (GAIC) and out-of-domain (FCDB) images.
Configuration
CPF ingredient
GAIC
FLMS
FCDB
Boxes ↑
multi-peak
continuity
sharpness
SRCC ↑
IoU ↑
LAION ↑
IoU ↑
Supervision strategies
(a)
CE on annotated GT
0.523
0.798
5.060
0.642
56
(b)
+ GRPO w/ IoU
0.564
0.822
5.074
0.671
92
(c)
CE on refined peaks
✓
0.538
0.827
5.076
0.704
154
(d)
+ GRPO w/ IoU
✓
0.559
0.835
5.082
0.701
199
Table 6: Ablation studies. Boxes counts the distinct crops the model emits over the 200 GAIC test images (at most 200 ), the diversity measure of Table 2 .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Zero-shot interpretable and interactive aesthetic cropping. Top: CPIC provides a concrete rationale grounded in the selected subject, spatial arrangement, contrast, and visual effect. Bottom: after its initial crop, CPIC follows free-form feedback to remove the distracting foreground pedestrian and tightly reframe the three musicians.
Output target
GAIC
FLMS
FCDB
ACC 10↑
IoU ↑
IoU ↑
Box only
0.905
0.842
0.714
Short text + box
0.910
0.822
0.723
Appendix
Table 7: Output-target ablation. Best bolded .
Coordinates
GAIC
FLMS
FCDB
12×12 anchors
1.0000
0.9099
0.8273
0 – 1000 integers
0.9994
0.9997
0.9989
Appendix
Table 8: Oracle IoU on three benchmarks.
ID
Config.
GAIC
Comp.
Subj.
GAIC
FLMS
FCDB
SRCC ↑
IoU ↑
IoU ↑
(a)
None
–
–
–
0.517
0.829
0.672
(b)
Direct
–
–
–
0.526
0.812
0.671
(c)
Refine
✓
–
–
0.563
0.827
0.704
(d)
Refine
✓
✓
–
0.578
0.828
0.709
(e)
Refine
✓
✓
✓
0.590
0.842
0.714
Appendix
Table 9: Refinement-score ablation.
Figure 6: CPF scores and IoU for pairs of crops. CPF distinguishes coherent compositions from visually flawed alternatives that IoU fails to adequately penalize, providing qualitative evidence of its closer alignment with human visual preference.
Figure 7: User study interface for cropping quality.
Aesthetic image cropping aims to enhance the aesthetic quality of an image by improving its composition through spatial cropping. Previous methods often rely on saliency prediction or retrieval augmentation, ignoring the task's core requirement: a deep understanding of composition and aesthetics. Consequently, saliency-based methods struggle to make compositional trade-offs in complex scenes, while retrieval-based methods blindly refer to similar cases, lacking adaptive reasoning for unique scenes. Both approaches fail to align their automated cropping results with those of human experts. To address the above issues, we propose a novel paradigm that reformulates aesthetic cropping as a multimodal reasoning task, aiming to activate the VLM's analytical and comprehension capabilities in aesthetics. We design a Compositional Reasoning and Optimizing Preference method (CROP) that directs the VLM to think like a professional photographer. It deconstructs a complex and subjective aesthetic problem into an "analysis-proposal-decision" process, reasoning step by step through the analysis of scene elements and compositional principles. Meanwhile, our expert preference alignment module makes the model's decision consistent with human expert aesthetics. Extensive experiments across multiple datasets validate our method's superiority and component effectiveness.
Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existing crop-and-explain methods largely treat explanation as post-hoc text generation and overlook composition, a key aesthetic factor that links crop decisions with interpretable reasoning. In this paper, we reformulate explainable aesthetic image cropping as a structured crop-composition-explanation problem. To support this setting, we introduce COMEX, a new benchmark built through image expansion and an IO-reversal pipeline. COMEX contains 33,161 quadruples, each consisting of an expanded image, a crop box, a composition category, and a composition-grounded explanation, enabling joint learning of crop localization, composition understanding, and explanation generation. We further propose a two-stage SFT+GRPO framework, where supervised fine-tuning establishes the structured output protocol and basic cropping ability, and GRPO further improves crop quality, composition prediction, and explanation faithfulness. We benchmark 15 large vision-language models and existing cropping methods on COMEX, establishing a comprehensive testbed for composition-grounded explainable aesthetic cropping. Experiments on both COMEX and prior benchmarks demonstrate the effectiveness and transferability of our framework, with strong performance across evaluation metrics.
Rui Yang, Wei Zhou, Dingyong Gou +5
State Key Laboratory of Mobile Network and Mobile Multimedia Technology, ZTE Corporation · School of Data Science and Institute of Artificial Intelligence, Chang’an University
Image cropping aims to improve image aesthetics by preserving important content within an appropriately composed region. However, most existing methods focus primarily on salient regions and therefore have limited sensitivity to the global relationships among the main image components. To address this limitation, we propose Global Attention-Fused Image Cropping (GAFIC), which consists of an Attention-Guided Feature Fusion (AGFF) and a Global-Aligned Crop Evaluator (GACE). AGFF aggregates the importance of local regions to construct a global representation that captures both image structure and local details. GACE aligns candidate crop features with this global representation, enabling crop evaluation to remain sensitive to boundary changes. We further combine three ranking losses across multiple scales to obtain accurate and stable crop scores. Extensive experiments on the GAIC and CPC datasets demonstrate that GAFIC outperforms existing image-cropping methods, particularly in terms of accuracy and stability. Unlike pixel-level retargeting methods such as seam carving, inpainting, and diffusion-based synthesis, GAFIC does not synthesize or modify the retained pixels; instead, it selects an aesthetically preferred crop from the source image, making it suitable for scenarios where pixel integrity and efficient batch processing are important. The source code is available at https://github.com/AIVRC/GAFIC.git.
Haotian Yang, Zhile Yang, Kin-Man Lam +2
Faculty of Data Science, City University of Macau, Macao, China · Shenzhen University of Advanced Technology, Shenzhen, China · Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China +2