FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
Organizations: Tongji University · University of Washington · Sun Yat-sen University · Microsoft · The Chinese University of Hong Kong
Abstract
Front-end engineering involves a complex workflow where engineers conceptualize designs, translate them into code, and iteratively refine the implementation. While recent benchmarks primarily focus on converting visual designs to code, we present FullFront, a benchmark designed to evaluate Multimodal Large Language Models (MLLMs) \textbf{across the full front-end development pipeline}. FullFront assesses three fundamental tasks that map directly to the front-end engineering pipeline: Webpage Design (conceptualization phase), Webpage Perception QA (comprehension of visual organization and elements), and Webpage Code Generation (implementation phase). Unlike existing benchmarks that use either scraped websites with bloated code or oversimplified LLM-generated HTML, FullFront employs a novel, two-stage process to transform real-world webpages into clean, standardized HTML while maintaining diverse visual designs and avoiding copyright issues. Extensive testing of state-of-the-art MLLMs reveals significant limitations in page perception, code generation (particularly for image handling and layout), and interaction implementation. Our results quantitatively demonstrate performance disparities across models and tasks, and highlight a substantial gap between current MLLM capabilities and human expert performance in front-end engineering. The FullFront benchmark and code are available in https://github.com/Mikivishy/FullFront.
Figures & tables
| Model | Gemini Visual Score | CLIP Score | DINOv2 Score | Human Score |
| GPT-4o | 5.0450 | 0.7445 | 0.5598 | 6.9600 |
| gemini-2.0-flash-exp- image-generation | 2.0190 | 0.6901 | 0.4798 | 6.0400 |
| Model | Real-world | Synthetic | Multi-window |
| Qwen2.5-VL-72B-Instruct | 0.4696 | 0.4950 | 0.4267 |
| InternVL2.5-78B | 0.4696 | 0.5050 | 0.4267 |
| InternVL3-78B | 0.4816 | 0.5375 | 0.4600 |
| LLaVA-Onevision-72B | 0.3296 | 0.3275 | 0.2733 |
| Claude 3.7 Sonnet | 0.5464 | 0.5325 | 0.4533 |
| Gemini 2.5 Flash | 0.4800 | 0.4250 | 0.3867 |
| Model | Code Score | Gemini Visual Score | CLIP Score | DINOv2 Score | ||||||||||||
| Ref | Img | Inter | Text | Ref | Img | Inter | Text | Ref | Img | Inter | Text | Ref | Img | Inter | Text | |
| Qwen2.5-VL-72B-Instruct | 0.53 | 0.40 | 0.40 | 0.49 | 6.21 | 4.48 | 6.22 | 5.18 | 0.79 | 0.72 | 0.76 | 0.73 | 0.63 | 0.49 | 0.67 | 0.47 |
| InternVL2.5-78B | 0.36 | 0.33 | 0.30 | 0.50 | 5.03 | 4.01 | 3.51 | 5.44 | 0.74 | 0.74 | 0.69 | 0.74 | 0.67 | 0.57 | 0.58 | 0.48 |
| InternVL3-78B | 0.49 | 0.42 | 0.38 | 0.50 | 5.87 | 4.47 | 4.48 | 4.91 | 0.77 | 0.73 | 0.73 | 0.72 | 0.61 | 0.53 | 0.60 | 0.44 |
| LLaVA-Onevision-72B | 0.33 | 0.14 | 0.06 | 0.41 | 4.90 | 1.89 | 0.45 | 5.13 | 0.73 | 0.65 | 0.58 | 0.74 | 0.66 | 0.42 | 0.17 | 0.50 |
| Claude 3.7 Sonnet | 0.63 | 0.64 | 0.55 | 0.64 | 8.36 | 8.93 | 9.18 | 8.14 | 0.88 | 0.89 | 0.86 | 0.87 | 0.80 | 0.81 | 0.82 | 0.78 |
| Model | Ref | Image | Inter | Text |
| Qwen2.5-VL-72B-Instruct | 6.18 | 5.72 | 7.02 | 4.90 |
| InternVL2.5-78B | 5.36 | 4.78 | 5.04 | 4.24 |
| InternVL3-78B | 6.32 | 5.56 | 5.44 | 4.62 |
| LLaVA-Onevision-72B | 5.64 | 2.96 | 0.58 | 4.42 |
| Claude 3.7 Sonnet | 8.00 | 8.48 | 8.80 | 8.10 |
| Gemini 2.5 Flash | 8.44 | 8.40 | 7.86 | 8.24 |
| Model | ||||
| Qwen2.5-VL-72B-Instruct | 22 | 28 | 10 | 40 |
| InternVL3-78B | 25 | 23 | 8 | 44 |
| GPT-4o | 32 | 13 | 21 | 34 |
| Claude 3.7 Sonnet | 44 | 3 | 36 | 17 |
| Gemini 2.5 Flash | 39 | 5 | 25 | 31 |
| GPT-5 | 48 | 6 | 31 | 15 |
| Model | Perc. | Code Score | G-Vis Score | CLIP Score | DINO Score |
| Qwen2.5-VL -72B-Instruct | Cor. (46) | 0.386 | 4.430 | 0.705 | 0.409 |
| Wrg. (48) | 0.405 | 5.117 | 0.753 | 0.529 | |
| InternVL3-78B | Cor. (59) | 0.418 | 4.444 | 0.736 | 0.554 |
| Wrg. (42) | 0.409 | 4.422 | 0.715 | 0.486 | |
| GPT-4o | Cor. (43) | 0.336 | 6.007 | 0.799 | 0.672 |
| Wrg. (56) | 0.342 | 5.896 | 0.809 | 0.697 |
| Model | Size | Blank | Isolation | ||||||
| Ref | Img | Inter | Text | Ref | Img | Inter | Text | Inter | |
| Qwen2.5-VL-72B-Instruct | 38 | 62 | 11 | 84 | 8 | 4 | 2 | 3 | 2 |
| InternVL2.5-78B | 5 | 20 | 2 | 76 | 2 | 14 | 12 | 4 | 11 |
| InternVL3-78B | 6 | 20 | 5 | 95 | 5 | 14 | 10 | 1 | 1 |
| LLaVA-Onevision-72B | 1 | 22 | 3 | 64 | 5 | 45 | 1 | 9 | 88 |
| Claude 3.7 Sonnet | 2 | 1 | 0 | 3 | 3 | 0 | 0 | 0 | 0 |
| Model | Structure | Text | Image | Form |
| Qwen2.5-VL-72B-Instruct | 0.51 | 0.19 | 0.54 | 0.41 |
| InternVL2.5-78B | 0.43 | 0.11 | 0.52 | 0.34 |
| InternVL3-78B | 0.49 | 0.15 | 0.61 | 0.42 |
| LLaVA-Onevision-72B | 0.29 | 0.06 | 0.38 | 0.23 |
| Claude 3.7 Sonnet | 0.72 | 0.39 | 0.65 | 0.53 |
| Gemini 2.5 Flash | 0.70 | 0.40 | 0.71 | 0.51 |
Appendix figures & tables50 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Interaction Success Rate | Interaction Success Rate (mini) |
| Qwen2.5-VL-72B-Instruct | 57.00 | 64.00 |
| InternVL2.5-78B | 47.00 | 54.00 |
| InternVL3-78B | 48.00 | 56.00 |
| LLaVA-Onevision-72B | 16.00 | 24.00 |
| Claude 3.7 Sonnet | 78.00 | 86.00 |
| Gemini 2.5 Flash | 70.00 | 72.00 |
| Model | C1 | C2 | C3 | C4 | C5 | C6 | H1 | H2 | H3 | H4 |
| Qwen2.5-VL-72B-Instruct | 9 | 7 | 6 | 7 | 5 | 8 | 3 | 2 | 2 | 8 |
| InternVL2.5-78B | 7 | 6 | 6 | 8 | 4 | 6 | 0 | 2 | 4 | 4 |
| InternVL3-78B | 7 | 3 | 5 | 7 | 6 | 4 | 1 | 7 | 1 | 7 |
| LLaVA-Onevision-72B | 8 | 0 | 7 | 0 | 0 | 0 | 0 | 0 | 0 | 1 |
| Claude 3.7 Sonnet | 9 | 8 | 5 | 10 | 9 | 9 | 6 | 3 | 9 | 10 |
| Gemini 2.5 Flash | 7 | 9 | 6 | 9 | 8 | 9 | 6 | 1 | 8 | 7 |
| Model | Gemini Visual Score | Clip Score | Code Score | DINOv2 Score |
| Qwen2.5-VL-72B-Instruct | 5.1750 | 0.7313 | 0.4854 | 0.4665 |
| InternVL2.5-78B | 5.4444 | 0.7366 | 0.4989 | 0.4824 |
| InternVL3-78B | 4.9144 | 0.7204 | 0.4969 | 0.4405 |
| LLaVA-Onevision-72B | 5.1273 | 0.7365 | 0.4129 | 0.5006 |
| Claude 3.7 Sonnet | 8.1358 | 0.8655 | 0.6436 | 0.7790 |
| Gemini 2.5 Flash | 8.0111 | 0.8660 | 0.5927 | 0.7929 |
| Metric | Ref | Img | Inter | Text | Average |
| Gemini Visual Score | 0.9394 | 0.9273 | 0.9394 | 0.9394 | 0.9364 |
| Code Score | 0.8909 | 0.9152 | 0.9515 | 0.9030 | 0.9152 |
| Clip Score | 0.9394 | 0.9152 | 0.9879 | 0.7333 | 0.8939 |
| DINOv2 Score | 0.9030 | 0.9030 | 0.9394 | 0.8061 | 0.8879 |
| Model | Parameter Setting | Source | URL |
| Qwen2.5-VL-72B-Instruct | temperature = 0.0 | local checkpoint | https://huggingface.co/Qwen/Qwen2.5-VL-72B-Instruct |
| InternVL2.5-78B | temperature = 0.0 | local checkpoint | https://huggingface.co/OpenGVLab/InternVL2_5-78B |
| InternVL3-78B | temperature = 0.0 | local checkpoint | https://huggingface.co/OpenGVLab/InternVL3-78B |
| LLaVA-Onevision-72B | temperature = 0.0 | local checkpoint | https://huggingface.co/llava-hf/llava-onevision-qwen2-72b-ov-hf |
| R1-Distill-Llama-8b | temperature = 0.0 | local checkpoint | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-8B |
| Claude 3.7 Sonnet | temperature = 0.0 | claude-3-7-sonnet-20250219 | https://www.anthropic.com/ |
| Model | Aes(w) | Aes(t) | Aes(l) | Inf(w) | Inf(t) | Inf(l) |
| Qwen2.5-VL-72B-Instruct | 15.2 | 34.4 | 50.4 | 13.8 | 38.4 | 47.8 |
| InternVL2.5-78B | 18.6 | 28.8 | 52.6 | 29.8 | 38.2 | 32.0 |
| InternVL3-78B | 23.6 | 33.4 | 43.0 | 34.6 | 34.0 | 31.4 |
| LLaVA-Onevision-72B | 6.0 | 10.8 | 83.2 | 7.6 | 25.0 | 67.4 |
| Claude 3.7 Sonnet | 80.0 | 15.6 | 4.4 | 62.8 | 29.0 | 8.2 |
| Gemini 2.5 Flash | 75.4 | 16.6 | 8.0 | 63.2 | 27.8 | 9.0 |
| Model | Code Score | Gemini Visual Score | CLIP Score | DINOv2 Score | Human Score | ||||||||||
| Ref | Img | Text | Ref | Img | Text | Ref | Img | Text | Ref | Img | Text | Ref | Img | Text | |
| Qwen2.5-VL-72B-Instruct | 0.27 | 0.35 | 0.34 | 4.05 | 4.85 | 5.83 | 0.69 | 0.69 | 0.76 | 0.48 | 0.47 | 0.57 | 6.16 | 4.16 | 5.62 |
| InternVL3-78B | 0.27 | 0.34 | 0.26 | 5.93 | 5.96 | 5.08 | 0.72 | 0.72 | 0.73 | 0.45 | 0.51 | 0.54 | 5.97 | 5.03 | 5.75 |
| Claude 3.7 Sonnet | 0.54 | 0.51 | 0.41 | 8.79 | 8.60 | 6.24 | 0.80 | 0.79 | 0.74 | 0.68 | 0.67 | 0.54 | 8.30 | 8.52 | 6.46 |
| Gemini 2.5 Flash | 0.48 | 0.48 | 0.38 | 8.07 | 8.27 | 6.20 | 0.80 | 0.79 | 0.78 | 0.69 | 0.67 | 0.62 | 7.87 | 7.62 | 6.15 |
| GPT-4o | 0.53 | 0.47 | 0.33 | 7.17 | 6.78 | 6.19 | 0.76 | 0.74 | 0.75 | 0.53 | 0.52 | 0.59 | 6.38 | 5.76 | 5.69 |
| Evaluated Model | Human Judge | Gemini 2.5 Flash Judge | Claude 3.7 Sonnet Judge | Grok 4 Judge | Qwen3-VL Judge |
| GPT-5 | 8.50 (R1) | 8.87 (+0.37, R1) | 7.67 (-0.83, R1) | 9.02 (+0.52, R1) | 8.92 (+0.42, R1) |
| Claude 3.7 Sonnet | 8.35 (R2) | 8.73 (+0.39, R2) | 7.51 (-0.84, R2) | 8.97 (+0.63, R2) | 8.83 (+0.49, R2) |
| Gemini 2.5 Flash | 8.23 (R3) | 8.61 (+0.38, R3) | 7.38 (-0.86, R3) | 8.92 (+0.68, R3) | 8.70 (+0.46, R3) |
| GPT-4o | 6.67 (R4) | 6.72 (+0.05, R4) | 5.93 (-0.74, R4) | 8.19 (+1.52, R4) | 7.34 (+0.67, R4) |
| Qwen2.5-VL-72B | 5.96 (R5) | 5.98 (+0.02, R5) | 5.31 (-0.65, R5) | 8.13 (+2.17, R5) | 6.51 (+0.55, R5) |
| InternVL3-78B | 5.49 (R6) | 5.34 (-0.15, R6) | 4.96 (-0.53, R6) | 7.98 (+2.50, R6) | 6.13 (+0.65, R6) |
| Judge Model | Avg. Score Std Dev | Inter-Run Correlation |
| Gemini 2.5 Flash | 7.38 0.080 | 0.965 |
| Claude 3.7 Sonnet | 6.46 0.300 | 0.876 |
| Grok 4 | 8.53 1.097 | 0.410 |
| Qwen3-VL-235B-A22B-Instruct | 7.74 0.079 | 0.972 |
| Experiment Condition | Gemini Visual Score | Code Score | CLIP Score | DINOv2 Score |
| (0-10 scale) | (0-1 scale) | (0-1 scale) | (0-1 scale) | |
| Baseline (Original HTML) | 10.00 | 1.00 | 1.00 | 1.00 |
| Exp 1: Valid Alternative | 9.266 | 0.952 | 0.969 | 0.942 |
| Penalty ( ) | -0.734 | -0.048 | -0.031 | -0.058 |
| Exp 2: Broken Implementation | 6.838 | 0.727 | 0.846 | 0.760 |
| Penalty ( ) | -3.162 | -0.273 | -0.154 | -0.240 |
| Intervention | Model | Eligible Cases | Target Outcome | Rate |
| QA-relevant crop ( ) | Qwen2.5-VL-72B-Instruct | 28 | 22 corrected | 78.6% |
| QA-relevant crop ( ) | InternVL3-78B | 23 | 18 corrected | 78.3% |
| Counterfactual local edit ( ) | Gemini 2.5 Flash | 25 | 18 edits ignored | 72.0% |
| Counterfactual local edit ( ) | GPT-5 | 31 | 21 edits ignored | 67.7% |
| Model | Perception QA Accuracy | Image-to-Code | Text-to-Code | ||||||||
| Real | Synth. | Multi. | Code | GVS | CLIP | DINOv2 | Code | GVS | CLIP | DINOv2 | |
| Qwen2.5-VL-72B-Instruct | 0.4696 | 0.4950 | 0.4267 | 0.40 | 4.48 | 0.72 | 0.49 | 0.49 | 5.18 | 0.73 | 0.47 |
| InternVL3-78B | 0.4816 | 0.5375 | 0.4600 | 0.42 | 4.47 | 0.73 | 0.53 | 0.50 | 4.91 | 0.72 | 0.44 |
| Claude 3.7 Sonnet | 0.5464 | 0.5325 | 0.4533 | 0.64 | 8.93 | 0.89 | 0.81 | 0.64 | 8.14 | 0.87 | 0.78 |
| GPT-5 | 0.5816 | 0.5775 | 0.5467 | 0.56 | 8.58 | 0.88 | 0.81 | 0.59 | 8.48 | 0.88 | 0.79 |
| Qwen3.5-122B-A10B | 0.5848 | 0.5925 | 0.5667 | 0.55 | 7.46 | 0.84 | 0.73 | 0.57 | 7.25 | 0.82 | 0.69 |