Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding and generation directly in pixel space. PixelUMM represents images as spatial patches and videos as spatiotemporal tubelets, connecting raw pixels to a shared multimodal backbone through single-layer linear projections. Its Mixture-of-Transformers architecture combines shared attention with task-specific parameters and extends clean-pixel prediction to video generation, jointly supporting autoregressive text prediction and pixel-space flow matching. Experiments show that PixelUMM achieves competitive performance across image and video understanding and generation tasks. We further conduct empirical studies of key design choices, including decoder design and spatial-temporal patch size, providing insights for future pixel-space unified multimodal models.
Figures & tables
Figure 1 : Image generation and understanding with PixelUMM. Top: PixelUMM text-to-image generation results. Bottom: benchmark inputs and raw model answers from TextVQA, ChartQA, AI2D, and MMMU.
Figure 2 : Video generation and understanding with PixelUMM. Top: PixelUMM text-to-video generation results. Bottom: PixelUMM responses on MVBench and LVBench.
Figure 3 : PixelUMM architecture. PixelUMM is a native unified multimodal model for image and video understanding and generation. It operates directly in pixel space without pretrained visual encoders or VAE-based latent tokenizers. PixelUMM adopts a Mixture-of-Transformers (MoT) architecture comprising an understanding expert for text prediction and visual understanding and a generation expert for image and video synthesis. We omit explicit timestep embeddings and timestep-conditioned normalization. The denoising network receives the noisy pixels without a separate timestep input. The two experts have symmetric structures with expert-specific normalization layers, projections, and FFNs, while all tokens interact through shared multimodal self-attention in every Transformer block.
Figure 5 : Attention patterns used by PixelUMM. Filled cells indicate visible query–key pairs (rows are queries and columns are keys). Image tokens form one bidirectional block, while sparse video frames form temporally ordered bidirectional blocks, so later frames can access earlier frames. Understanding text is causal and can attend to all preceding visual context. For generation, the noisy target attends to the causal prompt and all target tokens.
Joint Stage 1
Gen Stage 1
Und Stage 1
Und Stage 2
Und Stage 3
Joint Stage 2
Module settings Frozen Trainable
Understanding branch
Generation branch
Hyperparameters
Learning rate
1×10−4
1×10−4
2×10−5
2×10−5
2×10−5
2×10−5
LR scheduler
Constant
Constant
Constant
Constant
Constant
Constant
Table 1: Training recipe of PixelUMM. Stages are listed in training order. Gen/Und denote generation/understanding. Native denotes native-resolution images, and video resolutions are per frame. CE denotes cross-entropy; zero loss weights indicate inactive objectives. Per-step example ratios specify the relative numbers of training examples consumed by each task in a step and need not sum to one; a gray dash indicates an inactive task (zero ratio). Time shift applies only to generation tasks, with H and W denoting spatial dimensions. Elsewhere, dashes indicate inapplicable settings.
Figure 6 : Image patch-size ablation. Sample-mean T2I MSE (faint) and a 20-point EMA (opaque). Left: the full training trajectory with a logarithmic y -axis. Right: a zoom-in of 10K–15K steps with a linear y -axis. The outlined region and connecting lines indicate the enlarged interval.
Figure 7 : Qualitative comparison of image patch sizes. Rows show the 16×16 F1-R01 and 32×32 F1-R02 generation interfaces; within each prompt, columns follow training from 6K to 10K to 15K. All outputs use the same two prompts, DPM-Solver with 50 steps, timestep shift 1, and CFG 3.5 with global renormalization.
Figure 8 : Video patch-size ablation. Patch configurations of video_gen_linear_proj and video_linear_outproj : p32/t4 (F2-R01), p32/t2 (F2-R02), p16/t4 (F2-R03), and p32/t1 (F2-R04). All runs initialize the understanding branch from Joint Stage 1 and the generation branch from scratch; other training settings are identical. Faint curves show raw T2V losses; opaque curves show a 20-point EMA. The left panel covers 0–68K steps on a logarithmic scale; the right panel enlarges 40K–68K on a linear scale.
Figure 9 : Qualitative comparison of video patch configurations. Rows correspond to p32/t4 (F2-R01), p32/t2 (F2-R02), p16/t4 (F2-R03), and p32/t1 (F2-R04). For two matched prompts, columns show frames 0, 8, and 15 of each 16-frame, 192×320 clip.
Figure 10 : Patch artifacts with linear pixel heads. The left column shows an F3-R01 image output above a crop of its smooth sky; the right column shows frame 0 of the Upright piano video above a crop of the smooth wall behind the pianist. Red boxes identify the displayed regions. The subtle grid boundaries are clearer when the figure is enlarged.
Figure 11 : Pixel-space video output-head architectures. All three heads map the Transformer hidden grid at T/4×H/16×W/16 to RGB video at T×H×W . F3-R02 uses a Wan-style upsampling-and-convolution stack. F3-R03 and F3-R04 are PixelShuffle decoders that apply temporal and spatial shuffling in different orders; both end with the same joint temporal-spatial shuffle. T-Conv and S-Conv denote 3×1×1 temporal and 1×3×3 spatial convolutions, respectively. GFLOPs are the convolution-only forward cost for one 96×176×320 clip, using one multiply-accumulate as two FLOPs; normalization, activations, and parameter-free upsampling or shuffling are excluded. For reference, the F3-R01 linear projection requires 133 GFLOPs under the same accounting. At this clip size, F3-R01 has 12.6M parameters. F3-R02, F3-R03, and F3-R04 respectively cost 2,274, 1,279, and 747 GFLOPs ( 17.1× , 9.6× , and 5.6× the linear head), with 18.3M, 52.9M, and 15.1M parameters ( 1.5× , 4.2× , and 1.2× the linear head). Parameter counts are rounded to 0.1M from the shown kernel dimensions; bias and normalization terms do not affect this precision.
Figure 12 : Training loss for the pixel-space video output heads. Raw T2V pixel-flow losses for the Wan-style upsample-conv decoder (F3-R02), T-S PixelShuffle decoder (F3-R03), and S-T PixelShuffle decoder (F3-R04). The left panel shows the complete trajectories on a logarithmic scale; the right panel enlarges the interval after 2K training steps. Curves are shown exactly as logged, without EMA or other smoothing.
Figure 13 : Lossless spatial and temporal boundary probes over 72 evaluation prompts. For each output head, we measure all 96 frames and all spatial locations of the pre-clamp float32 output at 176×320 or 320×176 . The spatial probe compares RGB edge intensity at 16-pixel patch boundaries with non-boundary edges; the temporal probe compares adjacent-frame differences at 4-frame tubelet boundaries with within-tubelet transitions. We normalize each prompt by its corresponding video-wide adjacent-sample mean and then aggregate the 72 evaluation prompts with equal weight. Bars and error bars report mean ± SEM across prompts ( 72 prompts per run; 288 videos total).
Figure 14 : Patch artifacts across pixel-space video generation. Each row shows one output-head design under the same standard evaluation prompt: a linear baseline (F3-R01), a Wan-style upsample-conv decoder (F3-R02), a T-S PixelShuffle decoder (F3-R03), and an S-T PixelShuffle decoder (F3-R04). The red box in the full frame defines a fixed crop. The four panels to its right show that same spatial location in chronological order at frames 0, 32, 64, and 95, spanning the complete 96-frame video. The displayed frames are lossless PNGs captured before video encoding; the corresponding pre-clamp float32 tensors are retained separately for numerical analysis.
Figure 15 : Pixel-space versus VAE-space early training dynamics (F4). Left: T2I velocity MSE measured in pixel and VAE latent spaces; the difference in loss magnitude alone does not establish relative learning speed or generation quality. The inset enlarges a representative pixel-space spike at 7.9K–8.2K steps. Right: pre-clip global gradient norm. Both panels use logarithmic y -axes and show unsmoothed 10-step records.
Figure 16 : Model-size scaling in Joint Stage 1. We compare the 1.7B (F5-R01) and 8B (F5-R02) models using the same training recipe with a global batch size of 256.
Figure 17 : Distributed compute scaling under a matched MidJourney recipe. We compare F6-R01 and F6-R02, which use the same 1.7B model, data mixture, p32 tokenization, pixel embedder and head, loss weights, and optimizer settings. Opaque curves are exponential moving averages over 25 logged points; faint curves show raw measurements.
Figure 18 : Multimodal context conditioning. Text, two reference images, and an input video enter the understanding expert, while noisy target-video tokens enter the generation expert. Both routes interact through shared multimodal self-attention.
Figure 19 : Dual-encoder and encoder-free visual conditioning. (a) Clean visual conditions enter only through the understanding expert, while noisy targets enter through the generation expert. (b) BAGEL uses dual ViT/VAE representations of a clean image [ 26 ] . SenseNova-U1 introduced this encoder-free conditioning design for images [ 32 ] ; PixelUMM extends it to image-to-video generation and video editing. Rows are queries and columns are keys; colored cells indicate visible attention and white cells are masked. Block sizes are schematic.
Optimization
Data and objectives
Per-step example ratio
Understanding branch
Loss weight †
1:10:30
Text only
1
Generation branch
Seq length ‡
60K / 32,768
Image understanding (I2T)
2
Learning rate
2×10−5
Time shift
1
Image generation (T2I)
2
LR scheduler
Constant
Und resolution (image)
Native
Video understanding (V2T)
2
Optimizer
Fused Adam
Und resolution (video)
4482
Video generation (T2V)
3
(β1,β2,ϵ)
(0.9,0.99,10−8)
Gen resolution (image)
2562
Image-to-video (I2V)
3
Table 2 : Multi-task fine-tuning recipe. The model is initialized from an intermediate checkpoint in Joint Stage 2. Video resolutions are per frame. Per-step example ratios specify relative numbers of examples consumed per step and need not sum to one.
Checkpoint
MMMU
MMStar
RWQA
SEED-I
AI2D
DocVQA
ChartQA
InfoVQA
TextVQA
OCRBench
MME
F7-R01
40.44
55.81
70.98
78.10
80.18
90.21
82.32
64.59
78.72
77.60
1752.67
F7-R02
41.44
55.43
69.28
72.69
79.40
90.35
81.64
63.40
78.99
77.30
1805.29
Δ
+1.00
−0.38
−1.70
−5.41
−0.78
+0.14
−0.68
−1.19
+0.27
−0.30
+52.62
Table 3 : Understanding after multi-task fine-tuning with visual conditioning. F7-R01: before multi-task fine-tuning; F7-R02: after multi-task fine-tuning; Δ=F7-R02−F7-R01 . Video: dense_mode ( ≤ 384 frames) and sparse_mode ( ≤ 96 frames), without subtitles for Video-MME. LVB denotes LongVideoBench; scores retain their original scales.
Figure 20 : Image-to-video conditioning through the understanding expert. The clean first-frame condition is followed by frames 0, 47, and 95 of the generated 96-frame clip, in which the woman turns toward the camera and gives a small wave on a seaside pier at sunset. The output is sampled with DPM-Solver for 50 steps, timestep shift 10, text CFG 6, image-reference scale 1, and no CFG renormalization.
Figure 21 : Video editing through understanding-expert conditioning. The source and edited rows show frames 0, 47, and 60. The instruction changes the background to a waterfall valley with cliffs and replaces the short bob with long wavy hair while retaining temporal consistency. Sampling uses DPM-Solver for 50 steps, timestep shift 10, text CFG 6, video-reference scale 1, and no CFG renormalization.
Interface
MVBench
Video-MME
LongVideoBench
LVBench
sparse_mode
71.15
57.89
59.39
41.38
dense_mode
71.22
57.67
58.71
41.38
Δ
+0.07
−0.22
−0.68
0.00
Table 4: Video understanding interfaces at F8-R01. Video-MME is evaluated without subtitles. Δ=dense_mode−sparse_mode .
Unified Models
VLM
Benchmark
PixelUMM 8B MoT
BAGEL 7B MoT
TUNA 7B+5B
TUNA-2 7B+5B
Qwen2.5-VL 7B
LLaVA-OV-1.5 8B
LLaVA-OV-2 8B
Qwen3-VL-Inst. 8B
NEO-ov 8B
MMMU
41.67
55.30
49.80
50.70
51.30
55.40
–
69.60
68.10
MMStar
53.99
–
61.20
–
62.50
67.70
64.30
70.90
67.30
RWQA
71.63
72.80
66.10
67.70
68.50
68.10
69.70
71.50
67.80
SEED-I
70.39
–
74.70
–
77.50
77.30
–
–
76.60
AI2D
80.12
89.20
79.30
79.60
82.60
84.20
84.30
85.70
85.40
Table 5 : Image understanding benchmarks. Unified models are grouped first. PixelUMM is evaluated using the official LMMS-Eval protocol (64,750 generations over 21 tasks). Dark and light green denote the best and second-best results, respectively.
Unified Models
VLM
Benchmark
PixelUMM 8B MoT
Show-o2 1.5B+0.5B
TUNA 1.5B+?
Lance 3B MoT
LLaVA-OV-2 8B
Qwen3-VL 8B
Keye-VL-1.5 8B
InternVL-3.5 8B
PLM 8B
LLaVA-OV-1.5 8B
MVBench
70.53
49.80
54.40
62.00
66.20
69.00
56.90
72.10
77.10
51.20
Video-MME (w/o sub.)
57.33
48.00
49.10
–
71.90
71.40
73.00
65.90
60.50
61.10
LongVideoBench
59.61
49.20
49.70
–
66.90
68.00
66.00
62.40
59.60
56.20
LVBench
40.41
–
27.40
–
55.50
58.00
42.80
46.70
44.50
40.10
Table 6 : Video understanding benchmarks. Unified models are grouped first. PixelUMM is evaluated in sparse_mode with up to 96 sampled frames. Video-MME is evaluated without subtitles. Published models retain their original evaluation protocols. Dark and light green denote the best and second-best results, respectively.
Model
Size
GenEval
DPG-Bench
1-Obj.
2-Obj.
Count
Colors
Position
Col. Attr.
Overall
Global
Entity
Attribute
Relation
Other
Overall
T2I Models
SD3-M
2B
0.99
0.94
0.72
0.89
0.33
0.60
0.74
87.90
91.01
89.96
80.70
88.68
84.08
FLUX.1 [dev] †
12B
0.98
0.93
0.75
0.93
0.68
0.65
0.82
82.10
89.50
88.70
91.10
89.40
84.00
LongCat-Image
6B
0.99
0.98
0.86
0.86
0.75
0.73
0.87
89.10
92.54
92.00
93.28
87.50
86.80
Qwen-Image
20B
0.99
0.92
0.89
0.88
0.76
0.77
0.87
91.32
91.56
92.02
94.31
92.73
88.32
Table 7 : Image generation benchmarks. Results on GenEval and DPG-Bench. All multimodal systems are grouped as Unified Models ; generation-only systems are grouped as T2I Models . “Col. Attr.” denotes color attribute, and † marks reported use of an LLM prompt rewriter for GenEval. Baseline scores are taken directly from the original papers; the absence of † does not confirm that prompt rewriting was not used. GenEval uses both the original prompts and BAGEL-long rewritten prompts ( † ), with 553 prompts and four samples per prompt. DPG-Bench follows the official 1,065-prompt protocol with four samples per prompt. Dark and light green denote the best and second-best results among Unified Models, respectively.
Model
Size
Quality Score
Semantic Score
Subj. Consist.
Bkg. Consist.
Temp. Flicker
Motion Smooth.
Dynamic Degree
Aesthetic Quality
Imaging Quality
Object Class
T2V Models
ModelScope
1.7B
78.05
66.54
89.87
95.29
98.28
95.79
66.39
52.06
58.57
82.25
LaVie
3B
78.78
70.31
91.41
97.47
98.30
96.38
49.72
54.94
61.90
91.82
Show-1
6B
80.42
72.98
95.53
98.02
99.12
98.24
44.44
57.35
58.66
93.07
AnimateDiff-V2
1.3B
82.90
69.75
95.30
97.68
98.75
97.76
40.83
67.16
70.10
90.90
VideoCrafter-2.0
1.4B
82.20
73.42
96.85
98.22
98.41
97.73
42.50
63.13
67.22
92.55
Table 8 : Video generation benchmarks on VBench Part 1. Generation-only systems are grouped as T2V Models ; multimodal systems are grouped as Unified Models . Baseline scores are taken directly from the original papers; the absence of † does not confirm that prompt rewriting was not used. UniVideo and PixelUMM both use VBench GPT-enhanced prompts. Dark and light green denote the best and second-best results among Unified Models, respectively.
Model
Size
Multi. Objects
Human Action
Color
Spatial Relation
Scene
Appear. Style
Temp. Style
Overall Consist.
Total Score ↑
T2V Models
ModelScope
1.7B
38.98
92.40
81.72
33.68
39.26
23.39
25.37
25.67
75.75
LaVie
3B
33.32
96.80
86.39
34.09
52.69
23.56
25.93
26.41
77.08
Show-1
6B
45.47
95.60
86.35
53.50
47.03
23.06
25.28
27.46
78.93
AnimateDiff-V2
1.3B
36.88
92.60
87.47
34.60
50.19
22.42
26.03
27.04
80.27
VideoCrafter-2.0
1.4B
40.66
95.00
92.92
35.86
55.29
25.13
25.84
28.23
80.44
Table 9 : Video generation benchmarks on VBench Part 2. Generation-only systems are grouped as T2V Models ; multimodal systems are grouped as Unified Models . Baseline scores are taken directly from the original papers; the absence of † does not confirm that prompt rewriting was not used. UniVideo and PixelUMM both use VBench GPT-enhanced prompts. Dark and light green denote the best and second-best results among Unified Models, respectively.
Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creating misalignment between the two tasks and preventing fully end-to-end optimization from raw pixels. We introduce Tuna-2, a native unified multimodal model that performs visual understanding and generation directly based on pixel embeddings. Tuna-2 drastically simplifies the model architecture by employing simple patch embedding layers to encode visual input, completely discarding the modular vision encoder designs such as the VAE or the representation encoder. Experiments show that Tuna-2 achieves state-of-the-art performance in multimodal benchmarks, demonstrating that unified pixel-space modelling can fully compete with latent-space approaches for high-quality image generation. Moreover, while the encoder-based variant converges faster in early pretraining, Tuna-2's encoder-free design achieves stronger multimodal understanding at scale, particularly on tasks requiring fine-grained visual perception. These results show that pretrained vision encoders are not necessary for multimodal modelling, and end-to-end pixel-space learning offers a scalable path toward stronger visual representations for both generation and perception.
Zhiheng Liu, Weiming Ren, Xiaoke Huang +12
Meta AI · University of Waterloo · The University of Hong Kong
Unified multimodal models (UMMs) aim to handle perception and generation in a single model. Yet existing UMMs still rely on a frozen, separately pretrained VAE for image generation, imposing a structural bottleneck. Naively removing it introduces a quality gap, as the model must learn both high-level structure and low-level details from raw pixels. In this paper, we propose Representation Forcing (RF), a technique that closes this gap by making representation prediction a native capability of the model. Concretely, RF forces the decoder to autoregressively predict visual representations as intermediate tokens before pixels; these tokens then stay in context to guide pixel diffusion within the same backbone. By turning representations from perception outputs into generation targets, RF eliminates the need for any external generative latent space. We find that RF benefits both understanding and generation. On image generation, our pixel-space model with RF matches state-of-the-art VAE-based unified models. On image understanding, pixel-space RF generally outperforms its VAE-based variant. Together, these results offer an effective step toward end-to-end, bottleneck-free UMMs.
Yuqing Wang, Zhijie Lin, Ceyuan Yang +10
University of Hong Kong · ByteDance Seed · The Chinese University of Hong Kong +2
We present Lance, a lightweight native unified model supporting multimodal understanding, generation, and editing for both images and videos. Rather than relying on model capacity scaling or text-image-dominant designs, Lance explores a practical paradigm for unified multimodal modeling via collaborative multi-task training. It is grounded in two core principles: unified context modeling and decoupled capability pathways. Specifically, Lance is trained from scratch and employs a dual-stream mixture-of-experts architecture on shared interleaved multimodal sequences, enabling joint context learning while decoupling the pathways for understanding and generation. We further introduce modality-aware rotary positional encoding to mitigate interference among heterogeneous visual tokens and boost cross-task alignment. During training, Lance adopts a staged multi-task training paradigm with capability-oriented objectives and adaptive data scheduling to strengthen both semantic comprehension and visual generation performance. Experimental results demonstrate that Lance substantially outperforms existing open-source unified models in image and video generation, while retaining strong multimodal understanding capabilities. The homepage is available at https://lance-project.github.io.