From Pixels, Without Pre-training: Joint Generative and Self-Supervised Representation Learning in One Model
Authors: Vicente Balmaseda, Ching-Long Lin, Tianbao Yang
Organizations: Department of Computer Science and Engineering Texas A&M University College Station, TX, USA · Mechanical Engineering University of Iowa Iowa City, IA, USA
Strong image generation models are conditioned on class labels, aligned to frozen pretrained encoders, or built on separately trained autoencoders. While effective, generation then depends on supervision or pretraining: labels must be annotated, and encoders or autoencoders pretrained for the target domain. We study joint generative and self-supervised representation learning in a single model, enabling self-conditioned generation without labels or pretrained models. This is challenging because the objectives are mismatched: contrastive learning consumes clean augmented views and favors coarse, invariant semantics, while flow matching consumes noisy images and must preserve the fine detail and spatial layout that contrastive learning discards. We propose SCION (Self-conditioned Generation on Self-supervised representation), whose core is a single pixel-space encoder conditioned on the flow timestep and an embedding. For representation learning, this conditioning embedding is a learned global vector shared across images, with the encoder's [CLS] token yielding the semantic representation trained by the contrastive loss. For generative training, the conditioning embedding is the image's own [CLS] representation, while patch tokens pass through a decoder to predict the image. To sample without a reference image at inference, we jointly learn a prior over the embedding. Gradient-norm balancing and stop-gradient mechanisms enable joint optimization in one run. SCION is self-supervised and self-contained, with no labels or pretrained models. On ImageNet 256x256, with the JiT-B recipe and no representation guidance, SCION reaches 8.92 FID, surpassing class-unconditional iREPA, which aligns to pretrained DINOv2 (46.44), and RCG, which conditions on it (14.27). With JiT-L, SCION achieves 5.89 FID without guidance and 3.47 with representation guidance, outperforming RCG with the ADM recipe (6.24).
Figures & tables
four S requirements
capabilities
Method
Self-cont.
Self-sup.
Single-mod.
Single-stage
Repr.
Synth.
Pixel-space generation
JiT ( Li and He, 2026 )
✓
✗
–
✓
✗
✓
JiT + REPA ( Yu et al., 2025 )
✗
✗
✗
✗
✗
✓
JiT + Dispersive ( Wang and He, 2025 )
✓
✗
✓
✓
✗
✓
Semantic conditions
Table 1 : Comparison with related works under the four S requirements. Extended discussion on related works is presented in Appendix A . Self-contained means no pretrained component; self-supervised means no label supervision anywhere in the pipeline; single-model means that the representation is produced by the generator’s own network; and single-stage means that all components are learned jointly from scratch. Capabilities report a learned and evaluated Repr esentation and unconditional Synth esis with no source image to encode. All pipelines are as published, except REPA and Dispersive, which we combine with the pixel-space JiT recipe.
Table 2Table 3
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
four S requirements
capabilities
Method
Self-cont.
Self-sup.
Single-mod.
Single-stage
Repr.
Synth.
Generative recipes
DiT / SiT ( Peebles and Xie, 2023 ; Ma et al., 2024 )
✗
✗
–
✗
✗
✓
JiT ( Li and He, 2026 )
✓
✗
–
✓
✗
✓
JiT + REPA ( Yu et al., 2025 )
✗
✗
✗
✗
✗
✓
JiT + Dispersive ( Wang and He, 2025 )
✓
✗
✓
✓
✗
✓
Appendix
Table 6 : Comparison under the four S requirements , extending Table 1 to the comparable systems tabulated in this appendix. Self-contained means no pretrained component, including a latent autoencoder or tokenizer; self-supervised means no label supervision anywhere in the pipeline, including through pretrained components; single-model means that the representation is produced by the generator’s own network; and single-stage means that all components are learned jointly from scratch. Capabilities report a learned and evaluated Repr esentation and unconditional Synth esis with no source image to encode. All pipelines are as published, except REPA and Dispersive, which we combine with the JiT recipe removing the latent model dependency.
generation
representation
Rung
views
λCL
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
Acc. ↑
transfer-8 ↑
0. flow matching alone (JiT-B/16 unconditional)
1+0
—
61.17
9.05
17.53
.384
.559
10.65
26.19
1. add contrastive learning: which views?
1i
3+0
0.025
59.08
9.25
17.96
.385
.571
33.74
50.93
1ii
0+3
0.025
70.46
16.89
15.18
.237
.625
65.30
77.36
Appendix
Table 7 : Detailed ablation ladder. All rows use the JiT-B/16 architecture under the same recipe for 200 epochs and unconditional setting. Rungs 0 to 3 isolate the joint-training recipe before representation self-conditioning and the learned prior. Rung 4 adds both components, and rung 5 adds the conditioned-branch consistency loss LCBC to complete SCION. Rung 0 is our unconditional flow-matching reproduction. views gives the number of original and augmented inputs, respectively; bal. denotes gradient-norm balancing, which rescales the contrastive weight online.
Acc. ↑
transfer-8 ↑
Method
views
200
400
600
200
400
600
CL only
2 augmented
70.48
73.61
74.84
77.04
78.88
77.09
CL only
1 original +2 augmented
71.76
74.15
74.98
77.68
78.66
76.61
SCION (ours)
1 original +2 augmented
71.32
72.35
72.61
79.73
79.89
79.87
Appendix
Table 8 : Representation quality throughout training. We compare SCION with capacity-matched contrastive-only controls trained from scratch. All rows use the same 8 -layer encoder; the controls have no generative objective and are independently tuned for contrastive learning. During inference, the models use approximately the same number of parameters, 57.2 M for SCION and 57.1 M for the controls.
Method
FID ↓
IS ↑
Acc. ↑
transfer-8 ↑
JiT
61.17
17.5
10.65
26.19
JiT + Dispersive
58.68
18.0
8.88
24.06
JiT + REPA
53.07
23.0
63.94
73.59
JiT + iREPA
48.43
24.8
65.81
76.37
RCG (JiT; 400 total)
14.27
93.4
6.54
30.28
SCION (ours)
10.10
101.9
71.32
79.73
Appendix
Table 9: Controlled class-unconditional comparison at 200 epochs. All methods use the same JiT-B/16 flow-matching recipe on ImageNet ( 256×256 ). The RCG row is a reference: its 200-epoch generator is preceded by a frozen 200-epoch prior stage, for 400 total epoch-equivalents.
dependencies
generation
Method
Lat.
Ext.
Lab.
External data
Multi-stage
Params
FID ↓
IS ↑
latent-space
DiT-B/2 ∗ ( Peebles and Xie, 2023 )
SD-VAE
—
✓
✓
✓
179M
69.3
—
DiT-XL/2 ∗ ( Peebles and Xie, 2023 )
SD-VAE
—
✓
✓
✓
724M
44.6
—
ReDi (DiT-B/2) ( Kouzelis et al., 2025 )
SD-VAE
DINOv2
✓
✓
✓
179M
51.7
—
ReDi (DiT-B/2) (RG) ( Kouzelis et al., 2025 )
SD-VAE
DINOv2
✓
✓
✓
179M
47.3
—
Appendix
Table 10 : Unconditional ImageNet-1K ( Deng et al., 2009 ) 256×256 generation systems and their dependencies. FID and IS as reported by the corresponding work. Lat. names the pretrained VQGAN tokenizer ( Esser et al., 2021 ) or latent autoencoder, SD-VAE a or an LDM autoencoder (LDM-AE) ( Rombach et al., 2022 ) , and Ext. any other pretrained model. Lab. marks label supervision anywhere in the pipeline, including the supervised perceptual loss ( Zhang et al., 2018 ) inherited by VQGAN and SD-VAE. External data marks a dependency trained on data beyond ImageNet-1K. In particular, SD-VAE uses OpenImages ( Krasin et al., 2017 ) and, per its model card, is fine-tuned on LAION-Aesthetics ( Schuhmann et al., 2022 ) and an unreleased LAION-Humans subset, DINOv2 uses LVD-142M ( Oquab et al., 2024 ) , and CLIP uses WIT-400M ( Radford et al., 2021 ) . Multi-stage marks more than one separately trained component. Params count all inference-time models, including a latent autoencoder or tokenizer where applicable. † retrieves real images at sampling. ∗ denotes DiT values reported by ReDi and ‡ values reported by RCG. (CFG) and (RG) denote classifier-free and representation guidance, respectively. Training budgets are not matched. SCION (B/16) is reported at 800 epochs and SCION (L/16) at 400 epochs.
conditioning
reference
Prec. ↑
Rec. ↑
prior eω
conditioning set
.718
.524
disjoint set
.722
.510
oracle repr. †
conditioning set
.645
.828
disjoint set
.626
.664
Appendix
Table 11 : Oracle conditioning against the learned prior. SCION (B/16) at 400 epochs, unguided, 50 K samples per arm. The oracle arm ( † , semi-parametric) replaces the prior sample with the representation of a real training image, one sample per representation. Each arm is scored against 10 K images from the conditioning set and from a disjoint set of training images. With the same generator, real representations give higher recall and prior samples higher precision.
SCION (B/16)
SCION (L/16)
Architecture
Depth: total / enc. / dec.
12, 8, 4
24, 16, 8
Hidden dim
768
1024
Heads
12
16
In-context start block
4
8
[CLS] normalization
identity
BatchNorm
Appendix
Table 12 : Main hyperparameters for the reported SCION (B/16) and SCION (L/16) ImageNet-1K 256×256 runs. Rows spanning both columns are identical at the two scales. The baselines share this recipe apart from the objective and, for REPA and iREPA, the 4/8 encoder–decoder split, following their published split; see Appendix E . Guidance follows Ho and Salimans (2021) on the interval of Kynkäänniemi et al. (2024) .
Component
Symbol
Input
Output
Shared encoder
fθ
(xt∣t,c)∈R3×H×W×[0,1]×Rd
(e,h)∈Rd×RT×d
Noisy image, timestep and conditioning ↦ global embedding e and patch tokens h .
Generative decoder
gϕ
(h∣t,c)∈RT×d×[0,1]×Rd
x∈R3×H×W
Patch tokens, timestep and conditioning ↦ predicted clean image.
Contrastive projection
hψ
e∈Rd
z∈Rdz
Global embedding ↦ the space where LCL is optimized.
Appendix
Table 13 : Components of the shared model. T is the number of patch tokens, d the hidden width, dz the contrastive projection dimension, and c the conditioning vector of Section 3.1 .
Config
Value
Input / output width
representation dimension d
Hidden width C
808
Residual blocks N
14
Block SwiGLU width
2155
Timestep features
sinusoidal, width 256
Timestep embedding
808, added inside each block
Appendix
Table 14 : Prior architecture. The residual-MLP backbone of eω , shared by both scales; its input and output width follows the backbone, so the parameter count is the B/16 one.
Figure 4 : Preserving semantics while varying image details. Each pair shows a real reference and a generation conditioned on the representation extracted by SCION itself. Recognizable subjects and scene types persist across changes in pose, color, markings, background, and structure. The generator starts from image noise and bypasses the learned prior. All examples use SCION (JiT-L/16) at 400 epochs with representation guidance.
Figure 5 : Interpolation in the learned conditioning space. Real references appear at the outer edges. All five interior images are generated from interpolated representations, with α=0 and 1 using the endpoint representations directly. The first five rows connect images from the same ImageNet class, while the last four connect different classes. Image noise is fixed within each row. All paths use spherical interpolation, and SCION (JiT-L/16) at 400 epochs with representation guidance.
Figure 6 : Flower-bud representation interpolations. These selected paths illustrate representation changes rather than a recorded or predicted growth sequence. Flower names for the right-hand references follow their ImageNet class labels. Real photographs flank six generated images at α=0,0.2,0.4,0.6,0.8,1 , including generated endpoint reconstructions. All paths use spherical interpolation of the representations, fixed noise within each row, and SCION (JiT-L/16) at 400 epochs with representation guidance.