We present In-Token Learning, an image restoration framework that adapts a pretrained diffusion transformer using conditional rectified flow matching. Clean targets paired with degraded inputs supervise transport from Gaussian noise to restored images. Spatially aligned degraded-image tokens are fused with evolving latent tokens along the channel dimension, preserving the image-token count at a given resolution. Direct Low-Quality Guidance (DLG) combines frozen degraded-image embeddings with a fixed task prompt through the native conditioning pathway, without a trainable ControlNet-style branch or image captioning. We evaluate super-resolution and denoising on DIV2K, LSDIR, FFHQ, RealLQ250, and RealPhoto60, and automatic colorization on DIV2K and LSDIR. The tasks use separately trained checkpoints under the same framework. Results show competitive fidelity and perceptual quality under the evaluated protocols, with weaker generalization on RealLQ250. We report full-image QHD (2560×1440) inference and a tiled 12K restoration demonstration of Along the River During the Qingming Festival. Attention cost still increases with resolution. This technical report preserves the early broader study underlying Fill2SR, which subsequently developed the real-world super-resolution direction.
Figures & tables
Figure 1: In-Token Learning overview. Paired clean targets supervise a conditional velocity field sampled from noise toward restored images. In-token alignment preserves the spatial token count at a fixed resolution; DLG supplies task and image embeddings. SR/denoising and colorization use separately trained checkpoints. Full-image QHD and tiled 4K/8K/12K inference are distinct settings.
Figure 3: Qualitative comparisons on DIV2K-Val (synthetic). Instead of denoising from a degraded latent, our method generates from pure noise using a conditional velocity field guided by in-token alignment and DLG. The selected crops illustrate differences in eye shape, gaze direction, thin edges, and textures; they do not establish a universal fidelity ranking.
Figure 5: Qualitative comparisons on RealLQ250 (real-world). Left: our full-resolution results with crop locations marked (red boxes). Right: low-quality input and outputs from Real-ESRGAN, DiffBIR, FaithDiff, SeeSR, and our method on the corresponding crops. These selected examples complement the aggregate scores and failure cases; real-world performance is not uniformly best.
Figure 7: Qualitative comparisons on DIV2K-Val (colorization). Left: our full-resolution result with crop locations marked (red boxes). Right: grayscale input and outputs from InstColor, ColorFormer, BigColor, DDColor, and our method on the corresponding crops. Our method produces most material-consistent colors while preserving luminance and fine details.
Figure 9: Ablations for SR & denoising. Top: DLG variants show that combining text and LQ embeddings (Text & LQ Emb.) yields the best structure and detail. Middle: starting from pure noise performs better than denoising-initialization at varying strengths (0.9/0.6/0.3). Bottom: removing In-Token alignment (w/o In-Token) degrades both structure and perceptual quality.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11: Qualitative comparison under different inference precisions (qfloat8 vs. bf16). Each triplet shows the low-quality input, our model inferred with qfloat8 , and with bf16 . At QHD (1440p) inference, qfloat8 lowers the peak VRAM from 30 to 17 GB, but under challenging settings, it more often suppresses fine structures and introduces spurious “dirty” textures.
Figure 14: Qualitative comparisons of different degradation level on LSDIR-Val. Selected examples across D1–D3 illustrate changes in structure, texture, and residual noise as degradation increases. Visual differences complement the per-level numerical results rather than establish a universal fidelity ranking.
Figure 16: Qualitative comparisons on DIV2K-Val (D3). Selected examples illustrate structural and texture differences under severe degradation. The first example retains the duplicated SeeSR crop from the original comparison; no SUPIR crop is shown for that example. All original comparison image content is preserved.
Figure 18: Qualitative comparisons on RealPhoto60 (2 × ). Left: our full-resolution result with crop locations marked (red boxes). Right: low-quality input and outputs from DiffBIR, SUPIR, FaithDiff, SeeSR, and ours on the corresponding crops.
Figure 20: Failure cases on RealLQ250. These examples show residual degradation and lost details, consistent with limited generalization beyond the synthetic training recipe. Closeness to the low-quality input or lower sharpness does not certify correct geometry, reliability, or absence of hallucination. All generative restorations require caution when interpreting recovered content, especially historical artworks.
Figure 22: Sensitivity to input preprocessing on RealLQ250. Gaussian blur with a standard deviation equal to 1% of the image width changes the outputs in these selected examples. The corresponding no-reference scores improve, but original-detail fidelity cannot be established from these unpaired real-world images.
Figure 24: Qualitative comparisons on LSDIR (colorization). Left: our full-resolution result with crop locations marked (red boxes). Right: grayscale input and outputs from InstColor, ColorFormer, BigColor, DDColor, and ours on the corresponding crops.
College of Computer Science, Nankai University, Tianjin, China · Faculty of Information Science and Computing, University of Macau, Macau, China · CSIRO Data61, Australia +1
School of Artificial Intelligence, Shanghai Jiao Tong University, Shanghai, China · Alibaba Group, Beijing, China · School of Advanced Technology, Xi’an Jiaotong-Liverpool University, Suzhou, China +1