Language-aligned and self-supervised visual encoders offer complementary strengths in semantic abstraction and spatial detail. Harnessing this complementarity requires enriching local features while retaining distinctions between semantically related patches. We introduce PAIQ, a patch-aligned semantic injection framework that combines content-based cross-encoder matching with orthogonally constrained residual updates. Using DINOv3 patch features as the spatial base, PAIQ aggregates complementary SigLIP features through joint source allocation and injects the aggregate--base differences through a shared orthogonal transformation Q. This rotation adapts update directions while preserving residual norms and pairwise angles. For fixed projected features, we derive conditions for patch separability under similar semantic aggregates and show that rotation adds a nonnegative separation term over direct interpolation when the aggregate is shared. Only the projection and fusion parameters are trained; both visual encoders and the language model remain frozen, and fusion retains 196 visual tokens. Across diverse language backbones, PAIQ yields broad gains in judge-assessed correctness and reductions in hallucination severity over single-encoder interfaces on image description and visual question answering. On the 2B and 9B Qwen backbones, this compact interface outperforms the strongest evaluated fusion or token-compression baselines by about 2.9 correctness points on average.
Figures & tables
Figure 1: Complementary visual representations and the PAIQ fusion principle. DINOv3 and SigLIP both encode semantic and local information, but with different representational preferences. PAIQ takes the spatial features as a base and introduces language-aligned features through content matching and residual enrichment, producing a patch-wise fused representation.
Figure 2: PAIQ overview. Content matching aggregates SigLIP features for each DINOv3 patch, and a shared orthogonal transformation rotates the aggregate–base difference before adding it back. The two streams become 196 visual tokens; source allocation, update magnitude, and fixed-answer likelihood sensitivity describe how the representation is constructed and used.
Figure 3: Overview of PAIQ . (a) Features from two frozen visual encoders are fused into 196 tokens for a frozen language model. (b) Content matching and finite-step row–column normalization aggregate SigLIP features at DINOv3 output positions. (c) A shared Cayley-parameterized orthogonal transform rotates the aggregate–base residual; with row-stacked tokens, Z=D+λ(Sˉ−D)Q⊤ . (d) Aggregation weights πj , relative update norms ρj , and fixed-answer log-likelihood differences R(B) diagnose source mixing, update size, and the effect of local residual removal, respectively. Computing R(B) requires additional teacher-forced evaluations.
Method
FLUX-Reason
Caption
SEED
A-OKVQA
Acc. ↑
Hall. ↓
Acc. ↑
Hall. ↓
Acc. ↑
Hall. ↓
Acc. ↑
Hall. ↓
Table 1: Comparison across four language backbones and four tasks. Accuracy ( ↑ ) and Hallucination ( ↓ ) are mean Judge scores. The DINOv3, SigLIP, and PAIQ comparisons use 1,000 examples per setting. Bold and underlining mark the best and second-best displayed values, respectively, within each backbone and task metric. Colored arrows beside each baseline give the direction and size of PAIQ’s difference from that baseline in score points; green favors PAIQ and red favors the baseline.
Existing vision-language model (VLM)-based AI-generated image quality assessment (AIGIQA) methods suffer from a fundamental semantic-distortion dimensional conflict: monolithic representations optimized for semantic discrimination inherently entangle compositional understanding with low-level perceptual sensitivity, rendering them blind to fine-grained quality degradations. We introduce MST-CLIPIQA, a multi-scale two-stream framework that achieves hierarchical vision-language alignment through explicit representational decoupling. Our architecture leverages dual CLIP encoders with complementary patch granularities: coarse-grained streams capture global semantic coherence while fine-grained streams preserve textural signatures and artifact patterns. An information bottleneck-inspired gated fusion mechanism performs adaptive cross-scale distillation, with optional cross-attention enabling prompt-anchored correspondence evaluation when generation prompts are available. Extensive experiments across five benchmarks establish new state-of-the-art results, achieving average improvements of 1.11 percent SRCC on quality and 2.35 percent SRCC on text-image correspondence prediction, while maintaining efficiency with only 0.8M trainable parameters. Our project is available at https://github.com/YMlinfeng/MST-CLIPIQA.
A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, representing images as discrete signals in the same way as text inevitably introduces severe information loss. Existing work struggles to balance low-level details and high-level semantics in discrete representations: reconstruction-oriented representations often lack semantic information, whereas semantically stronger features typically suffer from severe loss of detail. We present ViQ, a Visual Quantized Representations framework, which is designed to balance semantics and details in discrete representations while supporting inputs at native resolutions, thereby enabling it to serve as a unified and general discrete representation for arbitrary visual inputs. Our approach structures quantization learning into two stages: text-aligned pre-training and feature discretization. With text-aligned pre-training, we enhance the visual encoder semantic-rich supervision from the pretrained language model and enable it to process native-resolution visual inputs. During discretization, we propose a proximal representation learning strategy to progressively compact the feature space, along with a position-aware head-wise quantization mechanism that enables flexible processing of arbitrary resolutions. Extensive experiments on multimodal tasks demonstrate that ViQ achieves competitive performance compared to state-of-the-art multimodal vision encoders with continuous and high-dimensional visual features, while maintaining high precision in low-level reconstruction. We also show that multimodal training with visual quantized representations largely improves efficiency, yielding up to 20%-70% acceleration with different base LLMs and training recipes.
Xumin Yu, Zuyan Liu, Zhenyu Yang +5
Tencent HY Vision Team · Tsinghua University · Institute of Automation, CAS +1
Representation autoencoders that reuse frozen pretrained vision encoders as visual tokenizers have achieved strong reconstruction and generation quality. However, existing methods universally extract features from only the last encoder layer, discarding the rich hierarchical information distributed across intermediate layers. We show that low-level visual details survive in the last layer merely as attenuated residuals after multiple layers of semantic abstraction, and that explicitly fusing multi-layer features can substantially recover this lost information. We propose DRoRAE (Depth-Routed Representation AutoEncoder), a lightweight fusion module that adaptively aggregates all encoder layers via energy-constrained routing and incremental correction, producing an enriched latent compatible with a frozen pretrained decoder. A three-phase decoupled training strategy first learns the fusion under the implicit distributional constraint of the frozen decoder, then fine-tunes the decoder to fully exploit the enriched representation. On ImageNet-256, DRoRAE reduces rFID from 0.57 to 0.29 and improves generation FID from 1.74 to 1.65 (with AutoGuidance), with gains also transferring to text-to-image synthesis. Furthermore, we uncover a log-linear scaling law (R2=0.86) between fusion capacity and reconstruction quality, identifying \textit{representation richness} as a new, predictably scalable dimension for visual tokenizers analogous to vocabulary size in NLP.
Xuanyu Zhu, Yan Bai, Yang Shi +4
Peking University · Meituan Inc · Tsinghua University +1