Visual Representation Learning

Latest papers 146

All topics
CardsList
  1. Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension

    Jul 6, 2026Lian Xu, Mohammed Bennamoun, Farid Boussaid +3Visual Representation LearningVision-Language Grounding

  2. Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

    Jul 3, 2026Chaofan Gan, Zicheng Zhao, Yuanpeng Tu +6Visual Representation LearningDiffusion Transformer

  3. Transformer Geometry Observatory TGO-II: Representational Similarity Observatory

    Jul 2, 2026Kaustubh Kapil, Kishor P. UplaIntrinsic DimensionalityRepresentational Similarity Analysis

  4. LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives

    Jul 1, 2026Lukas Kuhn, Giuseppe Serra, Randall Balestriero +1Visual Representation LearningVision-Language Pretraining

  5. Identifying Latent Concepts and Structures for Generalized Category Discovery

    Jul 1, 2026Boyang Dai, Chaoqi Chen, Yizhou YuRepresentation LearningVisual Representation Learning

  6. Cross-Resolution Semantic Transfer for Robust Text-to-Image Retrieval in Low-Resolution Surveillance

    Jun 29, 2026Wenjie Qian, Bin Yang, Xiao Wang +4Visual Representation LearningCross-Modal Retrieval

  7. ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

    Jun 25, 2026Xumin Yu, Zuyan Liu, Zhenyu Yang +5Visual Representation LearningMultimodal Pretraining

  8. Meta-learning as a principle for human-like visual representations

    Jun 24, 2026Can Demircan, Marcel Binz, Alireza Modirshanechi +1Visual Representation LearningMeta-Learning

  9. Semantic Allocation in Ordered Bottlenecks: Predictive Residual Inference for Visual Representation Learning

    Jun 23, 2026Erik Ayari, Manuel Traub, Martin V. ButzRepresentation LearningVisual Representation Learning

  10. Learning Diachronic Representations of Ancient Greek Letterforms

    Jun 23, 2026John Pavlopoulos, Spyros Barbakos, Lavinia Ferretti +7Representation LearningVisual Representation Learning

  11. Learning Entropy Signature for Image Representation and Classification

    Jun 21, 2026Jan Glaser, Ivo Bukovsky, Noriyasu Homma +1Visual Representation LearningImage Classification

  12. A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world

    Jun 19, 2026Steeven Janny, Leonid Antsfeld, Christian WolfVisual Representation LearningRobot Navigation

  13. LARE: Low-Attention Region Encoding for Text-Image Retrieval

    Jun 17, 2026Abdulmalik Alquwayfili, Faisal Almeshal, Jumanah Almajnouni +8Visual Representation LearningFine-Grained Image Retrieval

  14. Contrastive Action-Image Pre-training for Visuomotor Control

    Jun 15, 2026Yuvan Sharma, Dantong Niu, Anirudh Pai +16Contrastive LearningVisual Representation Learning

  15. Hierarchical Fine-Grained Aerial Object Detection

    Jun 15, 2026Yan Zhang, Fang Xu, Wen Yang +1Visual Representation LearningHierarchical Representation Learning

  16. Analyzing Visual Aircraft Representations with Sparse Autoencoders

    Jun 13, 2026Deepshik SharmaVisual Representation LearningSparse Autoencoders

  17. Rethinking Implicit Spatial Representation in Visuomotor Policy Learning

    Jun 13, 2026Xiangyu Chen, Yuxuan Hu, Chuhao Zhou +1Visual Representation LearningVisuomotor Policy Learning

  18. Beyond Scalar Distances: Semantic Attribute Gradients from Frozen MLLMs for Visual Embeddings

    Jun 13, 2026Shubhang Bhatnagar, Dheeraj Baiju, Narendra AhujaVisual Representation LearningFine-Grained Image Retrieval

  19. RATS! Patches Talk Through Registers: Emergent Parts in Register Attention Transformers

    Jun 12, 2026Timing Yang, Predrag Neskovic, Jansen Seheult +4Visual AttentionVisual Representation Learning

  20. IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder

    Jun 9, 2026Yitong Chen, Zijie Diao, Junke Wang +5Visual Representation LearningImage Tokenization

  21. Lighting-Aware Representation Learning under Controllable Lighting Variation

    Jun 5, 2026Lizhen Zhu, Charantej Reddy Pochimireddy, James Z Wang +1Contrastive LearningVisual Representation Learning

  22. Formal Concept Lattices are Good Semantic Scaffolds for Concept-Based Learning

    Jun 3, 2026Deepika SN Vemuri, Sayanta Adhikari, Ankit Saha +2Visual Representation LearningInterpretable ML

  23. Stateful Visual Encoders for Vision-Language Models

    Jun 3, 2026Zirui Wang, Junwei Yu, Adam Yala +3Vision-Language ModelsVisual Representation Learning

  24. Formalizing the Binding Problem

    Jun 2, 2026Lianghuan Huang, Yihao Li, Saeed Salehi +3Visual Representation LearningVision Transformer

  25. PRISM: Progressive Reasoning through Iterative Slot Memory for Vision

    May 29, 2026Ziyu Wang, Shuangpeng Han, Mengmi ZhangVisual Representation LearningVisual Reasoning

  26. DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

    May 28, 2026Jusuk Lee, Seungjae Lee, Jonghun Shin +6Cross-Modal LearningCross-Modal Representation Learning

  27. Deep Psychovisual Image Representations

    May 28, 2026Wendi Ma, Aryaman Sharma, Wei Dai +1Visual Representation LearningFrequency-Domain Feature Learning

  28. Misalignment Between Backpropagation and the Hierarchy of Brain Responses to Images

    May 27, 2026Joséphine Raugel, Maximilian Seitzer, Marc Szafraniec +6BackpropagationVisual Representation Learning

  29. Structure over Pixels: Learning Variable-Length Visual Programs

    May 26, 2026Piotr Wyrwiński, Kacper Dobek, Krzysztof KrawiecVisual Representation LearningImage Tokenization

  30. Uncertainty-DTW for Sequences and Visual Tokens

    May 24, 2026Lei Wang, Syuan-Hao Li, Yongsheng Gao +1Dynamic Time WarpingVisual Representation Learning

  31. Not Too Generative, Not Too Discriminative: The Human Alignment Sweet Spot

    May 22, 2026Jorge Chang Ortega, Bastien Le Lan, Thomas Serre +1Visual Representation LearningEnergy-Based Models

  32. TextTeacher: What Can Language Teach About Images?

    May 21, 2026Tobias Christian Nauen, Stanislav Frolov, Brian Bernhard Moser +3Cross-Modal Knowledge DistillationVisual Representation Learning

  33. RiT: Vanilla Diffusion Transformers Suffice in Representation Space

    May 21, 2026Le Zhang, Ning Mang, Aishwarya AgrawalFlow MatchingVisual Representation Learning

  34. Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts

    May 20, 2026Gene Tangtartharakul, Katherine R. StorrsVisual Representation LearningMixture of Experts

  35. Capability ≠\neq Interpretability: Human Interpretability of Vision Foundation Models

    May 19, 2026Julien Colin, Lore Goetschalckx, Nuria Oliver +1Visual Representation LearningVision Foundation Models

  36. PEIRA: Learning Predictive Encoders through Inter-View Regressor Alignment

    May 17, 2026Michael Arbel, Basile Terver, Jean PonceUnsupervised LearningVisual Representation Learning

  37. Characterizing the visual representation of objects from the child's view

    May 14, 2026Jane Yang, Tarun Sepuri, Alvin Wei Ming Tan +3Representation LearningVisual Representation Learning

  38. Rethinking the Good Enough Embedding for Easy Few-Shot Learning

    May 13, 2026Michael Karnes, Alper YilmazVisual Representation LearningFew-Shot Learning

  39. Characterizing Universal Object Representations Across Vision Models

    May 13, 2026Florian P. Mahner, Johannes Roth, Ka Chun Lam +3Visual Representation LearningRepresentation Alignment

  40. Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores

    May 12, 2026Alan Z. Song, Yinjie Chen, Mu Nan +3Visual Representation LearningVision Transformer

  41. WorldComp2D: Spatio-semantic Representations of Object Identity and Location from Local Views

    May 12, 2026SeongMin Jin, Doo Seok JeongRepresentation LearningVisual Representation Learning

  42. Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery

    May 12, 2026Chi-Nguyen Tran, Dao Sy Duy Minh, Huynh Trung Kiet +3Cross-View Geo-LocalizationVisual Representation Learning

  43. FeatMap: Understanding image manipulation in the feature space and its implications for feature space geometry

    May 11, 2026Elias B. Krey, Nils Neukirch, Nils StrodthoffNeural Representation GeometryVisual Representation Learning

  44. Learning to Align Generative Appearance Priors for Fine-grained Image Retrieval

    May 11, 2026Shijie Wang, Yadan Luo, Zijian Wang +2Visual Representation LearningFine-Grained Image Retrieval

  45. Lost or Hidden? A Concept-Level Forgetting in Supervised Continual Learning

    May 10, 2026Katarzyna Filus, Kamil Faber, Roberto Corizzo +1Visual Representation LearningContinual Learning

  46. SEMASIA: A Large-Scale Dataset of Semantically Structured Latent Representations

    May 10, 2026Mario Edoardo Pandolfo, Enrico Grimaldi, Lorenzo Marinucci +4Neural Representation GeometryVisual Representation Learning

  47. 3D MRI Image Pretraining via Controllable 2D Slice Navigation Task

    May 7, 2026Yu Wang, Qingchao ChenVisual Representation LearningSelf-Supervised Pre-Training

  48. Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search

    May 7, 2026Faisal Aljehrai, Mohammed A. Alkhrashi, Alreem Almuhrij +6Visual Representation LearningInformation Retrieval

  49. CRISP: Compositional Relations as Invariant Structural Priors for Domain Generalization

    May 7, 2026Dat Nguyen, Duc-Duy NguyenRepresentation LearningVisual Representation Learning

  50. MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality

    May 7, 2026Panqi Yang, Haodong Jing, Jiahao Chao +5Visual Representation LearningImage Tokenization

  51. Exploring Clustering Capability of Inpainting Model Embeddings for Pattern-based Individual Identification

    May 6, 2026Jens van Bijsterveld, Daniele Avitabile, Fons J. Verbeek +1Biometric IdentificationVisual Representation Learning