Cross-Modal Representation Learning

Momentum

14 papers in the last four weeks, up 100% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 136

All topics
CardsList
  1. 3D Masked Autoencoders are Robust Learners of Volumetric and Multimodal Cellular Representations for Microscopy

    Jun 22, 2026Amirhossein Kardoost, Lion Gleiter, Tingying Peng +1Cross-Modal Representation LearningMasked Autoencoders

  2. Retrieval-Augmented Multimodal Learning for Enzyme-Substrate Interaction Prediction Under Low-Homology Shift

    Jun 22, 2026Chen Liu, Bingxin Zhou, Xinyuan Wang +3Cross-Modal Representation LearningDomain Generalization

  3. T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition

    Jun 19, 2026Di Yang, Mahmoud Ali, Quan Kong +2Skeleton-Based Action RecognitionRepresentation Learning

  4. MV-WAM: Manifold-Aware World Action Model with Value Augmentation

    Jun 19, 2026Jintao Chen, Peidong Jia, Qingpo Wuwu +13RL ControlCross-Modal Representation Learning

  5. FUSE: Frequency-domain Unification and Spectral Energy Alignment for Multi-modal Object Re-Identification

    Jun 18, 2026Xuanhao Qi, Tom H. Luan, Yukang Zhang +4Cross-Modal Representation LearningMultimodal Robustness

  6. TactSpace: Learning a Physics-enriched Shared Latent Space for Tactile Sim-to-Real Transfer

    Jun 17, 2026Arunim Joarder, Arjun Bhardwaj, René Zurbrügg +6Cross-Modal Representation LearningTactile Sensing

  7. CAP: Towards PPG Universal Representation Learning with Patient-level Supervision

    Jun 13, 2026Chenyang He, Xinyi Shao, Shun Huang +4Representation LearningCross-Modal Representation Learning

  8. EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP--OCT Pretraining

    Jun 13, 2026Zhuo Deng, Ruiheng Zhang, Ziheng Zhang +23Medical Imaging Foundation ModelsCross-Modal Representation Learning

  9. HiST: A Hierarchical Sparse Transformer for Cross-Modal Spatial Transcriptomics Modeling

    Jun 12, 2026Weiyi Wu, Xinwen Xu, Xingjian Diao +4Computational PathologySpatial Transcriptomics

  10. When to Align, When to Predict: A Phase Diagram for Multimodal Learning

    Jun 9, 2026Ilay Kamai, Hugues Van Assel, Aviv Regev +2Cross-Modal Representation LearningCross-Modal Alignment

  11. AnyMod-LLVE: Low-Light Video Enhancement with Modality-Agnostic Inference

    Jun 9, 2026Hangfeng Liang, Yutao Hu, Yanhan Hu +3Missing-Modality LearningCross-Modal Representation Learning

  12. COMBINER: Composed Image Retrieval Guided by Attribute-based Neighbor Relations

    Jun 3, 2026Zixu Li, Yupeng Hu, Zhiwei Chen +3Cross-Modal Representation LearningCross-Modal Alignment

  13. Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation Models

    Jun 3, 2026Shuwen Yu, Zhanxuan Hu, Yi Zhao +2Cross-Modal Representation LearningVision Foundation Model Adaptation

  14. Building The Ph(ysical)AI Layer Of Machine Intelligence

    Jun 2, 2026Ulbert Jose Botero, Liam Smith, Brooks Olney +5Cross-Modal LearningCross-Modal Representation Learning

  15. SkelHCC: A Hyperbolic CLIP-Driven Cache Adaptation Framework for Skeleton-based One-Shot Action Recognition

    Jun 2, 2026Yanan Liu, Anqi Zhu, Jingmin Zhu +6Skeleton-Based Action RecognitionCross-Modal Representation Learning

  16. Channel-Oriented Design for EEG-to-Music Reconstruction

    Jun 2, 2026Jiaxin Qing, Junwei Lu, Lexin LiElectroencephalographyCross-Modal Representation Learning

  17. Spatial Representation Learning Beyond Pixels: Unifying Raster Data and Vector Semantics for Human-Centric Geospatial Foundation Models

    Jun 1, 2026Steffen Knoblauch, Hao Li, Gengchen Mai +3Remote Sensing Image UnderstandingRepresentation Learning

  18. Towards Resolving Optimization Conflicts Between Image- and Text-Based Person Re-Identification

    Jun 1, 2026Karina Kvanchiani, Timur MamedovCross-Modal Representation LearningMulti-Task Learning

  19. CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval

    May 30, 2026Md Aminur Hossain, Ayush V. Patel, Nitant Dube +1Cross-Modal Representation LearningJoint-Embedding Predictive Architecture

  20. HQ-JEPA: Hybrid Quantum Joint-Embedding Predictive Architecture for Cross-Modal Remote Sensing Representation Learning

    May 29, 2026Md Aminur Hossain, Ayush V. Patel, Sanjay K. Singh +1Representation LearningCross-Modal Representation Learning

  21. UniRTL: Unifying Code and Graph for Robust RTL Representation Learning

    May 29, 2026Yi Liu, Hongji Zhang, Lei Chen +2Representation LearningCross-Modal Representation Learning

  22. Variational Adapter for Cross-modal Similarity Representation

    May 29, 2026WenZhang Wei, Zhipeng Gui, Dehua Peng +2Cross-Modal Representation LearningVLM Adaptation

  23. ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

    May 29, 2026Jun-Hak Yun, Seung-Bin Kim, Seong-Whan LeeCross-Modal Representation LearningTTS Synthesis

  24. Improving Relative Representations with Learned Anchors and Whitened Inner Products

    May 28, 2026Oscar Thorsted Svendsen, Nikolaj Holst Jakobsen, Fabian Mager +1Representation LearningCross-Modal Representation Learning

  25. A Shared Valence Axis Across Modern LLMs and Human EEG: The Saturation Regularity

    May 28, 2026Yousef A. Radwan, Xuhui Liu, Kilichbek Haydarov +2Cross-Modal Representation LearningNeural Decoding

  26. DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

    May 28, 2026Jusuk Lee, Seungjae Lee, Jonghun Shin +6Cross-Modal LearningCross-Modal Representation Learning

  27. MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models

    May 25, 2026Kaixiang Chen, Pengfei Fang, Hui XueCross-Modal LearningCross-Modal Representation Learning