Discover cutting-edge AI research papers, automatically curated and categorized daily.
Latest Papers
[MASK] is All You Need
Vincent Tao Hu, Björn Ommer
Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis
M. Hamza Mughal, Rishabh Dabral, Merel C.J. Scholman, Vera Demberg, Christian Theobalt
Tactile DreamFusion: Exploiting Tactile Sensing for 3D Generation
Ruihan Gao, Kangle Deng, Gengshan Yang, Wenzhen Yuan, Jun-Yan Zhu
P3-PO: Prescriptive Point Priors for Visuo-Spatial Generalization of Robot Policies
Mara Levy, Siddhant Haldar, Lerrel Pinto, Abhinav Shirivastava
CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
Zhefei Gong, Pengxiang Ding, Shangke Lyu, Siteng Huang, Mingyang Sun, Wei Zhao, Zhaoxin Fan, Donglin Wang
Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation
Nicolas Dufour, David Picard, Vicky Kalogeiton, Loic Landrieu
Diverse Score Distillation
Yanbo Xu, Jayanth Srinivasa, Gaowen Liu, Shubham Tulsiani
AnyBimanual: Transferring Unimanual Policy for General Bimanual Manipulation
Guanxing Lu, Tengbo Yu, Haoyuan Deng, Season Si Chen, Yansong Tang, Ziwei Wang
Driv3R: Learning Dense 4D Reconstruction for Autonomous Driving
Xin Fei, Wenzhao Zheng, Yueqi Duan, Wei Zhan, Masayoshi Tomizuka, Kurt Keutzer, Jiwen Lu
Enhancing Robotic System Robustness via Lyapunov Exponent-Based Optimization
G. Fadini, S. Coros
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
Yi-Lun Lee, Yi-Hsuan Tsai, Wei-Chen Chiu
Visual Lexicon: Rich Image Features in Language Space
XuDong Wang, Xingyi Zhou, Alireza Fathi, Trevor Darrell, Cordelia Schmid
Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty
Meera Hahn, Wenjun Zeng, Nithish Kannen, Rich Galt, Kartikeya Badola, Been Kim, Zi Wang
Dynamic EventNeRF: Reconstructing General Dynamic Scenes from Multi-view Event Cameras
Viktor Rudnev, Gereon Fox, Mohamed Elgharib, Christian Theobalt, Vladislav Golyanik
Training Large Language Models to Reason in a Continuous Latent Space
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, Yuandong Tian
MAtCha Gaussians: Atlas of Charts for High-Quality Geometry and Photorealism From Sparse Views
Antoine Guédon, Tomoki Ichikawa, Kohei Yamashita, Ko Nishino
Ranking-aware adapter for text-driven image ordering with CLIP
Wei-Hsiang Yu, Yen-Yu Lin, Ming-Hsuan Yang, Yi-Hsuan Tsai
XRZoo: A Large-Scale and Versatile Dataset of Extended Reality (XR) Applications
Shuqing Li, Chenran Zhang, Cuiyun Gao, Michael R. Lyu
InstantRestore: Single-Step Personalized Face Restoration with Shared-Image Attention
Howard Zhang, Yuval Alaluf, Sizhuo Ma, Achuta Kadambi, Jian Wang, Kfir Aberman
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
Neel Jain, Aditya Shrivastava, Chenyang Zhu, Daben Liu, Alfy Samuel, Ashwinee Panda, Anoop Kumar, Micah Goldblum, Tom Goldstein