Efficient VLM Inference

VLM: Vision-Language Model

Latest papers 322

All topics
CardsList
  1. LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

    May 26, 2026Shihao Wang, Shilong Liu, Yuanguo Kuang +10Efficient VLM InferenceVision-Language Grounding

  2. MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

    May 26, 2026Runxi Huang, Liyu Zhang, Shengzhong Liu +1Mobile GUI AutomationEfficient VLM Inference

  3. MuCRASP: Multimodal Chain-of-thought Reasoning aware Structured Pruning

    May 25, 2026Aritra Dutta, Somak AdityaVision-Language ModelsEfficient VLM Inference

  4. CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception

    May 22, 2026Liupeng Li, Haoqian Kang, Zhenyu Lu +4Efficient VLM InferenceLarge Vision-Language Models

  5. CHASD: Language Increment-Calibrated Contrastive Decoding against Hallucination in LVLMs

    May 22, 2026Xiaoyi Huang, Kejia Zhang, Zhiming LuoHallucination Detection in VLMsEfficient VLM Inference

  6. Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving

    May 22, 2026Kewei Zhang, Jin Wang, Sensen Gao +9VLMs for Autonomous DrivingEnd-to-End Autonomous Driving

  7. GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation

    May 21, 2026Jiahao Yang, Zihan Wang, Xiangyang Li +4Efficient VLM InferenceGeometric Representation Learning

  8. Focus-then-Context: Subject-Centric Progressive Visual Token Reduction for Vision-Language Models

    May 20, 2026Yulin Zhao, Zheng ZhangVision-Language ModelsEfficient VLM Inference

  9. EventPrune: Cascaded Event-Assisted Token Pruning for Efficient First-Person Dynamic Spatial Reasoning

    May 19, 2026Pengtao Ma, Ziliang Zhou, Ciyu Ruan +7Visual Spatial ReasoningEfficient VLM Inference

  10. Wasserstein Equilibrium Decoding for Reliable Medical Visual Question Answering

    May 18, 2026Luca Hagen, Johanna P. Müller, Weitong Zhang +2HealthcareEfficient VLM Inference

  11. A More Word-like Image Tokenization for MLLMs

    May 18, 2026Hyun Lee, Hyemin Jeong, Yejin Kim +4Image TokenizationEfficient VLM Inference

  12. An Efficient Streaming Video Understanding Framework with Agentic Control

    May 18, 2026Jinming Liu, Jianguo Huang, Zhaoyang Jia +7Streaming Video UnderstandingAdaptive Model Routing

  13. SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening

    May 17, 2026Shahriar Kabir Nahin, Hadi Askari, Muhao Chen +1Adaptive InferenceEfficient VLM Inference

  14. t-gems: text-guided exit modules for decreasing clip image encoder

    May 17, 2026Alberto Presta, Grzegorz Stefanski, Michal Byra +1Cross-Modal LearningVision-Language Models

  15. FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing

    May 17, 2026Zihan Tang, Leqi Shen, Hui Chen +7Visual AttentionEfficient VLM Inference

  16. LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs

    May 17, 2026Jihwan Kim, Nikhil Parthasarathy, Danfeng Qin +5Efficient VLM InferenceVideo-Language Models

  17. Thinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern Induction

    May 16, 2026Yichang Jian, Boyuan Xiao, Zhenyuan Huang +2Efficient VLM InferenceLatent Visual Reasoning

  18. LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs

    May 15, 2026Hongyu Lu, Feng Zhang, Wenwei Jin +5Efficient VLM InferenceLarge Vision-Language Models

  19. RoiMAM: Region-of-Interest Medical Attention Model for Efficient Vision-Language Understanding

    May 15, 2026Jiayan Yang, Zhuoyu Wu, Wenqi FangHealthcareEfficient VLM Inference

  20. KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy

    May 14, 2026Yingbing Huang, Tharun Adithya Srikrishnan, Steven K. Reinhardt +1Vision-Language ModelsEfficient VLM Inference

  21. Do We Really Need External Tools to Mitigate Hallucinations? SIRA: Shared-Prefix Internal Reconstruction of Attribution

    May 14, 2026Tian Qin, Junzhe Chen, Yuqing Shi +3Efficient VLM InferenceHallucination in Language Models

  22. GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models

    May 13, 2026Mingzhe Huang, Weijun Wang, Xin Ding +7Efficient VLM InferenceVisual Token Pruning

  23. CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models

    May 13, 2026Sangin Lee, Yukyung ChoiEfficient VLM InferenceVision-Language Grounding

  24. AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding

    May 13, 2026Xiao Yang, Yingzhe Ma, Haoxuan Yu +2Efficient VLM InferenceLong-Video Understanding

  25. VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference

    May 12, 2026Hao Zhu, Shuo Jin, Wenbin Liao +4Image SegmentationVision-Language Models

  26. OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models

    May 12, 2026Minseok Kang, Minhyeok Lee, Jungho Lee +6Efficient VLM InferenceVideo-Language Models