Tool Use in VLMs

VLM: Vision-Language Model

Momentum

13 papers in the last four weeks, up 160% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 76

All topics
CardsList
  1. TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents

    May 11, 2026Bihui Yu, Caijun Jia, Jing Chi +6Tool Use in VLMsData Provenance

  2. DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents

    May 10, 2026Yixiong Chen, Wenjie Xiao, Pedro R. A. S. Bassi +7VLM EvaluationTool Use in VLMs

  3. Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models

    May 9, 2026Sagar Bharadwaj, Ziyong Ma, Anurag Ghosh +2Tool Use in VLMs3D Spatial Reasoning

  4. Act2See: Emergent Active Visual Perception for Video Reasoning

    May 3, 2026Martin Q. Ma, Yuxiao Qu, Aditya Agrawal +4Tool Use in VLMsVideo Reasoning

  5. See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection

    Apr 27, 2026Zhiheng Wu, Tong Wang, Shuning Wang +2Tool Use in VLMsVLM Reasoning

  6. Visual Reasoning through Tool-supervised Reinforcement Learning

    Apr 21, 2026Qihua Dong, Gozde Sahin, Pei Wang +4Tool Use in VLMsVLM Reasoning

  7. Q-DeepSight: Incentivizing Thinking with Images for Image Quality Assessment and Refinement

    Apr 18, 2026Xudong Li, Jiaxi Tan, Ziyin Zhou +6Tool Use in VLMsMultimodal CoT Reasoning

  8. Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering

    Apr 8, 2026Zhuohong Chen, Zhenxian Wu, Yunyao Yu +6Visual Question AnsweringTool Use in VLMs

  9. VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning

    Mar 26, 2026Zhe Gao, Shiyu Shen, Taifeng Chai +7Tool Use in VLMsEfficient VLM Inference

  10. Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs

    Mar 14, 2026Nimrod Shabtay, Moshe Kimhi, Artem Spector +5Tool Use in VLMsEfficient VLM Inference

  11. A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding

    Dec 24, 2025Christina Liu, Alan Q. Wang, Joy Hsu +2Tool Use in VLMsMedical Image Classification

  12. Ophiuchus: Incentivizing Tool-augmented "Think with Images" for Joint Medical Segmentation, Understanding and Reasoning

    Dec 16, 2025Yankai Jiang, Yujie Zhang, Peng Zhang +5Image SegmentationTool Use in VLMs

  13. REVEAL: Robust Evolution of Vision-Language Models for Explainable AI-Video Detection

    Feb 20, 2025Yun-Yun Tsai, Qingyuan Liu, Ruijian Zha +4Tool Use in VLMsVision-Language Models