Multimodal Tool Use

Momentum

3 papers in the last four weeks, against 2 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 41

All topics
CardsList
  1. OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

    Oct 1, 2026Haibo Wang, Jiteng Mu, Jialu Li +5Audio-Visual UnderstandingRL for Language Model Reasoning

  2. WeaveData: A Multimodal Data Analysis System with Self-Critiquing and Self-Evolving LLM Plans

    Sep 28, 2026Min Jia, Shihao Zhou, Jun-Peng Zhu +9LLM Self-RefinementLLM Planning

  3. MedTRACE: Tool-Augmented Multimodal Clinical Reasoning Agents for Evidence-Grounded Decision-Making

    Sep 13, 2026Ji Lu, Lifei Liu, Haoran Yu +5Medical DiagnosisClinical Decision-Making

  4. WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents

    Aug 28, 2026Zongkai Liu, Hui Zhang, Liqiang Niu +7Multimodal IRMultimodal Search Agents

  5. A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings

    Aug 13, 2026Hei Ting, Chan, Chenwei Wu +8Efficient Multimodal InferenceHealthcare

  6. Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents

    Aug 4, 2026Siqi Fan, Minghao Li, Xiaoqian Ma +6Computer-Use AgentsRL for LLM Agents

  7. See2Think: Do Multimodal Models Really Use Intermediate Visual States?

    Jul 29, 2026Siyu Yan, Zhuoran Yan, Haiying Xu +10Visual ReasoningMultimodal Reasoning

  8. ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition

    Jul 28, 2026Jooyeol Yun, Jintae Park, Hyesu Lim +3Agentic WorkflowsMultimodal Tool Use

  9. ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning

    Jul 17, 2026Binglin Zhou, Peng Shi, Ryo Kamoi +2Tool Use in VLMsScientific Claim Verification

  10. Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents

    Jul 13, 2026Ming Ma, Yi Zhu, Yiran Zhong +6Multimodal AgentsMultimodal Tool Use

  11. UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp

    Jul 12, 2026Xiyu Wei, Qingwei Zong, Zhuocheng Yu +1Multimodal QAAI Agent Benchmarks

  12. Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards

    Jul 6, 2026Tianhao Niu, Ziyu Han, Qiguang Chen +5Data VisualizationCode Generation Evaluation

  13. CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration

    Jul 6, 2026Hairui Zhu, Yiying Yang, Tengjin Weng +5Tool Use in VLMsText-Guided Image Editing

  14. Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories

    Jun 9, 2026Kevin Qinghong Lin, Batu EI, Yuhong Shi +5LLM Agent OrchestrationMultimodal Agents

  15. TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents

    Jun 4, 2026Chengqi Dong, Chuhuai Yue, Hang He +6Multimodal IRToken-Level Credit Assignment

  16. VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark

    Jun 2, 2026Amirhossein Dabiriaghdam, Shayan Vassef, Mohammadreza Bakhtiari +5Mathematical Reasoning BenchmarksMathematical Reasoning

  17. Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains

    Jun 1, 2026Garvin Guo, Donglei Yu, Yu Chen +6Tool-Use EvaluationLLM Agent Evaluation

  18. MetaForge: A Self-Evolving Multimodal Agent that Retrieves, Adapts, and Forges Tools On Demand

    Jun 1, 2026Shouang Wei, Houcheng Min, Xinpeng Dong +8LLM Agent Skill LearningMultimodal Agents

  19. Sandboxed Coding Agents are Competitive Omni-modal Task Solvers

    May 30, 2026Dongping Chen, Xuanao Huang, Zhihan Hu +3Audio-Visual UnderstandingAI Coding Agents

  20. Syll: Open-Source Personal Automation with Cross-Surface Execution

    May 28, 2026Bo Zhang, Borui Zhang, Chenghao Jiang +5LLM Agent Skill LearningHuman-in-the-Loop AI

  21. Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation

    May 28, 2026Chenghao Zhang, Guanting Dong, Yufan Liu +3Multimodal GenerationMultimodal Agents

  22. AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture

    May 21, 2026Zi Ye, Yibin Wen, Xiaoya Fan +10AI Agent BenchmarksTool-Using Vision-Language Agents

  23. Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning

    May 19, 2026Qinghe Ma, Zhen Zhao, Yiming Wu +3Multimodal Large Language ModelsMultimodal Reasoning

  24. Hallucination as Exploit: Evidence-Carrying Multimodal Agents

    May 18, 2026Guijia Zhang, Hao Zheng, Harry YangAI Agent SecurityMultimodal Hallucination