Data Analysis Agents

Momentum

9 papers in the last four weeks, up 200% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 39

All topics
CardsList
  1. DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis

    Oct 7, 2026Yizhi Song, Hang Ni, Weijia Zhang +1Data Analysis AgentsAI Agent Benchmarks

  2. Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses

    Sep 28, 2026Huachi Zhou, Yujing Zhang, Jiahe Du +7LLM Agent HarnessesLLM Agent Reliability

  3. StateGuard: Analytical-State Management with Validity-Aware Intervention for Long-Horizon Data Agents

    Sep 28, 2026Wenle Liao, Zhao Wang, Jingchao Zhang +3LLM Agent VerificationGraph-Based Agent Memory

  4. DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?

    Sep 27, 2026Nan Huang, Mario Tapia-Pacheco, Kun Zhou +4AI Agent ReliabilityScientific Hypothesis Generation

  5. TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent

    Sep 23, 2026Jie Yang, Yan Zheng, Jiarui Sun +8Time Series QALLM Agent Self-Improvement

  6. UniDataAgent: An Ontology-Grounded Agent for Enterprise Question-to-Report Automation

    Sep 23, 2026Yutai Duan, Yahui Zhao, Zhangti Li +5Tool-Using AgentsLLM Agents

  7. Data Agents: Agentic Data Systems

    Sep 21, 2026Guoliang Li, Peiyao Zhou, Xuanhe Zhou +3LLM AgentsData Analysis Agents

  8. Skill-based Agentic Evaluation for Real-time Data Science Tasks

    Sep 15, 2026Aniruddha Tamhane, Raghavendra Addanki, Ayushi Aggarwal +4LLM-as-a-JudgeLLM Agent Evaluation

  9. MasterControl Seventeen Every Time

    Sep 2, 2026MasterControl AI LabLLM Agent EvaluationText-to-SQL

  10. BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks

    Aug 26, 2026Zane Koch, Asmamaw T. Wassie, Javier Valdes-Aleman +5AI for ScienceAI Agent Benchmarks

  11. DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

    Aug 11, 2026Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub +3Computer-Use Agent BenchmarksAI Agent Evaluation

  12. SciDataSailor: Deep Scientific Data Exploring

    Jul 29, 2026Jiyong Rao, Yicheng Qiu, Chi Zhang +2Scientific QAAI Agent Benchmarks

  13. CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

    Jul 9, 2026Andrej Leban, Yuekai SunStructural Causal ModelsAI Agent Benchmarks

  14. AgenticDataBench: A Comprehensive Benchmark for Data Agents

    Jul 2, 2026Zhaoyan Sun, Shan Zhong, Daizhou Wen +10LLM Agent EvaluationAI Agent Benchmarks

  15. DA-Studio: An Agentic System for End-to-End Data Analysis

    Jun 30, 2026Yizhe Liu, Shaolei Zhang, Ju FanExecution-Guided Code GenerationAgentic Workflows

  16. Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

    Jun 23, 2026Tian Zheng, Kai-Tai HsuLLM-as-a-JudgeHuman-in-the-Loop Evaluation

  17. VeriGraph: Towards Verifiable Data-Analytic Agents

    Jun 15, 2026Jiajie Jin, Zhao Yang, Wenle Liao +5Data ProvenanceNeuro-Symbolic Reasoning

  18. Fantastic Scientific Agents and How to Build Them: AgentBuild for Rietveld Refinement

    Jun 11, 2026Woong Shin, Craig A. Bridges, Marshall T. McDonnell +1AI Agents for Scientific DiscoveryLLM Agent Self-Improvement

  19. GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models

    Jun 11, 2026Gabriel Diaz-Ireland, Diego Prieto-Herráez, Mario García Peces +2LLM Agent EvaluationGeospatial Reasoning

  20. Unsupervised Skill Discovery for Agentic Data Analysis

    Jun 4, 2026Zhisong Qiu, Kangqi Song, Shengwei Tang +4LLM Agent Skill LearningAgent Skill Learning

  21. LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis

    May 28, 2026Kewei Xu, Xiaoben Lu, Shuofei Qiao +4Long-Horizon Agent EvaluationLong-Horizon Agent Tasks

  22. AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery

    May 22, 2026Darek Kleczek, Fuheng Zhao, Alexander W. Lee +4AI Agent EvaluationAI Agent Benchmarks

  23. Toward AI VIS Co-Scientists: A General and End-to-End Agent Harness for Solving Complex Data Visualization Tasks

    May 20, 2026Haichao Miao, Zhimin Li, Kuangshi Ai +4Scientific VisualizationData Visualization

  24. Ambig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science Agents

    May 10, 2026Josefa Lia Stoisser, Marc Boubnovski Martell, Sidsel Boldsen +2LLM Agent EvaluationAI Agent Benchmarks