Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses
Organizations: The Hong Kong Polytechnic University, Hong Kong, China · Jilin University, Changchun, China
Abstract
Large language model agents are increasingly deployed for data-intensive work, yet reliable data analysis requires more than general-purpose reasoning and ad hoc tool augmentation. Data Agents, equipped with workflow harnesses, offer a promising paradigm for automating the end-to-end data science lifecycle. This paper examines Data Agents from a harness-centric perspective. First, we introduce a taxonomy of Data Agents and associated data environments, organizing the literature around five functional stages: perception, planning, execution, verification, and repair. Second, we analyze the key technical routes within each stage, identifying 15 distinct approaches ranging from data structure probing to data state reconstruction. Third, we identify four open reliability problems: inactive semantic calibration, missing clarification, missing experience transfer, and the missing verification-repair repository. These problems explain why silent failures can persist even when individual components function correctly, highlighting the need for rigorous workflow harnesses and shared reliability resources. Finally, we summarize the horizontal task families of Data Agents, examine their vertical application settings, and benchmarks for evaluation, while maintaining a companion repository at https://github.com/DEEP-PolyU/Awesome-Data-Agents.
Figures & tables
| Benchmark | Focus | Task & Process Coverage | Evaluation |
|---|---|---|---|
| StockGQL ( Liang et al., 2024 ) | Natural-language-to-GQL translation over structured financial knowledge. | Data Querying; Execution, Verification | Query correctness and retrieval of the required information. |
| TableBench ( Wu et al., 2025 ) | Table question answering covering fact checking, numerical reasoning, data analysis, and visualization. | Data Querying; Verification | TableQA accuracy across multiple reasoning categories. |
| Visual-TableQA ( Lompo and Haraoui, 2026 ) | Visual reasoning over rendered tables, including structure understanding and multi-step reasoning. | Visualization and Multimodal Analysis; Verification | Question-answering and reasoning accuracy on table images. |
| TopBench ( Ji et al., 2026 ) | Implicit predictive reasoning over tabular data, including prediction, decision making, and treatment-effect analysis. | Analysis and Prediction; Verification | Performance on analytical objectives beyond direct table lookup. |
| PrepBench ( Xu et al., 2026a ) | Natural-language-driven data preparation involving cleaning, restructuring, and table transformation. | Data Preparation; Execution, Verification | Correctness of generated output tables across preparation settings. |
| InfiAgent-DABench ( Hu et al., 2024 ) | End-to-end data analysis over CSV files requiring agents to interact with an execution environment. | Analysis and Prediction; Planning, Execution | Automatically evaluated answers across diverse analytical questions. |
| Survey / Study | Capability | Data | Lifecycle | Evaluation | Stage-wise | Reliability |
|---|---|---|---|---|---|---|
| Survey of Data Agents ( Zhu et al., 2025b ) | ✓ | ✓ | – | |||
| Autonomous Data Agents ( Fu et al., 2025b ) | ✓ | ✓ | – | |||
| LLM DATA ( Zhou et al., 2025 ) | ✓ | ✓ | – | – | ||
| LLM/Agent-as-Data-Analyst ( Tang et al., 2025b ) | ✓ | ✓ | – | |||
| LLM-Based Data Science Agents ( Rahman et al., 2025 ) | ✓ | ✓ | ✓ | ✓ | ||
| Clean Up Your Mess ( Zhou et al., 2026a ) | ✓ | ✓ | ✓ | – |