Exploratory data analysis (EDA) is rarely open-ended in practice: analysts work from high-level domain questions toward the concrete analyses that can answer them, prioritizing directions with domain knowledge and prior hypotheses. Large language models (LLMs) can supply such knowledge, but their responses are unstructured, leaving analysts no way to see what has been explored, what is missing, or why one direction was chosen over another. We present DAG-EDA, a system that lets analysts and an LLM co-navigate the space of possible analyses through two linked structures. An intent graph, governed by a grammar of analytical intent, decomposes an ambiguous natural-language question into progressively concrete analysis tasks, keeping alternative framings open and letting analysts branch, backtrack, and compare paths. A multi-layered knowledge graph externalizes the LLM's domain knowledge, linking domain concepts to the dataset variables that can measure them, so analysts can inspect and contest how their question is grounded in the data. Both graphs are constructed from only the dataset and the analyst's question, and the analyses the analyst reaches are rendered as interactive dashboards. We illustrate the system through a usage scenario and describe a user study design for examining whether the system scaffold analysts' reasoning and navigation.
Figures & tables
Figure 1. The DAG interface, early in Anna’s session. The question bar (A) holds the analytical question the session started from and the control that parses it. The state DAG canvas (B) shows the claims Anna has formed, one node per candidate analysis. Here it holds the interpretations the parser returned, each labeled with the claim it makes, a natural-language rendering of what it would ask, the confidence that it is the intended reading, and its slots as chips. The moves available from each node hang beneath it, grouped by the state they would produce and collapsed to a count, so that the shape of what is still open is visible without enumerating it; the toolbar names the four families of move the grammar admits. The active node panel (C) describes the selected claim in full, lists its slots, and offers to complete it to a dashboard once its required slots are grounded. The knowledge graph (D) shows the concept layer and the data layer joined by bridge edges, with filters for which layers and edge types to display.
Figure 2. Anna’s intent graph at the end of her session, with the dashboards she has open and the node she has selected. Across the top are the three readings the parser returned for “What makes a movie successful?” Two paths A and B are outlined. Each outline encloses the claims Anna formed while following that reading, and the two enclose one card in common, the comparison of Rotten Tomatoes rating across genre that both readings resolve to. The move marked reshape is the one that changes the shape of the claim as it binds, which turns a relation into a comparison. The node panel at the right describes the selected claim A node is marked renderable once its slots reach data and it can be completed to a dashboard.
Figure 3. The DAG architecture, instantiated on “What makes movies successful?” The analyst’s question and the dataset schema are the only inputs. \footnotesize1⃝ Parsing the question seeds the intent graph \footnotesizeA⃝ with the claims it could resolve into, and \footnotesize2⃝ graph construction builds the knowledge graph \footnotesizeB⃝ from the question and the schema together. In the intent graph, each node is a claim whose slots carry a type, measurable (Q) or categorical (C), and a binding that is open, an unresolved concept, or a set of columns; a chart icon marks a node whose required slots are grounded and which therefore renders to a dashboard. The two edge families appear here as one descend step, which resolves “critical success” to the columns beneath it, and one reshape step, which turns a relationship into a comparison by admitting a categorical slot. The knowledge graph splits into a concept layer \footnotesizeC⃝, built once per question, and a data layer \footnotesizeD⃝, built once per dataset, joined by operationalized by bridge edges. The binding edge is where the two structures meet: grounding a slot is a walk from the concept a slot holds to the columns that operationalize it. The dashed “award” concept has no column to operationalize it, which the system reports rather than silently dropping.
multiset
k
f : hypothesis
g : template
example claim
{Q}
1
Value
M1
The distribution of worldwide gross.
{Q, Q}
2
Relation
M3
Production budget against worldwide gross.
{Q, C}
2
Comparison
CAT1 M2
Rotten Tomatoes rating across genre.
{Q, T}
2
Trend
CH1
Worldwide gross over release year.
Table 1. The reference instantiation of the typology, as the two tables behind f and g . f is defined only up to arity two: these four multisets are every base state the grammar admits. g is multi-valued where more than one framing is defensible; choosing among them is a render-time decision, not a production. Swapping these tables swaps the analytical vocabulary. Every analysis over three or more variables is a composite, given in Table 2 . Examples are drawn from the dataset of Section 4 .
composite
well-formed when
compiles to
example
Condition(Spec,z)
z is categorical. It may be left open and bound afterwards.
A facet over z : small multiples or a colour encoding; the CAT2 crosstab where the conditioned spec is a Comparison.
Does the budget–gross relation hold across genres?
Chain(Spec,Spec,m)
Both legs are Relations, both bind m , and their remaining ends differ.
Two linked Relation panels, brushed on the shared m .
Budget → vote count → worldwide gross.
Compare(Spec,…)
At least two legs, sharing a hypothesis type, pairwise distinct.
Overlay on the shared axis for Relations and for Trends, the latter giving CH2 ; a slope chart for Comparisons; side-by-side otherwise.
The box-office reading against the critical reading.
Table 2. The three composites of the reference instantiation, which extend the base states of Table 1 and are the only route to an analysis over three or more variables. Composites are not states and so have no image under f or g . Each carries a well-formedness condition that is relational , a property of how its sub-specs fit together that a single state cannot express, and each compiles to a qualitative layout over templates rather than to one template. The layouts subsume the multi-variable templates directly: a Condition over a Comparison yields the two-way CAT2 crosstab, and a Compare over two Trends sharing a time axis yields the CH2 dual-series chart. Examples are drawn from the dataset of Section 4 .
Large Language Models (LLMs) are increasingly used in analytical workflows, but their suitability as exploratory data analysis (EDA) agents in business settings remains uncertain. In practice, a deployable EDA agent must provide not only useful average performance but also sufficient repeatability to support trust in its outputs. We evaluate this requirement in a controlled, business-relevant benchmark built on an agent-based supply chain simulation. The task is to identify supplier-product combinations responsible for low quality and downstream sales loss by reasoning from indirect operational traces rather than from explicit labels. Fifteen model-variant configurations from eight model families were evaluated under four experimental conditions that varied data representation, prompt clarity, and signal strength, with five trajectories per condition. Outputs were scored against deterministic ground truth using the Jaccard index and assessed through a framework that combines mean score (ms), coefficient of variation (CV), exploratory cross-condition significance tests, and Business utility, a risk-adjusted metric that we propose to summarise quality and repeatability in a single operational measure. The results show that most configurations are not reliable enough for autonomous EDA use, even when their average scores appear acceptable. GPT-5.4 with extra-high reasoning effort achieved the strongest overall profile, with an experiment-averaged ms of 0.8748 and an experiment-averaged Business utility of 0.6952, while the next-best configurations lost substantially more utility after variability discounting. Our findings suggest that evaluation of EDA agents should treat average quality, repeatability, and condition sensitivity as complementary dimensions of operational trustworthiness.
SGH Warsaw School of Economics · Bydgoszcz University of Science and Technology · At the time of contribution affiliated with deepsense.ai, present affiliation is Google
Real-world data analysis is a multi-step process over heterogeneous inputs rather than merely producing a final answer. A practical system should autonomously organize multi-step workflows, execute generated code in a sandboxed and controllable environment, and remain inspectable through visible action traces and intermediate artifacts. Existing LLM-based analysis tools, however, often emphasize isolated subtasks, leaving limited support for complete execution-grounded workflows. We present DA-Studio (Data Analysis Studio), an interactive web-based demo system for end-to-end data analysis that is autonomous, sandboxed, and inspectable. DA-Studio integrates an action-structured analysis backend, a sandboxed execution workspace, and a browser interface for task setup, streamed action traces, artifact preview, code editing and rerunning, and report export. Through iterative action generation, code execution, and feedback incorporation, it incrementally constructs executable analysis steps from raw files and natural-language requests while exposing intermediate results and artifacts throughout the process.
Autonomous data analysis agents are increasingly expected to conduct exploratory analysis with limited human guidance about data. However, existing benchmarks typically evaluate such agents in prior-guided settings, providing selected data sources, explicit data schemas, or cleaned data, thereby understating the exploratory burden. To evaluate this realistic exploratory data analysis task, we introduce DataClawBench, a benchmark built from financial think-tank consulting scenarios where agents must independently explore unfamiliar, noisy, cross-domain data and produce verifiable conclusions. DataClawBench provides a unified real-world data environment with approximately 2.06 million records across enterprise, industry, and policy domains, with native data noise preserved. On top of this data environment, it defines 492 multi-step cross-domain tasks, each annotated with intermediate milestones that diagnose exploration and reasoning failures beyond outcome accuracy. A systematic evaluation of eight advanced LLMs under the OpenClaw agent reveals that exploratory data analysis breaks agent reliability: more exploration does not reliably translate into task-relevant progress or correct final answers.
Qiaohong Zhang, Weihao Ye, Jialong Chen +7
School of Computer Science and Engineering, Sun Yat-sen University · School of Software Engineering, Sun Yat-sen University · Lingnan College, Sun Yat-sen University