physics.geo-phMar 22, 2026

TRACE: A Multi-Agent System for Autonomous Physical Reasoning for Seismology

Authors: Feng Liu, Xin Cui, Jian Xu, Xinghao Wang, Zijie Guo, Jiong Wang, S. Mostafa Mousavi, Xinyu Gu, +6 more

Organizations: School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, Shanghai, China. · Shanghai Artificial Intelligence Laboratory, Shanghai, China. · School of Earth and Space Sciences, University of Science and Technology of China, Hefei, China. · Department of Earth and Planetary Sciences,McGill University, Montreal, Canada. · Department of Earth and Planetary Sciences, Harvard University, Cambridge, United States of America. · Institute of Earthquake Forecasting, China Earthquake Administration, Beijing, China.

Abstract

Modern seismic networks resolve earthquake sequences in unprecedented detail, yet explaining how large earthquakes emerge from evolving fault systems remains difficult. We introduce TRACE, a seismology-guided artificial intelligence agent that plans and executes workflows while preserving auditable evidence chains from observations to physical interpretation. We evaluated TRACE through 104 benchmark tasks and two complementary earthquake sequences. For the well-studied 2019 Ridgecrest sequence, TRACE constructed a high-resolution catalog from continuous waveforms and retrospectively recovered delayed cascading activation between the Mw 6.4 and Mw 7.1 earthquakes without a prescribed target interpretation. In the less-understood 2025-2026 Sanriku sequence off northeastern Japan, TRACE developed a testable interpretation of progressive destabilization within a segmented megathrust. Its synthesis linked coupled seismic-aseismic activation around the MJ 6.9 sequence and subsequent persistent, spatially segmented shallow-interface activity to a megathrust patch that lay between regions of past large coseismic slip and later hosted the MJ 7.7 rupture. These results open a path from seismic observations to testable physical insight.

Figures & tables

Explore similar work

Jun 5, 2026cs.CL

TRACE: Trajectory Reasoning through Adaptive Cross-Step Evidence Aggregation for LLM Agents

Autonomous LLM agents can pursue hidden malicious objectives through sequences of individually benign actions, making sabotage difficult to detect using standard trajectory-level monitoring. Existing approaches either evaluate complete trajectories in a single pass or partition them into independently scored windows, limiting their ability to connect evidence across temporally distant actions. We propose TRACE, a monitoring framework for long-horizon LLM agent trajectories. TRACE operates through a TIJ (Triage-Inspect-Judge) loop that identifies high-signal regions, performs targeted inspection while maintaining accumulated evidence across reasoning steps, and synthesizes a trajectory-level verdict. We evaluate TRACE on ten task domains from SHADE-Arena against state-of-the-art baselines. TRACE achieves an aggregate F1 of 0.713 and recall of 0.844, with the largest gains on tasks requiring long-range evidence linking.
May 25, 2026cs.AI

VeriTrace: Evolving Mental Models for Deep Research Agents

Deep research agents face vast, interdependent, and pervasively uncertain information. Existing systems explore what evolving intermediate representations should look like, but leave their evolution to the LLM's implicit reasoning. Without explicit regulation, the intermediate layer is easily contaminated by mixed-quality information, and errors propagate along its dependencies, so model scale often ends up substituting for absent regulation. We argue that an agent's mental model should instead evolve through explicit feedback that continuously aligns task understanding with reality, and identify three regulatory loops: interpretive update, deviation feedback, and schema revision. We realise this in VeriTrace, a cognitive-graph framework that explicitly implements the three loops. Using matched Qwen3.5-27B backbones, VeriTrace improves over the strongest matched baseline by an average of 4.82 pp on DeepResearch Bench (DRB) Insight (1.83 pp Overall) and by 5.9 pp Overall win rate on DeepConsult. With Config-DeepSeek, it achieves the strongest reproducible open-source result on DRB.
Sep 10, 2026cs.AI

OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes, or distinguish systematic reasoning from fortunate guessing. We present \textbf{OpenDiscoveryTrace}, a public dataset of 558 complete AI scientific agent trajectories that captures how models reason, not just what they produce. Each trajectory records a structured 9-field-per-step trace---including thoughts, tool calls, observations, errors, revision triggers, and self-reported confidence---as models execute 124 scientific tasks spanning drug discovery, materials science, genomics, and scientific literature analysis. The dataset covers seven models: three frontier models (GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro; 124 trajectories each, fully balanced across domains and difficulty levels) and four open-weight models (Qwen2.5-7B, Mistral-7B-v0.3, Phi-3.5-mini, and Qwen2.5-1.5B; 30 each), plus 60 live-retrieval variant trajectories. Pilot analysis on 363 LLM-judged trajectories reveals that process traces expose behavioral differences invisible to output-only evaluation: all three frontier models achieve comparable success rates (84--89%), yet Claude Opus 4.6 produces 30×\times more errors than GPT-5.4 (2.5 vs. 0.08 per trajectory, p<0.0001p < 0.0001, Cliff's δ=0.613\delta = 0.613), with qualitatively different error profiles---66.7% tool misuse for Claude versus 83.6% reasoning errors for GPT-5.4. We define five benchmark tasks with baselines from logistic regression, random forests, LSTMs, and Transformer models. The dataset, trace schema, agent harness, and benchmark definitions are publicly available under CC BY 4.0 to support research on process-level evaluation, scientific agent auditing, and AI governance.