cs.CLAug 27, 2026

Research Design Tracking and Assessment for the Social Sciences

Authors: Marco Rovera, Sergiu Burlacu, Dominique Cappelletti, Alessio Tomelleri, Sonia Marzadro, Martina Bazzoli, Annalisa Tassi, Jessica Gagete-Miranda

Organizations: Fondazione Bruno Kessler Trento, Italy

Abstract

Reliable assessment of causal research designs in the social sciences is critical for evidence-based policy-making, yet has so far relied entirely on manual expert analysis. We introduce Automated Research Design Tracking and Assessment (ARDTrA), a task that involves detecting the research design used in a paper and assessing the quality of its application. We create an expert-annotated dataset of papers covering six families of counterfactual research designs and evaluate the task using a multi-turn RAG-based conversational pipeline. Across four retrieval strategies, four LLMs and six embedding models, we find that passage length is the main driver of performance, explaining 52-66% of the variance. A per-research-design analysis also shows that human and machine difficulty do not align: the designs that prove hardest for the system are not those on which expert annotators disagree most, pointing to two independent sources of task difficulty.

Explore similar work

CardsList
  1. How Is Automated Research Evaluated? A Survey of Benchmarks and Evaluation Practices

    Oct 8, 2026Liulei Zhang, Dejing Zhou, Chuyue Huang +5AI-Assisted Scientific ResearchBenchmark Design

  2. ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

    Aug 13, 2026Jiale Cui, Yueyao Yuan, Kaixi Zhong +3Human Preference EvaluationAI Agent Benchmarks