Reproducibility

Momentum

31 papers in the last four weeks, up 63% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 238

All topics
CardsList
  1. UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

    Oct 4, 2026Dolly Sah, Tanmay Sah, Harshul Jain +1Reproducibility

  2. AI-assisted mitotic counting improves reproducibility and efficiency across multiple tumour types

    Oct 1, 2026Simon Graham, Mostafa Jahanifar, Quoc Dang Vu +15Mitotic FiguresTumor

  3. Model validation in machine learning: A scenario-based guide from hold-out splits to nested group cross-validation in biomedical and applied research

    Oct 1, 2026Mehmet Baygin, Sengul Dogan, Turker TuncerCross ValidationValidation

  4. Code That Works, Environments That Don't: Measuring Environment Reproducibility in AI-Generated Software

    Sep 30, 2026Bhanu Prakash Vangala, Tanu MalikCoding AgentsCode Generation

  5. Retrieve, Reproduce, Reveal: Dissecting Retrieval-Augmented Software Vulnerability Detection

    Sep 29, 2026Sabrina Kaniewski, Tim Krämer, Julius Bächle +3Model VulnerabilitiesSoftware

  6. Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation

    Sep 29, 2026Mario Sanz-Guerrero, Minh Duc Bui, Manuel Mager +1ReproducibilityPrompt Engineering

  7. Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study

    Sep 24, 2026Showket Ahmad Khan, Mudasir Mohd, Nasrullah Sheikh +4SummarizationHindi

  8. RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?

    Sep 23, 2026Mithil Salunkhe, Haochen Ding, Samridhi Verma +1ReproducibilityMle-Bench Lite

  9. Driving Epidemic Models with AI Agents: the Epydemix Agent Framework

    Sep 23, 2026Nicolò Gozzi, Ciro Cattuto, Alessandro VespignaniEpidemiological ModelsAgent-Based Model

  10. Exact Feedback Is Not Control: Evaluating Text-based Closed-Loop Revision in LLMs

    Sep 23, 2026Haitong Jiang, Chunlin Liu, Yile Wang +1FeedbackReproducibility

  11. EvEMTBench: An Open Benchmark for Machine Learning in Power System Protection

    Sep 23, 2026Julian Oelhaf, Georg Kordowich, Christian Bergler +3Power SystemsFault Diagnosis

  12. The Illinois Social Attitudes Aggregate Corpus (ISAAC): An Open Tool and Reproducible Pipeline for Analyzing Social Group Discourse at Scale

    Sep 22, 2026Babak Hemmatian, Sarah Hadjarab, Jessica Chen +1Text CorporaDemographics

  13. Reproducible AI Requires Reproducible Randomness

    Sep 22, 2026Anthony Bertrand, Tom Schmitt, Engelbert Mephu Nguifo +1ReproducibilityRandom

  14. Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

    Sep 22, 2026Liam Cooper, Shinnung Jeong, Hyeran Jeon +2LLM Inference OptimizationFloating-Point

  15. GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

    Sep 21, 2026Yiran Wang, Xingyilang Yin, Junfu Pu +12Agentic BenchmarksVideo World Models

  16. LLJ Cards: Best practices for the Use of LLMs as Judges

    Sep 21, 2026Khaoula Chehbouni, Melina Medjdoub, Florian Carichon +2Llm-As-A-JudgeJudges

  17. RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

    Sep 18, 2026Shuai Bai, Jiayong Deng, Sicheng Fan +30Computer-Use AgentsSynthetic Environments

  18. Reproducibility is not construct validity: LLM measurement of institutionally situated communication

    Sep 17, 2026Veronika Batzdorfer, Carlo Romano Marcello Alessandro SantagiustinaInter-Annotator AgreementLarge Language Model Evaluation

  19. TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking

    Sep 16, 2026Giorgia Adorni, Michela Papandrea, Battista Rimoldi +1Text-To-MusicPerformance Metric

  20. OPEN-1B: A Fully Auditable Training Run

    Sep 15, 2026John Donaghy, Brian Wilcox, Oğuzhan Ersoy +6ReproducibilityTraining Data

  21. OpenAI4S: Code as Action, Science as Sessions

    Sep 14, 2026Gongbo Zhang, Hao Li, Yu Wang +15Scientific WorkflowsReproducibility

  22. Membership Inference via Pairwise Likelihood Ratios

    Sep 14, 2026Shengjie Niu, Zebin Yun, Yeheng Ge +1Membership Inference AttacksMembership Inference

  23. Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

    Sep 11, 2026Hanhua Hong, Yizhi Li, Luu Gia Huy +3Agentic BenchmarksReproducibility

  24. IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications

    Sep 9, 2026Yiling Ma, Yilun Zhao, Sihong Wu +2ReproducibilityResearch Automation

  25. APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents

    Sep 7, 2026Jintian Feng, Long Chen, Xiao Yu +7Graphical User Interface AgentsDevice

  26. Model Retirement Creates Reproducibility Risk in Biomedical AI Publications

    Sep 7, 2026Nathan Wolfrath, Meghan Conroy, Thomas Kosten +7Biomedical TextJournal

  27. Last Translation Benchmark

    Sep 3, 2026Vilém Zouhar, Niyati Bafna, Mukund Choudhary +257Machine Translation QualityReproducibility

  28. Concept of a Sensor Test Environment for Dusty Agricultural Conditions

    Sep 3, 2026Peter Buckel, Johannes Hermann, Jonas Wollmann +2Precision AgricultureAgricultural Robotics

  29. LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting

    Sep 1, 2026Yufei Chen, Yiran Zhao, Xiaogang Xu +3Large Language Model ForecastingProbabilistic Forecasting

  30. Perceptible or Not? Diagnosing Passive Fingerprints for Speech Deepfake Attribution

    Sep 1, 2026Yupei Li, Qiyang Sun, Emmanouil Benetos +2Audio Deepfake DetectionFingerprint

  31. Reliable Benchmarking of Artifact Detection in Computational Pathology: A Reproducibility and Uncertainty Analysis

    Aug 31, 2026Konstantinos Moutselos, Ilias MaglogiannisWhole-Slide ImagesComputational Pathology

  32. Aggregate Disambiguation Systems

    Aug 31, 2026José María Lago, Albert Castellana, Edgars NemšeDisambiguationMajority Voting

  33. Reifying Research Logic: AI-Assisted Workflow Construction and Incremental Refinement for Quantitative Syntax

    Aug 11, 2026He Wang, Jingbo Chen, Yuqiao Lai +3Agentic Workflow DesignNatural Language

  34. Fairness in Link Prediction Beyond Demographic Parity: A Reproducibility Study

    Aug 10, 2026Valentijn Oldenburg, Floris de Kam, Stef de Wildt +1Algorithmic FairnessDisparities

  35. Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

    Aug 10, 2026Diandian Zhang, Tingyu Song, Lin Fu +2Scientific DiscoveryVisual Question Answering Benchmarks

  36. XFeat Revisited: Reproducibility and Evaluation of a Lightweight Image Matcher

    Aug 10, 2026Lazar Đoković, Aimee LinHarder Better Faster Denser Feature MatchingCapability-Speed Trade-Offs

  37. Software Engineering for and with GUI Agent

    Aug 10, 2026Shengcheng Yu, Yuchen Ling, Junyang Xing +3Graphical User Interface AgentsSoftware Engineering

  38. EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility

    Aug 9, 2026Xudong Wu, Zeqing Wu, Jiarui Zhang +8Smart HomesParticipation

  39. JUMP-lite: Compact, reproducible benchmarking of cell representations

    Aug 7, 2026Alán F. Muñoz, Johan Fredin Haslum, Runxi Shen +2CellsPhenotypes

  40. When Does Consensus Mean Correctness? Measuring the Agreement-Accuracy Coupling with Semantics-Preserving Re-Rendering

    Aug 6, 2026Rasul Khanbayov, Hasan KurbanReproducibilityAgreement

  41. Sharding Prevents LLM Oversight Failures and Adversarial Exploitation

    Aug 5, 2026Victor Akinwande, J. Zico Kolter, Aran NayebiJudgesScalable Oversight

  42. CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications

    Aug 5, 2026Brendan Smith, Susana Lopez-Moreno, Eric Dolores-Cuenca +3Scientific WorkflowsAgentic Workflows

  43. CURATE: Leveraging LLM Agents to Compose, Catalog, and Deploy Reproducible Workflows

    Aug 4, 2026Nolan Cutler, Chia-Chen Kuo, Nanda Velugoti +2Agentic Workflow DesignCoding Agents

  44. Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

    Aug 4, 2026Mohsen Hariri, Weicong Chen, Nahal Shahini +11Test-Time ScalingInference-Time

  45. Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

    Aug 4, 2026William Bolton, Philip TorrScientific IdeationArtificial Intelligence Benchmarks

  46. Measuring Cross-Task Behavioral Consistency in Language Model Agents

    Jul 31, 2026Amritesh Banerjee, Pranil RaichuraLanguage-Model AgentsReproducibility