ML Reproducibility

ML: Machine Learning

Momentum

14 papers in the last four weeks, up 100% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 88

All topics
CardsList
  1. The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility

    May 19, 2026David Pape, Jonathan Evertz, Lea SchönherrLLM EvaluationLLM Inference

  2. ExECG: An Explainable AI Framework for ECG models

    May 19, 2026Jong-Hwan Jang, Yong-yeon JoExplainable Artificial IntelligenceHealthcare

  3. MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility

    May 15, 2026Sasi Kiran Gaddipati, Diyana Muhammed, Farhana Keya +2AI Agent BenchmarksML Reproducibility

  4. Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets

    May 14, 2026Rafi Al Attrach, Rajna Fani, Sebastian Lobentanzer +17ML Reproducibility

  5. Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling

    May 13, 2026Deepak Pandita, Flip Korn, Chris Welty +1Inter-Rater ReliabilityLLM Evaluation

  6. Rollout Cards: A Reproducibility Standard for Agent Research

    May 12, 2026Charlie Masters, Ziyuan Liu, Stefano V. AlbrechtAI Agent EvaluationAgent Evaluation

  7. BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

    May 7, 2026Jinge Wu, Hongjian Zhou, Mingde Zeng +8AI Agents for Scientific DiscoveryDeep Research Agents

  8. RamanBench: A Large-Scale Benchmark for Machine Learning on Raman Spectroscopy

    May 3, 2026Mario Koddenbrock, Christoph Lange, Robin Legner +6Benchmark DesignScientific ML

  9. SCARV: Structure-Constrained Aggregation for Stable Sample Ranking in Redundant NLP Datasets

    May 1, 2026Xu Zheng, Feiyu Wu, Linhong Wu +2Training Data SelectionML Reproducibility

  10. A Reproducibility Analysis of PO4ISR: Diagnosing and Mitigating Semantic Drift in LLM-Based Session Recommendation

    Apr 29, 2026Aditya Tiwari, Konduri Naga Lakshmi Rekha, Rajesh Kumar MundotiyaSequential RecommendationLLM Prompting

  11. Spreadsheet Modeling Experiments Using GPTs on Small Problem Statements and the Wall Task

    Apr 28, 2026Thomas A. Grossman, Yuan Chen, Sopiko DatuashviliML ReproducibilitySpreadsheet Automation

  12. NeuroClaw Technical Report

    Apr 27, 2026Cheng Wang, Zhibin He, Zhihao Peng +7Multi-Agent LLM SystemsAI Agent Benchmarks

  13. A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning

    Apr 25, 2026Qishi Zhan, Minxuan Hu, Liang He +2Bayesian Neural NetworksDeep Ensembles

  14. Introducing Background Temperature to Characterise Hidden Randomness in Large Language Models

    Apr 24, 2026Alberto Messina, Stefano ScottaLLM InferenceML Reproducibility

  15. Replicable Bandits with UCB based Exploration

    Apr 21, 2026Rohan Deb, Udaya Ghai, Karan Singh +1Multi-Armed BanditsContextual Bandits

  16. AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories

    Apr 21, 2026Xue Xia, Chengkai Yao, Mingyu Tsoi +10AI Coding AgentsML Reproducibility

  17. Improving reproducibility by controlling random seed stability in machine learning based estimation via bagging

    Apr 20, 2026Nicholas Williams, Alejandro SchulerDebiased MLEnsemble Learning

  18. A Random Matrix Theory Perspective on the Consistency of Diffusion Models

    Feb 2, 2026Binxu Wang, Jacob Zavatone-Veth, Cengiz PehlevanDiffusion Model SamplingML Reproducibility

  19. Glucose-ML: A collection of longitudinal diabetes datasets for development of robust AI solutions

    Jul 18, 2025Temiloluwa Prioleau, Baiying Lu, Yanjun CuiForecasting BenchmarksBlood Glucose Forecasting

  20. Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

    Jan 18, 2025Jialun Cao, Yuk-Kit Chan, Zixuan Ling +12Benchmark ConstructionBenchmark Design

  21. Replicability is Asymptotically Free in Multi-armed Bandits

    Feb 12, 2024Junpei Komiyama, Shinji Ito, Yuichi Yoshida +1Multi-Armed BanditsML Reproducibility

  22. No One Knows the State of the Art in Geospatial Foundation Models

    Date pendingIsaac Corley, Nils Lehmann, Caleb Robinson +6Geospatial Foundation ModelsML Reproducibility