ML Reproducibility

ML: Machine Learning

Momentum

14 papers in the last four weeks, up 100% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 88

All topics
CardsList
  1. Reproducible LLM Inference Benchmarking: A Sequential Isolation Protocol for Regression Testing

    Oct 7, 2026Arnold Olympio, Juan Manuel Servera Bondroit, Wael Abdelmalek +2LLM InferenceML Reproducibility

  2. Beyond the Model: The Critical Role of Data Filtering in Clinical Machine Learning

    Oct 5, 2026Noah Subedar, Colin Campbell, Wenjing Zhang +3Training Data CurationML Reproducibility

  3. Technical Report on the Turba Fertilizer Machine Learning Stack in Morocco

    Oct 5, 2026Abdelghani Belgaid, Zakaria Mahmoud, Fahd Chibani +3Surrogate ModelingML Reproducibility

  4. Measuring and Reducing Cross-Vendor Mismatch in Language Models

    Oct 4, 2026Erland Hilman Fuadi, Chong Tian, Xiaosong Ma +1Efficient Neural Network InferenceML Reproducibility

  5. Model validation in machine learning: A scenario-based guide from hold-out splits to nested group cross-validation in biomedical and applied research

    Oct 1, 2026Mehmet Baygin, Sengul Dogan, Turker TuncerData LeakageML Reproducibility

  6. Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation

    Sep 29, 2026Mario Sanz-Guerrero, Minh Duc Bui, Manuel Mager +1LLM EvaluationLanguage Model Generation Evaluation

  7. Making Cross-Continental Federated Learning Repeatable with FLIP: a Multi-Application Study

    Sep 28, 2026Rafael Garcia-Dias, Alexandre Triay Bagur, Chayanin Tangwiriyasakul +20HealthcareML Reproducibility

  8. Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models

    Sep 27, 2026Eduardo Ariño de la Rubia, Szilard PafkaAI Coding AgentsLLM Agent Evaluation

  9. RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?

    Sep 23, 2026Mithil Salunkhe, Haochen Ding, Samridhi Verma +1LLM Agent EvaluationAI Agent Benchmarks

  10. Reproducible AI Requires Reproducible Randomness

    Sep 22, 2026Anthony Bertrand, Tom Schmitt, Engelbert Mephu Nguifo +1ML ReproducibilityScientific Reproducibility

  11. Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

    Sep 22, 2026Liam Cooper, Shinnung Jeong, Hyeran Jeon +2GPU Kernel OptimizationLLM Inference Acceleration

  12. Exposing Blind Spots in Deep Imbalanced Regression Evaluation

    Sep 21, 2026Noah C. Puetz, Jens U. Brandt, Marc Hilbert +3Class-Imbalanced LearningML Reproducibility

  13. Rethinking How We Evaluate Methodological Progress in Health AI

    Sep 16, 2026Florent Pollet, Matthew McDermottML ReproducibilityLongitudinal EHR Modeling

  14. TuiML: Machine Learning for AI Agents

    Sep 16, 2026Nilesh Verma, Nick Lim, Albert Bifet +1Tool-Using AgentsML Reproducibility

  15. OPEN-1B: A Fully Auditable Training Run

    Sep 15, 2026John Donaghy, Brian Wilcox, Oğuzhan Ersoy +6Neural Network VerificationLLM Auditing

  16. Model Retirement Creates Reproducibility Risk in Biomedical AI Publications

    Sep 7, 2026Nathan Wolfrath, Meghan Conroy, Thomas Kosten +7ML ReproducibilityScientific Reproducibility

  17. Training seeds and model-selection stability in recommender-system evaluation

    Sep 2, 2026Juan Manuel Rodriguez, Oleg Lesota, Antonela TommaselRecommender SystemsML Reproducibility

  18. A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation

    Sep 1, 2026Hodong Lee, Sanghee Park, Dohoon Ryu +4ML ReproducibilityMultimodal Foundation Models

  19. MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines

    Aug 31, 2026Alireza Bayat Makou, Emirhan Böge, Phu Gia Hoang +5LLM InterpretabilityInterpretable ML

  20. Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data

    Aug 11, 2026Nicola Giuseppe Marchioro, Gabriele Padovani, Amal Gueroudji +5Data ProvenanceML Reproducibility

  21. Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR

    Aug 11, 2026Karamvir Singh Batra, Prathamjyot Singh, Ashima Sood +2ASR EvaluationML Reproducibility

  22. Why Ranking Anomaly Detection Algorithms Isn't as Reliable as You May Think

    Aug 5, 2026Simon Klüttermann, Jérôme Rutinowski, Frederik Polachowski +1Benchmark DesignML Reproducibility

  23. Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

    Aug 4, 2026Mohsen Hariri, Weicong Chen, Nahal Shahini +11LLM EvaluationTest-Time Scaling

  24. One Run Is Not an Idea: The Implementation Lottery in Automated Research

    Jul 29, 2026Jingjie Ning, Shanshan Zhong, Xiaochuan Li +2ML ReproducibilityAutonomous Research Agents