AI for Science

Latest papers 117

All topics
CardsList
  1. OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences

    Oct 8, 2026Zhiyi Li, Sihan Hu, Tianning Xiao +4AI for ScienceScientific Reasoning in Language Models

  2. SciExam for ENSO: Can AI Agents Build Climate Models?

    Oct 7, 2026Yinling Zhang, Langchen Liu, Dongbin Xiu +4AI for ScienceAI Agent Benchmarks

  3. OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

    Sep 30, 2026Dingyuan Dai, Heli Qi, Lei Liu +28Computer-Use Agent BenchmarksAI for Science

  4. Searching for BSM Experimental Signatures with Large Lagrangian Models

    Sep 29, 2026Ibrahim Elsharkawy, Victoria Knapp-Perez, Wahid Bhimji +1AI for ScienceHigh-Energy Physics

  5. ASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU Workstations

    Sep 26, 2026J. Paul Liu, Uthpala Herath, Andrew PetersenAI for ScienceHigh-Performance Computing

  6. PhyMo: A Physical-Field Modality for Multimodal AI4Physics

    Sep 23, 2026Henan Sun, Haitao Hu, Jin Liu +4AI for ScienceCross-Modal Representation Learning

  7. PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research

    Sep 15, 2026Xinle Yu, Fan Bai, Kaiser Sun +5AI for ScienceAutonomous Research Agents

  8. The AI-Enabled Scientific Frontier

    Sep 14, 2026Gabriel Manso, Emma Fu, Neil ThompsonAI for ScienceAI-Assisted Scientific Research

  9. AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era

    Sep 7, 2026Yunxiang Mo, Tianshi Zheng, Yisen Gao +7AI for ScienceAI Agent Evaluation

  10. BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks

    Aug 26, 2026Zane Koch, Asmamaw T. Wassie, Javier Valdes-Aleman +5AI for ScienceAI Agent Benchmarks

  11. K-Bench: measuring model performance on real scientific agent requests

    Aug 21, 2026Aubrey M. Brueckner, Darshil Patel, Yuhuan He +1AI for ScienceLLM Agent Evaluation

  12. The ethics of artificial intelligence in the life sciences: Universality, cultural diversity and an architecture of care

    Aug 5, 2026Jean-Pierre Changeux, Gustavo Deco, Morten L. KringelbachAI for ScienceHuman-Centered AI

  13. Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

    Aug 4, 2026Zejun Liu, Jian Wu, Ru Peng +4AI for ScienceBenchmark Design

  14. POMDPs for Autonomous Science Exploration

    Aug 4, 2026Daniel Guirguis, Nathan Wallace, Hanna Kurniawati +1Belief-Space PlanningAI for Science

  15. Towards a new paradigm of scientific discovery with socialized artificial intelligence

    Aug 3, 2026Xinjie Yao, Xingxin Xu, Xiyuan Gao +21AI for ScienceHuman-in-the-Loop AI

  16. Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable

    Aug 2, 2026Wenhui Chen, Jianlin Chen, Ziyao Lin +1Reward HackingAI for Science

  17. Artificial Intelligence and Modeling & Simulation: An Overview

    Aug 1, 2026Niclas Feldkamp, Philippe J. Giabbanelli, Istvan DavidAI for Science

  18. An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research Loop

    Jul 30, 2026Yiwen Zhang, Eloise Zeng, Jaeha Lee +1AI for ScienceHuman-in-the-Loop AI

  19. TCA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration Retrieval

    Jul 30, 2026Yuto Suzuki, Farnoush Banaei-KashaniAI for ScienceScientific Hypothesis Generation

  20. SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition

    Jul 30, 2026Yuqi Tang, Chenyi Zhou, Libin Wang +3AI for ScienceLLM Agent Skill Learning

  21. Can AI Follow In Einstein's Footsteps?

    Jul 30, 2026Michael Shalyt, Nathan Regev, Marin Soljačić +1AI for ScienceScientific Hypothesis Generation

  22. Agentic Autoresearch for CT Reconstruction

    Jul 24, 2026Andreas Maier, Lucas Kachelriess, Siming Bayer +4AI for ScienceSparse-View CT Reconstruction