We introduce SAIL, an open model with 35B total and 3B active parameters for literature research, scientific coding, and multi-step research workflows. SAIL is developed through a science-aware improvement loop: agents built on frontier AI models analyze its task failures and construct training tasks that address the underlying capability gaps. The diagnosis examines search and evidence selection in literature tasks, scientific assumptions and reasoning in coding, and planning and revision in longer investigations. The agents draw on paper collections and scientific code repositories to build problems, interaction trajectories, and executable tasks with the required environments and tools. We repeat this loop over multiple development cycles and train SAIL through supervised fine-tuning, specialist training, multi-teacher on-policy distillation, and agentic reinforcement learning. SAIL achieves competitive performance across scientific research tasks with substantially fewer parameters than leading open-weight models.
Figures & tables
Figure 1 : Performance across scientific benchmarks.
Figure 2 : Performance versus total parameters. Each logo represents a model, with total parameter count on a logarithmic horizontal axis and the unweighted mean score across twelve evaluations on the vertical axis. The aggregate excludes LitQA2-FullText and counts E2E-Bench Basic and Hard separately; all included scores are expressed on a 0–100 scale. Small dots and leader lines indicate the true coordinates of displaced logos.
Figure 3 : Scientific validity beyond successful execution. a, Schematic orbital simulation: forward Euler integration completes without runtime errors while accumulating energy drift. The comparison with a symplectic method illustrates the role of domain knowledge in selecting numerical methods and checking long-term behavior. b, Scientific knowledge guides method selection and the interpretation of computational results, which in turn guide subsequent actions.
Figure 4 : The science-aware improvement loop. a, Evaluation, diagnosis, task construction, and training repeat as the model improves. Agents built on frontier AI models drive diagnosis and construction. b, Examples of scientific reasoning failures and the problems, trajectories, and environments used to address them. c, An illustrative diagnosis: ignored energy drift motivates training on the relevant conservation principles and their use during execution.
Figure 5 : SAIL Training Recipe. The SFT checkpoint initializes both the specialist models and the student policy. Specialists undergo targeted SFT, RL, or both, and provide supervision for multi-teacher on-policy distillation (MOPD). The student then undergoes agentic reinforcement learning to obtain the final SAIL model. Solid arrows indicate model initialization or training progression; dashed arrows indicate teacher supervision.
Figure 6 : Infrastructure for Scalable On-Policy Training. The system decouples model rollout, task execution, and optimization. Agentic RL and MOPD share the same student trajectory collection pipeline, with policy weights synchronized between training rounds.
Model
Params. (B) Total / active
Paper Findings
LitQA search
ScholarQA CS2
LitQA2 FullText
Arxiv DIGESTables
Larger-scale models
Ling-3.0-flash inclusionAI (2026a)
124 / 5.1
18.03
26.67
67.88
91.23
25.45
DeepSeek-V4-Flash-0731 ( DeepSeek-AI, 2026 )
284 / 13
26.46
74.67
75.19
95.08
35.13
Hy3 ( Tencent Hy Team, 2026 )
295 / 21
28.90
73.33
85.63
94.12
32.42
MiMo-V2.5 ( Xiaomi MiMo Team, 2026a )
310 / 15
16.21
34.67
60.85
94.23
27.30
GLM-5.2 ( Z.ai, 2026 )
744 / 40
40.80
88.00
87.87
90.41
34.21
Table 1 : AstaBench results across literature understanding, code execution, data analysis, and end-to-end discovery. Higher scores are better.
Model
Params. (B) Total / active
SciCode
DeepResearch Bench II
Larger-scale models
Ling-3.0-flash ( inclusionAI, 2026a )
124 / 5.1
38.19
41.73
DeepSeek-V4-Flash-0731 ( DeepSeek-AI, 2026 )
284 / 13
39.17
43.22
Hy3 ( Tencent Hy Team, 2026 )
295 / 21
38.19
42.54
MiMo-V2.5 ( Xiaomi MiMo Team, 2026a )
310 / 15
27.64
27.46
GLM-5.2 ( Z.ai, 2026 )
744 / 40
47.57
45.51
Table 2 : Additional benchmark results on scientific coding and research. Higher scores are better.
Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain conventions, and specialized correctness criteria make this knowledge difficult to convert into reliable learning experience-a challenge we call the scientific experience bottleneck. We introduce ScienceIDE, infrastructure for turning the world's scientific code into programmable environments for scientific agents. Guided by expert-defined scientific cases and acceptance criteria, agents transform repositories into executable environments that support task generation, execution, and scientific verification. These environments provide a shared foundation for supervised fine-tuning, reinforcement learning, and evaluation. Using verified interaction trajectories, we train PhAI-IDE-72B, PhAI-IDE-9B, and PhAI-IDE-4B. The model family shows gains in held-out scientific-code repair and across selected general-purpose benchmarks in code, reasoning, and knowledge, providing evidence of positive transfer from scientific experience to broader capabilities. ScienceIDE lays the foundation for an integrated workspace for agent learning and scientific practice, making humanity's scientific software a shared substrate for developing scientific intelligence. Code: https://github.com/aitofound/ScienceIDE
Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answering, and therefore fail to assess key capabilities required in realistic scientific research workflows. We introduce SciExplore, a benchmark designed to evaluate scientific information-seeking and reasoning capabilities of LLMs and agents. SciExplore comprises four task types covering 103 expert-curated tasks across more than ten scientific disciplines: scientific database navigation, ambiguous literature retrieval, missing reference completion, and cross-source structured knowledge synthesis, which probe progressively higher-level abilities from entity-level reasoning and document-level identification to evidence-level grounding and domain-level synthesis. We evaluate over ten state-of-the-art LLMs and autonomous agents on SciExplore, revealing substantial performance gaps with performance degrading sharply as task complexity increases and extremely low accuracy on the most challenging structured synthesis tasks. These results highlight significant limitations of current models and agents in realistic scientific information-seeking scenarios.
Yinhao Tang, Youqing Fang, Yanan Sun +6
University of Science and Technology of China · Shanghai AI Laboratory
As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Across 16 benchmarks spanning real-world research, engineering, and digital work, Atria Dawn Preview is competitive with frontier agents and achieves the highest reported score on five of them. Beyond standalone performance, we examine the real research-and-development process behind this model as a case study of human--AI collaboration, analyzing 769 task records from 56 participants together with agent logs. When asked to evaluate completed tasks under comparable conditions, participants rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently propose methods and implement revisions, while humans retain most final decisions and guide exploration through judgment and feedback. These observations indicate a shift from task-level execution to project-level partnership, with human effort concentrating on what is worth pursuing and how evidence should guide research. Progress toward more autonomous AI research must therefore advance both the capacity for discovery and the capacity for meaningful human oversight, preserving accountable human authority over the risks and direction of continued development.