We introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tuning (SFT). Building on this SFT foundation, we propose Interleaved Distillation and Reinforcement Learning (IDRL), a novel post-training paradigm that tightly integrates on-policy distillation and reinforcement learning within a single training loop. By alternating between the two objectives, rather than optimizing either in isolation or combining them in a static joint loss, IDRL enables more effective knowledge transfer, greater optimization stability, and more precise credit assignment for long-horizon agentic trajectories, leading to stronger performance while mitigating common capability trade-offs. At both model scales, the framework produces two specialized variants: Pistis-Thinking, designed to strengthen deep multimodal reasoning, and Pistis-Agentic, which additionally incorporates agentic trajectory data to support long-horizon planning, iterative reasoning, and tool use. Pistis-Agentic is particularly strong in multimodal search. Both scales outperform their corresponding base models. Beyond model-parameter optimization, we further introduce Pistis-Auto-Harnessing (PAH), a system-level method that automatically improves the agent's inference harness through iterative optimization. Experiments demonstrate that PAH enhances the model performance without updating the model parameters or increasing the interaction budget.
Figures & tables
Figure 1: Overview of Pistis capabilities and performance. The understanding and reasoning panel compares Pistis-Thinking with the corresponding base models on MathVista-mini, MM-Vet, HallusionBench, OCRBench, AI2D-test, RefCOCO+ testA, Charades-STA (64 frames), and MVBench (8 frames) ( Lu et al., 2024 ; Yu et al., 2023 ; Guan et al., 2024 ; Liu et al., 2024 ; Kembhavi et al., 2016 ; Yu et al., 2016 ; Gao et al., 2017 ; Li et al., 2024 ) . The agentic and search panel compares Pistis-Agentic with the corresponding base models on MMSearch, BrowseComp-VL, LiveVQA, PinchBench, TreeBench, HRBench4K, LogicVista, and MathVerse-mini ( Jiang et al., 2024 ; Geng et al., 2025 ; Fu et al., 2025 ; Kilo Code, 2026 ; Wang et al., 2025a ; Wang et al., 2025c ; Xiao et al., 2024 ; Zhang et al., 2024b ) . All values match Tables 4 and 5 .
Figure 2: Overview of Interleaved Distillation and Reinforcement Learning (IDRL).
Figure 3
Reward scope
Task domain
Rule
Model
Binary
Reward design details
Format
All Domains
✓
✓
Score 1 if <think> and </think> tags are present and strictly paired; otherwise 0 .
STEM
Math
✓
✓
Numerical calculation & multiple-choice: Rule-based exact match; score 1 for correct, 0 for incorrect.
Physics
✓
✓
Chemistry
✓
✓
Long Document Chart & OCR
Long Document
✓
Open-ended QA: Model-based evaluation using a strong judge model.
Chart
✓
✓
Numerical calculation & multiple-choice: Rule-based exact match; score 1 for correct, 0 for incorrect.
Table 3: Task-aware reward design in the IDRL stage of Pistis, including a global format constraint and domain-specific correctness rewards.
Figure 4: Overview of Pistis-Auto-Harnessing (PAH). During development, an Optimization Agent runs a five-stage closed loop on a fixed development set: it attributes failures, proposes one falsifiable change, implements it, checks it with a small canary run, and evaluates it on the complete development set. A candidate replaces the current best harness only when the development metric improves; otherwise it is rolled back. The Optimized Harness is then frozen and evaluated once on a disjoint test set. At runtime, the Optimized Harness supports the frozen Pistis-Agentic model with a Candidate Ledger, conditionally loaded Search Skills, and bounded checkpoints with budget control, while the model itself makes every decision and writes the final answer.
Figure 5: Architecture of PistisEvalKit.
Figure 6: Effect of OPD objective design on training stability. We compare entropy (a) and gradient norm (b) across RKL, JSD-5, and JSD-50. Broader top- k JSD preserves higher policy entropy and yields smoother gradients, while RKL and narrow-support JSD exhibit entropy collapse or instability.
Figure 7: Effect of interleaving OPD and RL during post-training. Compared with vanilla RL, vanilla OPD, and the joint RL+OPD objective, IDRL shows the expected phase-wise entropy pattern (a) and more stable gradient norms (b), indicating that alternating the two objectives reduces optimization interference.
Category
SFT
RL
OPD
OPD → RL
RL+OPD
IDRL
Chart Understanding
81.0
81.5
79.6
81.2
81.3
82.0
Real-World Perception
77.5
77.8
78.5
78.5
77.7
78.4
Multimodal Reasoning
78.5
79.0
77.8
78.9
78.9
79.4
Search-Oriented
57.1
57.3
56.2
56.5
57.9
58.4
Claw-Style
77.1
77.5
74.4
77.2
77.5
77.6
AVG
73.7
74.0
73.4
74.1
74.1
74.7
Table 8: Comparison of IDRL against vanilla RL, OPD, the sequential OPD-then-RL pipeline (OPD → RL), and joint RL+OPD optimization on Pistis-9B-Agentic across agentic benchmark categories. All models start from the SFT checkpoint. AVG gives equal weight to each of the 18 individual benchmarks.
Modality
Capability
InternVL3.5 8B thinking
Keye-VL-1.5 8B thinking
Qwen3-VL 8B thinking
Qwen3.5 9B thinking
Pistis 9B thinking
Pistis † 9B thinking
Image
OCR
70.8
72.2
76.6
81.4
82.1
82.8
Grounding
52.6
56.8
71.2
69.7
73.8
80.3
Attribute
84.2
96.7
99.3
99.5
99.6
99.5
Multilingual
73.9
85.9
91.5
90.7
86.5
92.3
Video
Reasoning
83.1
84.3
87.3
86.0
92.5
89.5
Action recog.
91.5
86.0
82.9
94.5
92.3
97.0
Table 10: Evaluation of business-specific visual understanding capabilities for content safety and business integrity on Pistis Benchmark. † Pistis-9B is further fine-tuned on business-specific data, whereas all other models are evaluated in their original released form.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Representative failure cases of Pistis-Agentic in multimodal reasoning. (a) The age-group labels are read correctly, but the question intent (the gap between the highest and lowest plotted rates) is parsed literally as a gap between ages. (b) The drawn arrow, a salient symbolic cue, overrides the cross-frame evidence that the circle never moves. Model excerpts are abridged; red marks the faulty steps and final answers, and green marks the reference answers.
Figure 9: Representative failure cases of Pistis-Agentic in tool-integrated reasoning. In both rollouts, the code interpreter is used to confirm a preselected hypothesis rather than to measure: (a) only the hypothesized region is cropped and the candidate comparison is never performed; (b) the script prints a hard-coded count, and the printed output is then cited as confirmation. Excerpts are abridged; red marks the faulty steps.
Figure 10: Representative failure cases of Pistis-Agentic in agentic search. (a) Retrieval succeeds and the correct judgment (green) surfaces in the reasoning, yet the final answer contradicts it. (b) The model never invokes the available search tool and produces a hallucinated date with fabricated attribution. Excerpts are abridged; red marks the faulty steps.
The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinforcement learning with verifiable rewards (RLVR). However, SFT introduces distributional drift that neither preserves the model's original capabilities nor faithfully matches the supervision distribution. This problem is further amplified in multimodal reasoning, where perception errors and reasoning failures follow distinct drift patterns that compound during subsequent RL. We introduce PRISM, a three-stage pipeline that mitigates this drift by inserting an explicit distribution-alignment stage between SFT and RLVR. Building on the principle of on-policy distillation (OPD), PRISM casts alignment as a black-box, response-level adversarial game between the policy and a Mixture-of-Experts (MoE) discriminator with dedicated perception and reasoning experts, providing disentangled corrective signals that steer the policy toward the supervision distribution without requiring access to teacher logits. While 1.26M public demonstrations suffice for broad SFT initialization, distribution alignment demands higher-fidelity supervision; we therefore curate 113K additional demonstrations from Gemini 3 Flash, featuring dense visual grounding and step-by-step reasoning on the hardest unsolved problems. Experiments on Qwen3-VL show that PRISM consistently improves downstream RLVR performance across multiple RL algorithms (GRPO, DAPO, GSPO) and diverse multimodal benchmarks, improving average accuracy by +4.4 and +6.0 points over the SFT-to-RLVR baseline on 4B and 8B, respectively. Our code, data, and model checkpoints are publicly available at https://github.com/XIAO4579/PRISM.
Sudong Wang, Weiquan Huang, Xiaomin Yu +9
Hong Kong University of Science and Technology (Guangzhou) · Nanyang Technological University · Tsinghua University +3
Training language models (LMs) remains a highly human-intensive process, even as frontier language model agents become increasingly capable at software engineering and other long-horizon tasks. A central challenge is that autonomous post-training is not just a coding problem: it requires the agent to repeatedly plan iterations, construct benchmark-aligned data, run stable training jobs, evaluate checkpoints, and preserve experiment state across many hours of interaction. We present AutoTrainess, a LM agent that exposes these operations as a repository of agent-computer interfaces for planning, data preparation, training, evaluation, and logging. Rather than leaving the agent to operate in a raw CLI environment with an underspecified action space, AutoTrainess externalizes prior human experience as explicit workflows, rules, and execution constraints that guide the agent toward effective and reliable training behavior. On PostTrainBench, AutoTrainess consistently outperforms CLI-only baselines, achieving 26.94 average score with GPT-5.4 (Codex) versus 23.21 for CLI-only. It also generalizes across models and harnesses, improving DeepSeek-V4-Flash (OpenCode) from 12.13 to 19.58.
Zhaojian Yu, Penghao Yin, Shuzheng Gao +3
1Tsinghua University · 2The Chinese University of Hong Kong · 3Simple Agent Lab
Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research trajectories are generated, evaluated, and learned from. Existing pipelines typically train models under a fixed harness, including prompts, tools, skills, middleware, and memory, while leaving the data-generating process outside the optimization objective. This creates a mismatch between model updates and the static scaffolding that determines trajectory quality. We introduce Co-Harness, a framework that jointly optimizes the agent harness and model parameters during post-training. Co-Harness alternates between harness optimization and model optimization. An LLM-based HarnessCritic analyzes failed trajectories, identifies harness-level failure modes, and proposes validated local updates. The model is then fine-tuned on high-quality trajectories generated by the improved harness, distilling effective scaffolding into model parameters. A 200+ hour autonomous case study further shows that Co-Harness can recover from system crashes, improve inference efficiency, and discover ensemble strategies without human intervention. These results suggest that joint harness and model optimization is an effective way to improve agents beyond fixed-harness post-training.