cs.AISep 23, 2026

Pistis Technical Report

Authors: Heyun Chen, Xiaohan Lan, Jiaxi Li, Zhilin Lu, Qi She, Weiwen Xu, Fei Yu, Yujie Zhong, +12 more

Organizations: Pistis Team, ByteDance

Abstract

We introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tuning (SFT). Building on this SFT foundation, we propose Interleaved Distillation and Reinforcement Learning (IDRL), a novel post-training paradigm that tightly integrates on-policy distillation and reinforcement learning within a single training loop. By alternating between the two objectives, rather than optimizing either in isolation or combining them in a static joint loss, IDRL enables more effective knowledge transfer, greater optimization stability, and more precise credit assignment for long-horizon agentic trajectories, leading to stronger performance while mitigating common capability trade-offs. At both model scales, the framework produces two specialized variants: Pistis-Thinking, designed to strengthen deep multimodal reasoning, and Pistis-Agentic, which additionally incorporates agentic trajectory data to support long-horizon planning, iterative reasoning, and tool use. Pistis-Agentic is particularly strong in multimodal search. Both scales outperform their corresponding base models. Beyond model-parameter optimization, we further introduce Pistis-Auto-Harnessing (PAH), a system-level method that automatically improves the agent's inference harness through iterative optimization. Experiments demonstrate that PAH enhances the model performance without updating the model parameters or increasing the interaction budget.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL

    Apr 30, 2026Sudong Wang, Weiquan Huang, Xiaomin Yu +9Multimodal ReasoningEfficient On-Policy Distillation

  2. AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

    Jun 30, 2026Zhaojian Yu, Penghao Yin, Shuzheng Gao +3Large Language Model TrainingLanguage Modeling

  3. Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents

    Jul 17, 2026Zhengyu Chen, Teng Xiao, Huaisheng Zhu +3Agent HarnessLarge Language Model Agents