cs.CROct 8, 2026

Poster: A Preliminary Study of LLM Distillation Inference

Authors: Edward Chen, Yuntao Du

Organizations: Carmel High School Carmel, Indiana, USA · Purdue University West Lafayette, Indiana, USA

Abstract

Unauthorized model distillation, in which a model is trained on the outputs of a proprietary large language model (LLM), is a growing threat to model providers. We study distillation inference: determining whether a suspect model was distilled from another model or trained independently. We formulate this problem as a hypothesis test and estimate the behavior expected under each hypothesis by training shadow models: distilled shadow models learn from the teacher's reasoning traces, whereas independent shadow models learn only from reference answers. The auditor measures how closely each model predicts the teacher's reasoning outputs and then uses the shadow models to convert the suspect's score into a calibrated p-value. In a preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for the suspects, our test achieves a true positive rate of 1.0 at a significance level of 0.02. These results demonstrate the feasibility of using distillation inference to detect distillation attacks.

Figures & tables

Explore similar work

CardsList
  1. Reference-Based Distillation Detection in LLMs

    Jun 19, 2026Rajat Rawat, Sizhe Chen, Akshay Anand +3LLM AuditingTeacher-Student Learning

  2. Distillation Defenses Easily Break After Reinforcement Learning

    Sep 28, 2026Shidan Javaheri, Alexander Panfilov, Oliver Britton +2LLM SecurityLanguage Model Distillation

  3. Asking Back: Interaction-Layer Antidistillation Watermarks

    May 15, 2026Guang Yang, Amir Ghasemian, Fengchen Liu +3LLM AuditingLLM Watermarking