cs.CRMar 6, 2025

The Challenge of Identifying the Origin of Black-Box Large Language Models

Authors: Ziqing Yang, Yixin Wu, Yun Shen, Wei Dai, Michael Backes, Yang Zhang

Abstract

The tremendous commercial potential of large language models (LLMs) has heightened concerns over their unauthorized use. To address this, we focus on the task of identifying the origin of black-box LLMs. We further propose PlugAE, an effective and efficient identification method that proactively leverages LLM-specific adversarial embeddings and allows users to customize copyright tokens on a targeted query set. Extensive experiments demonstrate that PlugAE outperforms both state-of-the-art model watermarking and fingerprinting methods in accuracy and robustness. We further analyze its stealthiness and reliability from three complementary perspectives and conduct ablation studies under various configurations, confirming its practicality for real-world misuse detection.

Explore similar work

CardsList
  1. Detecting Data Contamination in Large Language Models

    Apr 21, 2026Juliusz Janicki, Savvas Chamezopoulos, Evangelos Kanoulas +1Membership Inference AttacksContamination