cs.CRSep 28, 2026

Similarity Is Not Validity: Defending LLM Semantic Caches Against Poisoning

Authors: Zihan Zhang, Shuangjie Yao, Zesen Liu, Zhixiang Zhang, Wai Ip Lai, Dung Hiu Hilton Yeung, Chun Kit Zhang, Fuchen Ma, +3 more

Organizations: The Hong Kong University of Science and Technology · Tsinghua University

Abstract

Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries. However, retrieval is based solely on embedding similarity between the incoming query and cached queries. This design enables cache poisoning: an attacker can cache a malicious response under a query with high cosine similarity to benign requests. The vulnerability stems from a gap between retrieval similarity and answer validity. From an information-bottleneck perspective, query embeddings can lose information needed to distinguish valid from invalid cache hits, which limits any matching algorithm that uses only these embeddings. We propose a novel defense that recovers this necessary information from the raw text of the cache key. Across poisoning attacks, adversarial queries share a rewrite-residual structure: they pair a rewrite of the target query with residual content. The rewrite maintains high similarity, while the residual elicits the malicious response. Deleting the residual makes the remaining rewrite more similar to the incoming query. We exploit this structure using Deletion Gain to search shortened variants of the cached query for similarity gains, and an Answer Check to test whether the removed text contributes to the stored answer. We prove that Deletion Gain stays positive when a deletion leaves text close enough to the rewrite, and we search for such deletions with a sliding window. Across three poisoning attack classes, our defense blocks 82.0% to 98.2% of poisoned entries at a 5% false-positive rate, with negligible serving overhead.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LaCache: Robust Semantic Caching for LLM Serving

    Aug 3, 2026Jiacheng Liang, Yuhui Wang, Tanqiu Jiang +1Attacker Large Language ModelLarge Language Model Serving

  2. From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching

    Jan 30, 2026Zhixiang Zhang, Zesen Liu, Yuchong Xie +2Attacker Large Language ModelLarge Language Model Memory

  3. MVR-cache: Optimizing Semantic Caching via Multi-Vector Retrieval and Learned Prompt Segmentation

    May 24, 2026Ali Noshad, Zishan Zheng, Yinjun WuCacheSimilarity and Metrics