cs.CROct 5, 2026

DP-ES: Differentially Private Evolution Strategies for Prompt Optimization

Authors: Ziniu Liu, Aiping Li, Yue Han, Han Yu, Junjian Zhang, Dong Zhu, Changjian Li, Shiqiang Zhang

Organizations: National University of Defense Technology · CRRC Zhuzhou Electric Locomotive Research Institute Co., Ltd., China Academy of Railway Sciences

Abstract

Token-level differentially private (DP) prompt optimization methods such as DP-OPT can become unstable under tight privacy budgets: on GSM8K, DP-OPT obtains 49.5±28.5%49.5\pm28.5\% across 30 runs, and a logged search trajectory reveals prompt-template drift and noise-sensitive irreversible choices. We diagnose these as structural consequences of greedy token-by-token construction over privately aggregated counts. We then propose DP-ES (Differentially Private Evolution Strategies), a structurally cleaner alternative that maintains a population of full prompts, mutates them via LLM calls that never access the private dataset, and spends privacy only on sampled-Gaussian evaluation; deterministic or Gumbel-smoothed selection is post-processing. Under a conservative (ε≤1.0,δ=10−5)(\varepsilon\leq1.0,δ=10^{-5}) guarantee, DP-ES achieves 88.1% on GSM8K (+38.6 pp over DP-OPT, approximately 9 times lower standard deviation), 99.7% on MedQA, 73.5% on BANKING77, and 86.8% on Alpaca. It is also 2.5 times faster in wall-clock time and uses 3.3 times fewer logged private-data call groups than DP-OPT. Selection and population ablations, implementation-level noise checks, and a 200-profile exact-match memorization stress test complement the formal guarantee. Scope: Our experiments establish optimization robustness under DP noise, especially where prompt structure is critical; end-to-end validation on genuinely sensitive, non-saturated deployment data remains future work.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MAPLE: Metadata Augmented Private Language Evolution

    Feb 26, 2026Eli Chien, Yuzheng Hu, Ryan McKenna +3Synthetic DataAlphaevolve

  2. DP-SelFT: Differentially Private Selective Fine-Tuning for Large Language Models

    May 17, 2026Haichao Sha, Zihao Wang, Yuncheng Wu +2Large Language Model Fine-TuningContinual Fine-Tuning