DP-ES: Differentially Private Evolution Strategies for Prompt Optimization
Organizations: National University of Defense Technology · CRRC Zhuzhou Electric Locomotive Research Institute Co., Ltd., China Academy of Railway Sciences
Abstract
Token-level differentially private (DP) prompt optimization methods such as DP-OPT can become unstable under tight privacy budgets: on GSM8K, DP-OPT obtains across 30 runs, and a logged search trajectory reveals prompt-template drift and noise-sensitive irreversible choices. We diagnose these as structural consequences of greedy token-by-token construction over privately aggregated counts. We then propose DP-ES (Differentially Private Evolution Strategies), a structurally cleaner alternative that maintains a population of full prompts, mutates them via LLM calls that never access the private dataset, and spends privacy only on sampled-Gaussian evaluation; deterministic or Gumbel-smoothed selection is post-processing. Under a conservative guarantee, DP-ES achieves 88.1% on GSM8K (+38.6 pp over DP-OPT, approximately 9 times lower standard deviation), 99.7% on MedQA, 73.5% on BANKING77, and 86.8% on Alpaca. It is also 2.5 times faster in wall-clock time and uses 3.3 times fewer logged private-data call groups than DP-OPT. Selection and population ablations, implementation-level noise checks, and a 200-profile exact-match memorization stress test complement the formal guarantee. Scope: Our experiments establish optimization robustness under DP noise, especially where prompt structure is critical; end-to-end validation on genuinely sensitive, non-saturated deployment data remains future work.
Figures & tables
| Method | Priv. | Disc. | Mechanism |
| TextGrad | × | ✓ | LLM gradients |
| OPRO | × | ✓ | LLM scoring |
| EvoPrompt | × | ✓ | Evolutionary search |
| PromptBreeder | × | ✓ | Evolutionary search |
| DP-SGD | ✓ | × | Gaussian noise |
| DP-OPT | ✓ | ✓ | Histogram + Exp |
| private data, size | candidates, parents, rounds | ||
| batch size, | Gaussian noise multiplier | ||
| per-record utility | clipping bound | ||
| privatized score | Gumbel smoothing scale |
| Method | GSM8K | MedQA | BANK77 | Alpaca † | Privacy |
| Non-DP | 95.0 0.9 | 100.0 0.0 | 75.0 2.8 | 91.7 0.6 | None |
| DP-OPT | 49.5 28.5 | 93.5 12.0 | 75.3 2.8 | 87.1 1.4 | (1.0, ) |
| DP-ES | 88.1 3.2 | 99.7 0.5 | 73.5 2.2 | 86.8 0.8 | ( 0.71, ) |
| (ES Opt) | +38.6 | +6.2 | 1.8 | 0.3 | – |
| (ES Non) | 6.9 | 0.3 | 1.5 | 4.9 | – |
| Selector | Accuracy (%) | Privacy |
| Gumbel-smoothed | post-processing | |
| Deterministic | post-processing |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | at | ||
| All reported sampled-score runs | 200 | 0.05 |
| Task | DP-OPT | DP-ES (Ours) | Speedup |
| DeepSeek API ( , 1,319 test examples, 3 iterations): | |||
| GSM8K | 45.7 hours | 17.9 hours | 2.5 |
| MedQA | 18.4 hours | 7.0 hours | 2.6 |
| Local Qwen2.5-7B ( , 200 examples): | |||
| GSM8K | 14.60 hours | 7.37 hours | 1.98 |
| MedQA | 3.24 hours | 1.87 hours | 1.73 |
| Cost component | DP-OPT | DP-ES | Ratio |
| Logged private-data call groups: | |||
| Mutation (token-level / privacy-free) | 3 (DP) | 0 (free) | — |
| Evaluation (Gaussian-noised) | 30 | 10 | 3.0 fewer |
| Total private-data groups | 33 | 10 | 3.3 fewer |
| Token usage: | |||
| Input tokens | 1,309 | 6,120 | 4.7 more |
| Method / diagnostic | Runs | Accuracy | Observed structure |
| DP-OPT aggregate | 30 | high run-to-run variance | |
| DP-ES aggregate | 30 | lower run-to-run variance | |
| DP-OPT seed 42 trajectory | 1 | 91.0% | 21/33 candidates omit {question} |
| Method | GSM8K | MedQA |
| DP-ES | ||
| DP-OPT | ||
| PromptDPSGD | ||
| Non-DP |
| Population | Iterations | Accuracy | vs Baseline |
| 3 | 4 | 78.5% | -8.0 pp |
| 6 | 3 | 86.5% | – |
| 9 | 2 | 82.0% | -4.5 pp |
| 12 | 2 | 82.0% | -4.5 pp |