PrivCert: Certifying Statement Support under Differential Privacy
Organizations: Acompany Co., Ltd.
Abstract
Differentially private (DP) text generation can protect individual records, but privacy alone does not specify what evidence a released statement carries about the underlying data. We identify this as an evidence gap: a private report may contain plausible claims without indicating whether they are strongly supported by the private dataset. We introduce PrivCert, a framework for privacy-preserving reporting that makes statement support explicit through privacy-preserving certificates and emit-or-abstain decisions. As a canonical instantiation, PrivCert-PF (Proposal-and-Filter) separates data-independent candidate discovery from private support certification, emitting only statements whose support passes a private evidence test. We provide theoretical grounding for this framework by characterizing the limits of implicit evidence under DP, deriving a sharp privacy--honesty frontier for single-statement certification, and establishing a worst-case cost for fine-grained multi-statement certification. Experiments on synthetic tasks and TAB, WildChat, and Yelp show that explicit certification maintains low unsupported emission, while free-text DP baselines frequently produce low-support claims under the same declared support semantics. We further show that the PrivCert contract can be realized with histogram, sparse-vector, and Gaussian mechanisms, and use DP synthetic data to illustrate an important boundary: support in a private proxy does not automatically certify support in the original data. Together, these results position privacy-preserving reporting as an evidence-design problem: not only how to generate private text, but what a private report can substantiate about its underlying data.
Figures & tables
| Dataset | Unfiltered | Doc-DP | InvisibleInk (TinyLlama) | InvisibleInk (Qwen) | PrivCert -PF |
|---|---|---|---|---|---|
| TAB | 1.000 | 0.894 | 0.958 | 0.994 | 0.010 |
| WildChat | 0.947 | 0.729 | 0.647 | 0.619 | 0.014 |
| Yelp | 0.943 | 0.803 | 0.877 | 0.856 | 0.000 |
| EmitRate | UnsupEmit | FalseEmit | MeanSupport | Interval | In? | |
|---|---|---|---|---|---|---|
| 1 | 0.051 | 0.0217 | 0.295 | 0.275 | [0.000, 0.335] | |
| 4 | 0.180 | 0.0085 | 0.033 | 0.327 | [0.135, 0.335] | |
| 10 | 0.220 | 0.0046 | 0.014 | 0.288 | [0.162, 0.335] | |
| 20 | 0.274 | 0.0014 | 0.004 | 0.250 | [0.189, 0.335] |
| Sampled | Exact-unique | ||||
|---|---|---|---|---|---|
| Dataset | PrivCert -Hist | PrivCert -PF | PrivCert -SVT | PrivCert -PF | PrivCert -PF (Gaus) |
| TAB | 0.199 | 0.404 | 0.400 | 0.422 | 0.444 |
| WildChat | 0.183 | 0.220 | 0.151 | 0.151 | 0.162 |
| Yelp | 0.390 | 0.749 | 0.796 | 0.816 | 0.833 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | DP type | Composition unit | Sensitivity | Final guarantee |
|---|---|---|---|---|
| PrivCert -PF | pure DP | support queries | -DP | |
| Document-level DP | pure DP | aggregates | -DP | |
| InvisibleInk | zCDP | private decoding | DClip sensitivity | -DP |
| Unfiltered LLM | none | no private-data access | n/a | reference only |
| Dataset | Proposer | #out | EmitRate | UnsupEmit | FalseEmit | MeanSupport |
|---|---|---|---|---|---|---|
| TAB | public LLM | 3 | 0.003 | 0.003 | 1.000 | 0.006 |
| tag-template | 404 | 0.404 | 0.008 | 0.010 | 0.545 | |
| WildChat | public LLM | 24 | 0.024 | 0.001 | 0.042 | 0.132 |
| tag-template | 220 | 0.220 | 0.004 | 0.014 | 0.288 | |
| Yelp | public LLM | 18 | 0.018 | 0.002 | 0.111 | 0.124 |
| tag-template | 749 | 0.749 | 0.000 | 0.000 | 0.239 |
| Method | EmitRate | UnsupEmit | Runs w/ any unsup. | SupportedRecall | |
|---|---|---|---|---|---|
| 0.05 | PrivCert -PF | 0.823 | 0.000 | 0.979 | |
| 0.05 | Aug-PE-style ( ) | 0.800 | 0.000 | 0.952 | |
| 0.10 | PrivCert -PF | 0.695 | 0.001 | 0.938 | |
| 0.10 | Aug-PE-style ( ) | 0.676 | 0.092 | 0.881 |
| Dataset | UnsupportedEmit | NLI–Qwen Jaccard | In interval |
|---|---|---|---|
| TAB | 0.0033 | 0.94 | 5/5 |
| WildChat | 0.0023 | 0.33 | 5/5 |
| Yelp | 0.0037 | 0.78 | 5/5 |
| Dataset | Method | NLI FalseEmit | Qwen FalseEmit |
|---|---|---|---|
| TAB | Document-level DP | 0.894 | 0.470 |
| PrivCert -PF | 0.010 | 0.040 | |
| WildChat | Document-level DP | 0.729 | 0.857 |
| InvisibleInk (Qwen) | 0.619 | 0.929 | |
| PrivCert -PF | 0.014 | 0.241 | |
| Yelp | Document-level DP | 0.803 | 0.239 |
| Dataset | Per-candidate EmitRate | Family-wise EmitRate |
|---|---|---|
| TAB | 0.404 | 0.357 |
| Yelp | 0.749 | 0.565 |
| WildChat | 0.220 | 0.169 |