cs.CRSep 26, 2026

Reading Is Not Leaking: Local, Auditable Measurement and Reduction of Inference Exposure from Public Footprints

Authors: Mahmudul Faisal Al Ameen

Abstract

Anyone with a public footprint leaks facts that were never stated, and language models make the inference cheap. We present a framework for measuring and reducing this inference exposure that runs on the owner's own CPU with no language model at analysis time, instantiated on organisations and on individuals. It starts from a measurement result: scoring an inference system against the target's private truth conflates how well the system reads the record with how much the record leaks. On a 128-question instrument over sixteen synthetic firms, almost half of the questions are never answered correctly by any of six readers, four of them language models, and a majority-class guess accounts for most of every reader's score. We therefore separate reading accuracy from leakage rate and introduce an injection protocol that creates cells with known support. Our analyser combines rules, statistical solvers and a 106M-parameter encoder trained from scratch that marks verbatim evidence and never generates text; every answer carries a graded certificate whose recorded proof replays. Its certified answers are correct in 93% of resolved cases, against 49-73% for the language models' quote-backed answers, whose citations are produced alongside the answer rather than deriving it; with plain-prose articles in the record, 70% of its evidence-bearing answers rest on evidence that establishes them, against 18-56% for the models. A constrained defence that rewrites each fact's carrier as a true but coarser statement hides every single-carrier fact from four language-model adversaries at 40% lower edit cost than deletion. On sixteen synthetic people the guessing term is larger still, and a decoy planner with no language model halves the correct answers of the estimator it targets without transferring to a second.

Explore similar work

Sep 18, 2026cs.AI

A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal

Large language models can hold knowledge they do not report. A model may sandbag on a capability evaluation, or answer against what it internally knows, and its outputs alone cannot tell whether it is hiding an answer or simply does not have one. We borrow the Concealed Information Test, a forensic method that identifies guilty knowledge by presenting a suspect with the true detail among plausible decoys and measuring a stronger response to the item they recognize. Our method, Probe of Internal Recognition (PIR), does the same inside a model. It presents a question with its candidate answers and reads, from the model's internal states, which candidate the model recognizes as correct. PIR is reference-free, needing no honest reference model and no labeled truth corpus. Across eight models from five families (Gemma, Qwen, Llama, Mistral, and Phi), PIR recovers the recognized answer at 0.70 to 0.87 balanced accuracy, well above the 0.28 to 0.40 unknown-item baseline and the 0.25 chance rate. It stays readable across every form of concealment we test, from prompted deception and trained sandbagging to external password-locked and circuit-broken checkpoints, with recognition between 0.85 and 0.93. When the model hides a known answer, recognition stays high. When unlearning removes the knowledge, recognition drops to the level of a question the model never knew. PIR therefore separates a model that will not answer from one that cannot, which supports sandbagging audits and unlearning verification. The signal is causal, adds information beyond black-box behavioral cues, and extends from multiple-choice questions to free-form generation.
Jul 31, 2026cs.LG

Leak It: Per-Document Extraction Beyond Aggregate Membership Inference

Membership inference (MIA) on language models is usually summarised by aggregate ROC-AUC, but such evaluations are confounded: model-free blind baselines can separate members from non-members using surface text alone. Building on probabilistic discoverable extraction, we study black-box training-data leakage using N samples from p_theta(. | x), placing mean overlap, extreme-value overlap, and self-concentration on a common functional-estimation footing. On WikiMIA, a blind bag-of-words classifier reaches AUC 0.97 (TPR 0.90 at 5% FPR) while sampling adds nothing. On an IID Pile split (MIMIR), neither self-concentration nor gold-continuation recovery significantly exceeds a blind baseline in aggregate. Aggregate metrics hide the real harm: sampling verbatim-extracts training data for a tail of documents no blind attack can reach. On Pythia-6.9B, 16.6% of 500 Pile documents bearing a real identifier (83 documents; 21.3% of those bearing an email address) have that identifier reproduced and not reproduced under a mismatched-prefix control. Each leak is attributable to that document rather than a globally common string. This per-document disclosure is invisible to aggregate AUC. Risk is uneven: identifier leakage is about 3x stronger in code than prose, though prose remains positive and grows with capacity (4.0% to 12.1% from 410M to 6.9B); recovery of arbitrary held-out continuations is essentially confined to code (+0.44 member gap on GitHub vs at most +0.014 on prose). Temperature and nucleus sampling have minor effect, a 16-token prefix suffices, and the sample-budget relationship corroborates prior probabilistic-extraction results. We detect no reduction from deduplication. Privacy audits should report per-document extraction, not only aggregate membership, and motivate differential privacy as the mitigation. We release leakit, a black-box tool implementing this probe and its control.
Sep 26, 2026cs.CR

Checking Leakage Witnesses versus Certifying Bounded Non-Leakage

When a language-model audit finds no leak, what is needed to certify non-leakage? We study guarantees over a declared prompt domain under an executable leakage criterion and decoding rule. For general bounded polynomial-time evaluators, a supplied leaking execution is polynomial-time checkable, while leak existence is \NP-complete and deterministic certification is \coNP-complete. Exact stochastic certification is \coNP\PP\coNP^{\PP}-complete at every fixed rational cutoff in (0,1)(0,1). Restricting the computation can change these bounds. For example, certification is in \coNP\ when all randomness is a terminal draw from an efficiently computed finite probability table. Attention models admit polynomial-time certification when local dependency windows of logarithmic length precede one global head, given deterministic decoding, fixed vocabulary, exact rational weighted means, a direct binary affine readout and finite-automaton prompt domains. A construction with two global layers instead makes certification \coNP-complete over template domains, with one head per layer, polynomial width, logarithmic precision and an inverse-polynomial logit margin. Planted-secret experiments measure what finite audits miss relative to complete references. Among 30 secret--model-state pairs that leak under greedy single-prompt execution on their secret's 4,096-prompt domain, uniformly selecting 256 recorded evaluations per pair misses every leak for an expected 41.06%41.06\% of these pairs. Batched and single-prompt checks disagree on one complete-domain decision among all 48 fine-tuned pairs, while a same-order repeat reproduces every single-prompt output. These results distinguish computational conditions for certification from the coverage and execution conditions needed to interpret a negative audit.