Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects
Organizations: Blossom AI Labs Blossom AI Tokyo, Japan
Abstract
Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.
Figures & tables
| Record | Observed evidence | Supports | Does not support |
|---|---|---|---|
| Assets | 113 IDs and linked audio; 1.552 h; 32 seeds; one empty Focus | cross-artifact identity and observed coverage | target correctness or representativeness |
| Realism | 75 final-rubric items; one label each; 58 usable, 17 minor; 72 natural | selected input/audio plausibility | paired-target validity or agreement |
| Correction | 12 notes; one rater each; eight changed | exploratory target defects | certified gold or error prevalence |
| Error cards | selective final records with source-linked dispositions and human additions; sparse overlap | review of displayed plant keys | detector precision/recall |