Rethinking Visual Provenance: Detection and Watermarking Across Direct Visual Generation and LLM-Driven Code Rendering
Organizations: UNSW Sydney
Abstract
AI systems create images and videos with image/video generation models or by writing code and graphics descriptions that are then rendered. These routes can produce similar visible artifacts but expose different representations, intervention points, and provenance evidence. We develop a production-centered framework that compares detection and watermarking across both routes. An explicit verification specification distinguishes passive inference, message recovery, and authenticated provenance. We organize image, video, source-code, and rendering-aware watermarks by production stage. We examine the different requirements of generated images and video, plots and SVG, programmable video, and agent-composed workflows. Documented Claude, OpenAI, and rendering-tool interfaces connect the framework to concrete systems. We pose ten scoped research questions on identifiability, observability, fair comparison across stages, recoverable payload, reconstruction, synchronization, composition, hybrid local contribution, and private production-event authentication. The result is a conceptual research agenda grounded in published methods, inspected interfaces, and elementary boundary examples. It reports no experiments and claims no new theorems; its appendix results are elementary calculations, and documentation and source inspection establish interfaces, not empirical robustness.
Figures & tables
| Part | Results | Serves |
|---|---|---|
| Appendix A : observation boundaries | Definition A.1 ; Propositions A.2 , A.3 and A.5 ; Lemma A.4 ; Remark A.6 ; Corollary A.7 ; Example A.8 . | Sections 2 and 6 ; Question 1 |
| Appendix B : rendering, task preservation, payload | Definitions B.1 and B.9 ; Remarks B.2 , B.8 and B.11 ; Corollary B.3 ; Examples B.4 and B.6 ; Propositions B.5 and B.10 ; Lemma B.7 . | Sections 2 and 6 ; Questions 2 to 5 |
| Appendix C : temporal search and editing | Definitions C.1 , C.4 and C.10 ; Examples C.2 , C.3 , C.7 , C.8 and C.9 ; Proposition C.5 ; Lemma C.11 ; Remarks C.6 and C.12 . | Sections 2 and 6.4 ; Questions 7 and 8 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Location | Image/video generation | Code-based rendering |
|---|---|---|
| Request / text | Conditions or generated textual content | Prompt, code tokens, identifiers, literals |
| Representation | Noise, latent states, model parameters | Program, component graph, scene parameters, assets |
| Emission stage | Decoder or generated-frame intervention | Renderer or captured-frame intervention |
| Final pixels | Pixels, transforms, codec-aware representation | The same output-level possibilities as for image/video generation |
| Side information | Signed manifests, service records, disclosed states | Signed manifests, source identity, execution receipts |
| Symbol | Meaning | Where |
|---|---|---|
| Production graph (nodes, edges); parents of node | Section 2.2 | |
| Artifact, node map, and randomness at node | Section 2.2 | |
| Task instance; joint law of the additional inputs | Section 2.2 | |
| Production state (random, realized); set of allowed states | Section 2.2 , Appendix B | |
| Verifier’s observation; realized value | Section 2.2 , Appendix A | |
| Verification specification | Section 2.2 |
| Family | Examples | Observation boundary |
|---|---|---|
| Video-asset marks | RivaGAN, DVMark, ItoV, Video Seal [ 23 , 24 , 25 , 9 ] | Marked frames are the received signal; this does not establish source authorship. |
| Video noise interventions | VideoShield, VideoMark [ 31 , 32 ] | Corresponding inversion paths and keys are auxiliary access. |
| Video decoder / model interventions | LVMark, Video Signature, SPDMark [ 35 , 36 , 37 ] | Learned emission changes differ from marking source tokens. |
| Graphical video payload | Safe-Sora [ 75 ] | A graphical message is the payload; abstract screening alone does not fix every embedding-stage detail. |
| Stage inheritance | LoT-Pass [ 45 ] | A marked input image passes through image-to-video generation; persistence needs its own channel model. |
| Localized / image neighbors | TrustMark, Watermark Anything [ 21 , 22 ] | Raster-image recovery and localization precede additional temporal questions. |