Generating the Wild: Individual-Consistent Image-to-Video Generation for Wildlife
Organizations: School of Computer Science, University of Auckland Auckland, New Zealand
Abstract
Individual-level wildlife identification often suffers from data scarcity, as varying observations of the same animal under diverse poses, viewpoints, and motions are rarely available. Image-to-video (I2V) generation offers a promising way to mitigate this limitation by synthesizing additional observations from a single reference image. However, existing I2V models mainly emphasize global layout, semantics, and motion, and therefore often fail to preserve fine-grained local appearance cues that distinguish one wildlife individual from another, such as fur texture, stripe boundaries, spot configurations, and contour transitions. We observe that these identity-critical cues are closely related to high-frequency information. To address this challenge, we propose WildIcon, a high-frequency-guided I2V framework for wildlife individual consistency. Specifically, WildIcon introduces a frequency-aware identity encoding branch that extracts individual-specific high-frequency cues from the reference image. Combined with isolated foreground information, the resulting identity tokens are then injected into cross-attention blocks as identity conditioning. Building on a frozen backbone with lightweight identity adaptation, WildIcon preserves fine-grained identity cues visible in the reference image while retaining the motion controllability and semantic fidelity of the base I2V model. In addition, to support the training and evaluation of wildlife individual-consistent I2V, we construct WildlifeVid, a wildlife-centric video dataset with high-quality, temporally coherent clips and individual-level identity labels. Experiments on I2V generation and downstream animal re-identification (ReID) show that WildIcon achieves stronger individual consistency than existing baselines, and that its filtered outputs can serve as useful candidate training augmentations for downstream ReID.
Figures & tables
| Method | I2V Subject | I2V Background | Subject Consistency | Motion Smoothness | Text Relevance | FVD |
| CogVideoX-5B [ 67 ] | 82.18 | 86.74 | 89.92 | 95.62 | 25.63 | 184 |
| Wan2.2-5B [ 55 ] | 86.02 | 88.01 | 90.16 | 95.94 | 26.57 | 175 |
| Wan2.2-A14B [ 55 ] | 86.23 | 88.27 | 90.79 | 96.43 | 26.58 | 152 |
| ConsistI2V [ 39 ] | 83.71 | 86.18 | 88.58 | 95.08 | 25.21 | 188 |
| ConsisID [ 68 ] | 84.06 | 86.29 | 89.72 | 95.56 | 25.98 | 168 |
| Phantom [ 31 ] | 85.64 | 87.53 | 91.02 | 96.11 | 26.25 | 172 |
| Setting | Subj. | Back. | Cons. | Mot. |
| FG-only | 86.81 | 88.39 | 91.33 | 96.53 |
| HF-only | 86.69 | 88.34 | 91.48 | 96.28 |
| FG+HF | 87.40 | 89.36 | 92.94 | 96.85 |
| Model | Setting | Tiger | Nyala | Cow | Panda | Stoat | |||||
| mAP | R1 | mAP | R1 | mAP | R1 | mAP | R1 | mAP | R1 | ||
| SigLIP-FT [ 69 ] | Real only | 55.13 | 90.21 | 25.84 | 41.83 | 62.57 | 91.53 | 28.97 | 45.43 | 61.16 | 92.06 |
| + Data Aug. | 55.41 | 90.08 | 26.11 | 42.06 | 62.74 | 91.47 | 29.22 | 45.31 | 61.38 | 92.14 | |
| + Wan synth. | 57.24 | 91.36 | 29.08 | 43.02 | 64.08 | 92.09 | 31.14 | 47.82 | 62.94 | 92.57 | |
| + Ours synth. | 59.72 | 93.41 | 32.01 | 46.36 | 65.84 | 92.11 | 32.26 | 49.36 | 65.18 | 94.71 | |
| DINO-FT [ 48 ] | Real only | 54.72 | 92.61 | 22.94 | 35.87 | 66.31 | 95.08 | 29.34 | 44.56 | 63.07 | 93.02 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Animal | Dataset | Train | Query | Gallery | Total | ||||
| Imgs | IDs | Imgs | IDs | Imgs | IDs | Imgs | IDs | ||
| Tiger | ATRW [ Li et al., 2020 ] | 3,678 | 148 | 424 | 33 | 521 | 33 | 4,623 | 181 |
| Nyala | Nyala [ Dlamini and Van Zyl, 2020 ] | 1,213 | 179 | 354 | 58 | 375 | 58 | 1,942 | 237 |
| Cow | CowDataset [ Fu et al., 2022 ] | 919 | 8 | 210 | 5 | 356 | 5 | 1,485 | 13 |
| Panda | IPanda50 [ Wang et al., 2021 ] | 4,114 | 31 | 1,022 | 19 | 1,738 | 19 | 6,874 | 50 |
| Stoat | Stoat (Proprietary) | 1,390 | 37 | 403 | 46 | 780 | 46 | 2,573 | 83 |
| Method | I2V Subject | I2V Background | Subject Consistency | Motion Smoothness | Text Relevance | FVD |
| Wan2.2-A14B | 86.23 | 88.27 | 90.79 | 96.43 | 26.58 | 152 |
| Identity branch input | ||||||
| Full image | 86.57 | 88.41 | 91.07 | 96.45 | 26.56 | 159 |
| Foreground only | 86.81 | 88.39 | 91.33 | 96.53 | 26.49 | 148 |
| High-frequency only | 86.69 | 88.34 | 91.48 | 96.28 | 26.43 | 150 |
| Injection strategy | ||||||
| Query source | Tiger | Nyala | Cow | Panda | Stoat |
| Real reference | 64.38 / 96.97 | 12.87 / 11.11 | 71.11 / 100.00 | 30.60 / 68.42 | 64.49 / 82.61 |
| Wan2.2-A14B | 42.73 / 78.79 | 4.18 / 3.97 | 54.62 / 85.71 | 21.87 / 47.37 | 42.36 / 60.25 |
| LoRA-FT | 43.91 / 82.68 | 4.95 / 5.29 | 54.74 / 88.57 | 21.68 / 49.62 | 46.83 / 67.08 |
| WildIcon | 53.72 / 89.61 | 10.84 / 10.32 | 64.27 / 94.29 | 26.14 / 58.65 | 54.61 / 75.78 |
| Selection filter | Tiger | Nyala | Cow | Panda | Stoat |
| DINOv3 | +1.95 / +0.93 | +2.86 / +2.65 | +1.47 / +0.24 | +1.09 / +1.88 | +1.61 / +0.94 |
| SigLIP | +1.81 / +0.68 | +2.94 / +2.11 | +1.12 / +0.08 | +1.57 / +2.03 | +1.44 / +1.06 |
| Method | GPUs | Train h/1K samples | Infer. s/video | Train mem. | Infer. mem. |
| Wan2.2 | 1 GPU | – | 693.3 | – | 51.36 |
| 4 GPUs | – | 300.5 | – | 45.03 | |
| WildIcon | 1 GPU | 5.2 | 701.5 | 51.67 | 55.73 |
| 4 GPUs | 1.3 | 315.8 | 54.48 | 46.47 |
| Dataset | Year | Domain | # Videos | Avg. Len (s) | Duration (h) | Resolution |
| UCF-101 Soomro et al. [2012] | 2012 | Human (Action) | 13.3K | 7.2 | 26.7 | 240p |
| NTU RGB+D Shahroudy et al. [2016] | 2014 | Human (Action) | 114K | – | 37 | 1080p |
| MSP-Avatar Sadoughi et al. [2015] | 2015 | Human (Action) | 74 | – | 3 | 1080p |
| Taichi-HD Siarohin et al. [2019] | 2019 | Human (Action) | 3K | – | – | 256p |
| SkyTimelapse Xiong et al. [2018] | 2018 | Sky | 35K | – | – | 360p |
| FaceForensics++ Rossler et al. [2019] | 2019 | Human (Face) | 1K | – | – | Diverse |