BossouChimpanzee: Long-term Chimpanzee Video Dataset
Authors: Daniel Schofield, Susana Carvalho, Vladimir Iashin, Andrew Zisserman, Max Bain, Arsha Nagrani, David Ng, Claudia Sousa, +5 more
Organizations: Visual Geometry Group, Department of Engineering Science, University of Oxford, UK · University of Oxford, UK · CIBIO–BIOPOLIS, University of Porto, Vairão, Portugal · Department of Science, Gorongosa National Park, Mozambique · Institut de Recherche Environnementale de Guinée (IREG), Bossou, Guinea · University of Rochester, USA · Chubu Gakuin University, Japan · Japan Monkey Centre, Japan · Northwest University, China
We describe the BossouChimpanzee video dataset, a unique long-term visual record of wild chimpanzees at an outdoor laboratory for field experiments in Bossou, Guinea, spanning three decades (1988-2018) and comprising over 1,200 hours of continuous video recordings collected through collaborative fieldwork and research. In this paper, we outline the history and scientific contributions of the experimental paradigm and video archive, provide key statistics and details on the structure of the main video dataset, and release an initial ~74h snapshot, BossouChimpanzee70h, covering 23 identified individuals focused on chimpanzee individual and action recognition, ahead of the full video resource. This dataset represents a valuable resource for cognitive and behavioural research in ethology and a rich benchmark for training and evaluating machine learning models on audiovisual data from the wild.
Figures & tables
Figure 1 : A sample of frames from a selection of years of the BossouChimpanzee dataset (1988–2018).
Face
Body
Action
Year
Boxes
Videos
Tracks
IDs
Boxes
Videos
Tracks
IDs
Videos
Tracks
IDs
2000
1,666,984
46
2,504
21
–
–
–
–
–
–
–
2004
1,436,429
23
1,877
12
–
–
–
–
–
–
–
2006
931,924
20
1,924
11
–
–
–
–
–
–
–
2008
1,607,059
49
2,819
13
–
–
–
–
–
–
–
2012
1,419,441
18
2,973
13
1,878,291
17
2,674
13
11
536
8
Table 1: BossouChimpanzee70h : Year-wise summary. Combined per-year counts across three identity annotation levels. For Face and Body we report the number of videos, tracks, bounding boxes, and individuals (IDs) that appear in a year. The Action columns give videos, tracks, and individuals with additional annotations for spatio-temporal (body-level) action recognition (nut-cracking). The Action Tracks column counts tracks from the earlier body annotation set (Section 3.2 ), not the Body tracks reported in the same row.
Figure 2 : A sample of frames from BossouChimpanzee70h showing Body and Face annotations. Green boxes show Body annotations (SAM 3 tracks [ 35 ] ; Section 3.2 ) and red boxes show Face annotations ( [ 25 , 26 ] ; Section 3.1 ); text labels indicate identity.
Face
Body
Action
Individual
Boxes
Vids
Tracks
Yrs
Boxes
Vids
Tracks
Yrs
Actions
Vids
Tracks
Yrs
Jeje (M, 1997)
1,154,000
113
1,803
6
384,289
29
355
2
2,287
16
87
2
Jire (F, 1958)
860,987
80
1,056
6
218,196
21
281
2
722
10
66
2
Fanle (F, 1997)
670,403
68
1,201
6
248,469
23
569
2
2,043
13
201
2
Peley (M, 1998)
635,376
81
1,182
5
220,987
11
239
1
1,111
9
33
1
Yolo (M, 1991)
593,501
82
683
4
–
–
–
–
–
–
–
–
Table 2: BossouChimpanzee70h : Individual-wise summary with Face , Body annotation levels, and annotations for spatio-temporal (body-level) Action recognition (nut-cracking). We report the name, sex Male ( M ) or Female ( F ), year of birth, number of annotated bounding boxes, videos, tracks, and years an individual appears in, and the number of action annotations. Body annotations were re-generated using a SAM 3 [ 35 ] + ByteTrack [ 36 ] pipeline; Unknown denotes an on-screen chimpanzee whose identity could not be verified. Actions counts action annotations attributed to a named individual; a further 517 annotations for which the actor could not be identified are not included. The Action Tracks column counts tracks from the earlier body annotation set (Section 3.2 ), not the Body tracks reported in the same row.
Bouts
Strikes
Failures
Successes
Total
Count
1,281
7,884 (7,367)
588
1,142
10,895
Table 3: BossouChimpanzee70h : Action (nut-cracking) fine-grained annotation counts. The strike count includes 7,367 strikes by identified individuals (in parentheses), and 517 ‘off-screen’ strikes where an actor could not be seen.
Figure 3 : Example of an on-screen nut-cracking action annotation. An on-screen individual, Peley, performs a successful cracking bout comprising 10 strikes. The bout includes one failure when the nut slips away. Another individual, Jeje, passes behind him.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Table A1 : BossouChimpanzee : Media attribute distributions. Count refers to individual digitised video files, each corresponding to a single archived tape. BossouChimpanzee70h (Table A2 ) instead counts clips excerpted from these recordings, so its per-year count may exceed the archive count while total duration remains a strict subset. Year denotes the field season a recording set is archived under, given by the date of its first video. Seasons spanning a new year (e.g. 2005–2006) are filed under the starting year; Table A2 assigns years by recording date, where such a season appears under the later year.
Table A2 : BossouChimpanzee70h : Media attribute distributions.
Leprosy (Mycobacterium leprae) has been confirmed in wild western chimpanzees (Pan troglodytes verus) in West Africa, presenting as clear and progressive visual symptoms. Manual review of camera-trap footage at landscape scale is infeasible, motivating the need for automated screening. We present the first deep learning pipeline for wildlife leprosy detection and contribute the PanLep300 dataset of 125,670 annotated bounding-box crops across 953 tracks from 303 camera-trap videos with ecologically-motivated splits that withhold whole individuals and camera installations. We benchmark spatial (2D), temporally aggregated (2.5D), and video-based (3D) classification approaches to investigate which approach is best suited to automated leprosy detection in wild apes. We find that simple aggregation of crop-level predictions consistently matches or outperforms both learned temporal models and end-to-end video architectures -- consistent with leprosy's static cutaneous presentation. We further find that performance is suppressed when tracklets contain frames of partially visible individuals -- as commonly occurs at the start and end of a track -- and demonstrate that this can be addressed through targeted construction and aggregation strategies.
Katie I. Murray, Anna C. Bowland, Marina Ramon +9
University of Exeter, Centre for Ecology and Conservation, Penryn, UK · Wild Chimpanzee Foundation, Leipzig, Germany · University of Bristol, School of Computer Science, Bristol, UK
Behavioural shifts in wild great ape populations, particularly the breakdown of social structures, can serve as an early indicator of population decline. Automating the detection of behaviours indicative of these shifts is therefore a critical task for conservation. Several valuable datasets have recently been introduced for the automated recognition of great ape behaviour, yet few include fine-grained social behaviour annotations, and those that do are captured either in captive settings or via aerial platforms such as UAVs. We address this gap by introducing PanAf-SBR, the first wild great ape camera trap dataset annotated with social behaviours. PanAf-SBR extends PanAf500 with 100 additional videos covering 36,063 frames. These come with 81,096 annotations including bounding boxes, segmentation masks, intra-video identities, and seven social behaviour classes defined under the action giver and receiver convention of ChimpACT. We use this data together with the AlphaChimp architecture to establish the first benchmarks for fine-grained social behaviour recognition in wild great apes from camera trap footage. We further conduct bidirectional transfer learning experiments between PanAf-SBR and the captive ChimpACT dataset, finding that cross-dataset pre-training is highly beneficial for specific classes rather than of uniform benefit. Finally, we examine the role of background context by inverting the segmentation masks to suppress non-ape pixels.
Maciej Braszczok, Otto Brookes, Xiaoxuan Ma +6
University of Bristol, School of Computer Science, Bristol, UK · Wild Chimpanzee Foundation, Leipzig, Germany · Carnegie Mellon University, Robotics Institute, Pennsylvania. USA +4
Recent advances in field technology have led to a massive influx of in-the-wild video data for ecological science. The primary bottleneck in leveraging this data is the high cost of expert annotation. While computer vision offers a potential solution, current models frequently fail when deployed in complex marine environments. To characterize these failures, we introduce WildFin, a novel benchmark for fish behavior recognition collected and annotated by ecologists. WildFin spans two critical real-world paradigms: stationary cameras monitoring groups of fish and dynamic divers following individual subjects. The dataset represents a massive curation effort, involving 1,350 hours of fieldwork and 600 hours of expert annotation to produce 9 hours of behavioral data with over 2 million frame-by-frame labels. We benchmark modern vision foundation models and quantify tradeoffs between static and spatiotemporal architectures, revealing the substantial gap that remains between current model capabilities and the demands of real-world underwater behavioral analysis. Project website: https://team-wildfin.github.io/.
Abigail G. Grassick, Jerome Tze-Hou Hsu, Ethan Lin +10
Cornell University, 14850 Ithaca, NY · University of Colorado Boulder, 80309 Boulder, CO · HHMI Janelia Research Campus, 20147, Ashburn VA