Automated Species Identification in Camera Trap Images for Wildlife Conservation
Organizations: Department of Computer Science and Engineering Brac University
Abstract
Wildlife conservation involves protecting, preserving, and managing wildlife species and their habitats. With today's rapid pace of human development, climate change, and other unsustainable practices, the need for wildlife conservation has heightened. Despite significant progress in species identification using deep-learning models, significant challenges still remain in effectively detecting small animals in low-contrast trap images due to limited feature extraction capabilities. This thesis presents a novel end-to-end framework integrating a self-attention mechanism to address these limitations. The proposed architecture involves a Swin-BiFPN backbone integrated in a Faster RCNN detection network, coupled with a visual semantic extraction module driven by the LLaVA v1.5 (13B) multimodal large language model. The detection framework, capable of extracting crucial features in challenging trap images, demonstrates consistently high results and robust generalization capabilities. Furthermore, the visual semantic extraction module provides zero-shot detection capability, as well as providing valuable insights and emergent cues of the animal's behavior, further supporting the conservation effort. The MLLM evaluation was conducted using both traditional NLP metrics (precision, recall, F1, and SBERT similarity) and subjective scoring by LLM-based judges (GPT-4.1 and GROK 3.0), across five MLLMs, demonstrating the model's strong performance in visual description generation. The proposed framework improves detection accuracy across low-contrast trap images and small animals while also demonstrating zero-shot detection capability leveraging the MLLM.
Figures & tables
| Model | Precision | Recall | Total Training Time (Hour) | |
|---|---|---|---|---|
| Classification model | ZF- Net | 10.93 | 16.29 | 0.65 |
| EfficientNetV2 | 92.52 | 91.45 | 1.6 | |
| Object Detection Model | YOLOv11 | 26.48 | 29.80 | 1.3 |
| F-RCNN | 68.07 | 72.12 | 3.34 |
| Model | Precision | Recall | Total Training Time (Hour) | |
|---|---|---|---|---|
| Classification model | ZF- Net | 11.98 | 16.45 | 0.65 |
| EfficientNetV2 | 86.07 | 82.90 | 1.6 | |
| Object Detection Model | YOLOv11 | 25.74 | 28.58 | 1.3 |
| F-RCNN | 60.96 | 64.67 | 3.34 |
| Model | Bear mAP | Deer mAP | Fox mAP | Frog mAP | Hedgehog mAP | Tiger mAP | Turtle mAP |
|---|---|---|---|---|---|---|---|
| YOLOv11 | 0.2358 | 0.4132 | 0.0222 | 0.2944 | 0.3964 | 0.0000 | 0.0000 |
| F-RCNN | 0.5206 | 0.5301 | 0.6477 | 0.6658 | 0.5628 | 0.8996 | 0.2879 |
| Model | Bear mAP | Deer mAP | Fox mAP | Frog mAP | Hedgehog mAP | Tiger mAP | Turtle mAP |
|---|---|---|---|---|---|---|---|
| YOLOv11 | 0.2476 | 0.4024 | 0.0111 | 0.2877 | 0.3333 | 0.0110 | 0.0000 |
| F-RCNN | 0.5032 | 0.4211 | 0.5213 | 0.6254 | 0.4416 | 0.8351 | 0.2576 |
| Object Detection Network | Precision | Recall | mAP |
|---|---|---|---|
| FRCNN (ResNet 50) | 0.6560 | 0.6694 | 0.6631 |
| ViT + FRCNN | 0.1121 | 0.1476 | 0.1358 |
| Swin + FPN + FRCNN | 0.8156 | 0.7915 | 0.7904 |
| Swin + Bi-FPN + FRCNN | 0.8343 | 0.8178 | 0.8161 |
| Object Detection Network | Precision | Recall | mAP |
|---|---|---|---|
| FRCNN (ResNet 50) | 0.5983 | 0.6103 | 0.6058 |
| ViT + FRCNN | 0.1121 | 0.1137 | 0.1129 |
| Swin + FPN + FRCNN | 0.7669 | 0.7478 | 0.7476 |
| Swin + Bi-FPN + FRCNN | 0.8059 | 0.7919 | 0.7889 |
| Object Detection Network | Bear mAP | Deer mAP | Fox mAP | Frog mAP | Hedgehog mAP | Tiger mAP | Turtle mAP |
|---|---|---|---|---|---|---|---|
| FRCNN (ResNet 50) | 0.5155 | 0.6868 | 0.5971 | 0.6523 | 0.5510 | 0.8502 | 0.0167 |
| ViT + FRCNN | 0.0000 | 0.0385 | 0.0000 | 0.1218 | 0.0000 | 0.1201 | 0.0000 |
| Swin + FPN + FRCNN | 0.8388 | 0.7917 | 0.8944 | 0.7901 | 0.5880 | 0.9000 | 0.5676 |
| Swin + Bi-FPN + FRCNN | 0.8351 | 0.8179 | 0.9176 | 0.7402 | 0.6475 | 0.8469 | 0.4935 |
| Object Detection Network | Bear mAP | Deer mAP | Fox mAP | Frog mAP | Hedgehog mAP | Tiger mAP | Turtle mAP |
|---|---|---|---|---|---|---|---|
| FRCNN (ResNet 50) | 0.4223 | 0.6088 | 0.5156 | 0.6505 | 0.4878 | 0.8041 | 0.0000 |
| ViT + FRCNN | 0.0000 | 0.0000 | 0.0000 | 0.0783 | 0.0000 | 0.1156 | 0.0000 |
| Swin + FPN + FRCNN | 0.8019 | 0.7083 | 0.8278 | 0.7500 | 0.5517 | 0.8963 | 0.4658 |
| Swin + Bi-FPN + FRCNN | 0.8213 | 0.7037 | 0.8385 | 0.7513 | 0.6163 | 0.8401 | 0.4459 |
| MLLM Model | BERTScore | SBERT Cosine Similarity | ||
|---|---|---|---|---|
| Precision | Recall | F1 Score | Mean Score | |
| LLaVa v1.5 7B | 0.9018 | 0.8930 | 0.8972 | 0.5788 |
| LLaVa v1.5 13B | 0.8952 | 0.9007 | 0.8978 | 0.6501 |
| LLaVa v1.6 Mistral | 0.8741 | 0.8971 | 0.8854 | 0.6283 |
| KOSMOS 2 | 0.8722 | 0.8841 | 0.8780 | 0.5513 |
| IDEFICS 9B | 0.8840 | 0.8547 | 0.8690 | 0.5718 |
| MLLM model | Avg Relevance | Avg Accuracy | Avg Depth | Avg Fluency | Total |
|---|---|---|---|---|---|
| LLaVa v1.5 7B | 0.6043 | 0.6429 | 0.6184 | 0.9836 | 2.8491 |
| LLaVa v1.5 13B | 0.65 | 0.8114 | 0.7484 | 0.9707 | 3.1806 |
| LLaVa v1.6 Mistral | 0.6286 | 0.5686 | 0.9489 | 0.9986 | 3.1446 |
| KOSMOS 2 | 0.46 | 0.81 | 0.7953 | 0.4993 | 2.5646 |
| IDEFICS 9B | 0.5643 | 0.0071 | 0.3479 | 0.3986 | 1.3179 |
| MLLM model | Avg Relevance | Avg Accuracy | Avg Depth | Avg Fluency | Total |
|---|---|---|---|---|---|
| LLaVa v1.5 7B | 0.85 | 0.80 | 0.75 | 0.82 | 3.22 |
| LLaVa v1.5 13B | 0.92 | 0.88 | 0.85 | 0.90 | 3.55 |
| LLaVa v1.6 Mistral | 0.88 | 0.84 | 0.80 | 0.87 | 3.39 |
| KOSMOS 2 | 0.90 | 0.82 | 0.70 | 0.80 | 3.22 |
| IDEFICS 9B | 0.87 | 0.83 | 0.78 | 0.89 | 3.37 |