cs.SDJul 16, 2026

Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026

Authors: Anthony MiyaguchiMurilo GustineliAdrian Cheung

Organizations: Georgia Institute of Technology, 225 North Ave NW, Atlanta, GA 30332 · 1Georgia Institute of Technology, 225 North Ave NW, Atlanta, GA 30332

Abstract

This paper details the DS@GT ARC team's approach to BirdCLEF+ 2026, multi-label detection of animal vocalizations in soundscapes from the Pantanal wetlands. The 2026 edition adds about an hour of labeled soundscapes, shifting the task toward supervised pipelines fit to the labeled set. First, we build a competitive supervised baseline that ensembles a frozen Perch v2 backbone, a trained HGNetV2-B0 sound-event-detection network, and a non-bird prototypical head, reaching a private leaderboard score of 0.936 at rank 1894 within a 90-minute CPU budget. Second, we ask whether token-based representations can compete, contrasting codec representations from neural audio codecs against semantic representations from foundational embeddings. We compare two bioacoustic specialist models against four token-based encoders trained on AudioSet. The repository for this work can be found at https://github.com/dsgt-arc/birdclef-2026.

Explore similar work

CardsList
  1. Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations

    Jun 10, 2026Chiara Semenzin, Faadil Mustun, Roberto Dessi +5Self-Supervised

  2. OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder

    Jul 18, 2025Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4Audio EncodersAudio Representation