Million-scale multimodal pollen microscopy with expert-guided foundation models
Organizations: Department of Physics of Complex Systems, ELTE Eötvös Lor´and University, Budapest, Hungary · 2The Palynological Laboratory at the Swedish Museum of Natural History, Stockholm, Sweden · 3National Centre for Public Health and Pharmacy, Budapest, Hungary · 4INRAE, UR 546 BioSP, Site Agroparc, Avignon 84914, France · 5National Kor´anyi Institute for Pulmonology, Budapest, Hungary · 6Health Data Science and AI Knowledge Centre, Health Services Management Training Centre, Faculty of Health and Public Administration, Semmelweis University, Budapest, Hungary · Department of Biological Physics, ELTE Eötvös Lor´and University, Budapest, Hungary
Abstract
Automated pollen identification from microscopy remains a bottleneck in aerobiology, palaeoecology and biodiversity monitoring, because scalable systems must generalise across specimen preparation, scanner settings and geographic origins while retaining palynological interpretability. To address this gap, we present a million-scale multimodal pollen microscopy resource, Pollen AI Atlas, assembled from pure-species whole-slide bright-field images spanning four geographic origins, four scanner settings and 46 taxon labels across 31 botanical families. Seeded by one manually selected exemplar per source slide, token-level mining and filtering produced 1,511,390 released grain detections with 99.6% proposal precision in expert-curated test regions. Each detection was paired with machine-generated grain-level morphological captions from five open-weight vision-language models, guided by expert-verified palynological anchors, yielding structured descriptions of aperture systems, wall ornamentation, shape and size. Among the evaluated models, Gemma4 provided the most controlled primary caption set, combining tight length control, no leakage and the strongest text-retrieval performance. Baseline benchmarks with frozen visual features reached 88.16% top-1 accuracy, while cross-regional retrieval showed that caption-derived text embeddings remained robust when image similarity degraded (mAP@20 0.811 versus 0.262). Released data, annotations, captions, splits, code, and weights provide a benchmark for pollen recognition, cross-regional domain adaptation and domain-specific multimodal microscopy learning.