Patch-based Querying Identifies Structures of Interest in Electron Microscopy
Authors: Niels Vyncke, Nicolas Nadisic, Yvan Saeys, Aleksandra Pižurica
Organizations: Department of Telecommunications and Information Processing, Sint-Pietersnieuwstraat 41, Ghent, 9000, Belgium · Royal Institute for Cultural Heritage (KIK-IRPA), Jubelpark 1, Brussels, 1000, Belgium · Department of Mathematics, Computer Science and Statistics, Krijgslaan 281, Ghent, 9000, Belgium
Volume electron microscopy (vEM) has emerged as an essential sensing technique in biomedical research, allowing the three-dimensional imaging of biological cells and tissues at nanometer-scale resolution. The ability to generate extensive datasets has reached the limitations of downstream analysis processes, which depend significantly on the intervention of human experts for preprocessing and annotation. We propose an efficient and reliable patch-based retrieval framework based on self-supervised learning of local image descriptors to locate self-similar structures in vEM datasets. Given a few manual annotations of a given cellular structure, our method can retrieve similar structures across the EM volume. Our framework is interactive, allowing the human expert to refine the search queries and retrieve relevant image patches quickly and using little labeled data. Experiments on real-world vEM images of biological tissues demonstrate that our framework can reliably identify relevant cellular structures, generalize across different organelles and acquisition modalities, and substantially reduce the search space for downstream analysis.
Figures & tables
Figure 1: An illustration of a volume electron microscopy dataset. The dataset contains 377 slices, each of size 508×966 .
Figure 2: Visualization of the proposed methodology. Ideally, the local image descriptor maps patches corresponding to different structures onto distinct manifolds in the latent space, which can be learned by extracting and encoding query patches.
Figure 3: Proof-of-concept visualization of the proposed descriptor. (Left) Color-coded overlay of the patches selected as references for each of the four structures (one color per structure). The pixel-level colors are manual annotations used only to identify which patches contain each structure; the proposed method neither predicts nor reconstructs these pixel-level masks, and no segmentation is performed. (Right) For every patch in the slice, a four-dimensional descriptor is computed as the average mean squared error (MSE) distance to each structure’s reference set; the resulting points are projected to three dimensions using principal component analysis (PCA).
Dataset
Number of slices
Slice dimensions
(h × w)
HeLa-EMBL
64
512 × 512
VIB
377
508 × 966
EPFL
330
768 × 1024
Platelet-EM
50
800 × 800
Table 1: Dimensions of the vEM datasets studied.
Figure 5: Illustration of the query selection strategy. The dataset is divided into equally sized groups (split size). From each group, the central slice is chosen as the query slice (blue), while the remaining slices together form a separate group.
Figure 6: t-SNE visualization of mitochondria and other patches from the HeLa-EMBL dataset for pretrained and fine-tuned classical autoencoder ( Figures 6a and 6b ) and variational autoencoder ( Figures 6c and 6d , β=10−5 ). Mitochondria patches ( >50% overlap) in green, other patches ( 0% overlap) in red.
Figure 7: t-SNE visualization of mitochondria and other patches for the classical autoencoder across groups of 10 slices (HeLa-EMBL dataset). Mitochondria patches ( >50% overlap) in green, other patches ( 0% overlap) in red.
Figure 8: t-SNE visualization of mitochondria and other patches for the variational autoencoder ( β=10−5 ) across groups of 10 slices (HeLa-EMBL dataset). Mitochondria patches ( >50% overlap) in green, other patches ( 0% overlap) in red.
Figure 9: t-SNE visualization of four different structures in the latent space for classical autoencoder ( Figure 9a ) and variational autoencoder ( Figure 9b , β=10−5 ), for the VIB dataset. Colors denote mitochondria (green), nuclear envelope (magenta), endoplasmic reticulum (blue), and vesicles (black).
Figure 10: Examples of retrieved patches based on specified positive/negative queries (left) from the VIB dataset. Right: AE/VAE results; bottom rows show the retrieved patches sorted by decreasing similarity score. Green indicates true positives; red indicates false positives.
Figure 11: Examples of retrieved patches based on specified positive/negative queries (left) from the HeLa-EMBL dataset. Right: AE/VAE results; bottom rows show the retrieved patches sorted by decreasing similarity score. Green indicates true positives; red indicates false positives.
Figure 12: Examples of retrieved patches based on specified positive/negative queries (left) from the EPFL dataset. Right: AE/VAE results; bottom rows show the retrieved patches sorted by decreasing similarity score. Green indicates true positives; red indicates false positives.
Figure 13: Examples of the retrieved patches corresponding to the specified positive and negative queries (100 from each class) taken from (pseudo-)labeling, from the VIB dataset. The query slice is shown on top, with positive (resp. negative) queries indicated in green (resp. red). Below are the 25 most similar patches for the classical autoencoder and the variational autoencoder, respectively. The patches are indicated in green (resp. red) if the patch is a true positive (resp. false positive).
Dataset
Split size
Pretrained
Fine-tuned
AE
VAE
AE
VAE
HeLa-EMBL
10
0.971
0.974
0.988
0.982
HeLa-EMBL
20
0.890
0.873
0.928
0.922
HeLa-EMBL
30
0.808
0.811
0.887
0.867
HeLa-EMBL
50
0.595
0.580
0.779
0.755
EPFL
10
0.916
0.923
0.932
0.930
Table 2: Overview of the Average Precision (AP) for the classical and variational autoencoder, and the pretrained and fine-tuned versions. The best-performing descriptor is indicated in bold .
Figure 14: Precision-recall curves for the HeLa-EMBL dataset (top), VIB (middle), and EPFL (bottom), and split sizes 10 (left) and 30 (right) for the fine-tuned classical autoencoder, which is the best-performing descriptor. Red points indicate reduction rates: the labeled value is the percentage of patches the user no longer needs to inspect; the corresponding retention rate (percentage of patches kept, sorted by similarity) is 100% minus the labeled value.
Figure 15: Operational retention analysis for the HeLa-EMBL dataset (fine-tuned classical autoencoder, split size 30 ). Patches are ranked by similarity and the top q fraction is retained. (Left) percentage of two-dimensional connected components that intersect at least one retained patch; (right) percentage of target pixels that lie in the union of retained patches. The dashed line marks 95% retention. Retaining 30% of patches already preserves 93.8% of mitochondrial instances and 78.4% of mitochondrial pixels; retaining 50% reaches 95.4% and 88.9% , respectively.
Figure 16: Retrieval quality as a function of the z -distance from each retrieved patch to the nearest query slice (VIB dataset, fine-tuned classical autoencoder, split size 30 , top- 10 % similarity threshold). Both metrics decay smoothly with distance but remain useful well beyond a single neighboring slice.
Figure 17: Long-range generalization. Queries are extracted from the first 20 slices (a) and retrieval is performed on a slice 323 slices away (b, c). Despite the very large z -gap, 8 out of 9 structures present in the target slice are recovered. The pixel-level coloring in (b) is the manually annotated ground truth, shown only to indicate where mitochondria are located; it is not produced by the method. In (c), colored rectangles mark whole retrieved patches (green: overlaps a mitochondrion; red: no structure), illustrating that the framework returns candidate regions rather than pixel-accurate segmentations.
Split size
Pretrained
Fine-tuned
AE
VAE
AE
VAE
10
0.961
0.964
0.969
0.963
30
0.851
0.859
0.882
0.872
Table 3: Average Precision (AP) for ER retrieval on the HeLa-EMBL dataset (structure label 2, positive-overlap threshold 15% ), for the pretrained and fine-tuned (V)AE descriptors. The best-performing descriptor is indicated in bold .
Figure 18: Precision-recall curves for ER retrieval on the HeLa-EMBL dataset (fine-tuned classical autoencoder), split sizes 10 (left) and 30 (right). Red points indicate reduction rates, as in Figure 14 .
Figure 19: Average Precision (AP) of the fine-tuned autoencoder against three handcrafted descriptors (HOG, LBP, dense SIFT) compressed to the same 32 dimensions by averaging contiguous blocks of their original feature vectors (split size 30 ), for every labeled structure across the three datasets (mitochondria in HeLa-EMBL, EPFL, and VIB; ER in HeLa-EMBL and VIB; nuclear envelope and vesicles in VIB). The AE attains the highest AP for every structure/dataset combination.
Structure
Threshold
Base rate
AP
Mitochondria
9%
7%
0.55
Alpha granule
20%
37%
0.78
Canalicular vessel
9%
46%
0.81
Table 4: Average Precision (AP) for retrieval on the SBF-SEM Platelet-EM dataset [ 27 ] (split size 10), with the base positive rate (AP of a random ranking) shown for context.
Figure 20: Precision-recall curves for retrieval on the SBF-SEM Platelet-EM dataset (fine-tuned classical autoencoder, split size 10). Red points indicate reduction rates, as in Figure 14 .
Dataset
Split
Patches
Retrieval
Manual
Speed-up
size
time (s)
vs. retrieval
HeLa-EMBL
10
2,793
2.3
21.3 min vs. 5.4 min
3.97×
30
2,989
2.1
21.3 min vs. 5.4 min
3.97×
EPFL
10
38,610
38.1
110.0 min vs. 28.1 min
3.91×
30
41,470
34.3
110.0 min vs. 28.1 min
3.92×
VIB
10
30,849
31.5
125.7 min vs. 31.9 min
3.93×
Table 5: Measured full-volume retrieval time and estimated time savings versus manual inspection, per dataset and split size (fine-tuned classical autoencoder, patch size 80×80 ; hardware and protocol described in the text). Manual time assumes 20 s/slice; retrieval-based time is the measured retrieval time plus 5 s/slice for candidate validation. Measured throughput ranges from 1,043 to 1,451 patches/s across all configurations.
Vision foundation models have substantially advanced computer vision, enabling state-of-the-art performance in zero- and few-shot settings. They have been successfully applied to biomedical imaging tasks ranging from organ segmentation in computed tomography to cell segmentation in light microscopy. Electron microscopy (EM) is a central modality for analyzing cellular ultrastructure due to its nanometer-scale resolution. However, the application of foundation models in EM has so far been limited to specific organelles, such as mitochondria, largely due to the diversity of segmentation tasks and the scarcity of comprehensively annotated data. As a result, EM segmentation still predominantly relies on supervised learning, requiring extensive manual annotation and limiting ultrastructural analysis. To address this gap, we propose μMatch, a framework for semi-supervised learning and domain adaptation that leverages foundation models. We implement state-of-the-art student-teacher-based methods and evaluate multiple foundation models (SAM, SAM2, μSAM, DINOv2/v3) on challenging EM tasks, including mitochondrion, nucleus, and neurite segmentation. Our results demonstrate consistent improvements over strong baselines and highlight a path toward substantially reducing the annotation effort in EM.
Quantitative microstructural characterization is fundamental to materials science, and electron micrographs (EMs) provide indispensable high-resolution insights. However, progress in deep learning-based analysis of EMs has been hampered by the scarcity of large-scale, expert-annotated public datasets. To address this issue, we introduce EM3M, a large-scale and multimodal dataset for instance-level understanding of EMs. EM3M comprises 5,091 high-quality EMs, approximately 3 million instance segmentation annotations, and image-level textual descriptions with disentangled attributes. The dataset is constructed through a rigorous multi-stage curation and validation pipeline, with comprehensive statistical analyses to ensure reliability and reproducibility. Building upon these curated image-text pairs, we further provide a text-to-image diffusion model that serves as a controllable data augmentation engine, demonstrating that synthetic augmentation consistently improves downstream segmentation performance. To establish a systematic benchmark, we evaluate representative instance segmentation methods on EM3M. Our results reveal that conventional detection-based and query-based methods struggle with the extreme instance densities and textural complexities inherent in EMs. We additionally provide an optimized flow-based baseline to facilitate fair comparison and future research. EM3M {Dataset: https://huggingface.co/datasets/UniParser/EM3M}, the generative engine {Generation: https://huggingface.co/UniParser/EM3M-Gen}, and an online demo {Segmentation demo: https://www.bohrium.com/apps/uni-aims} are publicly available to support future research in automated materials analysis.
Segmentation of cellular structures in electron microscopy (EM) images is fundamental to analyzing the morphology of neurons and glial cells in the healthy and diseased brain tissue. Current neuronal segmentation applications are based on convolutional neural networks (CNNs) and do not effectively capture global relationships within images. Here, we present DendriteSAM, a vision foundation model based on Segment Anything, for interactive and automatic segmentation of dendrites in EM images. The model is trained on high-resolution EM data from healthy rat hippocampus and is tested on diseased rat and human data. Our evaluation results demonstrate better mask quality compared to the original and other fine-tuned models, leveraging the features learned during training. This study introduces the first implementation of vision foundation models in dendrite segmentation, paving the path for computer-assisted diagnosis of neuronal anomalies.
Zewen Zhuo, Ilya Belevich, Ville Leinonen +4
A.I. Virtanen Institute for Molecular Sciences University of Eastern Finland Kuopio, Finland · Electron Microscopy Unit Institute of Biotechnology University of Helsinki Helsinki, Finland · Department of Medicine Faculty of Health Sciences University of Eastern Finland Kuopio, Finland