Siamese Neural Networks

Latest papers 10

Sep 11, 2026cs.CV

Brain-PACE: A Deep Siamese MRI Framework for Modelling Longitudinal Brain Acceleration

Brain age estimation has become a popular research proxy for assessing brain health and disease, yet longitudinal trajectories of brain ageing are still poorly defined, and clinical use is limited. Building on existing Siamese longitudinal frameworks, we develop Brain-Predicted Age Acceleration (Brain-PACE) to directly estimate the pace of structural brain ageing from paired T1-weighted MRI. Brain-PACE identified accelerated ageing in 42.642.6% of participants with mild cognitive impairment. Faster Brain-PACE was associated with greater functional and cognitive impairment (FAQ; r=0.35r=0.35, ADAS13; r=0.30r=0.30, CDR-SB; r=0.32r=0.32) and greater regional tau burden in the posterior cingulate (r=0.59r=0.59), precuneus (r=0.47r=0.47), and entorhinal cortex (r=0.37r=0.37). These associations were stronger than those observed when pace was calculated indirectly from repeated cross-sectional brain age estimates, suggesting that direct longitudinal modelling captures complementary information relevant to ongoing pathological change. Methodologically, Brain-PACE extends the LILAC framework by combining spatial attention with soft label distribution learning and a Cram'er distance objective, improving probabilistic performance and reducing prediction bias while providing measures of predictive uncertainty. Together, these findings support Brain-PACE as a complementary longitudinal imaging phenotype with sensitivity to relevant clinical and biological changes in early neurodegeneration.
Jul 6, 2026cs.LG

How Far is Too Far? Defining the Distance Threshold for Verification Siamese Networks

Siamese verification networks are widely used to compare items such as faces, cars, or signatures. In these scenarios, the network is trained to learn an embedding space in which similar objects are mapped closer together, while dissimilar objects are mapped further apart. Two objects are considered to belong to the same class (e.g., the same person in two different images) when the distance between their embeddings falls below a predefined threshold. Defining this threshold, however, is a non-trivial task and typically requires labeled data. In this work, we assume that the distribution of distances produced by a siamese verification network can be approximated by a bimodal function. Based on this assumption, we propose an unsupervised method to determine the verification threshold by identifying the minimum point between the two modes. The proposed approach does not require annotated samples, enabling the verification threshold to be updated directly in the deployment environment without the cost of manual labeling. We evaluate our method on four datasets: MNIST, CIFAR-10, LFW, and PKLot. The results indicate that the proposed approach achieves an average verification accuracy of 94%, comparable to the Equal Error Rate method, while eliminating the need for labeled data.
Jul 4, 2026cs.CV

SiamJEPA: On the Role of Siamese Student Encoders in JEPA

Joint Embedding Predictive Architectures (JEPAs) have emerged as a promising framework for self-supervised representation learning by predicting latent embeddings of masked regions rather than reconstructing pixels. Existing JEPA methods typically employ a single student encoder, leaving the role of Siamese student encoders largely unexplored. In this paper, we propose Siamese JEPA (SiamJEPA), a JEPA framework with masked Siamese student encoders and an exponential moving average (EMA) teacher, which can also be viewed as a JEPA formulation of the brain-inspired representation learning model PhiNet. We further introduce Random Shuffle Teacher (RST), which removes spatial correspondence in teacher targets to encourage semantic patch representations, and develop an RST-based semantic-to-spatial curriculum. Experiments on ImageNet show that Siamese student encoders effectively regularize the JEPA objective, improving representation separability and accelerating early-stage learning. Moreover, under RST, stronger Siamese regularization substantially increases class-discriminative information in individual patch tokens, suggesting that the semantic bias arises from the interaction between RST and the Siamese objective rather than from shuffling alone. SiamJEPA consistently outperforms comparable single-encoder JEPA variants under limited training budgets. With RST-based curriculum learning and a ViT-Base backbone, SiamJEPA achieves 74.2% linear-probing accuracy after 450 epochs using a substantially simpler masking strategy, compared with I-JEPA (72.9%) and DSeq-JEPA (73.5%) after 600 epochs. These results demonstrate that Siamese student encoders provide an effective inductive bias for predictive representation learning and can be further enhanced through semantic-to-spatial curriculum learning. The source code is publicly available at https://github.com/oist/SiamJEPA.
Jun 29, 2026cs.CV

UrbanCDNet: Appearance-Robust and Boundary-Aware Bitemporal Change Detection for Korean Urban Building Monitoring

Urban building change detection from bi-temporal aerial imagery is important for redevelopment monitoring, infrastructure management, and unauthorized-construction screening, but Korean urban scenes remain difficult because changed regions are often sparse, appearance varies strongly between acquisition dates, and useful outputs must follow building footprints rather than coarse blobs. This paper presents UrbanCDNet, a task specific Siamese CNN that combines appearance-robust multi-cue comparison, alignment-aware middle-scale differencing, lightweight context refinement, scene calibration, and auxiliary boundary supervision. Experiments use a corrected AIHub-based Korean benchmark with 3,998 training, 503 validation, and 499 test pairs, and report changed-class precision, recall, F1, and IoU. On the locked test split, UrbanCDNet achieves 0.7335 precision, 0.7696 recall, 0.7511 F1, and 0.6014 IoU, outperforming a strong Siamese U-Net baseline (0.7108 F1, 0.5514 IoU) and the strongest external competitor, ChangeFormer-MIT-B0 (0.7107 F1, 0.5512 IoU). Additional diagnostic slicing shows that the gain is concentrated in the operating regimes that motivated the design: on the sparse-change subset with less than 5% changed area, F1 improves from 0.4765 to 0.6175, and on the high photometric-gap subset it improves from 0.6349 to 0.7285. Boundary F1 at 3-pixel tolerance rises from 0.3445 to 0.4447, while object F1 at IoU 0.3 rises from 0.0690 to 0.2258. These results indicate that, on this Korean benchmark, task-shaped temporal comparison and boundary-aware supervision matter more than generic model scale alone
Jun 18, 2026cs.CV

Evaluation of Image Matching for Art Skills Assessment

While some individuals possess a natural talent for drawing, mastering this skill requires dedicated training and practice. Determining one's skill in the art of drawing requires proper comprehensive assessment. In this paper, we propose a method to measure drawing skill by by matching the hand-drawn image with the original template. Existing techniques often involve complex processes. However, advancements in computer vision allow us to train computers to perform these comparisons at a human-like level, thereby resolving the tedious and overwhelming traditional process. Using computer vision applications, determining image similarity involves identifying the level of similarities in an image with a reference image. We have implemented and analyzed the SIFT feature and Siamese network to measure image similarity. Our results indicate that it is feasible to assess art skill levels. Through feature analysis, we found that SIFT-based key point matching provides a more effective means of detecting drawing skills.
Jun 9, 2026cs.NI

A Unified Siamese Learning Framework for Zero-Day Anomaly Detection and Classification in Optical Networks

A multi-similarity Siamese neural network unifies zero-day anomaly detection and one-shot classification in optical networks, achieving over 99% accuracy and instant adaptability across lightpaths and unseen anomaly types without any retraining.
May 28, 2026cs.LG

Bridging the Gap Between Natural Language and Market Dynamics via High-Dimensional Representation Learning

Traditional multi-modal financial forecasting often relies on scalar sentiment scores, which fail to capture the nuances of financial news. To address this information loss, this paper explores high-dimensional representation learning by replacing discrete polarity ratings with dense FinBERT embeddings within a Transformer-based forecasting architecture. We benchmarked various embedding strategies on the FNSPID dataset, including raw embeddings, attention-weighted aggregation, and a custom Siamese network. While the attention-based mechanism struggled with the low signal-to-noise ratio typical of financial data, the integration of Siamese-optimized embeddings outperformed both the scalar baseline and raw embedding approaches, demonstrating that preserving high-dimensional narrative context yields improved predictive accuracy for short-term stock price movements.
May 27, 2026cs.LG

Learning Robust and Task-Invariant Functional Representation from fMRI through Siamese Self-Supervised Learning

Functional magnetic resonance imaging (fMRI) is a powerful tool for investigating human brain function. However, the high cost of data acquisition and the inherent subjectivity of psychiatric rating scales often lead to datasets with small sample sizes and variable label quality, especially when targeting a specific neurological condition. Combined with the inherently high dimensionality of fMRI data, these limitations substantially increase the risk of model overfitting. Recent years have seen growing interest in developing fMRI foundation models by combining multiple datasets; however, the computational resources needed for pretraining and fine-tuning are often prohibitive. We show that a lightweight self-supervised framework yields representations that generalize across diverse downstream tasks, outperforming fully supervised baselines and approaching the performance of large-scale models. We introduce BrainSimSiam, a data-efficient self-supervised representation learning framework that leverages positive-only data pairs to learn robust and generalizable features. We demonstrate that the learned representations achieve strong performance across multiple downstream classification and regression tasks, highlighting the potential of BrainSimSiam for data-limited neuroimaging applications.
May 12, 2026cs.CR

AccLock: Unlocking Identity with Heartbeat Using In-Ear Accelerometers

The widespread use of earphones has enabled various sensing applications, including activity recognition, health monitoring, and context-aware computing. Among these, earphone-based user authentication has become a key technique by leveraging unique biometric features. However, existing earphone-based authentication systems face key limitations: they either require explicit user interaction or active speaker output, or suffer from poor accessibility and vulnerability to environmental noise, which hinders large-scale deployment. In this paper, we propose a passive authentication system, called AccLock, which leverages distinctive features extracted from in-ear BCG signals to enable secure and unobtrusive user verification. Our system offers several advantages over previous systems, including zero-involvement for both the device and the user, ubiquitous, and resilient to environmental noise. To realize this, we first design a two-stage denoising scheme to suppress both inherent and sporadic interference. To extract user-specific features, we then propose a disentanglement-based deep learning model, HIDNet, which explicitly separates user-specific features from shared nuisance components. Lastly, we develop a scalable authentication framework based on a Siamese network that eliminates the need for per-user classifier training. We conduct extensive experiments with 33 participants, achieving an average FAR of 3.13% and FRR of 2.99%, which demonstrates the practical feasibility of AccLock.
Jul 5, 2025cond-mat.dis-nn

Siamese Neural Network for Label-Efficient Critical Phenomena Prediction in 3D Percolation Models

Predicting critical phenomena from limited labeled data remains a challenging task in statistical physics. As percolation theory provides a canonical model for phase transitions with well-established critical exponents, it serves as an ideal benchmark for validating new machine learning frameworks. Here, we introduce a label-efficient learning framework based on a Siamese Neural Network (SNN) to identify phase transitions in three-dimensional site and bond percolation models. Using only 22 labeled probability points drawn entirely from non-critical regions, the method locates percolation thresholds with percent-level accuracy and yields estimates of the critical exponent νν consistent with literature values within statistical uncertainty. Analysis of the learned representations clarifies what the network learns: although trained solely on binary similarity labels, the network autonomously converges to a statistic that coincides quantitatively with the normalized largest-cluster size Smax/L3S_{max}/L^3 (r>0.99r > 0.99), the finite-size order parameter of percolation. This underlies the framework's most distinctive capability -- a model trained solely on simple cubic lattices identifies the phase transition in face-centered cubic lattices without retraining. The framework thus offers a complementary route to criticality detection in settings where no quantitative order parameter is explicitly defined or labeled data is scarce.