Feature Fusion
Momentum
12 papers in the last four weeks, against 2 the four weeks before. 0.1% of all new papers.
Latest papers 77
Task-fMRI connectomes reveal state-dependent neural reconfigurations, yet conventional methods marginalize these signals by aggregating distinct conditions into static pairwise graphs, thereby obscuring condition-specific multi-ROI organization. We introduce CoHyFuse, a condition-aware ROI-centered hypergraph framework that constructs a task-state-specific incidence matrix from condition-wise functional connectivity (FC)-profile embeddings, allowing the same ROI to form different multi-ROI hyperedges across task phases. Condition-specific neighborhood sizes further adapt the hyperedge scale to each task state, and the resulting condition embeddings are fused with a complementary whole-session FC branch for prediction. In the AABC cohort (N=1,074), CoHyFuse achieved the best mean out-of-fold predictive performance among evaluated baselines on FACENAME Fluid Cognition Composite (FCC) prediction (7.830.10 MAE, 0.4390.026 ) and VISMOTOR age prediction (7.520.37 MAE, 0.5920.022 ). In an auxiliary CMI-HBN attention-deficit/hyperactivity disorder (ADHD) classification benchmark (N=223), CoHyFuse obtained 72.02.1% macro-AUC and 74.22.9% accuracy. Ablation studies support the contributions of condition-wise incidence construction and dual-view fusion, suggesting that state-resolved ROI-set structure provides complementary predictive information beyond whole-session FC alone. Occlusion analysis identifies the Distraction condition as the primary driver of model prediction, pointing toward the Salience/Ventral Attention Network (SAN)--FrontoParietal Network (FPN) and within-SAN hyperedge-defined ROI-set motifs as candidate model-relevant patterns. This framework provides an interpretable, state-resolved view of the connectome for downstream cohort analysis.
GateSPINE: Gated Cross-View Fusion for Lumbar Spine MRI Report Generation
Automated report generation can ease the burden radiolo gists face when interpreting multi-sequence MRI studies. Unlike CT, MRI examinations comprise multiple sequences and imaging planes, each con tributing complementary diagnostic information. Existing methods en code a study as a single volume and combine multiple acquisitions by fixed rules. Findings visible in only one plane are thus diluted and of ten missed, lowering recall on clinical efficacy metrics, where a missed abnormality is most costly. We propose GateSPINE, a vision-language framework that fuses sagittal T1 and T2 volumes with a training-free operator, encodes the fused sagittal and axial volumes with two parallel 3D encoders, and decodes their combined representation into a report. Its core mechanism is a gated cross view fusion module that predicts, per feature channel and token, how much of each view to admit, so the more informative view dominates at each spatial location. We evaluate GateSPINE on three lumbar MRI datasets, comprising two public bench marks and a private cohort collected from Phenikaa University Hospital, using both natural language generation (NLG) and clinical efficacy (CE) metrics. GateSPINE achieves the highest CE F1 through improved re call on all three datasets; on SPIDER, which lacks an axial sequence, this reflects the sagittal fusion component rather than the gated cross-view mechanism, which is validated on the two cohorts with both imaging planes. GateSPINE also remains competitive on standard NLG metrics.
Fusion techniques of time frequency-based images to predict the outcome of rTMS depression therapy
Depression is a mental condition that can lead to suicide and self-harm. Predicting the outcome of depression treatment is one of the most difficult tasks for clinicians. Among various treatment options, repetitive Transcranial Magnetic Stimulation (rTMS) is a widely used non-invasive method. Predicting rTMS response using Electroencephalogram (EEG) data is difficult because of high inter-subject variability and limited features from single-domain analysis. We introduce two fusion techniques, montage and blending, to overcome these limitations and extract richer features from EEG-derived Time-Frequency (TF) images. We then propose a lightweight custom Convolutional Neural Network (CNN) trained on fused TF representations. \textcolor{black}{We use a primary dataset of 15 patients and a secondary dataset of 46 patients. We run two sets of experiments. The first set uses segment-level 10-fold cross-validation. In this setup segments from the same patient can appear in both training and testing. The Montage CWT_ST fusion reaches 99.90% accuracy on the primary dataset and 91.90% on the secondary dataset. The second set uses strict subject-disjoint cross-validation. All segments of a patient stay in one fold and no patient appears in both training and testing. Performance collapses. We test four time-frequency methods, six fusion mechanisms, and fourteen model architectures. With one exception, every configuration on both cohorts falls between AUC 0.31 and 0.54 and every 95% confidence interval contains 0.5. A patient-level permutation test on the best standalone method returns . The best subject-level result is Montage CWT_ST on the primary cohort, which reaches AUC and 82.7% accuracy.
Training-Free Affinity Fusion of Neural and Embedding-Based Speaker Diarization
Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermediate representations are not directly compatible. We introduce Training-Free Affinity Fusion (TFAF), which integrates the speaker structure inferred by a neural diarizer into an embedding-based diarization system. The neural speaker partition is used to condition local speaker representations, from which we construct a continuous affinity matrix and combine it with the embedding-based acoustic affinity before a single global clustering step. The method requires no additional training, shared embedding space, speaker-label alignment, or hard transfer of the neural diarizer's speaker count. Experiments on AMI and CALLHOME show consistent DER improvements over both constituent systems; on AMI, fusion also improves speaker-attributed transcription. Ablations show that the neural speaker partition accounts for most of the gain, while retaining the continuous embedding-based affinities provides additional benefit over hard partition fusion.
Multichannel Audio Quality Assessment: Extending Pretrained Perceptual Models to Spatial Audio
Accurate perceptual quality assessment is essential for evaluating and optimizing spatial audio, where perceived quality depends on both signal fidelity and inter-channel spatial relationships. However, subjective evaluation is costly, while existing perceptual models are often trained for limited channel configurations and cannot be directly applied to higher-channel-count audio. This raises the question: how can pretrained perceptual knowledge be effectively reused for multichannel spatial audio? Using 5.1-channel audio, we study four levels of multichannel integration: signal, prediction, latent, and feature and propose two learned approaches: latent-level aggregation of spatial-group representations and the feature-level Feature-Band Group Attention (FGAtt), which adaptively fuses spatial groups at the feature level before perceptual processing. Across five 5.1-channel test sets, FGAtt achieves the strongest over- all performance, demonstrating the effectiveness of feature-level adaptation for reusing pretrained perceptual knowledge
Multi-Depth Temporal Fusion for Feedforward, Locally Trained Spiking Neural Networks
We propose a new spiking neural network (SNN) design to process static images and event streams using time-to-first-spike (TTFS) latencies. Our key research question is which architectural choices best accommodate local and online learning in multi-layer convolutional SNNs. This question is addressed via an original framework combining residual-like connections with multi-depth feature aggregation and consensus. The full SNN pipeline features an early-vision front end, to convert raw visual data into sparse spike latencies, a four-layer convolutional backbone trained layerwise with unsupervised spike-timing-dependent plasticity (STDP), a deterministic Multi-Depth Temporal Fusion (MDTF) and a final classifier trained with reward-modulated spike-timing-dependent plasticity (R-STDP). Rather than replacing early features in deeper layers, the proposed MDTF preserves early temporal evidence, adding sparse residual events from intermediate layers, and incorporating deeper features only when they agree in time with earlier representations. The resulting architecture is experimentally validated across MNIST, Fashion-MNIST, CIFAR-10, and N-MNIST, delivering strong classification performance under a fully local learning regime. Selective multi-depth fusion significantly outperforms traditional STDP/R-STDP baselines on higher-variability visual tasks (achieving +18.2 pp on Fashion-MNIST and +29.2 pp on CIFAR-10). Furthermore, activity-budget analyses show that the network retains high accuracy even when removing a large fraction of late or weak spike events, confirming its high data efficiency and reduced event-processing requirements. The codebase is publicly available at github.com/aidinattar/multi-depth-temporal-fusion-snn.
DBCF: Dual-Branch Complementary Fusion of Foundation Models for Generalized Deepfake Detection
As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulations. To address this limitation, researchers have turned to large-scale foundation models, which can provide richer representations and better generalization. Nevertheless, relying on a single foundation model alone remains insufficient for effective forgery detection. While models like CLIP offer robust global semantic cues, they lack the capacity to capture detailed local facial features. In contrast, DINO excels at capturing local structural features of faces, but provides weaker global semantic context. To fully utilize the synergies among multiple foundation models, we propose a hierarchical multi-granular framework that integrates complementary pretrained representations. Specifically, a Global Context Branch (GCB) based on CLIP captures holistic semantic cues, while a Fine-grained Cue Branch (FCB) built on DINOv3 captures localized structural irregularities. In addition, we design a feature fusion module that enables parameter-efficient adaptation of the frozen foundation backbones by adaptively extracting and integrating complementary features from the two models. By jointly leveraging global context and fine-grained cues, our method learns more comprehensive forgery representations and achieves strong cross-manipulation performance. Extensive experiments on multiple benchmarks demonstrate the benefit of the proposed design, particularly under cross-dataset and cross-manipulation settings.
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off: shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling preserves the full-layer latent mean in expectation while explicitly penalizing sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. Applying FuseReg to both stages also reduces unguided gFID by 29% on DiT-Base. The reconstruction and generation benefits also extend to other encoder families. FuseReg narrows the reconstruction-generation gap without additional training cost or architectural changes.
Complementary Roles of Radiomics and Foundation Representations in Renal Cell Carcinoma Classification: A Comparative Study of 2D and 3D CT Encodings
Accurate preoperative subtype classification of renal cell carcinoma (RCC) from contrast-enhanced computed tomography remains clinically challenging. Radiomics provides structured tumour descriptors, whereas foundation representations offer transferable image features. However, it remains unclear whether radiomics still adds value beyond pretrained representations, and how 2D and 3D MedVAE encoders compare in this setting. We compared handcrafted radiomics, 2D MedVAE, 3D MedVAE, and their fusion for binary clear-cell RCC versus non-clear-cell RCC classification on KiTS23 under a unified preprocessing pipeline. Concatenation, cross-attention, and gated fusion were evaluated as representative integration strategies, and radiomics feature importance was analysed to support decision-centric interpretability. Fusion consistently improved discrimination over image-only MedVAE branches. The best overall performance was achieved by 3D gated fusion, with an AUC of 82.7%, outperforming the best 2D fusion model (79.6%), the radiomics baseline (74.4%), and the single-modality MedVAE branches. Ablation analysis further showed clear gains of the full fusion model over both image-only and radiomics-only variants, indicating complementary contributions from radiomics and image representations. These findings suggest that radiomics remains relevant for RCC CT classification in the presence of foundation representations, and that its integration with MedVAE is more effective in the 3D setting. More broadly, the study supports a complementary role for radiomics and foundation representations in clinically meaningful imaging decision support.
Alliance Beats Isolation: Unifying Heterogeneous Allied Datasets Improves Classifier Performance
In many application domains, such as student dropout, insurance fraud, loan approval, and machine failures, several labelled public datasets are available where (i) data is about the same type of objects but the set of actual underlying objects are disjoint; and (ii) the class labels are same; and (iii) the feature spaces of the datasets are largely distinct (heterogeneous), with a few shared features. We call such datasets as allied. A single classifier cannot be trained on both datasets together, and one classifier trained on one dataset cannot be tested on the other. In this paper, we propose a method to merge the feature-spaces into a single feature-space for a pair of given allied heterogeneous datasets. We then use a matrix completion method to create a unified dataset based on the merged feature-space. The hypothesis is that the merged representation facilitates the transfer of classification knowledge from one dataset to another. We conduct experiments on several pairs of allied, heterogeneous datasets and several classifiers to demonstrate that any classifier trained on the unified representation always outperforms classifiers separately trained on the constituent allied datasets on several pairs of allied datasets. This work provides an easy way to substantially improve classifier performance by unifying and using multiple allied datasets together.
Predicting Collision Cross Sections with GRACE: Geometric Residual Adduct Conditioning via Early-fusion
Collision cross section (CCS), derived from ion mobility mass spectrometry, is a common descriptor for molecular annotation. Prediction is challenging for machine learning models because it reflects the size, shape, and ionization state of a gas-phase molecular ion. Most predictors either ignore explicit 3D structure or treat adduct identity as a late categorical feature, which limits their ability to capture adduct-dependent geometric effects. We present GRACE (Geometric Residual Adduct Conditioning via Early-fusion), a 3D CCS predictor that adapts a pretrained molecular geometry encoder using geometric residual adduct conditioning via early fusion. GRACE combines two inductive biases: a residual objective relative to an adduct-aware physical descriptor baseline and adduct conditioning within the encoder via a learned adduct token and low-rank attention adapters. We evaluate the model on a curated set of over 9,000 experimental molecule-adduct CCS records with random, scaffold, and adduct-sensitive splits designed to separate interpolation, scaffold generalization, and adduct-driven generalization. GRACE achieves the best mean percentage difference among the evaluated learned models on all three splits: 1.67% on the random split, 2.11% on the scaffold split, and 2.36% on the adduct-sensitive split. Diagnostic analyses suggest that residual learning stabilizes training by removing the dominant mass-CCS trend, while early fusion improves adduct-sensitive prediction relative to late fusion. Across four independent external test sets, GRACE shows consistently lower error than the other evaluated models. On a held-out set, GRACE also attains the lowest mean percent difference when compared with four previously reported physics-based workflows. These results support residual learning and encoder-level adduct conditioning as practical inductive biases for fast, accurate CCS prediction.
VertiFuseX: Generalizable Financial Forecasting via Multi-Stream Temporal Fusion
Stock price prediction remains challenging due to the non-stationary and noisy nature of financial time series. Existing deep learning models often rely on rigid decision-level fusion, ad hoc hyperparameter tuning, and compressed final-layer outputs, causing information loss, overfitting, and limited cross-market generalization. We propose VertiFuseX, a hybrid LSTM architecture using penultimate-layer vertical fusion of multi-scale temporal representations. VertiFuseX stacks and reweights penultimate features from LSTM, Bi-LSTM, and St-LSTM branches, integrates a parallel DNN stream, and jointly optimizes all components via backpropagation under a fixed hyperparameter configuration. This preserves richer intermediate temporal information across scales. Evaluated on 15 years (2010-2024) of closing prices from 10 global equity indices using strict chronological out-of-sample testing with the final 365 trading days held out, VertiFuseX achieves 30-54% MAPE reductions and over 40% improvements in MAE and RMSE versus LSTM-based baselines, and outperforms seven state-of-the-art models across 33 metric-dataset comparisons. Ablation studies confirm penultimate-layer fusion drives these gains over final-layer fusion and decision-level ensembling. Gradient-based saliency analysis shows consistent emphasis on mid-range dependencies at lags 9-15 days. Economic validation via algorithmic trading simulation under extreme market regimes shows reduced maximum drawdowns and superior risk-adjusted returns. With 675k parameters, a 2.6 MB memory footprint, and 1.5 ms/sample inference latency, VertiFuseX offers a lightweight, interpretable, deployment-ready framework for robust financial forecasting.
HiPerViT: A Hierarchical Perceiver-Vision Transformer Architecture for Multi-Scale Texture Recognition
Texture recognition remains challenging for modern vision models because discriminative evidence is often carried by higher-order spatial statistics rather than by object shape alone. While Vision Transformers provide strong long-range modeling capacity, their standard object-centric representations do not explicitly expose such statistical structure, which limits texture sensitivity in fine-grained recognition settings. We present HiPerViT, a compact vision-only architecture that injects an explicit second-order statistical prior into a transformer-based recognition pipeline. The method combines global and local image views with a compact bilinear descriptor encoded as a statistical token, and integrates this token with first-order spatial representations through Perceiver-style latent distillation. This design enables direct interaction between spatial tokens and second-order feature co-occurrence statistics, providing the model with explicit access to texture-relevant information without requiring multimodal pretraining or ensemble construction. Across six texture recognition benchmarks, HiPerViT achieves consistent improvements over strong vision-only baselines under the reported evaluation protocols, including gains of +3.05 percentage points on DTD, +10.48 on GTOS-Mobile, and +10.10 on 1200Tex. Beyond benchmark performance, our analyses show that these gains are largely invariant to the backbone depth used to extract second-order statistics and to the ordering of interaction and distillation stages. This pattern suggests that the primary source of improvement is not a specific fusion topology, but the explicit availability of second-order statistical information as a first-class representational signal. These results support explicit statistical tokenization as an effective and robust design principle for texture-centric visual recognition.
A Multi-Branch Feature Fusion Approach for Health Misinformation Detection and Propagation
This paper presents a multi-branch fusion framework for detecting and characterising the propagation of health misinformation in online social networks (OSNs). Grounded in the Elaboration Likelihood Model (ELM) and the Theory of Planned Behaviour (TPB), the model fuses transformer-based semantics with rhetorical cues, stance representations, and psychologically motivated proxies in a unified multi-task architecture. In addition to binary classification, we introduce the Cognitive Propagation Score (CPS), an interpretable post-hoc auxiliary score computed from psychologically motivated, text-derived cues capturing argument complexity, emotional intensity, and content-derived virality potential, to support diffusion-risk reasoning when engagement ground truth is incomplete or unavailable. Experiments on three benchmark datasets, Constraint, COVID--19_FNIR, and Monkeypox, show strong classification performance, achieving ROC--AUC up to 0.9999 on COVID--19_FNIR, while propagation-oriented ranking achieves near-perfect agreement when engagement-derived supervision is available (Monkeypox, Spearman's ) and similarly high ranking alignment under proxy-based supervision on COVID--19_FNIR (). Compared with representative literature baselines, the fusion model improves detection on Constraint and COVID--19_FNIR, while Monkeypox remains more challenging, reflecting domain- and signal-specific differences. Ablation analysis further indicates that psychological and rhetorical branches provide complementary gains beyond semantic embeddings. Overall, the framework bridges cognitive theory and neural modelling to improve transparency and to support scalable misinformation monitoring, with future work required to validate CPS against human-centred diffusion judgements.
Static in Frames, Dynamic in Events: Rethinking Features in Event Cameras as Motion Cues
Event cameras capture intensity changes asynchronously with high temporal resolution, requiring novel preprocessing methods for downstream tasks. Unlike static intensity snapshots, event data inherently encode information about scene dynamics and object motion, meaning that features derived from events can exhibit behaviors with no direct analogue in frame-based vision. In this paper, we analyze two features used in event-based corner detection---the eigenvalues of the structure tensor and the spatiotemporal density values---and show that they are \emph{motion cues}. We hypothesize that these features, combined with local geometric information, can enhance motion estimation tasks. To validate this, we first theoretically analyze how the eigenvalues of the structure tensor at moving corner points relate to the direction of motion. We then design controlled experiments on a synthetic dataset, confirming that extending local geometric features with eigenvalues and density values provides complementary motion information and is robust to texture and shot noise. Finally, we integrate the proposed features into a state-of-the-art event-based optical flow network and evaluate on the real-world DSEC benchmark, where the added features consistently improve accuracy, with the largest gains in data-scarce scenarios and for lower-capacity models. The code for this paper can be found at: https://github.com/hesamaraghi/static-in-frames-dynamic-in-events.
KnifeHunter: Structured Local Representation Learning for Fine-Grained Knife Image Retrieval in Law Enforcement
Knife-enabled violence presents a major public safety challenge, and law enforcement agencies require scalable tools for catalogue-level knife identification, intelligence analysis, and source attribution. Manual visual comparison is specialist, time-consuming, and difficult to scale under operational imaging conditions. We introduce KnifeHunter, an end-to-end forensic knife image retrieval system developed with UK law enforcement. The work contributes the KnifeHunter dataset, comprising 25,843 images across 543 knife classes from police evidence, retail catalogues, and border-force seizures, with structured metadata, Medium/Hard evaluation protocols, and large-scale distractor evaluation. We further propose CoRe-Net, a compact single-descriptor retrieval architecture that combines global context with spatially localised discriminative evidence. CoRe-Net introduces Structured Complementary Representation Learning (SCRL) to organise local evidence into complementary prototype-based representations, and Bi-Directional Reciprocal Fusion (BDRF) to integrate global and local evidence through residual projection and gated local-to-global injection. Using an EVA02-Base backbone and cosine-similarity retrieval, CoRe-Net achieves 88.0% mAP and 86.7% mP@10 on the Medium protocol, and 85.1% mAP and 83.8% mP@10 under distractor conditions. KnifeHunter was deployed by UK police forces during Operation Sceptre deployments from 2023 to 2025, achieving 99.2% mP@1 on field queries. These results demonstrate a practical and effective multimedia retrieval framework for fine-grained forensic knife matching in operational law-enforcement settings.
LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding
Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM features into action decoders remains underexplored. Existing designs either expose only a narrow part of the representation hierarchy or rigidly match each decoder block to one VLM layer, restricting access to complementary task evidence across depths. We introduce LIRA, a local cross-layer action-conditioning mechanism that formulates VLM-to-action conditioning as depth-aware information routing. LIRA operates on task-token features and LIRA Query features derived from intermediate VLM states, then assigns each Parallel Fusion Block a depth-aligned local window centered on its corresponding VLM layer. Parallel Fusion Blocks aggregate neighboring LIRA Query features and integrate them with task-token features and proprioceptive inputs before action prediction. This routing interface leaves the backbone architecture, action decoder, and supervised training recipe unchanged. Across LIBERO, LIBERO-Plus, CALVIN ABCD, and real-world manipulation, LIRA improves the principal aggregate metrics over the VLA-Adapter baseline under the same 0.5B-parameter configuration. In zero-shot transfer to LIBERO-Plus, LIRA increases average success from 59.1% to 78.0%, an 18.9-point gain indicating improved robustness under controlled distribution shifts.
Understanding Synergistic Interactions among Pathology Foundation Models via Adaptive Fusion
Pathology foundation models (PFMs) provide strong tile-level representations via self-supervised pre-training on large-scale pathology images. Yet, PFMs are developed under diverse and often opaque data, architecture, and objective choices, inducing latent representational biases that limit robustness and obscure what each model specialises in. We present AdaFusion, a lightweight adaptive fusion framework that integrates complementary signals from multiple frozen PFMs through (1) low-dimensional feature compression and (2) a sample-conditioned gating module that reweights model-wise (and optionally channel-wise) contributions. Beyond improving predictive accuracy, AdaFusion provides contribution-driven interpretation that offers evidence consistent with model-specific preferences and synergistic interactions across tissue phenotypes. We evaluate AdaFusion on three public benchmarks spanning treatment response prediction, prostate cancer grading, and spatial gene expression inference. AdaFusion consistently outperforms individual PFMs and other fusion baselines, while providing interpretable tissue visualisation which aligns model preferences with morphological patterns. Code is available at: https://github.com/xyx-98/PathoOracle.
Finding Change in Satellite Archives from Text: How to Combine Before-and-After Images Efficiently
Operational Earth observation increasingly calls for answering queries such as
find the image pairs where a new building appeared.'' This means searching an archive of before-and-after (bi-temporal) satellite image pairs and ranking each pair by how well it matches a natural-language description of the change. The component that performs this match, the fusion module that combines the before'' and ``after'' views, must be run at query time across many candidate pairs, so its speed largely sets the cost of every search. We present a controlled comparison of how to build that module. Using one fixed image encoder (a frozen CLIP model) and one training recipe for all variants, we evaluate eight designs drawn from three families: attention, state-space models (Mamba), and learned compression (our Temporal Bottleneck Fusion, TBF). Each design is tested on two benchmarks (LEVIR-CC and Dubai-CC) with ten random seeds, so the reported differences are statistically grounded. We outline three findings: first, a training-free two-stage search (a cheap difference model that shortlists candidates, followed by attention fusion that re-ranks them) matches or exceeds full-fusion recall on LEVIR-CC while cutting query cost -, with comparable R@1/R@5 on Dubai-CC; second, the linear-time scan of Mamba, attractive on paper, gives no speed benefit at the patch counts typical of vision transformers (): the scan is limited by memory bandwidth, whereas attention maps cleanly onto parallel hardware; and third, compressing the fused representation (TBF) reduces parameters by and latency by for a change-only BLEU-1 cost of , although more aggressive compression quietly discards change-relevant detail that aggregate metrics fail to reveal.SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification
Accurate cell segmentation and classification are foundational to digital pathology, enabling quantitative tissue analysis for diagnosis and treatment planning. Encoder-decoder architectures that fuse multi-scale features through skip connections have become the dominant paradigm for this task, yet standard direct skip connections treat every spatial location equally, which leads to redundant and potentially conflicting information reaching the decoder. To overcome this problem, various gating mechanisms have been introduced, but most of them operate solely on filtering encoder information, neglecting the benefit of global contextual information from the decoder. This study proposes replacing conventional skip connections in a CellViT-based model with a novel Spatial Attention Fusion (SAF) Gating module. Each SAF gate concatenates the encoder skip and upsampled decoder features, compresses them through two pointwise convolutions with an intermediate ReLU, and applies a channel-wise softmax to produce a per-pixel "heatmap of trust" that sums to unity at every spatial location, allowing the network to learn where each source is most trustworthy. The resulting fused features improve the model's ability to detect the minority "Dead" class, which in turn enhances the multi-class panoptic quality (mPQ) on the PanNuke dataset. SAF Gating is compared against six gating alternatives including no gating, attention gates, squeeze-and-excitation, CBAM, cross-attention, and attentional feature fusion on PanNuke and MoNuSeg datasets. SAF Gating achieves the highest mPQ (0.471), a gain driven primarily by a 14.5-point improvement in Dead-class F1 score compared to ungated CellViT baseline.
ReLATE: Reliability-Guided Evidence Fusion for Robust UAV--Satellite cross-view Geo-Localization
Unmanned aerial vehicle (UAV)-satellite cross-view geo-localization matches UAV images against satellite imagery and has achieved impressive accuracy on clean (non-degraded) image benchmarks. In real-world flights, however, UAV observations are frequently affected by adverse weather, illumination changes, platform motion, sensor noise, and compression, while the robustness of existing methods under such degradations remains largely unexamined. In this paper, we present UAVSat-Deg, a large-scale robustness benchmark for degraded UAV-satellite geo-localization, comprising University-1652-Deg and SUES-200-Deg. UAVSat-Deg covers 27 corruption types, including 19 core and 8 compound corruptions, at three severity levels, supports bidirectional drone-to-satellite and satellite-to-drone retrieval as well as multi-height UAV acquisition, and contains more than 11.7 million pre-generated corrupted test images. Benchmarking representative methods under this protocol reveals substantial robustness gaps, particularly under severe and compound corruptions. To address this problem, we propose ReLATE, a Reliable Evidence Learning framework with Adaptive Token Evidence Regulation, which realizes reliability-adaptive feature fusion during descriptor construction. ReLATE estimates a structure-smoothed reliability field over visual tokens, aggregates trustworthy local evidence, and adaptively integrates it into query-derived representations; the regulated query representations are then combined with the CLS-token and GeM-pooled branches to form the final cross-view descriptor. Across both test sets and retrieval directions, ReLATE achieves the best average corrupted-test performance among the compared methods while maintaining competitive accuracy on clean images. The code and dataset will be available at https://github.com/JHC626/ReLATE.
Beyond Appearance: A Multi-cue Framework and Large-scale Benchmark for Pedestrian Association and Tracking on Mobile Aerial-Ground Platforms
Multi-view Multi-object Association and Tracking (MvMoAT) associates objects across camera views and tracks them over time, supporting identity persistence and forensic trajectory reconstruction in multi-platform cooperative perception. Unlike conventional multiple object tracking, MvMoAT faces frequent viewpoint shifts that distort appearance and undermine cross-view association and temporal tracking. We propose FUSION, a viewpoint-robust Feature Unification framework for multi-view aSsociation and IdentificatiON. Its Multi-cue Adaptive Combination (MAC) module adaptively integrates viewpoint-invariant cues with appearance features to improve cross-view association, while Online Multi-view Feature Synchronization (OMFS) aggregates pedestrian features across historical and cross-view frames for temporally consistent tracking. We also introduce RealMvMoAT, a large-scale benchmark featuring substantial inter- and intra-camera viewpoint variation. It contains 504.9K frames from 7 cameras (5 UAV and 2 ground views) across 10 scenes, with over 7.3M identity-labeled bounding boxes. All cameras exhibit random and substantial motion. To the best of our knowledge, RealMvMoAT is the largest MvMoAT dataset to date. Its scale, viewpoint diversity, complex platform motion, and realistic trajectories provide a comprehensive resource for future research. Experiments on RealMvMoAT and six public benchmarks show that FUSION achieves state-of-the-art performance.
Rethinking Multi-Branch and Cross-Backbone Fusion for Vehicle Re-Identification under Foundation-Model Pretraining
Multi-branch architectures and CNN-Transformer fusion are widely believed to improve vehicle re-identification (Re-ID) by combining complementary representations. We revisit this for a DINOv3-pretrained backbone. A single DINOv3-pretrained ConvNeXt with a tuned recipe reaches 88.19 mAP on VeRi-Wild Small and 77.47 on Large from visual cues alone, within the combined evaluation and optimization noise of the strongest protocol-verified metadata-dependent multi-branch baseline, and 92.38/83.68 with training-free re-ranking. Using this baseline and retrieval-level branch diagnostics, we ask whether representational diversity still pays at this scale. In our runs, it does not. Across both benchmarks and every converged configuration, concatenating multiple heads over a shared backbone moves the best single head by under one mAP point in either direction while costing four times the embedding dimension; 99.7% of the concatenation's variance lies in 512 principal components, so the heads not only duplicate one another but each occupies a quarter of its nominal 2048 dimensions. Pushing diversity to its architectural limit, CNN versus Transformer, we grant fusion every advantage through an asymmetric frozen-anchor scheme. Every Transformer configuration still lands at least 13 mAP below the ConvNeXt backbone (13-15 for the two strongest, up to 46 for the weakest), and a paired per-query bootstrap bounds the fusion gain at +0.11 mAP (95% CI) even for the most favourable snapshot we obtained. One strong backbone with the right recipe and re-ranking is the efficiency frontier. All results use single-seed training and one foundation-model family; differences of this size are therefore reported as bounds rather than orderings, and we list falsifiers.
Improving the performance of an ASV system using hybrid speech features
The growing need for secure and convenient authentication methods has led to the increasing popularity of biometric solutions. In addition to traditional and popular methods, such as fingerprint or iris scanning, voice-based approaches are also employed. User identity verification based on voice is conducted using Automatic Speaker Verification (ASV) systems. Despite their many advantages, these systems are sensitive to various types of attacks and acoustic noises, which can reduce verification accuracy. This work examines the potential to improve the performance of ASV systems by using hybrid feature sets that combine different signal representations, starting with widely-used Mel-Frequency Cepstral Coefficients (MFCC), through Constant Q Cepstral Coefficients (CQCC) and ending with the innovative RAB descriptor. Experiments were conducted on recordings from the Google Speech Commands dataset under two scenarios: in clean conditions and in the presence of acoustic noise. Finally, the systems' performance was compared using the EER metric to determine whether hybrid feature sets decrease verification error. The results show that using a hybrid feature set (PNCC+RAB) improves speaker verification performance under noisy conditions.
How Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing Inspection
Housing-level urban physical examination is essential for identifying residential building problems and supporting targeted urban renewal. Existing automated inspection studies primarily rely on individual images and rarely examine whether surrounding urban functional context can provide supplementary information for building-level assessment. This study proposes a vision-POI fusion framework that combines multi-view visual inspection with POI-derived neighborhood context for residential building health assessment. The empirical dataset covers 92 old residential communities, 3,237 residential buildings, and 25,608 field-acquired inspection images in Qingdao, China, encompassing seven categories of housing-related issues. First, multiple object detection models are evaluated to extract issue locations, categories, and confidence scores from individual images. The image-level outputs are subsequently aggregated across multiple views to construct interpretable building-level representations. Second, POI features are extracted within 500m, 1,000m, and 1,500m neighborhood buffers to characterize surrounding functional environments. Pearson and Spearman correlation analyses, combined with false discovery rate correction, are used to identify candidate contextual features. Finally, visual and POI features are integrated using a cost-sensitive Random Forest classifier under community-isolated spatial cross-validation. The results show that multi-view aggregation provides the main performance improvement, increasing the building-level Macro-F1 from 60.84% under Direct Detection to 74.95%. Incorporating POI context further increases Macro-F1 to 76.79%, although the additional gain is modest and category-dependent. POI information therefore functions as a supplementary contextual prior rather than a substitute for direct visual evidence or a causal determinant of building condition.
CoGoal3D: Collaborative 3D Object Detection with 3D-Aware Fusion and Refinement
V2X collaborative object detection features overcoming the limitations of single-vehicle systems by aggregating environmental features from multiple collaborative agents. However, existing mainstream V2X perception methods mainly focus on 2D BEV object detection. When 3D detection task is concerned, inferior results are obtained because they ignore the 3D spatial misalignment caused by differing height and attitude among the collaborators. In this paper, we propose a novel collaborative 3D object detection framework called CoGoal3D, which extracts and refines the 3D feature gradually in a two-stage pipeline. In the first stage, a multiscale 3D-aware global fusion module is designed to mitigate the 3D spatial misalignment. The resulting proposals are then refined in the second stage with an auxiliary task of 3D point reconstruction. An effective multi-agent collaborative data augmentation strategy is further proposed to enrich the training data while minimizing information loss. Extensive experiments on public real-world datasets demonstrate that our CoGoal3D achieves new state-of-the-art performance, with 3D [email protected] improvements of 10.86%, 10.34%, and 10.18% on the DAIR-V2X, V2V4Real, and V2X-Real datasets, respectively. Code is available at https://github.com/Megalo-f/CoGoal3D.
UVFaceFusion: Fast Multi-view Topologically Consistent Face Reconstruction in the Wild via UV-space Neural Fusion
Reconstructing high-fidelity facial geometry with an assigned topology is essential for digital avatar creation and animation, yet existing automated methods often trade off geometric fidelity and in-the-wild generalization. We present UVFaceFusion, a feed-forward framework for multi-view, fixed-topology face reconstruction from daily images. Our key idea is to replace heuristic topological optimization with learnable neural fusion in a canonical UV space. Given multi-view images, we first obtain dense point maps and facial UV correspondences of each view using VGGT and Pixel3DMM, respectively. Then, the view-specific point maps are lifted into the canonical UV domain and fused with a novel mask-aware neural fusion network. The network predicts a complete UV-space point map, from which a fixed-topology mesh is directly sampled. Although trained only on Ava-256, UVFaceFusion generalizes well to multiple public benchmarks and in-the-wild captures, benefiting from its canonical UV-space geometry-to-geometry fusion that reduces dependence on dataset-specific appearance and capture conditions. Experiments on various benchmarks show that UVFaceFusion achieves state-of-the-art reconstruction accuracy while reconstructing a mesh from 16 input views in less than 3 seconds on a single RTX 4090. Code is available at https://github.com/grignarder/UVFaceFusion.
From Pixel to Prognosis: Convolutional and GLCM Feature Fusion for Automated Four-Class Cataract Severity Classification
Objective: To develop a low-cost automated cataract severity classification system operating on standard consumer-grade colour photographs of the eye, without specialised ophthalmic hardware. Methods: A hybrid framework was designed that fuses deep features from a Convolutional Neural Network (CNN) with five handcrafted Grey-Level Co-occurrence Matrix (GLCM) and intensity descriptors - mean intensity, uniformity, standard deviation, contrast, and energy - extracted from a Hough-circle-localised pupil Region of Interest (ROI). A multi-class Support Vector Machine (SVM) with Radial Basis Function (RBF) kernel classifies each image into one of four severity grades: normal, immature, mature, or hypermature cataract. Results: The proposed fused system achieved 95.0% accuracy, 93.8% sensitivity, and 96.1% specificity on an ophthalmologist-labelled test set drawn from 300 images (75 per class) collected at an ophthalmology clinic, outperforming texture-only (88.5%) and CNN-only (91.3%) baselines and surpassing recently published deep learning approaches. Conclusion: The CNN-GLCM-SVM fusion framework provides competitive four-class cataract grading without GPU acceleration or specialised cameras, making it suitable for primary-care and telemedicine deployment in resource-limited settings.
AffectFuse: Cross-Task Feature Fusion with Temporal Modeling for Multi-Task Affective Behavior Analysis
Affective behavior recognition in the wild requires joint prediction of continuous valence-arousal, categorical facial expression, and multi-label action units from unconstrained face images. We present our system for the Multi-Task Learning (MTL) track of the 11th Affective Behavior Analysis in-the-wild (ABAW) competition on s-Aff-Wild2, the static selected-frame version of Aff-Wild2. The method focuses on post-encoder adaptation: frozen AffectNet-supervised backbones provide multi-resolution features, while task-specific temporal heads and cross-task fusion modules select the useful signals for each target. For action-unit recognition, we adapt MAE-Face with Low-Rank Adaptation (LoRA) and use DISFA through per-unit expert routing rather than direct sequential transfer. Ablations over backbone, temporal, fusion, and AU-adaptation choices define the final configuration. The final system obtains P = 1.7302 on the official validation split, showing that post-encoder adaptation and task-wise modeling choices provide a strong MTL pipeline without training a new large-scale face foundation model.
Task-Specific Feature Fusion Method for Multi-Task Affective Behavior Analysis
The 11th Affective Behavior Analysis in-the-wild (ABAW11) Multi-Task Learning Challenge requires a unified system to predict valence-arousal, categorical expressions, and facial action units from the official s-Aff-Wild2 images. Although these tasks are naturally related through facial behavior, our validation experiments show that they benefit from different visual features, temporal processing strategies, fusion mechanisms, and calibration procedures. In this paper, we study task-adaptive feature fusion for ABAW11 multi-task affective behavior analysis. We first adapt two pretrained visual backbones, DINOv2 ViT-L and DINOv3 ConvNeXt-base, on an external expression-oriented facial image set and then freeze them to extract complementary frame-level features from the official ABAW11 data. On top of these frozen features, we systematically compare frame-level prediction heads, temporal convolutional heads, post-hoc temporal smoothing, LightGBM models, feature concatenation, gated fusion, residual fusion, late logit fusion, threshold calibration, and shared MTL structures. The final system selects task-specific fusion and prediction strategies rather than forcing all tasks to share a single architecture. On the ABAW11 validation set, the selected system achieves an EXPR macro-F1 of 0.4222, an AU macro-F1 of 0.5402, and a mean VA CCC of 0.6717, resulting in an overall validation score of 1.6341. The results suggest that task-adaptive fusion of frozen visual features is a simple and effective strategy for ABAW-style multi-task affective behavior analysis.