q-bio.QMJun 4, 2026

pp-adic Bi-Filtrations for Topological Machine Learning on Genomic Sequences

Authors: Tirtharaj Dash, Gunja Sachdeva

Abstract

We introduce pVR, a topological machine learning framework for alignment-free genomic sequence classification that combines pp-adic numbers with topological data analysis. Each DNA sequence is encoded along two complementary axes: a pp-adic distance on kk-mer prefixes, which captures hierarchical positional structure, and a compositional L1L_1 distance on kk-mer frequencies, which captures local sequence content. The two distances jointly parameterise a bi-filtered Vietoris--Rips complex, and per-sequence topological summaries from this bi-filtration serve as features for standard machine learning classifiers. We establish theoretical guarantees for the construction: stability under metric perturbations and invariance to the choice of prime, alongside a result that explains why a single pp-adic axis is topologically uninformative and why the bi-filtration recovers nontrivial homology. On twelve genomic benchmarks (2828 to 500500 sequences, 33 to 77 classes), pVR outperforms four established alignment-free baselines on three of six low-sample datasets, with gains of up to 2121 percentage points; it underperforms only on a SARS-CoV-2 variant benchmark whose point-mutation divergence violates the hierarchical assumption, and all methods saturate in the large-sample regime. pVR also outperforms zero-shot frozen embeddings from the 500M-parameter Nucleotide Transformer v2 by 6.76.7 to 11.411.4 percentage points on three low-sample benchmarks. The pVR codebase is publicly available at https://github.com/MAHI-Group/pVR.

Explore similar work

Jun 3, 2026cs.CL

LDARNet: DNA Adaptive Representation Network with Learnable Tokenization for Genomic Modeling

Genomic foundation models increasingly adopt large language model architectures, yet almost universally rely on fixed tokenization schemes such as kk-mers, BPE, or single nucleotides, which impose arbitrary sequence boundaries that may obscure biologically relevant structure. We present LDARNet, a 120M-parameter hierarchical genomic foundation model that adapts H-Net-style dynamic chunking from autoregressive generation to masked language modeling, combining BiMamba-2 state-space layers with local attention, bidirectional routing, and a ratio-based regularizer to induce adaptive token boundaries without supervision. Fine-tuned on 27 tasks from the Nucleotide Transformer and Genomic Benchmarks suites, LDARNet achieves 11/18 wins among compact models (<<300M parameters) and state-of-the-art results on 5 histone modification tasks, outperforming models up to 20×\times larger. A FLOPs-matched controlled experiment isolates learned routing as the source of these gains: learned boundaries beat fixed-grid boundaries by up to 14 percentage points on histone tasks at identical compute. Nucleotide-resolution analysis further shows that the learned boundaries align with canonical promoter motifs and splice junctions without supervision, providing a biological interpretation for adaptive tokenization in genomic foundation models.
Daria Ledneva, Denis Kuznetsov
Dec 29, 2025math.AT

Finite Topological Space Filtrations: A Topological Framework for Data Analysis

We introduce a data-analysis framework based on filtrations of finite topological spaces. Starting from a finite metric data set, we construct a sequence of coarsening topologies on the same set of points. These topologies give persistence modules and barcodes in the usual way, but they also retain information that is lost when the filtration is reduced to homology. At each level one can examine, for example, which points are topologically indistinguishable, how their minimal neighbourhoods overlap, how connected components merge, and how these features change from one level to the next. We develop the basic theory of these filtrations, establish stability results under suitable hypotheses, and give practical constructions starting directly from a distance matrix. We then study what can be learned from the resulting finite topologies. On synthetic data with known clusters of different shapes, sizes, and densities, we examine how these regions appear among the finite-topological structures and how they merge as the topology coarsens. We also study what happens when points that become uncovered early in the construction are removed and the analysis is repeated. For one-dimensional homology, we use paths in the finite-topological structure to locate cycles and to examine how their appearance is related to the geometry of the data. We finally apply these ideas to two real data sets with quite different structures. On the Paul15 single-cell data, we use the evolving finite topology to examine fine cellular states, their overlaps and relations, their assembly into larger groups, and the effect of removing points that connect these structures. On COIL20, where images of an object are sampled through a full rotation, we study how the cyclic organization of the images is reflected in the finite-topological evolution and in the associated one-dimensional homology.
Selçuk Kayacan
Date pendingcs.LG

When do cheap embeddings beat protein language models? A theoretically-grounded hashing sketch for biological sequence classification

\textbf{Motivation:} Pre-trained protein language models (PLMs) such as ESM-2 have become the default representation for biological sequence tasks, but they are computationally heavy and require GPUs both for embedding and for fine-tuning. Whether they are actually necessary for sequence \emph{classification}, as opposed to structure prediction, is rarely tested against strong, principled, lightweight alternatives. This question has direct practical stakes for large-scale genomic surveillance, where embedding millions of sequences on commodity hardware is a recurring bottleneck.\ \textbf{Results:} We introduce Murmur2Vec, an alignment-free, training-free embedding that aggregates kk-mer counts into a small hash table via the deterministic MurmurHash function, and we cast it as a randomized sketch of the classical kk-mer spectrum kernel. We provide a complete theoretical treatment: closed-form bias/variance of the inner product, an unbiased signed variant with a Johnson--Lindenstrauss-type concentration bound, an excess-risk bound for downstream linear classifiers that makes the bias--variance trade-off in the hash-table size explicit, and an implicit-regularization mechanism by which collisions damage frequent non-discriminative kk-mers more than rare lineage-defining ones. Across four classification tasks, SARS-CoV-2 spike lineage (22 classes), HIV-1 Env subtype (8 classes), and two protein-family benchmarks (8 and 6 classes), Murmur2Vec matches a LoRA-fine-tuned 650M-parameter ESM-2 model on the two tasks for which LoRA fine-tuning was run to convergence (SARS-CoV-2 and HIV-1) and ties frozen ESM-2 on the two protein-family tasks, and it \emph{outperforms} the fine-tuned model on the hardest task (SARS-CoV-2 lineage: 0.8540.854 vs.\ 0.8070.807 accuracy; macro-F1 0.6840.684 vs.\ 0.4010.401).
Sarwan Ali, Taslim Murad, Imdadullah Khan +1