stat.MLJul 9, 2025

Semi-parametric Functional Classification via Path Signatures Logistic Regression with Adaptive Order Selection

Authors: Pengcheng ZengSiyuan Jiang

Organizations: Institute of Mathematical Sciences, ShanghaiTech University, Shanghai, China

Abstract

We propose Path Signatures Logistic Regression (PSLR), a semi-parametric framework for classifying vector-valued functional data with scalar covariates. Classical functional logistic regression models rely on linear assumptions and fixed basis expansions, which limit flexibility and degrade performance under irregular sampling. PSLR leverages the well-established properties of path signatures - basis-free representation, cross-channel dependency capture, and robustness to sampling irregularity - as an enabling tool. The key novelty, however, lies in two distinctive contributions: (i) a semi-parametric additive structure that preserves interpretable linear effects for scalar covariates, and (ii) a fully data-driven procedure for adaptively selecting the signature truncation order via a penalized empirical risk criterion. This selection mechanism is supported by rigorous non-asymptotic guarantees, including the existence of an optimal truncation order, its consistent estimation from finite samples, convergence rates for the classifier risk, a finite computable search bound, and an error propagation framework that formally quantifies PSLR's robustness under irregular sampling. Experiments on synthetic and real-world datasets demonstrate that PSLR with adaptive order selection consistently outperforms traditional functional classifiers and fixed-order signature baselines in accuracy, robustness, and interpretability. Our results highlight the practical and theoretical value of integrating rough path theory with adaptive model complexity control.

Explore similar work

Sep 8, 2026stat.ML

Optimal estimation for Functional Linear Regression with Noisy Discretized Data

In this paper, we consider the scalar-on-function linear regression model under a realistic sampling scheme in which the functional covariates are observed on a regular grid and contaminated by additive noise. We propose a two-step estimation procedure: first, the underlying curves are reconstructed from the discrete noisy observations using a Fourier-based projection method; second, the slope function is estimated by a penalized least-squares criterion over finite-dimensional trigonometric spaces, with data-driven selection of the model dimension. We establish oracle-type inequalities for the prediction error, both with respect to the reconstructed curves and to the true latent curves. Under regularity assumptions on the slope function and polynomial decay of the eigenvalues of the covariate, we derive convergence rates for the prediction error and show that our estimator attains the minimax rate when the number of grid points is sufficiently large. Finally, the proposed method is illustrated on simulated data and on a real meteorological dataset.
Sixtine Sphabmixay
May 19, 2026stat.ML

Posterior Contraction of Lévy Adaptive B-spline Regression in Besov Spaces

We investigate the asymptotic properties of the Lévy Adaptive B-spline (LABS) regression model, a Bayesian nonparametric method that incorporates B-spline kernels into the Lévy Adaptive Regression Kernel (LARK) model. LABS applies splines of varying degrees with independently defined knots, yielding a flexible model class capable of adapting to irregular and locally structured features of the true function. Within the nonparametric regression framework with univariate random design and Gaussian errors, we establish that the LABS posterior contracts around the true function in Besov classes at nearly minimax-optimal rates, up to a logarithmic factor, while adapting automatically to unknown smoothness. This study contributes to filling a gap in the literature, where theoretical results on posterior contraction of the LARK model in Besov spaces remain scarce. Simulation experiments on standard test functions in Besov spaces, including Blocks, Bumps, HeaviSine, and Doppler, complement the theoretical results and demonstrate the practical utility of LABS.
Jeunghun Oh, Sewon Park, Jaeyong Lee
Jun 13, 2026cs.LG

Probabilistic Signature Inversion: Learning Conditional Distributions from Truncated Signatures

The signature transform is a principled feature map for continuous-time paths, valued for its uniqueness and universality. Recovering a path from its truncated signature is, however, structurally ill-posed because the truncated signature map is not injective. We therefore reframe truncated signature inversion as a probabilistic problem -- learning the conditional distribution of a path given its truncated signature -- and adopt a signature-conditioned flow matching model as a practical estimator. This probabilistic formulation elucidates the fundamental difficulty of inversion: Bayes reconstruction error quantifies the irreducible uncertainty remaining after conditioning on a statistic. We derive the Bayes-optimal error under linear statistics, obtaining a closed form for log-GBM and numerically tractable formulas for log-fBM and OU, yielding a concrete theoretical baseline for model validation. This baseline upper-bounds the Bayes error under truncated-signature conditioning, since truncated signatures provide richer information than linear statistics. Experiments show that empirical reconstruction errors under linear-statistics conditioning faithfully align with the theory-derived baseline, while errors decrease when the statistic is replaced with truncated signatures. Moreover, generated paths faithfully recover the conditioning signature while preserving key distributional and temporal structures, indicating that the estimator is well-calibrated to the target conditional distribution. Together, these results establish a well-posed probabilistic framework for truncated-signature inversion, with applicability demonstrated on real financial data beyond the parametric process families covered by theory.
Junoh Kang, Kiseop Lee, Bohyung Han