Causal discovery becomes particularly challenging when the available sample size is small relative to the number of variables. This challenge also arises in the linear non-Gaussian acyclic model (LiNGAM), an identifiable framework for causal discovery from observational data. DirectLiNGAM estimates a causal order, which arranges variables so that causes precede their effects, by sequentially identifying an exogenous variable and removing its linear effect from the remaining variables. We establish a structural limitation of this procedure: when the number of variables exceeds the sample size, repeated residualization necessarily becomes degenerate before the full causal order can be determined. Our analysis further reveals that each residual can be reconstructed using only a graph-determined subset of variables already placed earlier in the causal order, termed the active boundary. This result motivates AdaPS-LiNGAM (Adaptive Predecessor Selection LiNGAM), which reconstructs each residual directly from the original observations using an adaptively chosen sparse subset of those earlier variables. The same subset-selection principle is also applied to the final pruning step for edge estimation. Experiments on synthetic data demonstrate that AdaPS-LiNGAM provides accurate causal-structure recovery in sample-limited settings and degrades more gradually as the sample size decreases.
Figures & tables
Figure 1 : Results for the small-sample experiment with n=100 under uniform noise, with p∈{10,25,50,75,100,125,150,175,200} . Markers show the mean over 100 trials and error bars ±1 standard deviation. The SHD axis is truncated at 1000 .
Figure 2 : Results for the small-sample experiment with n=100 under Laplace noise, with p∈{10,25,50,75,100,125,150,175,200} . Markers show the mean over 100 trials and error bars ±1 standard deviation. The SHD axis is truncated at 1000 .
Figure 3 : Results for the reduced-sample experiment with p=100 under uniform noise, with n∈{100,500,1000,1500,2000} . Markers show the mean over 100 trials and error bars ±1 standard deviation. The SHD axis is truncated at 1000 .
Figure 4 : Results for the reduced-sample experiment with p=100 under Laplace noise, with n∈{100,500,1000,1500,2000} . Markers show the mean over 100 trials and error bars ±1 standard deviation. The SHD axis is truncated at 1000 .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Empirical mean active-boundary sparsity of the RG-indeg3 graphs under the fixed true causal order π(i)=i , together with the theoretical lower bound on E[smax(π)] from Example 2 . Error bars indicate one standard deviation across graph realizations.
Figure 6 : Active-boundary sparsity across the four graph families used in Section 6 . For each true graph realization, ten valid causal orders were sampled. Curves show the median across graph realizations and sampled orders, and shaded regions indicate the interquartile range.