Archetypal analysis (AA) was originally proposed in 1994 by Adele Cutler and Leo Breiman as a computational procedure for extracting distinct aspects, so-called archetypes, from observations, with each observational record approximated as a mixture (i.e., convex combination) of these archetypes. AA thereby provides straightforward, interpretable, and explainable representations for feature extraction and dimensionality reduction, facilitating the understanding of the structure of high-dimensional data and enabling wide applications across the sciences. However, AA also faces challenges, particularly as the associated optimization problem is nonconvex. This is the first survey that provides researchers and data mining practitioners with an overview of the methodologies and opportunities that AA offers, surveying the many applications of AA across disparate fields of science, as well as best practices for modeling data with AA and its limitations. The survey concludes by explaining crucial future research directions concerning AA.
Figures & tables
Fig. 1: AA computed on the subset of the MNIST handwritten digit dataset showing solely the digit 9 with three archetypes. The left part depicts the mixing weights of the nines whereas the right part depicts the actual handwritten digits.
Fig. 2: Citation overview of paper [ 1 ] in Google Scholar and SCOPUS until 5th November 2025.
Fig. 3: An example of AA with K=3 archetypes on two-dimensional toy data. After initializing the archetypes (purple squares), the optimization procedure pushes the archetypes (green squares) to lie on the boundary of the convex hull of data (see Theorem 1 ). The overall objective is to minimize the sum of projection errors (i.e., the sum of all dashed lines).
Fig. 4: An example of AA in two dimensions for various numbers ( K=2,3,4 ) of archetypes {a1,…,aK} . The archetypes are always located on the boundary of the convex hull of data.
Fig. 5: AA with K=3 archetypes compared to a nonnegative matrix factorization (NMF) with K=2 components, a principal component analysis (PCA) with K=2 components (first is solid, second dashed), k -means clustering with K=3 clusters, and k -maxoids clustering with K=3 clusters.
Algorithm
Restrictions on C and A
Restrictions on S
PCA
C
S
NMF [ 6 ]
A≥0
S≥0
Convex NMF [ 17 ]
C≥0
∥ck∥1=1
S≥0
Affine Coding [ 18 ]
0≤akm≤1
∥sn∥1=1
Convex Coding [ 18 ]
0≤akm≤1
S≥0
∥sn∥1=1
Conic Coding [ 18 ]
0≤akm≤1
S≥0
TABLE I: An overview of related methods and their restrictions.
Fig. 6: Only M+1 archetypes A′ are needed to define a data point. Hence, in the case of K>M+1 , the weights sn are not consistent. For the highlighted data point in the middle of the dataset, two different subsets A′ (blue shaded area) of M+1 archetypes can be used to define the point. Thus, the representation sn is not consistent. Here, M=2 and K=4 .
Method
Archetypes
Location
Representation space
Purpose / when to use
AA
Mixtures of observations
Boundary of the convex hull
Original data space
Standard choice for interpretable extremal profiles
Relaxed AA
Relaxed mixtures of observations
May extend outside the convex hull
Original data space
When extreme profiles may not be well represented by the observed convex hull
ADA
Actual observations
Boundary or interior of the convex hull
Original data space
When archetypes should correspond to tangible, observed cases
Kernel AA
Mixtures of mapped observations
Boundary of the convex hull in feature space
Nonlinear kernel space
When nonlinear structure or pairwise similarities are important
AAnet / DeepAA
Latent archetypal representations
Latent archetypal space
Learned nonlinear space
When the nonlinear representation should be learned jointly with the archetypal structure
Sparse AA
Sparse mixtures
May extend outside the convex hull
Original data space
For sparse representations in hyperspectral unmixing
TABLE II: Taxonomy of representative extensions of archetypal analysis.
Optimization algorithm
Abbreviation
Computational Cost OS
Computational Cost OC
Nonnegative Least Squares [ 81 , 1 ]
NNLS
OS(NK3)
OC(K∣A∣3)
Active Set [ 70 ]
AS
OS(KM+∣A∣2)
OC(KM+∣A∣2)
Principal Convex Hull Analysis [ 8 ]
PCHA
OS(NMK)
OC(NMK)
Frank–Wolfe [ 84 ]
FW
OS(NMK)
OC(NMK)
Softmax + Adam [ 56 , 66 ]
SA
OS(NMK)
OC(NMK)
Block Coordinate Descent [ 78 ]
BCD
OS(NMK)
OC(NMK)
TABLE III: Comparison of the principal optimization algorithms used in AA. Costs are per update and exclude the O(NMK) shared by every update for forming A=CX and SCX , which for the gradient-based procedures is already the dominant term.
Fig. 7: A comparison of the following initialization techniques of AA for K=3 : Uniform, FurthestFirst, FurthestSum, and AA++.
Name
Language
Description
Models
Optimization algorithms
Sparsity
Nonlinearity
Robustness
Weighted
LEAST-SQUARES AA
archetypes [ 111 ]
R
Package for AA
AA
NNLS [ 1 ]
✗
✗
✓
✓
adamethods [ 115 , 100 ]
R
Package for archetypoid analysis
ADA
NNLS [ 1 ] , ASA [ 19 ]
✗
✗
✓
✗
archetypes [ 46 ]
Python
Package for AA and visualization tools
AA, ADA, BiAA
PCHA [ 8 ] , NNLS [ 1 ] , PAA [ 3 ]
✗
✓
✗
✓
SPAMS [ 70 ]
R, Matlab, Python
Sparse modeling package including AA
AA
AS [ 70 ]
✗
✗
✗
✗
ParetoTI [ 110 ]
R
Package for AA and Pareto task inference
AA
PCHA [ 8 ]
✗
✗
✗
✗
TABLE IV: Summary of the main available packages and tools for archetypal analysis.
Fig. 8: Left panel: Simplex plot representing the scores, where each observation is a point, and the archetypes are positioned at the colored vertices. The location of each point indicates its relative contribution from the archetypes. Right panel: Stacked bar chart of the scores, where each bar represents an observation and the colored segments correspond to the contributions of each archetype.
Fig. 9: Bubble map with the main areas of AA applications.
Fig. 10: Analysis of data on Finches 6 6 6 Data taken from https://datadryad.org/dataset/doi:10.5061/dryad.g6g3h in which two features are measured based on beak length and depth for scandens (blue dots) and fortis (black dots) finches illustrated for a K=2 , K=3 , and K=4 AA model. When inspecting the variance explained we observe substantial improvements up until K=3 , and reproducibility as measured using NMI(S,S′) and sim(A,A′) we observe that a K=3 model is also reliably recovering S and A . Whereas the K=2 AA model mainly discriminates between the two types of finches, the K=3 characterizes in particular variability in the fortis finch class whereas the K=4 further subdivides the scandens finch class.
Fig. 11: Analysis of data on mixture designed NMR experiment containing systematically different fractions of propanol, butanol and pentanol 8 8 8 Data taken from https://ucphchemometrics.com/datasets/ . When inspecting the variance explained we observe substantial improvements up until K=3 and reproducibility as measured using NMI(S,S′) and sim(A,A′) we observe that a K=3 reliably recovers S and A . This three component model well corresponds to the three compounds whereas the estimated concentration fractions (given at the bottom) well correspond to the true concentration levels of the three compounds in each sample.
Fig. 12: Analysis of a remote sensing hyperspectral image data 10 10 10 Data taken from https://www.ehu.eus/ccwintco/index.php?title=Hyperspectral_Remote_Sensing_Scenes#Indian_Pines in which an image of a landscape in Indiana, US is measured across 200 wavelengths (removing bands pertaining to water absorption) in a 145×145 image. The image is turned into a matrix of 200 wavelength features by 1452 pixel observations. When inspecting the variance explained and reproducibility as measured using NMI(S,S′) and sim(A,A′) we observe that K=6 is the largest number of components that can be extracted while attaining reliable solutions. Inspecting the learned AA representation for K=6 we observe that the archetypes provide distinct spectral profiles with their convex combinations defining different regions of the image in which these profiles are present.
Fig. 13: On the left, we analyze a two-dimensional dataset where points follow a linear relationship, except for an outlier shown in red. The outlier appears clearly separated from the rest in the representation for K = 3 AA model, unlike the original space or PCA projection. The table shows AUC results (first row) and the rank of outlierness of the outlier (second row) for the outlier detection methods by [ 131 ] . As there are 11 points, the highest possible rank is 11, which corresponds to the lowest degree of outlierness. Projection in the archetypal space improves distance based outlier detection techniques such as k -NN. AA + k NN is the only method that detects the outlier. On the right, we reproduce an example by [ 132 ] where functional AA projections and k -NN are used to detect the anomaly in the landscape image.
Fig. 14: Analysis of data of congressional voting records 12 12 12 Data taken from https://archive.ics.uci.edu/dataset/105/congressional+voting+records considering 16 bills with votes missing by some congress members on some bills. The Kernel AA procedure is here used considering the Jaccard similarity between congress members in terms of bills they both voted for ignoring bills either had as missing value, i.e., KijJac.=∑m′:(xim′∈{0,1},xjm′∈{0,1})1−(1−xim′)(1−xjm′)∑m:(xim∈{0,1},xjm∈{0,1})ximxjm . The kernel has been double centered, i.e., TKJac.TT prior to analyses, corresponding to removing the mean. Black dots correspond to Republicans, whereas blue dots correspond to Democrats. The learned models are illustrated for a K=2 , K=3 , and K=4 AA model. When inspecting the variance explained and reproducibility as measured using NMI(S,S′) and sim(A,A′) we observe that up to K=4 the AA models can reliably be inferred whereas the variance explained steadily improves as the number of components is increased. Inspecting the model for K=2 , K=3 , and K=4 we observe that the models well separate Democrats (blue dots) from Republicans (black dots).
Figure 19
Irene Epifanio received her MS degree in mathematics and her PhD degree in statistics from València University, Spain. She is currently a full professor of Statistics at Jaume I University, Spain, and Senior Research Fellow at valgrAI. Her research interests include statistical learning, functional data analysis, computer vision and equality. She was the recipient of various awards in research, teaching, scientific dissemination and social commitment, including the Margarita Salas Prize of Talent Woman Spain.
Archetypal Analysis (AA) represents observations as convex combinations of extremal data-driven profiles, yielding interpretable low-dimensional descriptions of complex datasets. Classical AA relies on a least-squares objective, which is poorly suited to discrete observations such as binary, count, and categorical data. We introduce an efficient likelihood-based framework for AA supporting Bernoulli, Poisson, and multinomial observation models. Our optimization scheme employs local quadratic approximations of the negative log-likelihood, enabling constrained updates through sequential minimal optimization (SMO) and an active-set method. Scalability is improved by bounding the active set while preserving simplex feasibility. We further introduce a cross-validated predictive likelihood criterion for selecting the number of archetypes, providing a principled alternative to reconstruction-error heuristics and stability-based diagnostics. Synthetic experiments demonstrate computational efficiency and accurate recovery of model complexity. Applications to single-cell RNA sequencing, microbiome composition, and somatic mutation data show that the learned archetypes capture interpretable domain-specific structures while achieving competitive likelihood fits and stable solutions. Overall, the proposed framework enables efficient likelihood-based archetypal analysis of discrete data, complemented by predictive likelihood-based model selection.
A. Emilie J. Wedenborg, Jesper Løve Hinrich, Morten Mørup
Department of Applied Mathematics and Computer Science, Technical University of Denmark, 2800 Kgs. Lyngby, Denmark
Classical archetypal analysis is appealing for its interpretability, but its linear geometry can limit performance on data with strongly non-linear structure; at the same time, existing neural extensions improve flexibility while often weakening the geometric meaning of archetypes and interpolations. In this work, we develop a Riemannian version of archetypal analysis based on data-driven pullback geometry for real-valued data, with the goal of combining the interpretability of classical archetypal analysis with the expressive power of modern non-linear models. We introduce a class of deformed star distributions together with associated pullback Riemannian geometry to provide a statistical interpretation of the resulting manifold mappings, define the Riemannian archetypal mapping (RAM) as a projection onto the manifold of geodesically convex combinations of archetypes, and propose a practical optimization scheme based on convex relaxation followed by non-convex refinement. We further propose a learning scheme that yields reasonable, albeit generally suboptimal, deformed star distributions from data. Experiments on synthetic examples and MNIST show that the resulting framework produces meaningful geodesics, useful denoising projections, and geometry-aware classifications, while also clarifying where current optimization limitations remain.
Willem Diepeveen, Deanna Needell
Department of Mathematics University of California, Los Angeles Los Angeles, CA 90095, USA
We introduce biarchetype analysis for the first time in the context of univariate functional data. This unsupervised methodology extends archetype analysis by simultaneously identifying archetypal structures across both the cases (countries, in our application) and the temporal argument. Both cases and time points are expressed as mixtures of biarchetypes, yielding a concise and highly interpretable representation of complex functional observations. Although biarchetype analysis is not intended as a clustering technique, it offers superior interpretability compared with biclustering approaches, as it is based on extreme, representative patterns rather than average centroids, thereby enhancing human comprehension. We apply the proposed method to 10-year government bond yields of European countries over the period 2001-2025. The results identify three distinct time regimes (the pre-crisis period, the euro-area sovereign debt crisis, and the post-crisis period), and reveal Germany, Greece, and Hungary as country archetypes.
Aleix Alcacer, Rafael Benitez, Vicente J. Bolos +1
Jaume I University, Castelló 12071, Spain. · University of València, València 46022, Spain.