cs.CVNov 21, 2025

Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders

Authors: Jacob Beattie, Samuel Stevens, Neil Rosser, Yu Su, Tanya Berger-Wolf

Organizations: The Ohio State University

Abstract

Foundation models in several scientific domains, including visual domains, learn representations that capture complex semantics from their respective fields. Despite this, most existing applications focus on pre-specified concepts or targets, excelling at confirmation but not open-ended discovery of unknown patterns. We study whether sparse autoencoders (SAEs) with heuristic feature ranking can provide a practical methodology for surfacing candidate visual features from vision foundation models for scientific discovery without concept-specific supervision. We evaluate this approach in three stages: (1) general-purpose concept rediscovery, (2) domain-specific concept rediscovery, and (3) question-driven feature ranking. On ADE20K, SAEs outperform baseline methods (kk-means, PCA, SemiNMF) across most downstream metrics for concept rediscovery. On FishVista, the same approach recovers fine-grained fish anatomical features without concept supervision. Finally, on Heliconius butterflies, a heuristic based on the Mann-Whitney UU statistic consistently surfaces features corresponding to real diagnostic keys between mimic subspecies. Together, these results show that SAEs with heuristic ranking can efficiently learn and surface semantic structure in vision foundation model representations. We view this as a methodological prerequisite for open-ended visual scientific discovery: the present experiments validate feature surfacing and ranking through rediscovery, while prospective discovery of previously unknown scientific phenomena remains future work.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?

    Jun 22, 2026Nils Grandien, David Steinmann, Felix Friedrich +1HierarchicalUnsupervised

  2. Position: Use Sparse Autoencoders to Discover Unknowns

    Jun 30, 2025Kenny Peng, Rajiv Movva, Jon Kleinberg +2Improving Sparse AutoencodersModel Discovery

  3. Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders

    Jun 25, 2026Nathanaël Jacquier, Maria Vakalopoulou, Mahdi S. HosseiniImproving Sparse AutoencodersSparsity