cs.LGOct 1, 2026

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

Authors: Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, +2 more

Organizations: Stanford University

Abstract

Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforcement learning with sparse autoencoder features (RL-SAE), a post-training method that rewards the generation of sequences that activate specified feature sets. Across eight IDR design tasks, RL-SAE sequences activate, on average, 90% of 30 targeted features, compared to 24% for activation steering. We demonstrate that RL-SAE improves the predicted subcellular localization and transcriptional activity of generated IDRs compared to steering and supervised fine-tuning, and enables features associated with distinct biological functions to be combined within individual sequences. Thus, IDiom and RL-SAE enable interpretable and composable IDR design through explicit control of function-associated sequence features. More broadly, RL-SAE could extend to other protein design settings where interpretable features provide useful design targets. Code is available at https://github.com/rotskoff-group/idiom.

Figures & tables

Appendix figures & tables25 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Proteo-R1: Reasoning Foundation Models for De Novo Protein Design

    May 1, 2026Fang Wu, Weihao Xuan, Heli Qi +26Protein DesignMoleculenet

  2. ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning

    Jun 9, 2025Ziwen Wang, Jiajun Fan, Ruihan Guo +3Protein DesignProtein

  3. VFUSE: Virulent Feature Understanding with Sparse autoEncoders

    Jun 8, 2026Michael Yu, Matthew L. OlsonProtein DesignSparse Autoencoder Features