Generative Pretrained Transformers

Latest papers 42

All topics
CardsList
  1. Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

    Sep 29, 2026Zehao Jin, Ruixuan Deng, Junran WangTransformer ArchitecturesGenerative Pretrained Transformers

  2. ZonoGPT: Towards An Abstract Domain for Verifying Large GPT Models

    Sep 28, 2026Hai Duong, Thanh Le, ThanhVu NguyenGenerative Pretrained TransformersTransformer Architectures

  3. Pretraining Transformers with Quantized Softmax in Attention

    Sep 27, 2026Shangzhen Zhu, Muyan Hu, Tomasz KozlowskiTransformer AttentionGumbel-Softmax Relaxation

  4. Learning What to Remember: Test-Time Training via Context Distillation

    Aug 3, 2026Zixuan Wang, Xingyu Dang, Rui-Jie Zhu +4Efficient Long-Context InferenceTest-Time Training

  5. PolymerGPT: Multi-property Optimization with a Decoder-Based GPT Model for Generative Polymer Design

    Aug 2, 2026Charlie Pyle, Adarsh Gadari, C. Adrian Figg +3Polymer Property PredictionGenerative Pretrained Transformers

  6. Training nGPT

    Aug 2, 2026Ilya Loshchilov, Boris GinsburgGenerative Pretrained TransformersTransformer Architectures

  7. Verifier-Guided Model Discovery for Physical Dynamical Systems with Pretrained Symbolic Transformers

    Aug 1, 2026Farbod Faraji, Francesco BelardinelliVortexDynamical Systems

  8. StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis

    Jul 22, 2026Kaicheng Luo, Xuefei Gong, Yutao Sun +6Autoregressive Text-To-SpeechSpeech Synthesis

  9. Learning Standard Model structure from LHC data with Riemannian flow matching

    Jul 17, 2026Midori Kato, Kevin A. Urquía-Calderón, Inar Timiryasov +1Riemannian Flow MatchingGenerative Pretrained Transformers

  10. DNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks

    Jun 29, 2026Romain Karpinsky, Julien Mozziconacci, Mickaël DelceyGenerative Pretrained TransformersTransformer Architectures

  11. Muown Implicitly Performs Angular Step-size Decay

    Jun 22, 2026Florian Hübler, Kai Lion, Antonio Orvieto +1MuonAdaptive Optimizers

  12. RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild

    Jun 22, 2026Cheng Cui, Tingquan Gao, Xueqing Wang +11Document ParsingOrder Matters

  13. A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition

    Jun 18, 2026Nabil Mosharraf Hossain, Riasat Islam, Unaizah ObaidellahDiscrete Speech RepresentationsGenerative Pretrained Transformers

  14. HLS-GPT: A Generative Pretrained Transformer (GPT) for Continental-Scale NASA Harmonized Landsat and Sentinel-2 (HLS) Reflectance Reconstruction Across All Bands on Arbitrary Dates

    Jun 16, 2026Junjie Li, Hankui K. Zhang, David P. RoySentinel-2Remote Sensing

  15. Toward Controllable Catalyst Inverse Design via Large-Scale Autoregressive Pretraining

    Jun 16, 2026Dong Hyeon Mok, Jonggeol Na, Seoin BackHeterogeneous CatalysisGenerative Pretrained Transformers

  16. Simplifying the Modeling of Arbitrary Conditionals in Natural Language

    Jun 12, 2026Yinhan Lu, Eric Elmoznino, Léo Gagnon +3Generative Pretrained TransformersHigh-Fidelity Conditional Generation

  17. SpikeDecoder: Realizing the GPT Architecture with Spiking Neural Networks

    Jun 10, 2026Claas Beger, Florian Walter, Alois KnollSpiking Neural NetworksTransformer Encoder

  18. Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking

    Jun 2, 2026Zekun Qi, Xuchuan Chen, Dairu Liu +10Generative Pretrained TransformersSequential Scaling

  19. LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback

    May 28, 2026Jiwon Kim, Maya Ajit, Sherry Gong +4Mental HealthAi-Assisted Writing

  20. Argo: Efficient Importance Labeling for Enterprise Email Systems

    May 20, 2026Siddhant Ray, Ganesh Ananthanarayanan, Kevin Chian +5Inference CostImportance

  21. LLM Pretraining Shapes a Generalizable Manifold: Insights into Cross-Modal Transfer to Time Series

    May 19, 2026Alexis Roger, Prateek Humane, Zhenghan Tai +4Large Language Model PretrainingGenerative Pretrained Transformers

  22. MiniGPT: Rebuilding GPT from First Principles

    May 17, 2026Jibin JosephAutoregressive Language ModelsGenerative Pretrained Transformers

  23. From Sparsity to Simplicity: Enabling Simpler Sequential Replacements via Sparse Attention Distillation

    May 15, 2026Yuxin Ren, Maxwell D Collins, Miao Hu +1Transformer AttentionAttention Layers

  24. Masked Generative Transformer Is What You Need for Image Editing

    May 11, 2026Wei Chow, Linfeng Li, Xian Sun +14Diffusion-Based Image EditingImage Editing

  25. Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining

    May 11, 2026Jinchang Zhu, Jindong Li, Yuwen Hao +3Large Language Model PretrainingPretraining

  26. Priming: Hybrid State Space Models From Pre-trained Transformers

    May 8, 2026Aditya Chattopadhyay, Elvis Nunez, Prannay Kaul +6Generative Pretrained TransformersPriming

  27. Spreadsheet Modeling Experiments Using GPTs on Small Problem Statements and the Wall Task

    Apr 28, 2026Thomas A. Grossman, Yuan Chen, Sopiko DatuashviliSpreadsheetsReproducibility

  28. RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian

    Apr 21, 2026Mircea Timpuriu, Mihaela-Claudia Cercel, Dumitru-Clementin CercelGrammatical Error CorrectionLegal Domain

  29. Freeze, Diffuse, Decode: Task-Aware Adaptation of Transformer Embeddings for Antimicrobial Peptide Design

    Nov 28, 2025Pankhil Gawade, Adam Izdebski, Myriam Lizotte +4Generative Pretrained TransformersBehavioral Embeddings

  30. RheOFormer: A generative transformer model for simulation of complex fluids and flows

    Oct 1, 2025Maedeh Saberi, Amir Barati Farimani, Safa JamaliFluid DynamicsSurrogate Models

  31. Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture

    Jun 24, 2025Shuchen Xue, Tianyu Xie, Tianyang Hu +5Diffusion Language ModelsMasked Diffusion Language Models

  32. MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning

    Jun 10, 2025Javier Lopez-Piqueres, Pranav Deshpande, Archan Ray +3Parameter-Efficient Fine-Tuning MethodsModel Fine-Tuning

  33. Ensemble Learning for Large Language Models in Text and Code Generation: A Survey

    Mar 13, 2025Mari Ashiga, Wei Jie, Fan Wu +5Large Language Model GenerationCode Generation

  34. Attention is All You Need Until You Need Retention

    Jan 15, 2025M. Murat YasliogluRetentionRecurrent Model