Pretraining

Latest papers 428

All topics
CardsList
  1. Pretraining Shapes Spectral Structure: Architecture- and Strategy-Conditional Prediction of OOD Robustness in Foundation Models

    Oct 7, 2026Sangyoon Bae, Sk Miraj Ahmed, Shinjae Yoo +1Pretraining

  2. Sequential Pretraining Favors Large Models

    Oct 7, 2026Mohnish Harwani, Yujia ZhengPretrainingModel Size

  3. Atom-JEPA: Joint-Embedding Predictive Architecture for 3D Atomistic Systems

    Oct 6, 2026Kasper Helverskov Petersen, Rasmus Hannibal Tirsgaard, François R J Cornet +2Joint-Embedding Predictive ArchitecturesPretraining

  4. Differentially Private Mixing of Public Datasets Improves Private Learning

    Oct 5, 2026Yufei Chen, Tejumade Afonja, Anvith Thudi +1Pretraining

  5. EMG-FM-Bench: A Comprehensive Benchmark for Foundation Model Transfer and Adaptation on Electromyography

    Oct 5, 2026Tianhao Wu, Xu Wu, Amirmohammad Radmehr +4Electroencephalography Foundation ModelsElectrocardiogram Foundation Models

  6. Learning Decision-Stump Thresholds in Context: Dynamics of Softmax Attention

    Oct 5, 2026Hong Ha Le, Jackie Lok, Atsushi Nitanda +1Gumbel-Softmax RelaxationPretraining

  7. RepICL: Reusable In-Context Prediction Across Heterogeneous Representation Spaces

    Oct 5, 2026Yu-Hsiang Liu, Kuan-Yu Chen, Chih-Sheng Chen +3Few-Shot LearningIn-Context Learning

  8. What Is a Repeated Token Worth? The Scaling Geometry of Multi-Epoch Pretraining

    Oct 4, 2026Yekun Chai, Haoyi XiongEpochPretraining

  9. Local Support Learning

    Oct 1, 2026Assaf Ben-Kish, Akarsh Kumar, James Glass +1Catastrophic ForgettingLarge Language Model Memory

  10. Artifact Annotations Partially Substitute for Per-User Calibration: SAFE-EDA and a Normalization-Controlled Evaluation of Wrist-EDA Affect Recognition

    Oct 1, 2026Haochen Chai, Xinbi Luo, Zining Liu +1Surface ElectromyographyNormalization

  11. Localizing Transfer Between Memorization Tasks

    Sep 30, 2026Yimiao Yu, Florentin GuthTransfer LearningPretraining

  12. Replay on Demand: An Emergent Curriculum for Balancing Adaptation and Forgetting in Continued Pretraining

    Sep 30, 2026Lukas Thede, Shengzhuang Chen, Stefan Winzeck +3ReplayPretraining

  13. How Does Local Landscape Geometry Evolve in Language Model Pre-Training?

    Sep 30, 2026Zhanpeng Zhou, Yuhan Sun, Bingrui Li +4Large Language Model PretrainingBatch

  14. Comparative study of adapting pre-trained models for driving behavior video captioning

    Sep 30, 2026Sayak Mallick, Philipp Geiger, Augustin KelavaAutonomous DrivingVision-Language Model Adaptation

  15. Lasting Effects of Abstract Pretraining Beyond Perplexity

    Sep 30, 2026Zachary Shinnick, Hemanth Saratchandran, Damien Teney +1Large Language Model PretrainingPretraining

  16. What Pretraining and Midtraining Make Learnable from Rewards?

    Sep 29, 2026Chiwun Yang, Xiaoyu LiReinforcement Learning Post-TrainingPretraining

  17. Rethinking Representations for World-Action Modeling

    Sep 29, 2026Haoyi Jiang, Liu Liu, Xinjiang Wang +12Efficient World-Action ModelPretraining

  18. Pretraining Latent Information Feedback Transformers with Teacher Supervision

    Sep 29, 2026Dor Tirosh, Ido Amos, Mor GevaTransformer ArchitecturesPretraining

  19. Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing

    Sep 29, 2026Guannan Lai, Han-Jia YeLarge Language Model RoutingLLM Inference Optimization

  20. Emergent Specialization in Populations of Self-Supervised Collaborative Vision Experts Without a Shared Gate or Cross-Agent Gradients

    Sep 29, 2026Aram Davtyan, Pablo Acuaviva, Sebastian Stapf +1Decentralized LearningPretraining

  21. FLOORA: A Human-Aligned Domain-Specific Language Model for Architectural Design

    Sep 28, 2026Sahand Rezaei-Shoshtari, Patryk Wozniczka, Shu Ishida +14Floor Plan GenerationLayouts

  22. Temporal Heterogeneous Graph Pretraining for Relational Deep Learning

    Sep 28, 2026Yixin Peng, Er Jin, Diego Collarana +1Relational Deep LearningTemporal Graph Neural Networks

  23. Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks

    Sep 28, 2026Alexandre Declèves, Etienne Boursier, Nicolas FlammarionPretrainingModel Fine-Tuning

  24. Two-Timescale Fine-tuning Provably Learns New Features for Two-Layer ReLU Networks

    Sep 28, 2026Etienne Boursier, Nicolas FlammarionFeature LearningModel Fine-Tuning

  25. SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis

    Sep 28, 2026Hangyul Yoon, Hyungyung Lee, Edward Choi +1Chest X-RayRecent Vision-Language Models

  26. Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training

    Sep 28, 2026Junlin Chen, Daize Dong, Huanwei Di +9Block Sparse Flash AttentionMatched Fp16 Intermediate

  27. What masking geometry works best for EEG foundation models?

    Sep 27, 2026Pierre Guetschel, Bruno Aristimunha, Yassine El Ouahidi +3Electroencephalography Foundation ModelsPretraining

  28. Self-Play Pretraining with Zero Data

    Sep 24, 2026Aditya Cowsik, Kfir Dolev, Michael Y. Li +4PretrainingSelf-Play

  29. Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions

    Sep 23, 2026Michael Lawrence Castanares, Princess Ventures, Allan TanComputerized Adaptive TestingPedagogical Frameworks

  30. A generalizable structural brain MRI foundation model built through dual-priority federated pretraining

    Sep 23, 2026Zhen Yu, Yang Liu, Xiahai Zhuang +1Functional Magnetic Resonance ImagingWireless Foundation Models

  31. A Scaling Study for fMRI Foundation Models

    Sep 23, 2026Wenhao Ye, Xuanye Pan, Junfeng Xia +3Model SizeFunctional Magnetic Resonance Imaging

  32. MMAP: Multimodal Missing-Aware Pretraining for Longitudinal Alzheimer's Prediction

    Sep 22, 2026Fiona Kekwick, Matthew Baugh, Bernhard Kainz +2Multimodal Clinical DataMultimodal Pretraining

  33. Foundation model embeddings capture pre-diagnostic changes on screening mammograms

    Sep 22, 2026Kalina P. Slavkova, Eric Brattain, Aditya Gowd +6MammographyMultidisciplinary Tumor Boards

  34. Learning to Fluctuate: Statistical Foundations for Causal Tabular Pretraining

    Sep 22, 2026Zhiheng ZhangCausalPretraining

  35. Signed Graph Pre-Training and Prompt Learning

    Sep 22, 2026Zihan Mei, Rong Pan, Yuzhou Chen +1Graph Representation LearningPersistent Homology

  36. A Deployment Study of Identity-Gated Drone Gesture Control

    Sep 22, 2026Diyari Mohammed Salih, Ilyes Chaabeni, Naima Ait OufroukhAutonomous DronesAerial Navigation

  37. PermuFormer: Multi-Task Pretraining for Permutation Representation in Algebraic Combinatorics

    Sep 21, 2026Henry KvingePretrainingPermutation

  38. Toward a foundation model for forest point clouds

    Sep 21, 2026Yuanwen Yue, Stefano Puliti, Damien Robert +7ForestsForest Biomass

  39. Video-STLayout Pre-training

    Sep 21, 2026Akash Abdu Jyothi, Greg MoriFine-Grained Video UnderstandingVision Encoders

  40. Time-Incremental Continued Pretraining of LLMs: Knowledge Updates Without Catastrophic Forgetting

    Sep 20, 2026Fırat Öncel, Salman Hussain Ali, Mirco Ravanelli +2Replay-Based Continual LearningCatastrophic Forgetting

  41. QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability Synchronization

    Sep 17, 2026Yujie Li, Zezhi Shao, Chengqing Yu +7Zero-Shot ForecastingPretraining

  42. Procedural Pretraining for Molecular Property Prediction

    Sep 15, 2026Moritz Friedemann, Zachary Shinnick, Philip Torr +1Molecular Property PredictionPretraining

  43. Measuring Annotation Efficiency for Handwritten Devanagari Recognition: Sample-Complexity Curves for Four Pretraining Regimes

    Sep 15, 2026Manglesh Kumar Pandey, Sumit Kumar BanshalHandwritten Text RecognitionPretraining

  44. Which Pretext Task Transfers? Self-Supervised Pretraining Objectives for Lung Ultrasound

    Sep 15, 2026Moein Heidari, Junbo Rao, Jai Choraria +3UltrasoundSelf-Supervised Learning

  45. Compute-Optimal Pretrain--Fine-tune in Ridge Gradient Descent

    Sep 14, 2026Alex Buna, Fanghui Liu, Patrick RebeschiniPretrainingModel Fine-Tuning

  46. Parameter-Efficient Adaptation of Pretrained Language Models for Time-Series Forecasting

    Sep 14, 2026Tamanna Kumavat, Georg Brunner, Kyriakos FlourisForecasting BackboneTime Series Forecasting

  47. BRIDGE-EEG: Bridging Self-Supervised Pretraining and Efficient Deployment for Cross-Dataset EEG Classification

    Sep 14, 2026Meghna Roy Chowdhury, Chengwei Zhou, Haotian Yu +2Electroencephalography Foundation ModelsElectroencephalography

  48. Self-supervised Pre-training Helps Retinal Disease Progression Modelling Most When Data Is Scarce

    Sep 14, 2026Ifeoma Veronica Nwabufo, Julius Gervelmeyer, Sarah Müller +1Self-Supervised LearningLongitudinal Imaging

  49. Pre-Trained Low-Rank Tensor Decomposition for Multi-Dimensional Image Recovery

    Sep 14, 2026Bing-Zhang Fu, Zhi-Long Han, Ting-Zhu Huang +2Tensor DecompositionLow-Rank Structure

  50. Towards a knowledge-enhanced single-cell foundation model

    Sep 14, 2026Hanqing Zhang, Jie Bao, Mei Ma +7Regulatory NetworksPretraining

  51. Real-World Reinforcement Learning with MPC Scaffolding for Dexterous Manipulation

    Sep 14, 2026Emek Barış Küçüktabak, Karankumar Patel, Zhaodong Yang +3Reinforcement Learning ControlModel Predictive Control