Gumbel-Softmax Relaxation

Latest papers 58

All topics
CardsList
  1. Two-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion

    Oct 7, 2026Walid Bendada, Guillaume Salha-GalvanGumbel-Softmax RelaxationOptimal Sample Complexity

  2. Learning Decision-Stump Thresholds in Context: Dynamics of Softmax Attention

    Oct 5, 2026Hong Ha Le, Jackie Lok, Atsushi Nitanda +1Gumbel-Softmax RelaxationPretraining

  3. Underscoring the Problem: Why Softpick Fails at Initialization

    Oct 4, 2026Aryan Sood, Jaikaran Singh, Ishaan BansalGumbel-Softmax Relaxation

  4. Does a Shared Temperature Imply a Shared Angular Scale in Probabilistic Contrastive Learning?

    Sep 30, 2026Ningkang Peng, Qianfeng Yu, Jingyang Mao +4Contrastive LearningCosine Similarity

  5. When a Flatness Proxy Is Not a Function: Robustness Certificates and Training Interventions

    Sep 29, 2026Vicente Opazo, Jose Calatayud-Mateu, Cristobal Rojas +1Gumbel-Softmax Relaxation

  6. Dual-Stream Simultaneous Translation via 2D Grid Attention

    Sep 28, 2026Yu Pu, Wei-Qiang ZhangBidirectionalFast Inference

  7. Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks

    Sep 28, 2026Manya Singh, Arjun PakrashiAmbiguityMulti-Label Classification

  8. Pretraining Transformers with Quantized Softmax in Attention

    Sep 27, 2026Shangzhen Zhu, Muyan Hu, Tomasz KozlowskiTransformer AttentionGumbel-Softmax Relaxation

  9. Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration

    Sep 27, 2026Shangzhen Zhu, Muyan Hu, Tomasz KozlowskiGumbel-Softmax RelaxationTransformer Architectures

  10. Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

    Sep 18, 2026Richard Zhe WangGumbel-Softmax RelaxationGating

  11. Mixed-Integer Nonlinear Differentiable Predictive Control for Underground Pumped Hydro Energy Storage Systems

    Sep 16, 2026Honghui Zheng, Ján Boldocký, Yury Dvorkin +1Model Predictive ControlNeural Policies

  12. EFQ-Softmax: Exp-Free Quantization for Softmax

    Sep 9, 2026Haohui Han, Yuming Wan, Hongni Wang +4Gumbel-Softmax RelaxationBlock Sparse Flash Attention

  13. Soft-Argmax for the Projective Plane via the Veronese Embedding

    Sep 1, 2026Benjamin El-Zein, Dominik Eckert, Paul Zech +4ProjectionSpherical Latent Space

  14. Intrinsic Interaction Geometry Controls the Low-Rank Complexity of Softmax Attention

    Aug 28, 2026Yuhe Sui, Jianing Zhang, Yingzhi TangLow-Rank StructureSpectral Norm

  15. SoftWater: Class-Aware Rate Allocation for Softmax Quantization

    Aug 12, 2026Joao V. Cavalcanti, Ashia C. WilsonGumbel-Softmax RelaxationKullback-Leibler Divergence

  16. A Quantum Roadmap for Softmax Attention: Exact Born-Rule Analogs for Softmax Attention on the Probability Simplex

    Aug 11, 2026Eric A. F. Reinhardt, Adam J. HauserGumbel-Softmax RelaxationQubit

  17. Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training

    Aug 11, 2026Artyom Sabitov, Daniil Volkov, Alexey ZaytsevReal-World Content Recommendation ProblemBounded-Memory

  18. Fast LapSum: Exact Differentiable Top-kk at Million Scale

    Aug 7, 2026Jakub Antczak, Joanna Wojciechowicz, Kamil Książek +3Top-KDecode Speedup

  19. Matrix Zonotopic Attention: A Context-Adaptive Value Projection for Set Transformers

    Aug 5, 2026Zhen Zhang, Amr AlanwarMulti-Head AttentionTransformer Architectures

  20. Stochastic Sequential Search in Very-High-Dimensional Feature Selection

    Aug 2, 2026Petr Somol, Jiří GrimFeature SelectionHigh-Dimensional

  21. Latent Softmax for Data-Efficient Phoneme-Based Multilingual ASR Across Tonal and Non-Tonal Languages

    Aug 2, 2026Saierdaer Yusuyin, Nanling Jiang, Hao Huang +1Multilingual Automatic Speech RecognitionGrapheme-To-Phoneme

  22. Generalised Balanced Softmax: A Finite-Data Perspective on Logit Adjustment for Long-Tailed Recognition

    Jul 24, 2026Yi-Hang Zhu, Rajeev Raman, Shiqi Su +4Imbalanced ClassificationGumbel-Softmax Relaxation

  23. Temperature Scaling Is Not Enough: Calibration Gaps Under Human Label Distributions

    Jul 15, 2026Wisdom DogahGumbel-Softmax Relaxation

  24. A Strong Balanced-Softmax Classifier-Retraining Baseline for Long-Tailed Recognition

    Jul 10, 2026Juan Terven, Diana Margarita Córdova Esparza, Julio Alejandro Romero Gonzalez +4Few-Shot LearningImagenet

  25. Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection

    Jul 6, 2026Víctor Yeste, Paolo RossoMulti-Label ClassificationIntermediate Decoder Layers

  26. Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains

    Jun 9, 2026Zilai Wang, Natarajan Balaji Shankar, Mohan Shi +2Speech EncoderDomain Adaptation

  27. Federated Variational Preference Alignment with Gumbel-Softmax Prior for Personalized User Preferences

    May 29, 2026Jabin Koo, Hoyoung Kim, Minwoo Jang +1Preference LearningPreference Alignment Learning

  28. A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification

    May 21, 2026Marcel Kühn, Yoon Thelge, Bernd RosenowBatchGumbel-Softmax Relaxation

  29. Reading Calibrated Uncertainty from Language Model Trajectories

    May 19, 2026Aliai Eusebi, Alexander Herzog, Xiaoyu Liang +3Large Language Model UncertaintyCalibrated Uncertainty

  30. The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought

    May 18, 2026Moritz Brösamle, Stephan EcksteinTransformer ArchitecturesGumbel-Softmax Relaxation

  31. Fitting Multilinear Polynomials for Logic Gate Networks

    May 9, 2026Youngsung KimDifferentiable Logic Gate NetworksPolynomials

  32. Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization

    May 8, 2026Navid Rezazadeh, Arash Gholami DavoodiGumbel-Softmax RelaxationTransformer Attention

  33. Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning

    May 8, 2026Zixuan Xie, Xinyu Liu, Claire Chen +3In-Context LearningLinear Attention

  34. Streaming Adversarial Robustness in Fuzzy ARTMAP: Mechanism-Aligned Evaluation, Progressive Training, and Interpretable Diagnostics

    May 7, 2026Shane Cairns, Leonardo Enzo Brito da Silva, Sasha Petrenko +2Adversarial TrainingGumbel-Softmax Relaxation

  35. Linearizing Vision Transformer with Test-Time Training

    May 4, 2026Yining Li, Dongchen Han, Zeyu Liu +3Self-Supervised Vision TransformersVision Transformer

  36. Stochastic Sparse Attention for Memory-Bound Inference

    May 3, 2026Kyle Lee, Corentin Delacour, Kevin Callahan-Coray +5Dynamic Sparse AttentionAutoregressive Decoding

  37. Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models

    Apr 29, 2026Vijay Sadashivaiah, Georgios Dasoulas, Judith Mueller +1Wireless Foundation ModelsGumbel-Softmax Relaxation

  38. Transformer Approximations from ReLUs

    Apr 27, 2026Jerry Yao-Chieh Hu, Mingcheng Lu, Yi-Chen Lee +1Gumbel-Softmax RelaxationRectified Linear Unit

  39. ELSA: Exact Linear-Scan Attention for Fast and Memory-Light Vision Transformers

    Apr 26, 2026Chih-Chung Hsu, Xin-Di Ma, Wo-Ting Liao +1Kimi Delta AttentionGumbel-Softmax Relaxation

  40. Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices

    Apr 26, 2026Dawon Choi, Hana Kim, Ji-Hoon KimBatch NormalizationGumbel-Softmax Relaxation

  41. GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling

    Apr 20, 2026Alireza Dadgarnia, Soroush Tabesh, Mahdi Nikdan +4Large Language Model QuantizationPost-Training Quantization

  42. TabICLv2: A better, faster, scalable, and open tabular foundation model

    Feb 11, 2026Jingang Qu, David Holzmüller, Gaël Varoquaux +1Tabular Foundation ModelsPretraining

  43. Entropy-Generated Attention Beyond Softmax and Entmax: Kaniadakis and Reciprocal-Symmetric Abe Operators

    Feb 9, 2026Gunn KimEntropyGumbel-Softmax Relaxation

  44. FlexAct: Why Learn when you can Pick?

    Jan 10, 2026Ramnath Kumar, Kyle Ritscher, Junmin Judy +2Activation FunctionsGumbel-Softmax Relaxation

  45. The EM-algorithm and the Method of Moments in Softmax Mixture Models

    Sep 16, 2024Xin Bing, Florentina Bunea, Jonathan Niles-Weed +1Mixture ModelsExpectation-Maximization

  46. Smoothing the Score Function to Enhance Generalization in Diffusion Models

    Date pendingXinyu Zhou, Jiawei Zhang, Stephen J. WrightScore-Based Diffusion ModelGumbel-Softmax Relaxation