Fp8

Momentum

1 paper in the last four weeks, against 1 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 21

All topics
CardsList
  1. Concurrent Split Learning Through Stable Client Clustering

    Sep 24, 2026Mohammad Kohankhaki, Valentin Rentschler, Anke SchmeinkClientsFundamental Limits

  2. FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling

    Aug 3, 2026Xianglong Yan, Hong Liu, Chengzhu Bao +4Large Language Model QuantizationFp8

  3. Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention

    Jul 5, 2026Siyu Ding, Mingchuan Ma, Jiabo Tong +3Fp8Large Language Model Pretraining

  4. W4A4 Quantization for Inference on Wan2.2-I2V-A14B

    Jun 28, 2026Yidong Chen, Chengyu Shi, Jiahao LiuLarge Language Model QuantizationFp8

  5. Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe

    Jun 18, 2026Qian Zhao, Kunlong Chen, Changxin Tian +9Fp8Shrink

  6. ReQAT: Achieving Full-Precision Reasoning Accuracy with 4-bit Floating-Point Quantization-Aware Training

    Jun 14, 2026Janghwan Lee, Sihwa Lee, Jinseok Kim +4Quantization-Aware TrainingFp8

  7. ReSET: Accurate Latency-Critical NVFP4 Reasoning via Step-Aware Temperature Scaling

    Jun 11, 2026Sihwa Lee, Janghwan Lee, Donghoon Yoo +4Fp8Autoregressive Decoding

  8. Characterizing the Impact of NVFP4 Quantization for Low-Power Edge AI Deployment

    Jun 3, 2026Ovishake Sen, Venkata Nithin Kamineni, Daniel Lobo +3Fp8Neural Network Inference

  9. ScaleSweep: Accurate NVFP4 Post-Training Quantization of LLMs via Block Scale Initialization

    May 30, 2026Li Lin, Xiaojun WanFp8Full-Precision Performance

  10. Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2

    May 26, 2026Zhanfeng Feng, Shuai Guo, Xin Di +3Post-Training QuantizationFp8

  11. Grid Games: The Power of Multiple Grids for Quantizing Large Language Models

    May 12, 2026Vage Egiazarian, Erik Schultheis, Andrei Panferov +3Large Language Model QuantizationFp8

  12. Pretraining large language models with MXFP4 on Native FP4 Hardware

    May 11, 2026Musa Cim, Poovaiah Palangappa, Miro Hodak +3Large Language Model QuantizationFp8

  13. FP8 is All You Need (Part 1): Debunking Hardware FP64 as the HPC Holy Grail (Sep 3rd version)

    Date pendingSatoshi MatsuokaHigh-Performance ComputingFp8

  14. FP8 is All You Need (Part 2): Full-FP64 3-D FFT on FP8-Generation Tensor CoresThe Integer-Epilogue Wall and the Minimal Hardware That Would Remove It

    Date pendingSatoshi MatsuokaFp8Nvidia