Conditional Computation

Momentum

5 papers in the last four weeks, with none the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 21

All topics
CardsList
  1. QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing

    Oct 7, 2026Yueying Li, Zhongle Xie, Ke Chen +1Efficient Transformer InferenceConditional Computation

  2. Overcoming Kernel Redundancy for Scaling Logic Gate Networks

    Oct 1, 2026Sejin Park, Hongjae Lee, Changwoo Han +1Neural Scaling LawsConditional Computation

  3. Shared Weights, Selected Computations: How Looped Transformers Route What Each Loop Does

    Sep 30, 2026Jiaju Wu, Yi Hu, Muhan ZhangTransformer InterpretabilityTransformer Attention

  4. X-MoD: Practical Scaling Laws for Sparse-Depth Routing Beyond Mixture-of-Depths

    Sep 28, 2026Bowen Dong, Yilong Fan, Tengyu Pan +6Language Model Scaling LawsMixture-of-Experts Models

  5. MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup

    Sep 14, 2026Muchen Li, Leonid Sigal, Renjie LiaoMemory-Augmented Language ModelsPersistent Memory for Language Models

  6. An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS

    Sep 12, 2026Roberto Campbell, Momin Abbass, Muneeza Azmat +5LLM Safety AlignmentLow-Rank Adaptation

  7. PromptPath: Prompt-Adaptive Computational Pathways for In-Context Learning

    Aug 3, 2026Hangrui Zhang, Feifei Shao, Yawei Luo +6Mixture of ExpertsIn-Context Learning

  8. SecondOpinion: Anatomy-Aware Gated Reasoning for Efficient Medical Image Analysis

    Aug 3, 2026Siam Tahsin Bhuiyan, Rashedur Rahman, Sefatul Wasi +4Chest X-Ray ClassificationMedical Image Analysis

  9. Sparse by Command: Task-Conditional Compute Skipping for Multi-Task Inference Accelerators

    Jul 24, 2026Afzal Ahmad, Gaoyu Mao, Shoubo Hu +4Structured SparsityFPGA-Based Neural Network Acceleration

  10. SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs

    Jul 20, 2026Huzaifa Shaaban Kabakibo, Eric Schniedermeyer, Artem Burchanow +1On-Device Language Model InferenceMemory-Efficient Inference

  11. Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning

    Jul 12, 2026Sudipto Ghosh, Tanmoy ChakrabortyGroup Relative Policy OptimizationMulti-Agent Reasoning

  12. TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation

    Jul 7, 2026Andrii Balashov, Olena PonomarovaSparse Mixture-of-ExpertsLLM Routing

  13. End-to-End Dynamic Sparsity for Resource-Adaptive LLM Inference

    Jun 26, 2026Yuhang Chen, Jinhao Duan, Ruichen Zhang +11LLM InferenceStructured Sparsity

  14. Sigma-Branch: Hierarchical Single-Path Network Reconstruction for Dynamic Inference with Reduced Active Parameters

    Jun 7, 2026Kohga Tanaka, Hiroaki NishiEfficient Neural Network InferenceDynamic Neural Networks

  15. Linear-Time Global Visual Modeling without Explicit Attention

    May 3, 2026Ruize He, Dongchen Han, Gao HuangTransformerTransformer Attention

  16. evMLP: An Efficient Event-Driven MLP Architecture for Vision

    Jul 2, 2025Zhentan ZhengEfficient Neural Network InferenceConditional Computation

  17. Tracing Computation Density in LLMs

    Date pendingCorentin Kervadec, Iuliia Lysova, Iuri Macocco +2Transformer InterpretabilityLLM Interpretability