Instruction Following

Momentum

11 papers in the last four weeks, down 8% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 85

All topics
CardsList
  1. Task Competence Is Not Instruction Following: Evaluating Instruction-Conflicting Behavior in Small Language Models

    Jul 21, 2026Mahdiyeh Farajidizaji, Vatsal RainaLLM EvaluationInstruction Following

  2. Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents

    Jul 13, 2026Chenglin Yu, Li Yin, Ying Yu +4Instruction FollowingRuntime Enforcement for AI Agents

  3. Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

    Jul 2, 2026Congrui Du, Yang Zhang, Kaizhi Qian +1Language Model PretrainingInstruction Following

  4. CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes

    Jun 30, 2026Yuchen Huang, Xiang Li, Zhenqing Ling +5LLM EvaluationInstruction Following

  5. Indi-RomCoM: Code-Mixed Benchmark for Evaluating LLMs on Romanized Indic-English Instructions

    Jun 29, 2026Avisha Das, Mihir Parmar, Mohana Ramnath +1Multilingual Language Model EvaluationInstruction Following

  6. IHDec: Divergence-Steered Contrastive Decoding for Securing Multi-Turn Instruction Hierarchies

    Jun 29, 2026Nicole Geumheon Liu, Haeun Jang, Yonghyun Jun +1Instruction FollowingPrompt Injection Defense

  7. Masked Language Flow Models

    Jun 26, 2026Iskander Azangulov, Kianoosh Ashouritaklimi, Leo Zhang +2Masked Diffusion ModelsInstruction Following

  8. FBK's Long-form SpeechLLMs for IWSLT 2026 Instruction Following

    Jun 25, 2026Zhihang Xie, Marco Gaido, Sara Papi +2Audio-Language Model EvaluationSpeech Processing

  9. A Framework for Evaluating Agentic Skills at Scale

    Jun 16, 2026Maksim Shaposhnikov, Nicolas Fortuin, Simon Stipcich +3LLM Agent EvaluationInstruction Following

  10. Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation

    Jun 10, 2026Selen Erkan, Bastian Boll, Kristian Kersting +2LLM EvaluationInstruction Following

  11. ComplexConstraints and Beyond: Expert Rubrics for RLVR

    Jun 8, 2026Sushant Mehta, Liudas Panavas, Suhaas Garre +1LLM-as-a-JudgeRubric-Based RL

  12. RECAP: Regression Evaluation for Continual Adaptation of Prompts

    Jun 4, 2026Harsh Deshpande, Kushal Chawla, Sangwoo Cho +2Continual Learning for LLMsLLM Evaluation

  13. MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following

    Jun 4, 2026Mohammad Mahdi Salmani-Zarchi, Zahra Rahimi, Heshaam Faili +1Group Relative Policy OptimizationInstruction Following

  14. Multilingual Long-Form Speech Instruction Following: KIT's Submission to IWSLT 2026

    Jun 3, 2026Enes Yavuz Ugan, Maike Züfle, Yuka Ko +5MBR DecodingInstruction Following

  15. Bridging Auxiliary Constraints to Resolve Instruction Following in Large Reasoning Models

    Jun 2, 2026Zhengyi Zhao, Shubo Zhang, Huimin Wang +7LLM PromptingInstruction Following

  16. AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following

    Jun 2, 2026Haitao Li, Tian Tan, Yuguang Yang +2Audio-Language Model EvaluationLLM-as-a-Judge

  17. Label-Free Reinforcement Learning via Cross-Model Entropy

    May 27, 2026Matt Gorbett, Hossein ShiraziReinforcement LearningInstruction Following

  18. Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards

    May 27, 2026Saurabh Dash, Pierre Clavier, John Dang +4Reinforcement LearningInstruction Following

  19. The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates

    May 27, 2026Shaobo Wang, Guo Chen, Ziyue Wang +5LLM EvaluationLanguage Model Decoding

  20. Disentangling Language Roles in Multilingual LLM Task Execution

    May 26, 2026Qishi Zhan, Minxuan Hu, Seoyeon Jang +7Multilingual Language ModelsMultilingual Language Model Evaluation

  21. MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation

    May 26, 2026Haoyu Zheng, Yun Zhu, Shu Yuan +5Instruction FollowingData Contamination in Language Models

  22. ContextGuard: Structured Self-Auditing for Context Learning in Language Models

    May 26, 2026Hongbo Jin, Chi Wang, Haoran Tang +5LLM AuditingInstruction Following

  23. Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

    May 19, 2026Carolina Camassa, Derek ShillerLLM EvaluationLanguage Model Self-Assessment