Instruction Following

Momentum

11 papers in the last four weeks, up 10% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 83

All topics
CardsList
  1. When Verifiable Counts Depend on Wording: Auditing Wording Robustness in Instruction Following

    Oct 4, 2026Qishi Zhan, Seoyeon Jang, Zihan Dong +3LLM EvaluationInstruction Following

  2. Beyond Leaderboards: Tokenomics of Agentic Small Language Model Ensembles

    Oct 1, 2026Alexei N. Skurikhin, Emily M. Taylor, Nathan A. DeBardelebenLLM Inference EfficiencyLanguage Model Generation Evaluation

  3. Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness

    Sep 30, 2026Pardis Sadat Zahraei, Janvijay Singh, Gokhan Tur +1LLM AlignmentInstruction Following

  4. Coverage Before Control: Route-Instruction Grounding and Steering for Controllable Retrosynthesis

    Sep 30, 2026Xuemin Chen, Xiaozhuang Song, Xinjian Zhao +2Instruction FollowingInstruction Tuning

  5. A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

    Sep 30, 2026Seonho Lee, Wonryeol Jeong, Alberto Cereser +4AI Coding AgentsAI Agent Benchmarks

  6. IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

    Sep 12, 2026Kainan Zhou, Zhaoyi Li, Janet Sung +2LLM EvaluationBenchmark Design

  7. In-Place Instruction Following in Diffusion Language Models

    Sep 7, 2026Zheng Nie, Zherui Li, Jiaming Zhang +3Instruction FollowingDiffusion Language Models

  8. ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults

    Sep 7, 2026Weide Zhan, Qumu Shaqu, Yuanqing Liu +6Mobile GUI AutomationComputer-Use Agent Benchmarks

  9. Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning

    Sep 1, 2026Jinhu Qi, Minda Hu, Wentao Zhang +4Instruction FollowingIn-Context Learning

  10. Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues

    Aug 12, 2026Haoyuan ZhuInstruction FollowingLLM Reliability

  11. Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

    Aug 12, 2026Zining Huang, Haoran Que, Hong Zeng +8AI Coding AgentsLLM Agent Evaluation

  12. LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs

    Aug 9, 2026Minhan Cho, Soyoung Park, Kihyeon Jeong +3LLM Inference EfficiencyInstruction Following

  13. Task Competence Is Not Instruction Following: Evaluating Instruction-Conflicting Behavior in Small Language Models

    Jul 21, 2026Mahdiyeh Farajidizaji, Vatsal RainaLLM EvaluationInstruction Following

  14. Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents

    Jul 13, 2026Chenglin Yu, Li Yin, Ying Yu +4Instruction FollowingRuntime Enforcement for AI Agents

  15. Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

    Jul 2, 2026Congrui Du, Yang Zhang, Kaizhi Qian +1Language Model PretrainingInstruction Following

  16. CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes

    Jun 30, 2026Yuchen Huang, Xiang Li, Zhenqing Ling +5LLM EvaluationInstruction Following

  17. Indi-RomCoM: Code-Mixed Benchmark for Evaluating LLMs on Romanized Indic-English Instructions

    Jun 29, 2026Avisha Das, Mihir Parmar, Mohana Ramnath +1Multilingual Language Model EvaluationInstruction Following

  18. IHDec: Divergence-Steered Contrastive Decoding for Securing Multi-Turn Instruction Hierarchies

    Jun 29, 2026Nicole Geumheon Liu, Haeun Jang, Yonghyun Jun +1Instruction FollowingPrompt Injection Defense

  19. Masked Language Flow Models

    Jun 26, 2026Iskander Azangulov, Kianoosh Ashouritaklimi, Leo Zhang +2Masked Diffusion ModelsInstruction Following

  20. FBK's Long-form SpeechLLMs for IWSLT 2026 Instruction Following

    Jun 25, 2026Zhihang Xie, Marco Gaido, Sara Papi +2Audio-Language Model EvaluationSpeech Processing

  21. A Framework for Evaluating Agentic Skills at Scale

    Jun 16, 2026Maksim Shaposhnikov, Nicolas Fortuin, Simon Stipcich +3LLM Agent EvaluationInstruction Following

  22. Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation

    Jun 10, 2026Selen Erkan, Bastian Boll, Kristian Kersting +2LLM EvaluationInstruction Following

  23. ComplexConstraints and Beyond: Expert Rubrics for RLVR

    Jun 8, 2026Sushant Mehta, Liudas Panavas, Suhaas Garre +1LLM-as-a-JudgeRubric-Based RL

  24. RECAP: Regression Evaluation for Continual Adaptation of Prompts

    Jun 4, 2026Harsh Deshpande, Kushal Chawla, Sangwoo Cho +2Continual Learning for LLMsLLM Evaluation

  25. MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following

    Jun 4, 2026Mohammad Mahdi Salmani-Zarchi, Zahra Rahimi, Heshaam Faili +1Group Relative Policy OptimizationInstruction Following

  26. Multilingual Long-Form Speech Instruction Following: KIT's Submission to IWSLT 2026

    Jun 3, 2026Enes Yavuz Ugan, Maike Züfle, Yuka Ko +5MBR DecodingInstruction Following

  27. Bridging Auxiliary Constraints to Resolve Instruction Following in Large Reasoning Models

    Jun 2, 2026Zhengyi Zhao, Shubo Zhang, Huimin Wang +7LLM PromptingInstruction Following

  28. AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following

    Jun 2, 2026Haitao Li, Tian Tan, Yuguang Yang +2Audio-Language Model EvaluationLLM-as-a-Judge

  29. Label-Free Reinforcement Learning via Cross-Model Entropy

    May 27, 2026Matt Gorbett, Hossein ShiraziReinforcement LearningInstruction Following

  30. Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards

    May 27, 2026Saurabh Dash, Pierre Clavier, John Dang +4Reinforcement LearningInstruction Following

  31. The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates

    May 27, 2026Shaobo Wang, Guo Chen, Ziyue Wang +5LLM EvaluationLanguage Model Decoding

  32. Disentangling Language Roles in Multilingual LLM Task Execution

    May 26, 2026Qishi Zhan, Minxuan Hu, Seoyeon Jang +7Multilingual Language ModelsMultilingual Language Model Evaluation

  33. MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation

    May 26, 2026Haoyu Zheng, Yun Zhu, Shu Yuan +5Instruction FollowingData Contamination in Language Models

  34. ContextGuard: Structured Self-Auditing for Context Learning in Language Models

    May 26, 2026Hongbo Jin, Chi Wang, Haoran Tang +5LLM AuditingInstruction Following

  35. Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

    May 19, 2026Carolina Camassa, Derek ShillerLLM EvaluationLanguage Model Self-Assessment

  36. Dimension-Level Intent Fidelity Evaluation for Large Language Models: Evidence from Structured Prompt Ablation

    May 14, 2026GAng PengMultilingual Language Model EvaluationLLM Evaluation

  37. When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction

    May 13, 2026Vardhan Dongre, Joseph Hsieh, Viet Dac Lai +3Self-AttentionLLM Interpretability