Coding Agents

Momentum

65 papers in the last four weeks, up 195% on the four weeks before. 0.6% of all new papers.

Jul 13Week of Sep 28

Latest papers 336

All topics
CardsList
  1. AECP: Artifact-Exclusive Communication Protocol for Multi-Agent Code Generation

    Oct 5, 2026Jiaqi Xue, Yanjun Wang, Xiangci Li +5Coding Agents

  2. Correct Code, Broken Contributions? SWE-CC: Benchmarking Repository Policy Compliance for Coding Agents

    Oct 5, 2026Hai Dang Truong, Rayner Goh, Thanh Le-Cong +1Code QualityCoding Agents

  3. AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

    Oct 1, 2026Xuan Zhang, Longtao Zheng, Cunxiao Du +2Coding AgentsCompaction

  4. Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents

    Oct 1, 2026Junyu Guo, Shangding Gu, Ming Jin +1Code QualityReviewer

  5. Cross-Benchmark Transfer from RL on Agentic Coding Tasks

    Oct 1, 2026Sushant Mehta, Logan Ritchie, Edwin ChenCoding AgentsAgentic Benchmarks

  6. AuraForge: Scaling Security Supervision for Training Coding Agents

    Oct 1, 2026Danqing Wang, Songwen Zhao, Harsh Sharma +4Coding AgentsForge

  7. Self-Evolving Coding Rules for AI Coding Agents

    Sep 30, 2026Zhengyuan Jiang, Reachal Wang, Yuepeng Hu +3Coding AgentsSelf-Evolving Agents

  8. Incident-Arena: Getting agents to the last nine of reliability

    Sep 30, 2026Andre Fu, Malik Drabla, Leon Liu +5Incident ResponseCoding Agents

  9. How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

    Sep 30, 2026Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel AbbéAgent HarnessCoding Agents

  10. Code That Works, Environments That Don't: Measuring Environment Reproducibility in AI-Generated Software

    Sep 30, 2026Bhanu Prakash Vangala, Tanu MalikCoding AgentsCode Generation

  11. Learning When and How to Intervene: A Hindsight-Distilled Sentinel for Coding Agents

    Sep 30, 2026Jiangrui Zhao, Chenglong Li, Meng Zhang +1Coding AgentsHindsight

  12. EngramBench: A Capability-Grounded Benchmark for Skill-Evolution Harnesses

    Sep 30, 2026Zhixuan Tan, Pengjie Gu, Zhao Li +4Coding AgentsAgentic Benchmarks

  13. Coding Agents for Coding Theory

    Sep 30, 2026Abraham YeungCoding AgentsSource-Channel Coding

  14. From Verification Failures to Reusable Guidance for Coding Agents

    Sep 30, 2026Yuqing Zhai, Xiaohong Chen, Lingming Zhang +2Coding AgentsAgentic Evaluations

  15. SimEX: Simulation-Integrated Robotics AutoResearch

    Sep 30, 2026Jiaheng Hu, Roberto Martin-Martin, Peter Stone +3Robotic SystemsMulti-Agent Simulations

  16. OpenCollab: A Multi-Agent Coding Framework with Programmable Collaboration and Controllable Runtime

    Sep 29, 2026Chun-Wah Hsu, Kai Gong, Yu Wu +12CollaborationCoding Agents

  17. Zero2Repo: Can Coding Agents Build Repositories from Scratch?

    Sep 29, 2026Pei Yang, Tianyu Shi, Yuhang Yao +23Coding AgentsSoftware

  18. LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

    Sep 29, 2026Yun Peng, Zihan Wu, Zeyang Zhuang +6CodebasesCoding Agents

  19. Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes

    Sep 29, 2026Xiaokang Ye, Siddhant Hitesh Mantri, Zimeng Chen +63D Scene Generation3D Editing

  20. WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses

    Sep 29, 2026Haomin Qi, Xiangzhe Xu, Yiming Huang +2BugCoding Agents

  21. LEGO-Anything: Coding Agents for 3D Scene Reconstruction

    Sep 28, 2026Xirui Li, Peng Shi, Mingwen Dong +73D ReconstructionScene Understanding

  22. Towards an AI Software Factory for Data Systems

    Sep 28, 2026Anna Pavlenko, Bogdan Crivat, Brandon Haynes +23Software EngineeringAi-Based

  23. StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

    Sep 28, 2026Ziyang Yu, Liang Zhao, Bowen Zhu +1Long-Horizon AgentsCoding Agents

  24. GPUPhysBench: Benchmarking Coding Agents for Correct and Efficient GPU Physics Simulation

    Sep 28, 2026Yuchen Sun, Jinjin He, Sinan Wang +1Physics SimulationCoding Agents

  25. Do Coding Agents Reuse Existing Code or Reinvent the Wheel?

    Sep 28, 2026Dongsheng Ma, Sizhe Wang, Xinyi Huang +5Coding AgentsReuse

  26. Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero

    Sep 28, 2026Zipei Yu, Yue-Jiao Gong, Zeyuan Ma +2Agentic OptimizationBlack-Box Optimization

  27. Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut

    Sep 28, 2026Gunin Gupta, Nirmit Arora, Pavan Kalyan TankalaVideo EditingVideo Agent

  28. RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers

    Sep 28, 2026Haitong Ma, Chenxiao Gao, Rushi Qiang +2Coding AgentsRobot Systems

  29. SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities

    Sep 27, 2026Xiaonan Luo, Yue Huang, Kehan Guo +9Coding AgentsItem Response Theory

  30. SWE-Game: Can Coding Agents Build the Games We Want?

    Sep 27, 2026Xiaoyu Chen, Lai Wei, Jin Wang +8Swe-Bench VerifiedCoding Agents

  31. Graph-Guided Repository Environment Construction

    Sep 27, 2026Jianying Pan, John Zhang, Hongyu ZhangSynthetic EnvironmentsCoding Agents

  32. Coding Agents for Generalized Task and Motion Planning Problems

    Sep 24, 2026Matteo Merler, Bowen Li, Josh Roy +4Classical PlanningCoding Agents

  33. Coding Agents Aren't Enough! Evaluating an Enterprise Security Brain for Agentic Cloud Investigations

    Sep 24, 2026Leon Goldberg, Gal Engelberg, Eden Yavin +3Security EvaluationSecurity

  34. Rufus-Air: An Open LLM Post-Training Recipe

    Sep 24, 2026Chia-Yuan Chang, Renyuan Cheng, Rui Feng +19Post-TrainingOffline Reinforcement Learning

  35. Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise

    Sep 24, 2026Arian Abbasi, Alan Aqrawi, Ted KwartlerAgent HarnessCoding Agents

  36. Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior

    Sep 23, 2026Chuyi Wang, Xiaohui Xie, Tongze Wang +2FingerprintCoding Agents

  37. Generalizing Manipulation Skills with a Local Coding Agent

    Sep 22, 2026Raman Talwar, Elias Nijs, Andreas Verleysen +1Scalable Robot LearningCoding Agents

  38. Evaluating Coding Agents on Kernel Exploit Generation

    Sep 22, 2026Junyoung Jang, Gwanhyun Lee, Hwiwon Lee +4ExploitationCoding Agents

  39. VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

    Sep 20, 2026Liyang Fan, Yingcheng Shi, Yongbin Li +7Repository-Level Code UnderstandingMemoryagentbench

  40. Quantifying Overclaiming Propensity in Frontier LLM Agents

    Sep 17, 2026Nolan Smyth, Yorguin-Jose Mantilla-Ramos, Pascal Jr Tikeng Notsawo +6Frontier Large Language Model AgentsCoding Agents

  41. An Empirical Study of Harness Design for Coding Agents

    Sep 17, 2026Run-Ze Fan, Zihao Zhang, Simin Ma +6Coding AgentsAgent Harness

  42. The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

    Sep 17, 2026Zhexi Feng, Ruiyi Zhang, Yongbo Yang +1Incomplete EvidenceCoding Agents

  43. Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents

    Sep 17, 2026Alex Remedios, Simon Storf, Fabien Roger +1Red-TeamingCoding Agents

  44. Self Improvement via Fast Tree-search

    Sep 17, 2026Xinghong Fu, Aravinth Kulanthaivelu, Yutaro YamadaRecursive Self-ImprovementTree Search

  45. Higher-order pruning of experts in mixture-of-experts language models

    Sep 16, 2026Alex M. Tseng, Prannay Kaul, Luca Zancato +2Mixture-Of-Experts Large Language ModelsMixture-Of-Experts

  46. ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

    Sep 16, 2026Jeonghye Kim, Minseon Kim, Young Jin Kim +5Coding AgentsSoftware Engineering

  47. WetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories

    Sep 16, 2026Yuna Oikawa, Kei Endo, Takanori Uzawa +5LaboratoryTeleoperation

  48. Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks

    Sep 15, 2026Franziska Roesner, Tadayoshi KohnoCoding AgentsPoisoning