AI Coding Agents

Momentum

28 papers in the last four weeks, up 115% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 212

All topics
CardsList
  1. BrickBench: Evaluating Agentic Brick Design

    Oct 8, 2026Peter Kulits, Yiqing Xu, R. Kenny Jones +2AI Agent BenchmarksAI Coding Agents

  2. Forms of LLM-Integrated Applications from LLM-Chats to Autonomous AI Agent System

    Oct 8, 2026Irene WeberLLM AgentsLLM Agent Orchestration

  3. Closed-loop evaluation of LLM agents for embedded software development

    Oct 8, 2026Jorge García-Carrasco, Sergio García-Carrasco, Alejandro Maté +1AI Coding AgentsLLM Agent Evaluation

  4. Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

    Oct 7, 2026Tan Yu, Alexander Bukharin, Khushi Bhardwaj +19AI Coding AgentsAI Agent Evaluation

  5. ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding

    Oct 6, 2026Hanjun Luo, Xiucheng Zhang, Zhuoning Xu +4AI Coding AgentsAI Risk Management

  6. A Case Study in Assuring AI-Written Software

    Oct 6, 2026Lindsey Ferris, Sierra BonillaAI Agent ReliabilityAI Assurance

  7. Small Agents with Semantic Search: Efficient Multilingual Code Localization

    Oct 4, 2026Maxence Lasbordes, Aarush Sinha, Raphael Sourty +2AI Coding AgentsSemantic Search

  8. AuraForge: Scaling Security Supervision for Training Coding Agents

    Oct 1, 2026Danqing Wang, Songwen Zhao, Harsh Sharma +4AI Coding AgentsSoftware Security

  9. Incident-Arena: Getting agents to the last nine of reliability

    Sep 30, 2026Andre Fu, Malik Drabla, Leon Liu +5Agent ReliabilityAI Coding Agents

  10. How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

    Sep 30, 2026Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel AbbéAI Coding AgentsLLM Agent Evaluation

  11. A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

    Sep 30, 2026Seonho Lee, Wonryeol Jeong, Alberto Cereser +4AI Coding AgentsAI Agent Benchmarks

  12. LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

    Sep 29, 2026Yun Peng, Zihan Wu, Zeyang Zhuang +6Software EngineeringAI Coding Agents

  13. Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes

    Sep 29, 2026Xiaokang Ye, Siddhant Hitesh Mantri, Zimeng Chen +6AI Coding Agents3D Scene Editing

  14. Can Agents Design Libraries for Agents?

    Sep 29, 2026Gabriel Orlanski, Alex L. Zhang, Avi Trost +4Software Engineering AgentsAI Coding Agents

  15. StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

    Sep 28, 2026Ziyang Yu, Liang Zhao, Bowen Zhu +1AI Coding AgentsLLM Agent Memory

  16. Opera: A Verbal Critic Framework for Long-horizon Coding Agents

    Sep 27, 2026Kai Mei, Zhiyuan Hu, Yutong Dai +7AI Coding AgentsAI Agent Auditing

  17. Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models

    Sep 27, 2026Eduardo Ariño de la Rubia, Szilard PafkaAI Coding AgentsLLM Agent Evaluation

  18. SWE-Game: Can Coding Agents Build the Games We Want?

    Sep 27, 2026Xiaoyu Chen, Lai Wei, Jin Wang +8Software Engineering AgentsAI Coding Agents

  19. Evaluating Coding Agents on Kernel Exploit Generation

    Sep 22, 2026Junyoung Jang, Gwanhyun Lee, Hwiwon Lee +4AI Coding AgentsSoftware Security

  20. Quantifying Overclaiming Propensity in Frontier LLM Agents

    Sep 17, 2026Nolan Smyth, Yorguin-Jose Mantilla-Ramos, Pascal Jr Tikeng Notsawo +6AI Agent ReliabilityAI Coding Agents

  21. The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

    Sep 17, 2026Zhexi Feng, Ruiyi Zhang, Yongbo Yang +1AI Coding AgentsAgentic Retrieval