AI Coding Agents

Momentum

28 papers in the last four weeks, up 115% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 212

All topics
CardsList
  1. BrickBench: Evaluating Agentic Brick Design

    Oct 8, 2026Peter Kulits, Yiqing Xu, R. Kenny Jones +2AI Agent BenchmarksAI Coding Agents

  2. Closed-loop evaluation of LLM agents for embedded software development

    Oct 8, 2026Jorge García-Carrasco, Sergio García-Carrasco, Alejandro Maté +1AI Coding AgentsLLM Agent Evaluation

  3. Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

    Oct 7, 2026Tan Yu, Alexander Bukharin, Khushi Bhardwaj +19AI Coding AgentsAI Agent Evaluation

  4. ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding

    Oct 6, 2026Hanjun Luo, Xiucheng Zhang, Zhuoning Xu +4AI Coding AgentsAI Risk Management

  5. A Case Study in Assuring AI-Written Software

    Oct 6, 2026Lindsey Ferris, Sierra BonillaAI Agent ReliabilityAI Assurance

  6. Small Agents with Semantic Search: Efficient Multilingual Code Localization

    Oct 4, 2026Maxence Lasbordes, Aarush Sinha, Raphael Sourty +2AI Coding AgentsSemantic Search

  7. AuraForge: Scaling Security Supervision for Training Coding Agents

    Oct 1, 2026Danqing Wang, Songwen Zhao, Harsh Sharma +4AI Coding AgentsSoftware Security

  8. Incident-Arena: Getting agents to the last nine of reliability

    Sep 30, 2026Andre Fu, Malik Drabla, Leon Liu +5Agent ReliabilityAI Coding Agents

  9. How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

    Sep 30, 2026Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel AbbéAI Coding AgentsLLM Agent Evaluation

  10. A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

    Sep 30, 2026Seonho Lee, Wonryeol Jeong, Alberto Cereser +4AI Coding AgentsAI Agent Benchmarks

  11. LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

    Sep 29, 2026Yun Peng, Zihan Wu, Zeyang Zhuang +6Software EngineeringAI Coding Agents

  12. Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes

    Sep 29, 2026Xiaokang Ye, Siddhant Hitesh Mantri, Zimeng Chen +6AI Coding Agents3D Scene Editing

  13. Can Agents Design Libraries for Agents?

    Sep 29, 2026Gabriel Orlanski, Alex L. Zhang, Avi Trost +4Software Engineering AgentsAI Coding Agents

  14. StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

    Sep 28, 2026Ziyang Yu, Liang Zhao, Bowen Zhu +1AI Coding AgentsLLM Agent Memory

  15. Opera: A Verbal Critic Framework for Long-horizon Coding Agents

    Sep 27, 2026Kai Mei, Zhiyuan Hu, Yutong Dai +7AI Coding AgentsAI Agent Auditing

  16. Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models

    Sep 27, 2026Eduardo Ariño de la Rubia, Szilard PafkaAI Coding AgentsLLM Agent Evaluation

  17. SWE-Game: Can Coding Agents Build the Games We Want?

    Sep 27, 2026Xiaoyu Chen, Lai Wei, Jin Wang +8Software Engineering AgentsAI Coding Agents

  18. Evaluating Coding Agents on Kernel Exploit Generation

    Sep 22, 2026Junyoung Jang, Gwanhyun Lee, Hwiwon Lee +4AI Coding AgentsSoftware Security

  19. Quantifying Overclaiming Propensity in Frontier LLM Agents

    Sep 17, 2026Nolan Smyth, Yorguin-Jose Mantilla-Ramos, Pascal Jr Tikeng Notsawo +6AI Agent ReliabilityAI Coding Agents

  20. The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

    Sep 17, 2026Zhexi Feng, Ruiyi Zhang, Yongbo Yang +1AI Coding AgentsAgentic Retrieval