Code Generation Evaluation

Momentum

17 papers in the last four weeks, up 113% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 147

All topics
CardsList
  1. TaoD2C-Bench: Benchmarking MLLMs for Industrial UI Code Generation Beyond Visual Fidelity

    Oct 7, 2026Chengwei Shi, Yunnong Chen, Tingting Zhou +5Code GenerationCode Generation Evaluation

  2. AeroEval: Staged Program and Execution Validation for AI-Generated Drone Missions

    Oct 7, 2026Kautuk Astu, Naina Rabha, Yogesh SimmhanCode Generation EvaluationAerial Robotics

  3. Verification Trap: Understanding Test-Time Selection Failures under False Premises in Code Generation

    Oct 4, 2026Feng He, Hejia Wang, Linghao Meng +2Execution-Guided Code GenerationCode Generation Evaluation

  4. Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents

    Oct 1, 2026Junyu Guo, Shangding Gu, Ming Jin +1AI Agent AuditingCode Generation Evaluation

  5. Code That Works, Environments That Don't: Measuring Environment Reproducibility in AI-Generated Software

    Sep 30, 2026Bhanu Prakash Vangala, Tanu MalikCoding AgentsCode Generation

  6. A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

    Sep 30, 2026Seonho Lee, Wonryeol Jeong, Alberto Cereser +4AI Coding AgentsAI Agent Benchmarks

  7. Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen

    Sep 22, 2026Om Nepal, Sushant Aryal, Oluseyi Olukola +1Automated Program RepairLLMs for Cybersecurity

  8. ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

    Sep 16, 2026Jeonghye Kim, Minseon Kim, Young Jin Kim +5Coding AgentsSoftware Engineering Benchmarks

  9. An Empirical Evaluation of Cost-Efficient Large Language Models on Algorithmic Programming Tasks

    Sep 16, 2026Chandimal Adikari, Nandika HerathCode GenerationLLM Reliability

  10. DepthBenchCAD: When Does Deeper Auditing Yield More Reliable Conclusions?

    Sep 14, 2026Hongye Yang, Zhihao Xie, Shengjun Xiong +1Benchmark AuditingText-to-CAD Generation

  11. What is the Difference Between Me and You? Benchmarking the Quality Gap Between Human-Written and AI-Generated Code

    Sep 14, 2026Cristina Improta, Pietro Liguori, Domenico CotroneoLLM SecurityCode Generation Evaluation

  12. The Vibe Shift in Software Engineering: Evaluating AI-Led Conversational Programming for Performance, Cognition, and Responsible Adoption

    Sep 10, 2026Sales G. Aribe Jr., Louie Jay S. LabastidaCode Generation EvaluationHuman-AI Collaboration

  13. It Is Not My Code Anymore

    Sep 8, 2026Augusto CamargoAlgorithmic AccountabilityCode Generation Evaluation

  14. CodeTD: Topology of Attention Detects Hallucinations in Code LLMs

    Sep 7, 2026Daria Voronkova, Ilya Trofimov, Anton Dmitriev +3LLM Hallucination DetectionCode Generation Evaluation

  15. Rubric-to-Code Credit Assignment for Reinforcement Learning

    Aug 28, 2026Rui Jin, Jikai Chen, Yihan Chen +6RL for Code GenerationCode Generation

  16. RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

    Aug 28, 2026Gyuhyeong Kim, Hyojung Gwon, Jeonghyeon Kim +2LLM PromptingCoding Agents

  17. QuoteBench: How Matched Scores Can Hide Command-Path Failures

    Aug 13, 2026Shangao Li, Yao Zhang, Volker Tresp +1LLM Agent EvaluationAgent Evaluation

  18. VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?

    Aug 11, 2026Mizanur Rahman, Arshia Azimlu, Shadikur Rahman +4VLM EvaluationData Visualization