Code Generation Evaluation

Momentum

17 papers in the last four weeks, up 113% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 147

All topics
CardsList
  1. Specification Grounding Drives Test Effectiveness for LLM Code

    Jul 7, 2026Amin Haeri, Mahdi GhelichiLLM GroundingAutomated Software Testing

  2. Evaluating Neural Decompilation of Dart AOT Binaries: Fine-Tuning, Metric Validity, Specification Leakage, and Reliability

    Jul 7, 2026Raafat Abualazm, Ayman AboElhassan, Amr G. WassalSoftware EngineeringLLM Evaluation

  3. Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards

    Jul 6, 2026Tianhao Niu, Ziyu Han, Qiguang Chen +5Data VisualizationCode Generation Evaluation

  4. AlgoBench: Benchmarking Algorithmic Adaptation in Code Generation

    Jun 30, 2026Xinyuan Song, Zekun Cai, Liang ZhaoSynthetic Benchmark GenerationCode Generation Evaluation

  5. LibEvoBench: Probing Temporal Knowledge Stratification in Code Generation Models

    Jun 24, 2026Daniele Cipollone, Sergey Titov, Maliheh Izadi +2Software EngineeringTemporal Reasoning in Language Models

  6. SoK: AI Secure Code Generation: Progress, Pitfalls, and Paths Forward

    Jun 23, 2026Rupam Patir, Keyan Guo, Haipeng Cai +1AI Coding AgentsCode Generation Evaluation

  7. Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent

    Jun 21, 2026Ishaan Bhola, Adithyan Krishnan, Sravanth Kurmala +1Software Engineering AgentsCode Generation Evaluation

  8. NL2Scratch: An Executable Benchmark and Evaluation for Block-Based Programming

    Jun 20, 2026Heejin Do, Alexandre Ballenghien, Yang Wu +1Code GenerationCode Generation Evaluation

  9. Is Agent Code Less Maintainable Than Human Code?

    Jun 19, 2026Shaswat Patel, Betty Li Hou, Arun Purohit +4Software EngineeringCoding Agents

  10. Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages

    Jun 18, 2026Maria Ivanova, Pavel Zadorozhny, Rodion Levichev +5Code GenerationCode Generation Evaluation

  11. GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

    Jun 16, 2026Tongxu Luo, Rongsheng Wang, Jiaxi Bi +22AI Agent BenchmarksAgentic Code Generation

  12. Unlocking LLM Code Correction with Iterative Feedback Loops

    Jun 16, 2026Le Zhang, Suresh KothariLLM Self-CorrectionExecution-Guided Code Generation

  13. Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment

    Jun 12, 2026Kirill Vasilevski, Ximing Dong, Benjamin Rombaut +8Software Engineering AgentsLLM-as-a-Judge

  14. Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive Programming

    Jun 11, 2026Tingqiang Xu, Hangrui Zhou, Tianle Cai +2Automated Program RepairCompetitive Programming

  15. FASE: Fast Adaptive Semantic Entropy for Code Quality

    Jun 8, 2026Shizhe Lin, Ladan TahvildariSemantic EntropyLLM Uncertainty Estimation

  16. Beyond Pass Rate: A Multilingual, Execution-Grounded Evaluation of Open Code LLMs

    Jun 7, 2026Sayed Erfan ArefinCode Generation EvaluationCode Language Models

  17. Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests

    Jun 5, 2026Thanawat Lodkaew, Johannes Ackermann, Soichiro Nishimori +3Reward HackingCoding Agents