cs.SEMay 5, 2026

ProgramBench: Can Language Models Rebuild Programs From Scratch?

Authors: John YangKilian LieretJeffrey MaParth ThakkarDmitrii PedchenkoSten SootlaEmily McMilinPengcheng Yin+4 more

Organizations: 1Meta FAIR · 4Harvard University · 2Meta TBD · 3Stanford University

Abstract

Turning ideas into full software projects from scratch has become a popular use case for language models. Agents are being deployed to seed, maintain, and grow codebases over extended periods with minimal human oversight. Such settings require models to make high-level software architecture decisions. However, existing benchmarks measure focused, limited tasks such as fixing a single bug or developing a single, specified feature. We therefore introduce ProgramBench to measure the ability of software engineering agents to develop software holisitically. In ProgramBench, given only a program and its documentation, agents must architect and implement a codebase that matches the reference executable's behavior. End-to-end behavioral tests are generated via agent-driven fuzzing, enabling evaluation without prescribing implementation structure. Our 200 tasks range from compact CLI tools to widely used software such as FFmpeg, SQLite, and the PHP interpreter. We evaluate 9 LMs and find that none fully resolve any task, with the best model passing 95% of tests on only 3% of tasks. Models favor monolithic, single-file implementations that diverge sharply from human-written code.

Explore similar work

CardsList
  1. MirrorCode: AI can rebuild entire programs from behavior alone

    Jun 29, 2026Tom Adamczewski, David Owen, David Rein +4Coding AgentsAi-Ran