Organizations: Northeastern University Bellevue, USA · Columbia University New York, USA · University of Pennsylvania Philadelphia, USA · University of Southern California San Jose, USA · Independent Researcher Mukilteo, USA
Large language models for code generation often fail on execution, multilingual coverage, and contamination control, especially under frozen backbone constraints. We present CodeForge-MA, a unified framework that improves code synthesis through a multi-agent data forge, execution verified reinforced instruction tuning, and a language conditioned mixture of LoRA adapters. Four specialized agents, Composer, Reviewer, Executor, and Curator, iteratively refine instruction code pairs, validate them with tests, and filter duplicates and benchmark leakage. During training, we combine masked supervised fine tuning with a test driven reinforcement objective to align generations with executable correctness. For the larger model, we use sparse expert routing over low rank adapters to improve cross language transfer while keeping the base model unchanged at inference. Experiments show that joint data, objective, and adapter design yields robust gains across programming languages.
Figures & tables
Fig. 1: Overview of the CodeForge-MA framework. Four sequential stages (Multi-Agent Data Forge, EV-RIT Training, LC-MoLoRA Adaptation, Inference) operate above a frozen Qwen-1.8B/72B backbone, faithfully respecting the competition’s no-modification rule.
Fig. 2: The Multi-Agent Data Forge. Four specialized agents {ACmp,ARev,AExe,ACur} cooperate in a closed Reflexion-style loop; a sample is accepted only when πk=1 , qk≥τq , and k≤Kmax .
Fig. 3: Execution-Verified Reinforced Instruction Tuning (EV-RIT). The upper lane implements masked cross-entropy on the assistant span; the lower lane implements a PPO surrogate whose reward is the unit-test pass ratio. The two streams are combined by the warm-up coefficient λt .
Symbol
Value
Description
τq
0.75
Forge acceptance threshold
Kmax
3
Max forge revisions per sample
M
5
Reviewer self-consistency samples
α
0.3
Language sampling temperature
τJ
0.8
MinHash deduplication threshold
βKL
0.02
KL penalty in PPO surrogate
TABLE I: Key hyperparameters of CodeForge-MA.
Fig. 4: Chain-of-Repair prompting schema. The model emits an ordered triple (Φloc,Φrat,Φfix) ; only Φfix is forwarded to the hidden unit tests at evaluation time, while LCoR ties the three fields together during training.
Model / Configuration
pass@1 / Score
CodeBLEU
CR
BMPS
1.8B-class backbones
Qwen-1.8B (base)
13.6
22.1
48.3
10.9
CodeGen-2B-mono
16.6
25.8
55.7
13.1
StarCoderBase-3B
21.0
29.6
62.4
17.4
Qwen-1.8B + Evol-Instruct
21.9
30.2
63.8
17.9
DeepSeek-Coder-1.3B-Instruct
27.1
34.7
71.2
22.8
TABLE II: Consolidated results across four metrics. Top: 1.8B-class backbones, pass@1 averaged over HumanEvalSynthesize / HumanEvalFixTests / MBPP. Middle: 70B-class backbones, where the first column is Scorefinal from ( 32 ). Bottom: ablation of CodeForge-MA on Qwen-1.8B. All numbers in %.
State-of-the-art code generation frameworks rely on mental simulation, where LLMs internally trace execution to verify correctness. We expose a fundamental limitation: the Mental-Reality Gap -- where models hallucinate execution traces and confidently validate buggy code. This gap manifests along two orthogonal dimensions: the Specification Gap (overlooking edge cases during planning) and the Verification Gap (hallucinating correct behavior for flawed code). We propose SolidCoder with a simple principle: don't imagine -- execute. The S.O.L.I.D. architecture addresses both dimensions by forcing edge-case awareness before algorithm design and replacing imagined traces with sandboxed execution using property-based oracles. With GPT-4o, SolidCoder achieves state-of-the-art pass@1 performance: 95.7% on HumanEval (+0.6%p), 77.0% on CodeContests (+4.3%p), and 26.7% on APPS (+3.4%p). Ablation reveals that edge-case awareness provides the largest individual gain, while execution grounding catches categorically different errors that specification improvements cannot address. These gains generalize to RL post-trained models, validating that bridging both gap dimensions is essential for robust code synthesis. We release our code and framework to facilitate future research.
Woojin Lee, Jin-Xia Huang
Electronics and Telecommunications Research Institute, Republic of Korea
Large Language Models achieve strong code generation for high resource languages like Python and Java but suffer sharp performance drops on Low-Resource Programming Languages~(LRPLs) such as Julia. Improving Small Language Models~(SLMs) for these languages faces a trilemma: Supervised Fine-Tuning~(SFT) is bottlenecked by data scarcity, inference-time scaling is too expensive for deployment, and Reinforcement Learning from scratch yields near zero advantages. We propose a three-phase pipeline that resolves this trilemma by decoupling syntax acquisition from algorithmic reasoning. First, we \emph{left-shift} inference-time compute to an offline data synthesis engine that uses iterative compiler and test feedback to generate verified training examples. Second, we fine-tune an SLM on this synthetic, verified data to embed strong syntactic priors. Third, we apply Reinforcement Learning with Verifiable Reward~(RLVR) grounded by language-agnostic Input/Output tests, where the SFT prior constrains exploration away from syntax errors. Applied to Qwen3-8B, our pipeline improves pass@1 by up to +7.6 points on MultiPL-E and +14.2 points on the Agnostics LiveCodeBench for Julia compared to SOTA results. Furthermore, the pipeline only used 31 data and 61 cost over the previous state-of-the-art. We further demonstrate that the pipeline generalizes to Ballerina achieving 49.7% MultiPL-E Pass@1, a language with near-zero pretraining representation. Ablations confirm that both the SFT phase and execution-grounded rewards are necessary for stable training.
Large language models (LLMs) already excel at writing code in high-resource languages such as Python and JavaScript, yet stumble on low-resource languages that remain essential to science and engineering. Besides the obvious shortage of pre-training data, post-training itself is a bottleneck: every new language seems to require new datasets, test harnesses, and reinforcement-learning (RL) infrastructure. We introduce Agnostics, a language-agnostic post-training pipeline that eliminates this per-language engineering. The key idea is to judge code solely by its externally observable behavior, so a single verifier can test solutions written in any language. Concretely, we (i) use an LLM to rewrite existing unit-test datasets into an I/O format, (ii) supply a short configuration that tells the verifier how to compile and run a target language, and (iii) apply reinforcement learning with verifiable rewards (RLVR) in a robust code execution environment. Applied to five low-resource languages--Lua, Julia, R, OCaml, and Fortran--Agnostics (1) improves Qwen-3 4B to performance that rivals other 16B-70B open-weight models; (2) scales cleanly to larger and diverse model families (Qwen-3 8B, DeepSeek Coder 6.7B Instruct, Phi 4 Mini); and (3) for ≤16B parameter models, sets new state-of-the-art pass@1 results on MultiPL-E and a new multi-language version of LiveCodeBench that we introduce. We release the language-agnostic training datasets (Ag-MBPP-X, Ag-Codeforces-X, Ag-LiveCodeBench-X), training code, and ready-to-use configurations, making RL post-training in any programming language as simple as editing a short YAML file.
Aleksander Boruch-Gruszecki, Yangtian Zi, Zixuan Wu +4
1Northeastern University · University of Texas · 3Wellesley College