cs.CLSep 28, 2026

ReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code Generation

Authors: Huifei Wang, Xinying Huang, Yiheng Sun, Yifan Yuan

Organizations: College of Computer Science and Software Engineering, Shenzhen University · School of Artificial Intelligence, Shenzhen University

Abstract

Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search framework. It organizes program candidates as tree states, retains branch-local debugging context, retrieves failure experience across branches, and distinguishes failed checks from unavailable evidence. On HumanEval and MBPP-Sanitized, visible-test ReMCTS improves over direct generation in 8 of 10 model-dataset pairs under held-out evaluation, whereas proxy-only search is less stable. Controlled tree-search, sampling, repair, and memory ablations characterize the source and limits of these gains. A 30-task HumanEval-X C++ pilot further demonstrates compatibility with compiler-backed execution, but does not constitute a broad multilingual evaluation.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Probabilistic Programs of Thought

    Apr 19, 2026Poorva Garg, Renato Lui Geh, Daniel Israel +3Code GenerationLLM Reasoning Strategies

  2. Understanding Scattered Forest Search: A Version-Space Perspective on Multi-Turn Program Correction

    Apr 27, 2026Yuto Tanaka, Issei SatoEfficient InferenceIterative Refinement

  3. Playing Psychic: Using Thought Trees to Predict Reasoning Models Accuracy on Coding Tasks

    Apr 18, 2026Jiaxin Fang, Runyuan He, Sahil Bhatia +2Reasoning TracesLarge Reasoning Models