ReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code Generation
Organizations: College of Computer Science and Software Engineering, Shenzhen University · School of Artificial Intelligence, Shenzhen University
Abstract
Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search framework. It organizes program candidates as tree states, retains branch-local debugging context, retrieves failure experience across branches, and distinguishes failed checks from unavailable evidence. On HumanEval and MBPP-Sanitized, visible-test ReMCTS improves over direct generation in 8 of 10 model-dataset pairs under held-out evaluation, whereas proxy-only search is less stable. Controlled tree-search, sampling, repair, and memory ablations characterize the source and limits of these gains. A 30-task HumanEval-X C++ pilot further demonstrates compatibility with compiler-backed execution, but does not constitute a broad multilingual evaluation.
Figures & tables
| Model | Dataset | Direct no_tests | ReMCTS no_tests | ReMCTS known | TG– no_tests (pp) |
|---|---|---|---|---|---|
| Qwen3-1.7B | HumanEval | 60.37% | 69.51% (+9.14) | 83.54% (+23.17) | +14.03 |
| Qwen3-1.7B | MBPP | 53.40% | 51.29% (-2.11) | 63.47% (+10.07) | +12.18 |
| Qwen3-8B | HumanEval | 79.88% | 86.59% (+6.71) | 93.90% (+14.02) | +7.31 |
| Qwen3-8B | MBPP | 64.17% | 66.74% (+2.57) | 74.24% (+10.07) | +7.50 |
| Qwen3-14B | HumanEval | 82.93% | 78.66% (-4.27) | 82.93% (+0.00) | +4.27 |
| Qwen3-14B | MBPP | 66.98% | 60.42% (-6.56) | 65.57% (-1.41) | +5.15 |
| Model | Method | Hidden solve | Tok/task | Calls/task | Wall (s) | Cost (CNY) |
|---|---|---|---|---|---|---|
| Qwen3-8B | ReMCTS | 66.74 | 10,295 | 19.7 | 1,980 | 3.72 |
| Qwen3-8B | Best-of-32 | 59.95 | 8,290 | 44.7 | 1,502 | 3.18 |
| Qwen3-8B | Reflexion-16 | 62.30 | 15,777 | 17.1 | 903 | 4.54 |
| Qwen3-32B | ReMCTS | 70.49 | 9,735 | 19.7 | 3,053 | 13.15 |
| Qwen3-32B | Best-of-32 | 64.17 | 8,600 | 33.0 | 2,232 | 13.79 |
| Qwen3-32B | Reflexion-16 | 69.09 | 11,072 | 12.4 | 1,147 | 12.94 |
| Model | Dataset | Setting | Full | W/O global | W/O path |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash | HumanEval | known | 98.78 | 99.39 | 99.39 |
| DeepSeek-V4-Flash | HumanEval | no_tests | 95.12 | 94.51 | 94.51 |
| DeepSeek-V4-Flash | MBPP | known | 81.03 | 80.56 | 80.09 |
| DeepSeek-V4-Flash | MBPP | no_tests | 72.83 | 73.07 | 72.83 |
| Qwen3-8B | HumanEval | known | 93.90 | 92.07 | 95.12 |
| Qwen3-8B | HumanEval | no_tests | 86.59 | 85.37 | 86.59 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Type | Definition |
|---|---|---|
| Search and semantic-heuristic hyperparameters | , , ; ReMCTS uses , while pure-MCTS/MCTS-style ablations set the memory heuristic to zero. | |
| Reward and acceptance hyperparameters | known uses unit/exec/semantic weights ; no_tests uses proxy/exec/semantic/risk weights . The no_tests acceptance penalties are , and its thresholds are , , , , and . | |
| Proxy-evidence weights | Contract/probe/meta/static weights are ; fixed before evaluation. | |
| Runtime evidence quantities | Produced by visible tests, sandbox execution, contract checks, probes, property checks, static checks, risk scans, and LLM reviews; is computed from the confidence of available evidence groups. | |
| Rule- or LLM-assisted scores | combines path-reflection confidence and high-severity global-memory similarity; depends on nonblank code length; is a confidence- and risk-weighted LLM review. | |
| Numerical stability term | Used only to avoid division by zero; not a tunable experimental parameter. |
| Evidence | Meaning | Role |
|---|---|---|
| Unit | Visible unit tests or assertions | Provides the highest-confidence functional feedback in known |
| Exec | Compilation, loading, entry call, timeout, and crash status | Checks whether the candidate is runnable and callable |
| Contract | Function name, argument count, return form, or schema | Prevents interface mismatches and entry-point errors |
| Probe | Four LLM-generated weak assertions from task text; fixed confidence 0.45 | Provides low-cost behavioral signals without explicit tests after sanitizer checks |
| Meta | Properties or metamorphic relations extracted from the requirement; fixed confidence 0.35 | Checks task properties such as sorting, length, idempotence, and reversal |
| Static | Compilation/execution status; fixed confidence 0.80 | Filters obviously non-runnable code |
| File | Tasks | Use | Source or Description |
|---|---|---|---|
| HumanEval_feedback.jsonl | 164 | known input | Visible feedback parsed from doctest examples in HumanEval prompts; some tasks have no visible examples |
| HumanEval_no_tests.jsonl | 164 | no_tests input | Explicit tests removed; only the prompt and entry point are retained |
| HumanEval_eval.jsonl | 164 | hidden eval | Original HumanEval test fields retained only for final offline evaluation |
| mbpp_sanitized_feedback_split.json | 427 | known input | Converted from Google Research sanitized-mbpp.json ; only the first assertion per task is used as search feedback |
| mbpp_sanitized_no_tests.json | 427 | no_tests input | Explicit assertions removed; only task description and entry point are retained |
| mbpp_sanitized_eval_split.json | 427 | hidden eval | Setup/imports and remaining assertions retained for final evaluation |
| Model | Method | Tasks | Visible pass | Hidden solve | Tok/task | Cost (CNY) |
|---|---|---|---|---|---|---|
| Qwen3-8B | Direct | 30 | 73.33 | 73.33 | 335 | 0.0094 |
| Qwen3-8B | ReMCTS | 30 | 93.33 | 90.00 | 3,217 | 0.0818 |
| Qwen3-32B | Direct | 30 | 83.33 | 83.33 | 337 | 0.0381 |
| Qwen3-32B | ReMCTS | 30 | 96.67 | 96.67 | 3,245 | 0.3260 |
| DeepSeek-V4-Flash | Direct | 30 | 60.00 | 60.00 | 330 | 0.0123 |
| DeepSeek-V4-Flash | ReMCTS | 30 | 96.67 | 96.67 | 2,996 | 0.1063 |
| Item | Setting |
|---|---|
| Model alias | Qwen3-1.7B , Qwen3-8B , Qwen3-14B , Qwen3-32B , DeepSeek-V4-Flash |
| Invocation | OpenAI-compatible API; main experiments were run in May 2026 and the budget-up/C++ pilots in July 2026 |
| Direct, ReMCTS, and tree-search baselines | Code-generation temperature 0.2, maximum output length 1024 tokens, and API request timeout 90 seconds |
| Best-of-8 | 8 independent candidates per task, generation temperature 0.7, and 5 self-generated assertion tests for selection |
| Best-of-32 | 32 independent candidates per task with the same generation and proxy-selection protocol as Best-of-8 |
| Reflexion-style | Up to 8 Actor–Evaluator–Self-Reflection rounds; Actor, Evaluator, and reflection temperatures are all 0.2 |
| Parameter | Value |
|---|---|
| max_iterations | 8 |
| branch_factor | 5 |
| max_depth | 8 |
| no_improve_patience | 8 |
| min_accept_iterations | 3 |
| complete_branch_before_accept | true |
| Metric | Definition | Applicable Settings |
|---|---|---|
| Hidden Solve Rate | Fraction of tasks whose final candidate passes hidden tests | All experiments |
| Accepted Rate | Fraction of tasks whose candidate is automatically accepted | Methods/settings with an internal acceptor |
| Acceptance Precision | Fraction of automatically accepted candidates that pass hidden tests | Methods/settings with an internal acceptor |
| False Acceptance Rate | Fraction of automatically accepted candidates that fail hidden tests | Methods/settings with an internal acceptor |
| Avg. Iterations | Average search iterations per task | Search efficiency analysis |
| Avg. Tree Size | Average search-tree nodes per task | Search efficiency analysis |
| Model | Dataset | Method | Hidden Solve | Accepted | Acc. Prec. | False Acc. | Oracle Pool |
|---|---|---|---|---|---|---|---|
| Qwen3-8B | HumanEval | Best-of-8 | 82.93% | 95/164 | 98.95% | 1.05% | 82.93% |
| Qwen3-8B | HumanEval | Reflexion | 81.10% | 99/164 | 96.97% | 3.03% | 84.15% |
| Qwen3-8B | HumanEval | pure-MCTS | 84.76% | 163/164 | 84.66% | 15.34% | 87.20% |
| Qwen3-8B | HumanEval | MCTS-style | 85.98% | 90/164 | 90.00% | 10.00% | 86.59% |
| Qwen3-8B | MBPP | Best-of-8 | 60.42% | 190/427 | 80.53% | 19.47% | 61.83% |
| Qwen3-8B | MBPP | Reflexion | 61.59% | 234/427 | 74.36% | 25.64% | 64.17% |
| Model | Dataset | Method | Hidden Solve | Accepted | Acc. Prec. | False Acc. | Avg. Iter. |
|---|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash | HumanEval | Reflexion-style | 96.34 | 96.34 | 100.00 | 0.00 | 1.360 |
| DeepSeek-V4-Flash | HumanEval | MCTS-style | 96.34 | 96.34 | 100.00 | 0.00 | 3.183 |
| DeepSeek-V4-Flash | MBPP | Reflexion-style | 94.15 | 93.21 | 100.00 | 0.00 | 1.684 |
| DeepSeek-V4-Flash | MBPP | MCTS-style | 75.88 | 74.24 | 100.00 | 0.00 | 4.288 |
| Qwen3-8B | HumanEval | Reflexion-style | 86.59 | 86.59 | 100.00 | 0.00 | 2.098 |
| Qwen3-8B | HumanEval | MCTS-style | 86.59 | 86.59 | 100.00 | 0.00 | 3.445 |
| Model | Dataset | Setting | ReMCTS Tok. (M) | MCTS-style Tok. (M) | Ratio |
|---|---|---|---|---|---|
| Qwen3-8B | HumanEval | no_tests | 2.52 | 2.26 | 1.12 |
| Qwen3-8B | HumanEval | known | 1.07 | 1.74 | 0.61 |
| Qwen3-8B | MBPP | no_tests | 4.40 | 3.96 | 1.11 |
| Qwen3-8B | MBPP | known | 3.93 | 3.77 | 1.04 |
| Qwen3-32B | HumanEval | no_tests | 2.32 | 2.23 | 1.04 |
| Qwen3-32B | HumanEval | known | 0.80 | 1.72 | 0.46 |
| Model | Dataset | Direct | ReMCTS | ReMCTS-only | Direct-only | Exact | ReMCTS 95% CI |
|---|---|---|---|---|---|---|---|
| Qwen3-1.7B | HumanEval | 99/164 | 137/164 | 40 | 2 | 77.11–88.43 | |
| Qwen3-1.7B | MBPP | 228/427 | 271/427 | 62 | 19 | 58.80–67.89 | |
| Qwen3-8B | HumanEval | 131/164 | 154/164 | 25 | 2 | 89.14–96.65 | |
| Qwen3-8B | MBPP | 274/427 | 317/427 | 57 | 14 | 69.89–78.16 | |
| Qwen3-14B | HumanEval | 136/164 | 136/164 | 15 | 15 | 1.00 | 76.43–87.92 |
| Qwen3-14B | MBPP | 286/427 | 280/427 | 40 | 46 | 0.59 | 60.95–69.92 |
| Model | Dataset | MCTS-style | ReMCTS | ReMCTS-only | MCTS-only | Exact | ReMCTS 95% CI |
|---|---|---|---|---|---|---|---|
| Qwen3-8B | HumanEval | 142/164 | 154/164 | 13 | 1 | 0.00183 | 89.14–96.65 |
| Qwen3-8B | MBPP | 285/427 | 317/427 | 33 | 1 | 69.89–78.16 | |
| Qwen3-32B | HumanEval | 154/164 | 156/164 | 5 | 3 | 0.727 | 90.67–97.51 |
| Qwen3-32B | MBPP | 304/427 | 333/427 | 30 | 1 | 73.82–81.66 | |
| DeepSeek-V4-Flash | HumanEval | 158/164 | 162/164 | 5 | 1 | 0.219 | 95.66–99.66 |
| DeepSeek-V4-Flash | MBPP | 324/427 | 346/427 | 27 | 5 | 77.04–84.47 |
| Model | Dataset | Method | Hidden Solve | Accepted | False Acc. | Avg. Iter. | Avg. Tree |
|---|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash | HumanEval known | No global | 99.39 | 99.39 | 0.00 | 3.043 | 7.165 |
| DeepSeek-V4-Flash | HumanEval known | No path | 99.39 | 99.39 | 0.00 | 3.030 | 7.238 |
| DeepSeek-V4-Flash | HumanEval no_tests | No global | 94.51 | 99.39 | 5.52 | 3.030 | 6.835 |
| DeepSeek-V4-Flash | HumanEval no_tests | No path | 94.51 | 100.00 | 5.49 | 3.000 | 6.695 |
| DeepSeek-V4-Flash | MBPP known | No global | 80.56 | 79.39 | 0.00 | 4.082 | 10.761 |
| DeepSeek-V4-Flash | MBPP known | No path | 80.09 | 78.45 | 0.00 | 4.082 | 11.354 |
| Model | Dataset | Setting | Avg. Iter. | Avg. Tree | Median | Range | Avg. Depth | Eff. Branch |
|---|---|---|---|---|---|---|---|---|
| Qwen3-1.7B | HumanEval | no_tests | 3.098 | 5.165 | 5 | 2–17 | 1.396 | 2.858 |
| Qwen3-1.7B | HumanEval | known | 3.835 | 5.762 | 5 | 2–21 | 1.384 | 2.861 |
| Qwen3-1.7B | MBPP | no_tests | 3.068 | 5.433 | 5 | 2–15 | 1.412 | 2.967 |
| Qwen3-1.7B | MBPP | known | 4.974 | 7.651 | 7 | 2–29 | 1.609 | 2.820 |
| Qwen3-8B | HumanEval | no_tests | 3.500 | 3.671 | 3 | 2–12 | 1.341 | 1.904 |
| Qwen3-8B | HumanEval | known | 3.354 | 3.537 | 3 | 2–21 | 1.366 | 1.741 |
| Model | Dataset | Setting | Score@1 | Score@3 | Score |
|---|---|---|---|---|---|
| Qwen3-1.7B | HumanEval | no_tests | 64.02 | 69.51 | +5.49 |
| Qwen3-1.7B | HumanEval | known | 82.32 | 83.54 | +1.22 |
| Qwen3-1.7B | MBPP | no_tests | 51.99 | 51.29 | -0.70 |
| Qwen3-1.7B | MBPP | known | 63.70 | 63.47 | -0.23 |
| Qwen3-8B | HumanEval | no_tests | 87.20 | 86.59 | -0.61 |
| Qwen3-8B | HumanEval | known | 93.29 | 93.90 | +0.61 |
| Model | Dataset | Setting | Iter@1 | Iter@3 | Tree@1 | Tree@3 | Tree |
|---|---|---|---|---|---|---|---|
| Qwen3-1.7B | HumanEval | no_tests | 1.293 | 3.098 | 2.591 | 5.165 | +2.573 |
| Qwen3-1.7B | HumanEval | known | 2.317 | 3.835 | 3.927 | 5.762 | +1.835 |
| Qwen3-1.7B | MBPP | no_tests | 1.237 | 3.068 | 2.506 | 5.433 | +2.927 |
| Qwen3-1.7B | MBPP | known | 3.864 | 4.974 | 6.129 | 7.651 | +1.522 |
| Qwen3-8B | HumanEval | no_tests | 3.677 | 3.500 | 4.073 | 3.671 | -0.402 |
| Qwen3-8B | HumanEval | known | 1.585 | 3.354 | 2.677 | 3.537 | +0.860 |
| Case | Observation |
|---|---|
| Qwen3-1.7B, HumanEval/5, known | Direct fails hidden evaluation on intersperse . In the ReMCTS run, the search tree has 10 nodes. Three root candidates fail all visible tests. Several sibling candidates pass all 3 visible tests. The selected candidate inserts the delimiter only between adjacent elements. It passes hidden evaluation. This supports the view that weaker models often benefit from sandbox filtering and evidence-based selection. |
| Qwen3-8B, HumanEval/26, known | Direct fails hidden evaluation on remove_duplicates . The ReMCTS run needs 7 iterations and reaches depth 3. Early nodes repeatedly fail visible tests. The reflection memory records two issues. The model confuses “keep first occurrence” with “keep elements that occur once”. It also misses a List import. The final node counts element frequencies. It keeps only elements with count one. It passes all 3 visible tests and hidden evaluation. This shows how path reflections can prevent repeated repair mistakes. |
| Qwen3-32B, HumanEval/83, known | Direct fails hidden evaluation on starts_one_ends . It overcounts numbers ending in 1. Qwen3-32B Direct is already strong overall. Thus, this is a residual error rather than a broad failure mode. ReMCTS explores 10 nodes over 8 iterations. The reflection memory points to an inclusion–exclusion error and the special case . The selected candidate passes all 7 visible tests and hidden evaluation. This supports the claim that stronger models mainly benefit by recovering a smaller set of missed implementations. |
| Qwen3-14B, MBPP task 9, no_tests | The task asks for the minimum positive rotation needed to obtain the same string. In no_tests , ReMCTS accepts a candidate that returns -1 for strings with no smaller period. Hidden evaluation expects full-length rotations, e.g., find_Rotations("ab")==2 and find_Rotations("abc")==3 . Probe evidence partially fails. Contract and static checks pass. The LLM semantic review remains optimistic. This produces a false acceptance. This case shows why proxy evidence can be insufficient on MBPP. It also shows why no_tests improvements are less stable. |