Direct Preference Optimization (DPO) treats all constraint violations equally: a 1budgetovershootanda1,000 overshoot induce the same training signal. It is also susceptible to length and style bias when preference pairs come from different model families. We introduce Constraint-Margin DPO (CM-DPO), which replaces DPO's binary preference signal with a continuous margin derived from a deterministic symbolic verifier and scaled by violation severity. Hard and soft constraints are separated through a lexicographic objective, ensuring hard constraints are never traded off against preferences. To supply CM-DPO with bias-reduced training pairs, we generate preference data through procedurally generated constraint profiles (DCCG) and minimal-edit distillation from a reasoning teacher (RT-MED), within a framework we call SynPlan-R. On TravelPlanner, NaturalPlan, and out-of-distribution PlanBench, an 8B model fine-tuned with CM-DPO achieves 89.2% pass rate and 93.4% solve rate, matching multi-agent systems at 13x lower latency while outperforming GPT-4o on unseen Blocksworld by 9.2 points.
Figures & tables
Figure 1: The SynPlan-R architecture. Phase 1: DCCG decomposes each profile into hard constraints Chard (red), soft preferences Csoft (blue), and a procedural knowledge
Figure 2: RT-MED qualitative example The hard constraint is total cost ≤600 , and the soft constraint is a preference for boutique hotels. Left: the student draft violates the budget constraint (hotel at 195/night,yieldingatotalcostof650). Center: RT-MED swaps only the hotel ( ∼12 tokens). Right: a full rewrite changes ∼245 tokens and introduces style bias. Bottom: token-count and PR comparison.
Figure 3: CM-DPO vs. Standard DPO. Standard DPO assigns the same penalty regardless of violation severity. CM-DPO scales the margin γmc proportionally to the budget overshoot, producing stronger repulsion for larger violations.
Table 4
Method
TravelPlanner
NaturalPlan
Eff.
PR ↑
DR ↑
Sˉ↑
SR ↑
Lat. ↓
External baselines
Zero-Shot (Llama-3-8B)
1.3
35.4
—
18.5
2.1s
Zero-Shot (GPT-4o)
39.8
88.3
0.61
68.4
5.4s
MAS (LRPlan w/ GPT-4o)
89.4
98.1
0.72
95.1
28.5s
SynPlan-R ablations (Llama-3-8B)
Table 3: Main results. SynPlan-8B reaches a TravelPlanner pass rate close to MAS performance (89.2% vs. 89.4%) while reducing latency from 28.5s to 2.1s. The soft-preference score Sˉ is computed only over feasible plans.
Standard DPO
CM-DPO
PR
Tok Δ
PR
Tok Δ
Full Rewrite
64.1
245
72.6
245
RT-MED
78.3
32
89.2
32
Table 4: 2×2 ablation on TravelPlanner pass rate (%).
Table 7
Domain
∣VChard∣
∣VCsoft∣
∣VE∣
avg. ∣Egrounds∣
Travel
3.4
1.8
41.2
14.7
Retail
3.2
2.0
28.4
11.9
Table 7: Constraint-profile graph statistics (averaged across instances).
Multi-constraint planning involves identifying, evaluating, and refining candidate plans while satisfying multiple, potentially conflicting constraints. Existing large language model (LLM) approaches face fundamental limitations in this domain. Pure reasoning paradigms, which rely on long natural language chains, are prone to inconsistency, error accumulation, and prohibitive cost as constraints compound. Conversely, LLMs combined with coding- or solver-based strategies lack flexibility: they often generate problem-specific code from scratch or depend on fixed solvers, failing to capture generalizable logic across diverse problems. To address these challenges, we introduce the Scalable COde Planning Engine (SCOPE), a framework that disentangles query-specific reasoning from generic code execution. By separating reasoning from execution, SCOPE produces solver functions that are consistent, deterministic, and reusable across queries while requiring only minimal changes to input parameters. SCOPE achieves state-of-the-art performance while lowering cost and latency. For example, with GPT-4o, it reaches 93.1% success on TravelPlanner, a 61.6% gain over the best baseline (CoT) while cutting inference cost by 1.4x and time by ~4.67x. Code is available at https://github.com/DerrickGXD/SCOPE.
Derrick Goh Xin Deik, Quanyu Long, Zhengyuan Liu +2
Nanyang Technological University, Singapore · Agency for Science, Technology and Research (A*STAR), Singapore
Large language model reasoning is often treated as a monolithic capability, relying on binary preference supervision that fails to capture partial progress or fine-grained reasoning quality. We introduce continuous utility direct preference optimization (CU-DPO), a framework that aligns models to a portfolio of prompt-based cognitive strategies by replacing binary labels with continuous scores that capture fine-grained reasoning quality. We prove that learning with K strategies yields a Theta(K log K) improvement in sample complexity over binary preferences and that DPO converges to the entropy-regularized utility-maximizing policy. To exploit this signal, we propose a two-stage pipeline: (i) strategy selection, which optimizes the model to choose the best strategy via best-vs-all comparisons, and (ii) execution refinement, which trains correct execution using margin-stratified pairs. The framework is domain-agnostic: any task admitting cognitively distinct solution strategies and a decomposable continuous utility signal can be incorporated into the portfolio. On mathematical reasoning benchmarks, CU-DPO improves strategy selection accuracy from 35-46% to 68-78% across seven base models, yielding downstream reasoning gains of up to +6.6 points on in-distribution datasets with effective out-of-distribution transfer. CU-DPO demonstrates consistent gains on code generation and causal reasoning benchmarks, confirming generalization beyond the mathematical domain.
Offline preference optimization has become a practical substitute for reinforcement learning from human feedback, but pairwise objectives such as Direct Preference Optimization (DPO) and its variants use only the chosen and rejected responses stored in a static dataset. This leaves a useful signal unused: the response that the reference model itself would generate for the same prompt. We propose Direct Preference Optimization with Penalization (DPOP), a simple extension of DPO that augments the base preference loss with a gated penalty on reference-greedy responses. DPOP activates this penalty only when the current policy still assigns a lower likelihood to the preferred response than to the rejected response. On AlpacaEval 2.0, DPOP improves length-controlled win rate over DPO, SimPO, and AlphaDPO on both Llama-3-8b-it and Gemma-2-9b-it, achieving relative gains of 5.3% and 4.4% over baselines on the two models, respectively. Ablations further show that a SimNPO-style length-normalized penalty is stronger than NPO and token-level unlikelihood in this setting.