cs.AIOct 6, 2026

CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning

Authors: Rabimba Karanjai, Qun Gu, Hemanth Hegadehalli Madhavarao, Wenhuan Sun, Xiaojiao Yu, Suryabhan Singh Hada, Libin N. George, Uma Kona, +3 more

Organizations: PayPal AI Lab

Abstract

Direct Preference Optimization (DPO) treats all constraint violations equally: a 1budgetovershootanda1 budget overshoot and a 1,000 overshoot induce the same training signal. It is also susceptible to length and style bias when preference pairs come from different model families. We introduce Constraint-Margin DPO (CM-DPO), which replaces DPO's binary preference signal with a continuous margin derived from a deterministic symbolic verifier and scaled by violation severity. Hard and soft constraints are separated through a lexicographic objective, ensuring hard constraints are never traded off against preferences. To supply CM-DPO with bias-reduced training pairs, we generate preference data through procedurally generated constraint profiles (DCCG) and minimal-edit distillation from a reasoning teacher (RT-MED), within a framework we call SynPlan-R. On TravelPlanner, NaturalPlan, and out-of-distribution PlanBench, an 8B model fine-tuned with CM-DPO achieves 89.2% pass rate and 93.4% solve rate, matching multi-agent systems at 13x lower latency while outperforming GPT-4o on unseen Blocksworld by 9.2 points.

Figures & tables

Explore similar work

CardsList
  1. Programming over Thinking: Efficient and Robust Multi-Constraint Planning

    Jan 14, 2026Derrick Goh Xin Deik, Quanyu Long, Zhengyuan Liu +2Large Language Model PlanningClassical Planning

  2. Continuous-Utility Direct Preference Optimization

    Jan 31, 2026Muhammad Ahmed Mohsin, Muhammad Umer, Emily Fox