cs.ROSep 28, 2026

SAGE: Symbolic Action-Gating and Editing for LLM Task Planners

Authors: Trung Minh Bui, JongSul Moon, YoungOuk Kim, Quang-Ngoc Phung, Se-Woong Jun, Dongin Shin

Organizations: Intelligent Robotics Research Center, Korea Electronics Technology Institute (KETI), Seongnam 13449, Republic of Korea

Abstract

Large language models (LLMs) are now the default cognitive core of embodied household agents, yet the plans they emit are rarely checked against a grounded model of the environment before execution, and the task-success they report is often measured on benchmarks so saturated that no method can be separated from another. We present SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate (~250 lines of Python, zero tokens, O(∣π∣)O(|π|)) that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and untouched work intact; a hybrid seed+live memory store supports cold-start coverage. We evaluate under a leak-free protocol (leave-one-out retrieval) over five open-weight models and a 75-task AI2-THOR benchmark. On the standard benchmark goal-completeness saturates (52% of instances trivially solved) and SAGE ties strong hierarchical baselines. On a harder, method-agnostic multi-goal composition, SAGE's completeness lead re-emerges large (+0.06 to +0.23 across four models). Under injected mid-execution failures, SAGE recovers as reliably as whole-plan replanners at 2.4-3.3x fewer LLM calls. As a verify-before-execute gate, the symbolic monitor blocks unsafe actions before actuation and raises simulator-reported step-success for every planner tested (up to +0.11), a signal the verifier never sees (non-circular). Because the gate calls no model (0.008 ms/plan), it is a safety layer that runs essentially free on the edge: SAGE planning reproduces its quality on a Jetson AGX Orin, where small-model verification helps most. We release the benchmark, the leak-free protocol, the recovery and safety-gate harnesses, and a verifier-portability study (auto-induced on ALFWorld, 0.89 held-out).

Figures & tables

Explore similar work

CardsList
  1. SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model

    Jun 12, 2026Xiaoxin Lu, Ranran Haoran Zhang, Rui ZhangLarge Language Model PlanningLarge Language Models Fail

  2. Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution

    Sep 29, 2026Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera +2Large Language Model PlanningLarge Language Model Agents

  3. Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

    Jun 3, 2026Haoyu Sun, Wenxuan Wang, Mingyang Song +5Large Language Model PlanningAgentic Benchmarks