cs.ROOct 8, 2026

SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning

Authors: Weizhe Xu, Jialiang Fan, Mengyu Liu, Fanxin Kong

Organizations: University of Notre Dame · Washington State University

Abstract

Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation. We develop an inference-time monitor that exposes and verifies intermediate plans without disrupting the original decoding trajectory. Building on this monitor, we propose SafeInferCom, a formal verifier-guided framework that preserves valid intermediate plans and directs error correction during generation. Experiments across multiple LRLMs and planning domains reveal reasoning-response inconsistency and limited self-correction under one-shot inference. SafeInferCom improves planning success and accelerates error correction relative to one-shot inference. When combined with iterative refinement, it further improves success while reducing token usage compared with refinement alone. We additionally evaluate SafeInferCom in VirtualHome and provide a real-world robotic-arm demonstration.

Figures & tables

Explore similar work

CardsList
  1. Using large language models for embodied planning introduces systematic safety risks

    Apr 20, 2026Tao Zhang, Kaixian Qu, Zhibin Li +4Large Language Model-Based Robot PlanningRobot Safety

  2. Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy

    Aug 10, 2026Rohan Bhagra, Mahantesh Halapannavar, Uddhav BhattaraiLarge Language Model-Based Robot PlanningAdversarial Robustness

  3. SafeRun: Enabling Determinism in LLM Planning for Running

    Jun 8, 2026Meilin Chen, Zepeng Zhai, Jiaxuan Zhao +1LLM PlanningLLM Safety Evaluation