cs.ROOct 7, 2026

iAm.md: Robot Skill Self-Assessment through Agentic Introspection for Unknown Open-Vocabulary Domains

Authors: Vincenzo Guarino, Emanuele Musumeci, Vincenzo Suriani, Daniele Nardi

Organizations: Department of Computer, Control and Management Engineering “Antonio Ruberti”, Sapienza University of Rome, Via Ariosto 25, 00185 Rome, Italy

Abstract

Agentic AI based on Large Language Model generalization capabilities offers a wide range of potential applications, including planning for embodied tasks. For example, embodied agents based on Foundation models can generate plausible plans in autonomous robotics scenarios. Due to limited context windows or hallucinatory phenomena in the next-token prediction formulation, behaviors may be generated without establishing whether the deployed robot and the observed environment actually support the requested operation, in what we call a "grounding failure". Thanks to the recent improvements in reasoning capabilities of foundation models, autonomous robot behavior generation problem can be formulated as a code generation problem. We present iAm.md, a Markdown standard and generation framework, that allows anchoring this process in complementary forms of deployment evidence. Through open-vocabulary semantic mapping, we combine local vision-language detections and object segmentation and refer them to persistent object records in this intermediate standardized representation, allowing agentic introspection. We then study this new technique on a simulated TIAGo, on navigation-and-manipulation tasks, showing how this standardized representation jointly supports skill self-assessment and executable task generalization.

Figures & tables

Explore similar work

CardsList
  1. Inspect Robots: Evaluating the Capabilities and Safety of Embodied AI

    Oct 5, 2026Christopher Leet, Achu Menon, Sravanthi Machcha +12AI Agent EvaluationRobotic Policy Evaluation

  2. ETA: A New Agentic Paradigm for Embodied Tasks

    Aug 4, 2026Yitong Chen, Zezheng Huai, Sixian Li +7Tool-Using AgentsRobot Task Planning

  3. Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

    Oct 20, 2025Yulin Luo, Chun-Kai Fan, Menghang Dong +19Long-Horizon Robotic ManipulationMultimodal Large Language Models