cs.AISep 28, 2026

A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations): A Visual-Symbolic Framework for Virtual Humans

Authors: Alessandro Emmanuel Pecora, Stefano Calzolari, Francesco Strada, Andrea Bottino

Organizations: DAUIN, Politecnico di Torino, Torino, Italy

Abstract

Creating believable vh requires the coherent integration of perception, reasoning, and action mediated by language. A central challenge is to combine these components into a control loop grounded in interactive 3D environments. To this end, we present A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations), a visual-symbolic framework for language-driven vh that leverages a pretrained vlm with tool calling to unify perception, reasoning, and action within a single control loop. A.D.A.M.O. maintains a dual visual-symbolic world model that combines egocentric visual input and synchronized symbolic state to support grounded task-oriented behavior from natural language prompts. To support diagnostic evaluation, we introduce a controlled task suite organized by a cd taxonomy that breaks down spatial tasks into procedural and linguistic complexity. Experiments in controlled scenes show that semantic labeling strongly influences task completion and failure modes, reducing perceptual ambiguity while shifting failures toward downstream execution, whereas reasoning errors remain comparatively rare.

Figures & tables

Explore similar work

CardsList
  1. DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model

    Jun 10, 2026Pankhuri Vanjani, Zhuoyue Li, Jakub Suliga +4Diffusion-Based Vision-Language-ActionsAction Generation

  2. AR-WAM: A Visual-Conditioned Agent-Ready World Action Model for Robotic Manipulation

    Sep 20, 2026Yicheng Jiang, Zesen Gan, Xiaobo Wang +8Efficient World-Action ModelRobotic Manipulation

  3. ChunkVLA-AM: Parallel Action Chunking for Vision-Language-Action Robot Control in Additive Manufacturing

    Oct 1, 2026Zhugang Liu, Kaichuang Zhang, Jinman Zhang +7Vision-Language-Action FrameworkRobot Systems