Toward Trustworthy Physical AI for Human Interaction
Organizations: The BioRobotics Institute, Scuola Superiore Sant’Anna, Pontedera, Italy. · CSAIL, Massachusetts Institute of Technology, Cambridge, MA, USA. · CHARM Lab, Stanford University, Stanford, CA, USA. · The University of Tokyo, Tokyo, Japan. · Boston University, Boston, MA, USA. · Department of Mechanical Engineering, College of Design and Engineering, and Advanced Robotics Centre, National University of Singapore, Singapore.
Abstract
Robots are entering human spaces faster than we can establish when they deserve trust. We propose a framework for trustworthy physical AI that integrates Safety, Behavioral Intelligibility, and Perceptual Alignment across embodiment, control, cognition, and design. Trustworthiness emerges from aligning physical capabilities, observable behavior, and expectations people form during interaction.
Figures & tables
| Dimension | Evaluation target | Figure of merit and operationalization |
| Safety | Maintain safety | Generalized safety margin (positive satisfies the specified criterion; zero marks its boundary; negative violates it) and trajectory minimum , operationalized via the contact-force margin or task-specific separation, pressure, energy, compression, or acceleration margins. |
| Safety | Control physical risk under uncertainty | is a conservative lower bound on the minimum safety margin predicted across possible outcomes of action ; is the tolerated frequency with which the actual margin may fall below this bound. Report this frequency, the typical gap between the bound and actual margin, and how often the robot selects an action with or invokes a defined fallback. Also report actual violations. |
| Safety | Recover from an unsafe state | Minimum margin during recovery and time to return to and remain at ; report unrecovered trials separately. |
| Behavioral Intelligibility | Produce predictable motion | Trajectory predictability , including fallback behavior, operationalized via or ; validate against human anticipation accuracy, surprise, intervention, or correction rates as appropriate. |
| Behavioral Intelligibility | Reveal the intended goal early | Time-weighted legibility ; report the goal hypothesis set, observation horizon, and time to correct goal inference. |
| Perceptual Alignment | Align perceived and actual capability | Signed discrepancy between users’ elicited estimates of limits relevant to the interaction, such as task success rate, sensing range, and contact tolerance, and the validated values reported in the Safety rows [ 61 , 60 ] . Report overestimation and underestimation separately, since they differ in consequence and remedy. State the user population, prior exposure, and elicitation instrument [ 66 ] . |
| Technical Pillars Trustworthiness Dimensions | Embodied Intelligence | Motor-Cognitive Intelligence | Communicative Intelligence |
| Safety | Passive harm mitigation. Mechanical compliance, damping, mass distribution, contact area, surface texture, and mechanical stops shape contact severity. Soft or compliant bodies can reduce peak forces, pressures, and transferred energy when perception, planning, or control fail. | Active risk control. Human-aware sensing, human motion prediction, stable control, collision avoidance, safe motion generation, fail-safe recovery, and safety filters actively bound risk under uncertainty in closed-loop interaction. Controller gains, impedance/admittance design, latency, and passivity directly affect safety. | Behavior-shaping safety cues. Mode indicators, motion previews, lights, sounds, displays, transparent warnings, and role-consistent appearance shape human behavior around the robot, including proximity, handover timing, workspace entry, and appropriate reliance. |
| Behavioral Intelligibility | Intelligible embodied dynamics. Passive dynamics, compliant or underactuated morphology, soft contact, and visible shape adaptation create physical regularities that humans can learn. Hidden deformation, hysteresis, nonlinear material response, or oscillation can make behavior harder to predict. | Predictable and legible motion. Low-level stability and high-level motion policies determine whether actions are smooth, bounded, repeatable, goal-directed, and consistent with human expectations. Perception failures, abrupt replanning, jitter, hidden timing references, or unexplained pauses can make motion appear unpredictable. | Observable state, intent, and upcoming action. Appearance, affordances, gaze-like orientation, body posture, lights, sounds, displays, transparent covers, and motion previews help humans infer the robot’s state, goal, mode, turn-taking structure, and upcoming action. |
| Perceptual Alignment | Physical compliance calibration. Visual, tactile, material, geometric, and scale cues should make perceived softness, stiffness, yielding, inertia, and contact tolerance match the robot’s actual passive mechanical compliance. This calibrates the embodied component of perceived safety: what the body itself can absorb or limit during contact. | Competence, reliability, and active-safety calibration. Observable motion and autonomy cues, such as smoothness, responsiveness, consistency, recovery, hesitation, interruption, and failure handling, should make perceived confidence, competence, reliability, and active risk control match the robot’s safety-aware control capabilities. | Aesthetic-social approachability. Appearance, morphology, expressiveness, surface-material style, perceived cuteness, and behavioral congruence should make people feel appropriately comfortable approaching, being near, and socially interacting with the robot. These cues should match the robot’s intended role and actual social-interaction capabilities. |