Natural face-to-face human--robot interaction requires a robot to understand an evolving social situation, decide when to engage, and express its intent through coordinated physical behavior. Yet existing approaches rarely close this loop: foundation-model agents provide increasingly capable multimodal reasoning and memory but remain largely disembodied, while expressive virtual agents do not face the physical constraints of real robots, and physical social robots typically address social reasoning and embodied expression only partially. We present ARISE, a unified framework that bridges Agentic Reasoning and Interactive Social Embodiment on the Sophia humanoid robot. ARISE integrates multimodal context understanding, long-term memory, and reactive and proactive interaction to determine when and what to communicate, and translates social intent into robot-native gestures coordinated with speech and mechanical facial expressions through streaming execution. Extensive evaluations on Sophia demonstrate strong perceived interaction quality, expressive and well-coordinated embodied behavior, and substantial latency reductions through streaming execution. These results highlight the importance of jointly reasoning about what to communicate, when to engage, and how to physically express social intent for natural interaction with humanoid robots. Project Page: https://robosocial.github.io/
Figures & tables
Method
Physical robot
Long-term memory
Proactive engagement
Facial expression
Co-speech gestures
EMAGE [ 1 ]
✗
✗
✗
✓
✓
Social Agent [ 2 ]
✗
✗
✗
✓
✓
MIBURI [ 3 ]
✗
✗
✗
✓
✓
Aly and Tapus [ 4 ]
✓
✗
✗
✓
✓
Attentive Support [ 5 ]
✓
✗
✓
✗
✗
ExFace [ 6 ]
✓
✗
✗
✓
✗
TABLE I: Capability comparison with previous works. ✓ : demonstrated; ✗ : not reported.
Fig. 2: Overview of ARISE . Given conversational audio and real-time RGB observations, the Multimodal Social Agent performs reactive and proactive planning with long-term memory and online search, generating a verbal response and symbolic gesture plan. Streaming Multimodal Orchestration synchronizes speech, facial animation, and gestures for progressive execution, and Robot Deployment maps the coordinated outputs to robot’s actuators.
Fig. 3: Multi‑round Interaction Sequences (a) Birthday Gift : Sophia preserves confidential information for individual users and remembers the gift’s specific location. (b) Whose Bag : Leveraging visual perception to distinguish different users and backpacks, Sophia recalls object‑user associations and answers identity‑related queries. (c) Clean up Desk : Sophia can offer appropriate contextual guidance for object placement, and give positive feedback proactively after the user finishes the task.
Method
OIQ( ↑ )
CA( ↑ )
CC( ↑ )
IE( ↑ )
Full System
4.38 ± 0.41
4.34 ± 0.43
4.25 ± 0.46
4.31 ± 0.44
w/o Visual Input
3.72 ± 0.58 ***
2.90 ± 0.69 ***
4.08 ± 0.50
4.06 ± 0.52
w/o Long-Term Memory
3.54 ± 0.61 ***
3.78 ± 0.57 **
3.29 ± 0.68 ***
3.62 ± 0.64 **
w/o Proactive Interaction
3.76 ± 0.51 **
4.23 ± 0.44
4.16 ± 0.47
3.12 ± 0.71 ***
TABLE II: Results for Multi-Round Social Interaction.
Fig. 4: Gesture Generation Sequences (a) Group Photo : The robot actively responds to social invitations and cooperates for group‑photo scenarios. (b) Job Interview : Detecting user nervousness, Sophia delivers comforting verbal feedback and appropriate expressive gestures.
Method
Motion Natural.
Motion Express.
Speech-Motion Coord.
Overall Quality
Full System
3.92 ± 0.34
4.22 ± 0.34
4.03 ± 0.60
4.09 ± 0.26
Single Candidate
3.63 ± 0.41 **
3.88 ± 0.43 *
3.79 ± 0.41
3.92 ± 0.29
w/o R-R Agent
3.00 ± 0.39 ***
3.35 ± 0.46 ***
3.21 ± 0.44 **
3.58 ± 0.24 ***
EMAGE
3.37 ± 0.27 ***
3.67 ± 0.38 ***
3.89 ± 0.33
3.98 ± 0.25
TABLE III: Results for the Gesture-Generation Experiment.
Latencies
Definition
LMSA
The latency of Multimodal Social Agent .
LES
The latency of embodied synthesis.
Lrobot
The latency of Robot Deployment .
LAF
The end-to-end latency from the end of user’s speech to the beginning of audio-facial playback.
TABLE IV: Latency definitions.
Method
LMSA ( ↓ )
LES ( ↓ )
Lrobot ( ↓ )
LAF ( ↓ )
Streaming Multimodal Orchestration (Ours)
3.306
1.906
0.003
5.214
Serial Embodiment Processing
3.227
10.621
0.003
13.851
TABLE V: Results for System latency evaluation. Values are reported as mean.
Fig. 5: Effect of gesture library size on user-rated motion quality. Library size is measured by the number of semantic keyframes. The scores increase substantially as the library grows from 8 to 37 keyframes, while the gain becomes marginal when further expanded to 45 keyframes.