GLARE: Generating Listening Heads with Appropriate Reactions
Organizations: Department of Computer Science, Stony Brook University · Atmanity Inc.
Abstract
While talking head generation has advanced rapidly, generating natural listener behavior in dyadic conversations, which know when to react, how to react, and with what type of response, remains underexplored. Existing dyadic datasets lack fine-grained listener reaction annotations, and prevailing evaluation metrics inherited from talking-head and video generation measure visual realism rather than whether a listener reacted appropriately. We address these gaps along three aspects. First, we curate a listening-head-specific dataset built from RealTalk and Seamless Interaction, comprising approximately 147 hours of paired speaker-listener videos with 64,557 event-level reaction annotations across six categories: nodding, head shaking, smiling, laughing, frowning, and surprised. Second, we introduce an audio-driven baseline built on a flow-matching transformer, namely GLARE, with prosody conditioning derived from Qwen2-Audio and a temporal reaction loss that explicitly supervises frame-wise reactions. Third, we propose a reaction-oriented evaluation protocol that jointly measures reaction occurrence (R-F1), temporal alignment (R-tIoU), asymmetric temporal deviation (R-ATD), and reaction-region visual quality (R-FID), giving a more behaviorally grounded assessment than visual-quality-only metrics. Experiment results show consistent gains over prior listening-head methods in both visual fidelity and reaction-level metrics, suggesting that reaction-aware data, modeling, and evaluation are critical for natural listening behavior.
Figures & tables
| Reaction type | nodding | head shaking | smiling | laughing | frowning | surprised | Total |
| Count | 10,230 | 12,790 | 15,767 | 8,320 | 9,986 | 7,464 | 64,557 |
| Dataset | Method | PSNR | SSIM | FID | FVD | Var | LPIPS | rPCC | DI-Sync | R-F1 | R-tIoU | R-ATD | R-FID |
| RealTalk | L2L | 14.684 | 0.575 | 45.782 | 202.798 | 1.885 | 0.637 | 0.323 | 0.182 | 0.454 | 0.505 | 78.265 | 24.683 |
| DIM | 16.223 | 0.495 | 37.717 | 188.266 | 2.765 | 0.585 | 0.289 | 0.190 | 0.576 | 0.551 | 133.082 | 22.971 | |
| ViCo | 14.932 | 0.602 | 44.089 | 185.040 | 2.560 | 0.579 | 0.261 | 0.178 | 0.429 | 0.572 | 92.454 | 26.105 | |
| ListenFormer | 17.454 | 0.582 | 36.173 | 165.290 | 1.625 | 0.525 | 0.256 | 0.221 | 0.334 | 0.650 | 73.379 | 23.097 | |
| DyStream | 17.894 | 0.611 | 37.416 | 147.537 | 2.802 | 0.467 | 0.248 | 0.208 | 0.535 | 0.694 | 63.171 | 15.337 | |
| Ours | 17.972 | 0.601 | 35.697 | 142.454 | 2.916 | 0.454 | 0.227 | 0.245 | 0.594 | 0.704 | 57.388 | 15.192 |
| Dataset | Reaction | Data Amount | R-F1 | R-tIoU | R-ATD | R-FID | ||||
| Dystream | Ours | Dystream | Ours | Dystream | Ours | Dystream | Ours | |||
| RealTalk | nodding | 3033 | 0.562 | 0.602 | 0.672 | 0.686 | 66.168 | 60.375 | 15.114 | 14.680 |
| head shaking | 3588 | 0.521 | 0.581 | 0.655 | 0.652 | 67.447 | 61.894 | 15.626 | 15.316 | |
| smiling | 4804 | 0.554 | 0.633 | 0.718 | 0.753 | 60.396 | 54.882 | 14.883 | 14.371 | |
| laughing | 2397 | 0.569 | 0.656 | 0.735 | 0.769 | 57.026 | 48.189 | 16.005 | 15.697 | |
| frowning | 2745 | 0.511 | 0.549 | 0.705 | 0.707 | 61.518 | 54.793 | 15.010 | 15.024 | |
| Reaction | Reaction Naturalness | Contextual Appropriateness | Timing Plausibility |
| nodding | 0.954 | 0.975 | 0.982 |
| head shaking | 0.969 | 0.914 | 0.937 |
| smiling | 0.897 | 0.868 | 0.944 |
| laughing | 0.852 | 0.836 | 0.955 |
| frowning | 0.902 | 0.874 | 0.884 |
| surprised | 0.824 | 0.898 | 0.905 |
| Reactions | Prosody Cond. | R-F1 | R-tIoU | R-ATD | R-FID |
| nodding | 0.422 | 0.626 | 84.732 | 23.284 | |
| 0.474 (+12.32%) | 0.583 | 97.428 (+14.98%) | 21.032 | ||
| head shaking | 0.398 | 0.584 | 91.645 | 24.226 | |
| 0.441 (+10.08%) | 0.539 | 103.763 (+13.22%) | 21.445 | ||
| smiling | 0.421 | 0.640 | 82.519 | 22.571 | |
| 0.517 (+22.80%) | 0.595 | 93.314 (+13.08%) | 19.164 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Source | Raw scale | Resolution | Role in our dataset |
| RealTalk | 694 videos | 1080P | Dyadic conversational footage for speaker–listener pairing |
| Seamless Interaction | 65k+ videos | 4K | Large-scale dyadic audiovisual data for diverse interactions |
| Curated dataset | 107,149 pairs | 512 512 crops | 147 hours with 64,557 reaction instances |
| Reaction | Primary visual cues | Temporal rule | Post-processing |
| Nodding | Vertical head-center displacement, pitch-related dynamics, velocity and acceleration patterns | Peak/valley cycles with vertical zero-crossings and sufficient normalized amplitude | Smooth scores, me-rge overlapping intervals, threshold at 0.5 |
| Head shake | Horizontal head-center displacement, yaw-related dynamics, horizontal velocity and acceleration patterns | Alternating left-right extrema with horizontal zero-crossings and sufficient normalized amplitude | Smooth scores and merge adjacent intervals |
| Smiling | Happy-expression confidence, smile-related action units, lip-corner movement | Sustained high smile confidence over consecutive frames | Remove isolated p-eaks and merge sh-ort gaps |
| Laughing | Smile confidence, mouth opening, high-intensity happy-expression cues, stronger facial dynamics | Sustained expression with lar-ger mouth/facial motion than ordinary smiling | Merge nearby high-conf-idence intervals |
| Frowning | Negative-expression confidence, Brow-lowering cues (AU04), mouth-corner depression cues | Sustained negative facial expression over a short temporal window | Remove brief neutral fluctuations |
| Surprised | Surprise-expression confidence, eyebrow raising, eye opening, mouth opening | Short high-confidence peaks or short sustained surprise intervals | Allow shorter eve-nts than other expression classes |
| Item | Setting |
| Training precision | Mixed precision |
| Distributed training | Accelerate |
| GPUs | 4 L40s |
| Optimizer | AdamW |
| Learning rate | |
| Learning-rate schedule | Cosine decay with warmup |
| R-ATD | |||||
| 2.0 | 0.5 | 22.864 | 17.312 | 17.212 | 57.388 |
| 1.0 | 1.0 | 18.120 | 14.615 | 14.315 | 47.050 |
| 0.5 | 2.0 | 22.436 | 19.238 | 18.563 | 60.237 |
| Audio Length (s) | Metrics | |||||||||||
| PSNR | SSIM | FID | FVD | Var | LPIPS | rPCC | DI-Sync | R-F1 | R-tIoU | R-ATD | R-FID | |
| 2 | 16.684 | 0.584 | 37.780 | 164.74 | 2.686 | 0.496 | 0.212 | 0.224 | 0.577 | 0.678 | 68.697 | 17.126 |
| 3 | 16.987 | 0.592 | 36.158 | 152.08 | 2.874 | 0.492 | 0.236 | 0.242 | 0.588 | 0.684 | 64.563 | 16.554 |
| 5 | 17.371 | 0.599 | 35.796 | 144.32 | 2.952 | 0.466 | 0.201 | 0.237 | 0.596 | 0.695 | 60.285 | 15.724 |
| 10 | 17.973 | 0.601 | 35.691 | 142.456 | 2.913 | 0.454 | 0.227 | 0.245 | 0.594 | 0.704 | 57.386 | 15.192 |
| 20 | 17.884 | 0.597 | 35.824 | 147.67 | 2.949 | 0.444 | 0.233 | 0.240 | 0.590 | 0.712 | 58.597 | 16.055 |
| Channel number | Metrics | |||||||||||
| PSNR | SSIM | FID | FVD | Var | LPIPS | rPCC | DI-Sync | R-F1 | R-tIoU | R-ATD | R-FID | |
| 1 | 17.797 | 0.596 | 35.704 | 148.925 | 2.783 | 0.475 | 0.235 | 0.242 | 0.602 | 0.689 | 70.474 | 16.345 |
| 2 | 18.021 | 0.603 | 35.832 | 144.796 | 2.928 | 0.457 | 0.234 | 0.233 | 0.582 | 0.695 | 60.178 | 16.502 |
| 4 | 17.973 | 0.601 | 35.691 | 142.456 | 2.913 | 0.454 | 0.227 | 0.245 | 0.594 | 0.704 | 57.386 | 15.192 |
| 8 | 18.792 | 0.584 | 35.989 | 145.776 | 2.862 | 0.460 | 0.230 | 0.247 | 0.582 | 0.699 | 59.404 | 15.220 |
| 16 | 18.464 | 0.581 | 35.794 | 146.435 | 2.848 | 0.457 | 0.211 | 0.250 | 0.556 | 0.704 | 58.691 | 15.348 |
| Channel number | Metrics | |||||||||||
| PSNR | SSIM | FID | FVD | Var | LPIPS | rPCC | DI-Sync | R-F1 | R-tIoU | R-ATD | R-FID | |
| BCE | 17.695 | 0.597 | 35.702 | 159.266 | 2.884 | 0.461 | 0.240 | 0.249 | 0.598 | 0.712 | 59.185 | 17.996 |
| Smooth L1 | 17.973 | 0.601 | 35.691 | 142.454 | 2.913 | 0.454 | 0.227 | 0.245 | 0.594 | 0.704 | 57.386 | 15.192 |