Facial movements convey subtle and important information that is critical for human social communication. Optical methods for face capture are difficult or impossible to use when the face is occluded by head-mounted devices (HMDs), such as VR headsets. Even with a clear line of sight, such methods raise privacy concerns and require head-mounted capture rigs that offset cameras and lighting from the face. We show that high-density surface electromyography (HD-sEMG) provides a viable non-optical alternative that addresses these challenges. We measured 64 EMG channels, using two textile EMG grids, with 32 from the forehead (typically occluded by an HMD) and 32 from the side of the face. EMG data were digitized at 2048 Hz and filtered. Facial movements were simultaneously recorded and used to estimate 478 3D facial landmarks using MediaPipe's Face Landmarker. A major challenge in such multimodal recordings is synchronizing EMG and video data, which have different sampling frequencies and independent clocks. We developed a novel synchronization method using analog audio bursts that is capable of sub-millisecond synchronization. We also developed a staged fitting method that fits a recent high-resolution parametric head model (GNM), with 253 identity blendshapes and 383 expression blendshapes, to the MediaPipe landmarks as participants performed different facial expressions. We trained a deep neural network comprising per-grid spatial encoders followed by a dilated temporal convolutional network (TCN) to predict blendshape parameters from HD-sEMG signals at 100 Hz. Once trained, the network can predict expression blendshapes solely from HD-sEMG recordings. The output can be rendered using standard real-time blendshape animation methods. We demonstrate the methods using recordings from 25 participants, and direct expression transfer to a variety of human faces and non-human characters.
Figures & tables
Figure 1: (a) Capture setup. Left: the participant sits facing a monitor that displays instructions for each task, while a front camera records the face, which is lit by light boxes. Right: seatbelts attached to the chair stabilize the torso, and the EMG amplifiers are placed behind the chair. (b) Electrode grids placed on a participant’s face (bottom right), with a visualization of the raw EMG signals recorded by the forehead grid (top) and the cheek grid (left). The numbers show the channel at each electrode.
Figure 2: Synchronization and segmentation. (a) Tones detected in the audio track of the video. Top: the three tones that mark one trial. Below: each onset at millisecond scale, showing the raw audio sampled at 44.1 KHz, a band-passed copy (gray), and the detected onset (dashed). Note that the first cycle has some artifacts, likely due to the camera’s gain control hardware, but our method is immune to such artifacts. (b) The same tones recorded on an auxiliary channel of the EMG amplifier, sampled at 2048 Hz. (c) One recording of eight expressions, preceded by saccade trials marked by start tones only: tone envelopes in the camera audio (top) and in the EMG channel (bottom), and the timeline of events recovered from them (middle). Note that the detected onset marks the same phase of the first cycle in both recordings, so any fixed bias cancels in the offset between the two clocks.
Figure 3: Staged fitting of GNM to MediaPipe landmarks, for a frame in which the participant blinks while speaking. Spheres show the landmarks, colored by their distance from the fitted mesh (0–5 mm). Left: stage 1 fits only the rigid pose of the head. Middle: stage 2 adds eyelid closure and gaze, which fits the closed eyes. Right: stage 3 adds the remaining expression blendshapes, which fits the open mouth and reduces the residuals over most of the face.
Figure 4: Network architecture. Each grid is a 4×8 image of EMG envelopes at every frame, and has its own spatial encoder of two 3×3 convolutions followed by average pooling. The pooled features of both grids are concatenated and projected to 128 features per frame. A dilated temporal convolutional network (TCN) then combines 2.5 s of context centered on each frame, and two 1×1 convolutions map the result to the 383 GNM expression weights. Note that the dilations double with depth, so six layers of kernel size 3 span the whole context.
Figure 5: A closed eyelid before (left) and after (right) the correction, on the same frame. Left: the eyeball emerges through the lid, which penetrates it by up to 3.4 mm. Right: after Equation 1 , nothing penetrates, and the margin of the lid reads as a closed crease.
Figure 6: Prediction performance across participants, on the held-out cued expressions. Each box shows the distribution over the 25 participants, with open circles marking the seven recordings of lower signal quality. Left: correlation between predicted and fitted blendshape weights, for the eye-region blendshapes of GNM (in GNM, the eye region includes the forehead) and for the rest of the face (higher is better). Right: mean vertex error of the reconstructed mesh, over the whole face, in the eye region, and elsewhere, next to the error of holding each participant’s mean face (gray; lower is better). Note that the error is about half that of the mean face in every region.
Figure 7: Predicted (solid) and fitted (dashed) facial motion during four held-out trials of one participant, whose accuracy on these expressions is close to the median over all participants. Each trace is the displacement of the face along the main mode of motion of that trial, and r is computed over the whole trial. For most participants brow raises, eye closure, and blinks are tracked closely, including their timing, while the lip and jaw motion of speech is underestimated, as Büchner et al. [ BAGLD25 ] also observed.
Figure 8: Where the errors are. Top: mean vertex error (mm) on the held-out cued expressions, averaged over the 25 participants and shown on the GNM template from the front and from both sides. Bottom: the same error relative to the error of holding each participant’s mean face (lower is better). Vertices that expression cannot move are gray. Note that the relative error is lowest at the brows and the cheeks and highest at the chin, and that both maps are nearly symmetric.
Figure 9: Decoding facial expressions with the face covered by a sleep mask covering the forehead. (a) Upper face occluded. (b) GNM faces at the peak of one brow raise under the mask, fitted to the landmarks tracked by MediaPipe and predicted from HD-sEMG, colored by displacement from rest. (c) Brow height over the same trial; the dashed line and gray band show the median and range of the participant’s 16 brow raises without the mask, and the dotted line the rest cue. Note that the tracked brows rise only half as far, while the predicted brows rise as far as without the mask.
Figure 10: Decoding from the forehead grid alone. (a) Fraction R2 of the variance of facial motion explained for each participant, with both grids and with the forehead grid alone, over the whole face, the eye region (which in GNM includes the forehead), and the rest of the face. Open circles mark the seven recordings of lower signal quality. (b) Change in R2 over the frames of each cued expression when the cheek grid is removed (median and interquartile range). (c) Mean vertex error with the forehead grid alone, relative to that with both grids, from the subject’s right, the front, and the subject’s left. Note that the error grows by a similar fraction over the whole face.
Figure 11: Clenching the teeth, an expression that barely moves the face. One participant clenched the teeth (orange) and smiled (blue) in seven repetitions each. (a) EMG of the cheek grid: envelope of the differences between neighboring electrodes, median over the pairs of electrodes. Lines show the median over repetitions and bands their range; the movement cue is at 0 s and the rest cue at the dotted line. Panels (b)–(d) are relative to the second before the cue. (b) Motion of the fitted GNM face: mean distance between corresponding vertices. (c) EMG of each electrode during the hold, laid out as in Figure 16 f (Appendix A ); the two faulty electrodes (gray) are excluded. (d) Displacement of the fitted face during the hold. Note that the clench produces strong EMG, in a pattern of its own, while the face hardly moves.
Figure 12: The same expression weights applied to nine GNM heads: the mean head (center, outlined) and eight other identities.
Figure 13: Expressions transferred to non-human characters. Expression weights from the GNM head (left) are transferred to multiple non human characters: a squirrel (middle) and a stylized character (right). Current limitations on registration for extremely different facial types result in artifacts for the squirrel character while smiling.
Figure 14: Without (left) and with (right) expression-driven wrinkles, on the same frame. Observe the highlighted region, as the subject raises their eyebrows furrows are naturally created along the forehead, through Equation 2 .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 15: Closure of the eye at the deepest fitted closure of each held-out blink and eye closure, fitted against predicted at the same frame. Trials in which the fitted eye closes by less than 3 mm are omitted, and open circles mark the seven recordings of lower signal quality. Note that most of the blinks that the prediction misses come from these recordings.
Figure 16: Processing of the raw EMG, for one forehead electrode during a held brow raise. The movement cue is at 0 s and the rest cue at the dotted line. (a) The raw signal is dominated by a large offset and a slow drift. (b) Band-pass and notch filtering isolate the muscle activity. (c) Rectification and low-pass filtering produce the envelope (blue). (d) Resampled to 100 Hz and log-compressed, the envelope is the input to the network. (e) Power spectra of the same electrode: filtering removes the low-frequency drift, which dominates the raw signal, and the mains interference at 60, 120, and 180 Hz. (f) Network input of all 64 electrodes on the forehead and cheek grids, averaged over rest and over the movement; the outlined electrode is the one shown in (a)–(d). Note that the activity rises over the forehead grid, while the cheek grid changes little.
Head-mounted displays (HMDs) fundamentally limit emotion recognition in virtual reality (VR): by occluding the upper face, they render conventional image-based facial expression analysis incomplete, particularly for applications requiring real-time affective assessment. We address this challenge by fusing lower-face video with facial electromyography (EMG) from the occluded upper face to classify seven emotional categories (six basic emotions plus neutral). We introduce a synchronized multimodal dataset from 20 participants, pairing lower-face video with seven-channel upper-face EMG elicited by validated emotion stimuli. Under subject-independent test, our proposed late-fusion architecture merging convolutional visual embeddings with RBF-kernel EMG representations achieves 51% macro-F1, outperforming both image-only (41%) and EMG-only (43%) baselines. These results demonstrate that upper-face EMG provides robust complementary information under HMD-induced visual occlusion and establish a foundation for multimodal emotion recognition in naturalistic VR environments. This approach facilitates affect-adaptive applications, including communication training and therapeutic interventions. The dataset will be shared upon request under an ethical-use agreement.
Birgit Nierula, Karam Tomotaki-Dawoud, Mert Akguel +5
Fraunhofer HHI, Berlin, Germany · Humboldt University of Berlin, Berlin, Germany
Facial expression recognition (FER) is crucial for social interaction in mixed reality environments that employ head-mounted displays (HMD). However, collecting FER data from head-mounted cameras (HMC) is challenging due to privacy concerns and the diversity of HMD platforms. Moreover, existing FER datasets are not directly applicable due to the unique perspectives of HMCs. The lack of sufficient data hinders the development of neural network-based HMC FER methods. To address data scarcity, we propose a data synthesis framework that generates HMC-view images from frontal-view images, leveraging abundant existing annotated datasets. Specifically, we first reconstruct 3D textured meshes from images and then apply a configurable camera system to render images from the HMC perspective. Additionally, we introduce a texture-space alignment network (TSAN) that enables accurate texture sampling from images to preserve detailed facial expressions. To evaluate the proposed method, we conduct extensive experiments on both simulated and real HMC datasets. Experimental results demonstrate that models trained on our synthetic dataset outperform those trained on existing datasets and exhibit better generalization across different camera configurations.
We present RegHead, a framework for constructing semantic blendshape sets for animatable non-humanoid head avatars. With a fixed expression vocabulary, semantic blendshapes provide a low-dimensional and interpretable animation interface and support cross-identity retargeting. Building such blendshape sets remains expensive because (i) expression-consistent supervision is scarce, (ii) generated 4D assets typically lack correspondence, and (iii) facial motion is highly localized. We propose (1) a large-scale dataset of non-humanoid identities paired with a shared expression vocabulary, obtained by expanding a small artist-rigged library via fine-tuned image editing; (2) a dense stochastic anchor motion representation tailored to localized facial deformations; and (3) a fast feed-forward registration model that converts unregistered expression meshes into a corresponded blendshape basis by predicting anchor-based deformations from the neutral shape. Experiments show that our approach produces higher-fidelity expression meshes than baselines, while running orders of magnitude faster than optimization. We further demonstrate real-time retargeting from human face tracking signals to non-humanoid characters, capturing both head pose and localized facial motions. Our project page is available at https://snap-research.github.io/RegHead/.