Organizations: Signal Analysis and Interpretation Lab (SAIL), University of Southern California, USA · Center for Language and Speech Processing, Johns Hopkins University, USA
Neural codecs encode continuous signals into compact sequences of discrete tokens, providing an interface for efficient transmission, storage, and token-based sequence modeling. This paradigm has been widely adopted in modern speech and audio frameworks; however, the biosignal domain still lacks a neural codec designed specifically for low-bitrate streaming and generalization across diverse downstream tasks. We present MyoCodec, a streaming neural codec designed for electromyography (EMG). Inspired by recent neural audio codecs, MyoCodec combines causal Transformers with residual vector quantization to encode continuous EMG signals into different levels of EMG representations spanning from continuous latent features to discrete tokens operating at 50 Hz. Trained on twelve public EMG datasets, MyoCodec achieves favorable performance in both intrinsic codec quality and representative downstream tasks, including typing (emg2qwerty), hand-pose (emg2pose), speech decoding (emg2speech), and speech-to-EMG synthesis (speech2emg). Across these tasks, MyoCodec exhibits strong performance against prior models while providing a compact and causal EMG representation. During streaming inference, it requires compute time of only 0.482 ms for each 20 ms frame, enabling real-time streaming. Also, the discrete token representation provided by MyoCodec has the potential to support integration into language-model based approaches, creating a path toward LLM-based interactive systems, where tokenized EMG representations are directly processed into such language or speech models. Code and model weights are released.
Figures & tables
Model
Params
Hz
Nq
bitrate/ch
reconstruction ↑
codebook
SI-SDR
SNR
rwav
renv
util.
perp.
TinyMyo
3.6M
100.0
—
≤ 614k †
0.46
1.70
0.735
0.445
—
—
BioCodec
12.3M
27.8
6
1333
−1.51
2.20
0.643
0.907
1.000
203.4
MyoCodec (Ours)
12.7M
50.0
2
800
3.62
4.99
0.832
0.933
1.000
172.1
3
1200
5.89
6.75
0.881
0.949
1.000
177.8
6
2400
9.80
9.98
0.933
0.967
1.000
186.0
Table 1: Intrinsic codec quality. Reconstruction quality is evaluated using SI-SDR and SNR (dB), waveform Pearson correlation rwav , and envelope Pearson correlation renv . Nq is the number of codebooks used. MyoCodec outperforms TinyMyo and BioCodec in signal reconstruction by using only two codebooks ( 800bits/s ).
CB 1
CB 2
CB 3
CB 4
CB 5
CB 6
mean
Utilization
1.000
1.000
1.000
1.000
1.000
1.000
1.000
Perplexity
154.9
189.2
189.2
192.6
194.1
196.0
186.0
Table 2: Per-codebook utilization and perplexity for MyoCodec . All six codebooks are fully utilized, with higher perplexity in later RVQ stages.
EMG Representation
Hz
bitrate/ch
CER (%) ↓
greedy
LM beam
seen
unseen
seen
unseen
Log-spectrogram ( Sivakumar et al., 2024 )
125.0
≤ 132k †
24.02
52.98
16.86
48.57
TinyMyo ( Fasulo et al., 2026 )
100.0
≤ 614k †
21.16
50.02
14.48
44.55
BioCodec ( Avramidis et al., 2025 )
27.8
1.3k
18.52
51.87
13.07
47.78
MyoCodec (Ours)
50.0
2.4k
18.43
50.47
12.74
45.77
Table 3: emg2qwerty. Character error rate (CER) is reported with and without language-model beam search, across seen and unseen subjects. The best and second-best are bolded and underlined , respectively.
EMG encoder
joint-angle error ( ∘ ) ↓
land. (mm) ↓
fingertip (mm) ↓
val
test
Log-spectrogram ( Salter et al., 2024 )
13.13
15.23
20.76
35.07
TinyMyo ( Fasulo et al., 2026 )
13.15
14.83
20.13
33.97
BioCodec ( Avramidis et al., 2025 )
13.11
14.75
20.12
34.00
MyoCodec (Ours)
12.75
14.24
19.04
32.02
Table 4: emg2pose. Mean joint-angle error (degrees) and mean landmark and fingertip distances (mm).
EMG Encoder
CER ↓
WER ↓
PER ↓
PFER ↓
UTMOS ↑
Ground truth
1.43
3.76
7.13
4.19
3.41
Gaddy & Klein ( Gaddy and Klein, 2021 )
22.95
38.49
31.62
14.98
1.90
TinyMyo ( Fasulo et al., 2026 )
16.39
27.77
26.47
13.29
1.91
BioCodec †
16.56
28.33
24.63
12.59
1.93
MyoCodec (Ours)
12.15
21.12
21.61
11.54
2.00
Table 5: emg2speech results. We report character, word, and phoneme error rates (CER, WER, and PER), phone feature error rate (PFER), and UTMOS. All of the downstream models are causally trained with 100ms of lookahead.
Env. CC ↑
System
lookahead
50ms window
20ms window
STE-GAN ( Scheck and Schultz, 2023 )
∞ (offline)
0.705± 0.014
0.578± 0.012
TinyMyo ( Fasulo et al., 2026 )
∞ (offline)
0.577± 0.011
0.474± 0.010
BioCodec ( Avramidis et al., 2025 )
0ms
0.511± 0.010
0.434± 0.008
144ms
0.600± 0.009
0.522± 0.008
MyoCodec (Ours)
0ms
0.587± 0.010
0.503± 0.009
Table 6: speech2emg. Envelope correlation coefficients (Env. CC) with 95% confidence intervals are reported using 20ms and 50ms smoothing windows. We compare offline and causal systems with different lookahead. The best and second-best are bolded and underlined , respectively.
Compute time (ms) ↓
Real-time throughput ( × ) ↑
System
basis
Encoder
Decoder
Total
Encoder
Decoder
Total
MyoCodec (Ours)
per second
0.782
0.692
1.474
1279
1445
678
per frame
0.253
0.229
0.482
79
87
41
BioCodec
per second
1.091
0.941
2.032
917
1062
492
per frame
N/A †
Table 7: Inference speed. All measurements are obtained on a single NVIDIA RTX PRO 6000 Blackwell GPU. Per-second speed is measured using 5s input windows and normalized by signal duration, while MyoCodec ’s per-frame speed is measured using its compiled CUDA Graph implementation.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Corpus
ch.
subj.
ch.-hours
content
emg2qwerty ( Sivakumar et al., 2024 )
32
108
11,074
typing
emg2pose ( Salter et al., 2024 )
16
193
5,925
hand pose
Hyser ( Jiang et al., 2021 )
256
20
2,593
finger force
putEMG ( Kaczmarek et al., 2019 )
24
44
1,045
hand gestures
EMG-EPN-612 ( Benalcázar et al., 2020 )
8
612
1,014
hand gestures
MeganePro ( Cognolato et al., 2020 )
12
45
610
grasping
Appendix
Table 8: The twelve public EMG corpora. Channel-hours count only the channels and recordings retained after filtering.
Neural audio codecs are a fundamental component of modern speech generation systems. While recent codecs achieve increasingly low bitrates, reducing frame rate remains challenging, as each token must preserve more information while maintaining reconstruction quality. We present ZipCodec, a streaming neural speech codec operating at 6.25 Hz and 0.80 kbps with a theoretical latency of 160 ms. Our approach combines large-scale WavLM distillation with a redesigned transformer-based architecture, a scalar spherical quantizer, and a latency-aware streaming decoder. Experiments show that ZipCodec substantially outperforms existing streaming codecs at comparable bitrates in both reconstruction and downstream tasks, while operating at a significantly lower frame rate. Despite its 842M parameters, ZipCodec achieves real-time single-stream inference on a consumer-grade CPU. Demo samples, code and checkpoints are available at https://lucadellalib.github.io/zipcodec-web/.
Luca Della Libera, Cem Subakan, Mirco Ravanelli
Concordia University · Mila-Quebec AI Institute · Universit´e Laval
Neural speech codecs provide discrete representations for speech language models, but emotional cues are often degraded during quantization. Existing codecs mainly optimize acoustic reconstruction, leaving emotion expressiveness insufficiently modeled at the representation level. We propose an emotion-guided neural speech codec that explicitly preserves emotional information while maintaining semantic fidelity and prosodic naturalness. Our framework combines emotion-semantic guided latent modulation, relation-preserving emotional-semantic distillation, and emotion-weighted semantic alignment to retain emotionally salient cues under compression. Extensive evaluations across speech reconstruction, emotion recognition, and downstream text-to-speech generation demonstrate improved emotion consistency and perceptual quality without sacrificing content accuracy.
Jiacheng Shi, Hongfei Du, Xinyuan Song +3
College of William & Mary · Emory University · George Mason University
Neural speech codecs have become the discrete interface between raw audio and speech language models, yet they remain optimized primarily for acoustic reconstruction fidelity, which leaves emotion-relevant cues vulnerable to being discarded during quantization, limiting the affective capacity of downstream models. We trace this degradation to two mechanisms: reconstruction-driven bit allocation under limited bitrate and cross-stream leakage in concatenation-based codecs, where acoustic gradients can overwrite nominally emotion-reserved dimensions. We propose AffectCodec, an emotion-preserving neural speech codec built on Block-Diagonal Residual Finite Scalar Quantization (BD-RFSQ). By imposing block-diagonal input and output projections over emotion and acoustic subspaces, BD-RFSQ transforms bit allocation from implicit and loss-driven to explicit and structurally guaranteed, while still preserving a flat token interface for downstream speech language models. AffectCodec further combines this structurally constrained quantizer with multi-granularity emotion conditioning and multi-rate training, enabling robust affect preservation at low bitrates. Experiments across multiple emotional speech benchmarks show that AffectCodec substantially improves emotion preservation, especially in the low-bitrate regime, while maintaining competitive acoustic quality and intelligibility. These results suggest that structurally protected quantization is an effective principle for preserving emotion-relevant information and may provide a general route toward attribute-aware neural speech compression.
Zhaoyang Meng, Zhengyao Ma, Kecan Mao +2
Beijing University of Posts and Telecommunications