Electromyography-to-Speech (ETS) synthesis is typically a non-linear, chaotic dynamical system. However, no prior work has studied the chaotic behavior of ETS synthesis to date. Yet, prior works strictly rely on standard reconstruction metrics with parameter-heavy transformers that systematically over-smooth natural acoustic dynamics. To close this gap, for the first time, we propose a chaos-inspired Lyapunov-derived activation function (LyRIC) with two novel chaotic loss functions, Lyapunov Exponent Regularization and Multi-Scale Detrended Fluctuation Analysis, to explicitly capture the deterministic chaos of human phonation. In addition, we introduce a compressed novel encoder, GLAME, which synergizes global Mamba state-space modeling with localized GLARE attention. We comprehensively perform frame-level acoustic evaluation in a multilingual and multi-speaker setup using English and Mandarin datasets. The proposed system outperforms the established baseline with a 4.69x increase in objective intelligibility (STOI: 0.61 vs. 0.13) and a 2.08x improvement in spectral reconstruction (LSD: 1.08 vs. 2.25). Importantly, this improvement is achieved with 73.49% fewer parameters (14.34M vs. 54.10M), establishing a new baseline for ETS synthesis. To the best of our knowledge, this is the first work demonstrating that integrating non-linear chaotic physics into neural networks yields superior yet compact inductive biases for real-time ETS synthesis.
Figures & tables
Figure 1: Detailed architecture of CGS-ETS with chaos-driven LyRIC activation and GLARE attention.
Figure 2: Architecture of GLARE attention.
Figure 3: (Left) Comparing different activation functions (Eqn. 2 ). (Right) Effect of α on log-divergence.
Figure 4: LyRIC activation with α∈{−1,−1.5} .
Function
FLOPs/element ↓
P0
α=0.5
11
P1
α=1
1
P2
α=1.5
12
P3
α=2
2
P4
α=2.5
13
P5
α=3
3
Table 1: FLOPs footprint analysis of hyperparameter α .
Figure 5: (Left) Comparison of different activation functions. (Right) Effect of α on LyRIC activation.
Model
Cr
L ↓
ST ↑
P ↑
SD ↑
N ↑
W ∗↓
C ↓
Pr ↓
P0
Baseline 1
EN
3.23
0.04
1.17
-52.5
3.28
68.00
38.2
60.21
P1
Baseline 1
ZH
3.50
0.03
1.05
-58.1
1.31
×
59.1
60.21
P2
Baseline 2
EN
2.25
0.13
1.18
-42.0
3.30
42.20
23.7
54.10
P3
Baseline 2
ZH
2.67
0.10
1.08
-39.5
1.84
×
39.8
54.10
P4
Baseline 3
EN
×
×
×
×
×
46.14
×
×
P5
Baseline 4
ZH
×
×
×
×
×
×
38.0
×
Table 2: Comparison against SOTA methods. Here, Cr=Corpus, EN=English, ZH=Mandarin, L=LSD, ST=STOI, P=PESQ, SD=SI-SDR, N=NISQA-MOS, W=WER %, C=CER %, Pr=Parameters in Millions. ∗ As Mandarin text lacks spaces, corpora rely on CER rather than WER. P4 and P5 codebases are not public.
Model
CFE ↓
Enc ↓
Total ↓
IT ↓
Fs ↓
RTF ↓
Gaddy and Klein (2021)
10.17
44.03
54.10
5.97
2400
0.0057
Proposed CGS-ETS
4.49
9.85
14.34
5.61
926.6
0.0057
Table 3: Computational efficiency. Here, Enc=Encoder Size, Total = Total Parameters in Millions, ITs = Inference Time in Milliseconds, and Fs = FLOPs in Millions.
Model Config.
L ↓
ST ↑
P ↑
SD ↑
N ↑
W ↓
Pr ↓
P1
w/o LER+MSDFA
1.14
0.60
1.12
-35.5
3.25
51.56
14.34
P2
w LER w/o MSDFA
1.10
0.60
1.18
-33.7
3.29
48.34
14.34
P3
w MSDFA w/o LER
1.09
0.60
1.19
-32.6
3.31
45.65
14.34
P4
w LER+MSDFA
1.08
0.61
1.19
-32.6
3.33
41.36
14.34
Table 4: Ablation study on chaotic losses. Here, metrics acronyms are provided in Tables 2 and 3 .
CFE
Mam.
FFN
SD ↑
N ↑
W ↓
IT ↓
Fs ↓
RTF ↓
P0
ReLU
SiLU
SwiGLU
-31.5
3.25
51.16
7.12
1038
0.0075
P1
ReLU
SiLU
GeLU
-31.3
3.28
47.06
5.92
929.8
0.0059
P2
LyRIC
SiLU
GeLU
-33.2
2.90
71.86
5.92
929.8
0.0059
P3
ReLU
SiLU
LyRIC
-31.9
3.23
52.63
5.70
928.4
0.0057
P4
ReLU
LyRIC
SwiGLU
-32.6
3.30
45.71
6.21
1035
0.0062
P5
ReLU
LyRIC
GeLU
-32.6
3.33
41.36
5.61
926.6
0.0057
Table 5: Ablation on activations. Here, metrics acronyms are provided in Tables 2 and 3 .
Layers
SD ↑
N ↑
W ↓
Pr ↓
IT ↓
Fs ↓
RTF ↓
P0
None
-34.4
2.91
73.22
12.68
4.23
848.3
0.0046
P1
L=1,2
-32.4
3.30
46.14
14.34
5.91
926.6
0.0059
P2
L=3,4
-33.0
3.25
50.25
14.34
5.97
926.6
0.0059
P3
L=1,4
-33.5
3.30
49.01
14.34
5.91
926.6
0.0061
P4
L=1,2,3
-33.3
3.26
46.38
15.05
6.86
966.5
0.0068
P5
L=2,3,4
-32.8
3.28
44.91
15.05
7.11
966.5
0.0071
Table 6: Ablation study on GLARE attention on different layer depths. Metrics LSD, STOI, and PESQ showed negligible variance, hence are omitted here. Metrics acronyms are provided in Tables 2 and 3 .
α
L ↓
ST ↑
P ↑
SD ↑
N ↑
W ↓
IT ↓
Fs ↓
RTF ↓
P0
0
1.09
0.60
1.18
-34.1
3.30
48.22
5.98
927.0
0.0058
P1
1.5
1.09
0.61
1.18
-30.8
3.26
44.73
5.95
929.9
0.0060
P2
2
1.08
0.60
1.18
-32.3
3.30
45.40
5.95
927.2
0.0059
P3
3
1.09
0.60
1.18
-33.0
3.30
49.26
6.00
927.3
0.0059
P4
1
1.08
0.61
1.19
-32.6
3.33
41.36
5.61
926.6
0.0057
Table 7: Ablation study on hyperparameter α of LyRIC activation. Acronyms are provided in Tables 2 and 3 .
Attention Type
SD ↑
N ↑
W ↓
Enc ↓
IT ↓
Fs ↓
RTF ↓
P0
MHA, All Layers
-32.4
3.30
43.32
14.49
7.42
1166
0.0074
P1
MHA, L=1,3
-32.4
3.27
44.73
12.39
5.92
1061
0.0059
P2
MQA, All Layers
-31.7
3.33
44.85
10.56
7.10
966.5
0.0071
P3
MQA, L=1,3
-31.8
3.31
44.90
9.38
5.69
907.4
0.0058
P4
GLARE, L=1,3
-32.6
3.33
41.36
9.85
5.61
926.6
0.0057
Table 8: Ablation study on different attention mechanisms. Here, Enc=Encoder Size in Millions. Other metrics acronyms are provided in Tables 2 and 3 .
Agg.
Model
n
Mean ↑
SDv
SEM
95% CI ↑
Utter.
Gaddy 2021
99
3.301
0.593
0.060
[3.193, 3.409]
Utter.
CGS-ETS
99
3.329
0.560
0.056
[3.217, 3.440]
Sess.
Gaddy 2021
7
3.303
0.195
0.074
[3.132, 3.474]
Sess.
CGS-ETS
7
3.328
0.059
0.022
[3.273, 3.383]
Table 9: MOS (NISQA) descriptive statistics.
Agg.
n
d
Median(d)
W+
W−
T
p
Utter.
99
0.028
-0.064
2405.0
2545.0
2405.0
0.807
Sess.
7
0.025
-0.008
13.0
15.0
13.0
0.938
Table 10: Paired Wilcoxon signed-rank test results.
Lyapunov (LER)
Fractal (MSDFA)
Phase
MACs
FLOPs
Lat.
MACs
FLOPs
Lat.
Training
31.81M
63.62M
1.28ms
0.358M
0.716M
0.16ms
Inference
0
0
0
0
0
0
Table 11: Code-derived computational complexity for deterministic chaotic objectives per 200-frame utterance. These modules are exclusively active during the training phase. Here, MACs, FLOPs are in Millions and Lat. = Latency in Milliseconds.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
CFE
Mam.
FFN
L ↓
ST ↑
P ↑
SD ↑
N ↑
W ↓
IT ↓
FLOPs ↓
RTF ↓
P0
ReLU
SiLU
SwiGLU
1.10
0.59
1.16
-31.47
3.25
51.16%
7.12ms
1.038G
0.0075
P1
ReLU
SiLU
GeLU
1.09
0.60
1.18
-31.30
3.28
47.06%
5.92ms
929.76M
0.0059
P2
LyRIC
SiLU
GeLU
1.16
0.54
1.13
-33.22
2.90
71.86%
5.92ms
929.76M
0.0059
P3
ReLU
SiLU
LyRIC
1.10
0.59
1.17
-31.85
3.23
52.63%
5.70ms
928.43M
0.0057
P4
ReLU
LyRIC
SwiGLU
1.08
0.61
1.19
-32.63
3.30
45.71%
6.21ms
1.035G
0.0062
P5
ReLU
LyRIC
GeLU
1.08
0.61
1.19
-32.61
3.33
41.36%
5.61ms
926.55M
0.0057
Appendix
Table 12: Ablation Study on Activation Functions.
Layers
L ↓
ST ↑
P ↑
SD ↑
N ↑
WER ↓
Para. ↓
IT ↓
FLOPs ↓
RTF ↓
P0
No Attention
1.15
0.53
1.14
-34.42
2.91
73.22%
12.68M
4.23ms
848.33M
0.0046
P1
L=1,2
1.08
0.60
1.18
-32.41
3.30
46.14%
14.34M
5.91ms
926.55M
0.0059
P2
L=3,4
1.10
0.59
1.17
-31.96
3.25
50.25%
14.34M
5.97ms
926.55M
0.0059
P3
L=1,4
1.10
0.59
1.17
-33.54
3.30
49.01%
14.34M
5.91ms
926.55M
0.0061
P4
L=1,2,3
1.08
0.61
1.19
-33.31
3.26
46.38%
15.05M
6.86ms
966.45M
0.0068
P5
L=2,3,4
1.09
0.60
1.19
-32.81
3.28
44.91%
15.05M
7.11ms
966.45M
0.0071
Appendix
Table 13: Ablation Study on GLARE Attention.
Attention Type
L ↓
ST ↑
P ↑
SD ↑
N ↑
W ↓
Enc. ↓
IT ↓
FLOPs ↓
RTF ↓
P0
MHA (All Layers)
1.08
0.61
1.19
-32.43
3.30
43.32%
14.49M
7.42ms
1.166G
0.0074
P1
MHA ( L=1,3 )
1.08
0.61
1.18
-32.36
3.27
44.73%
12.39M
5.92ms
1.061G
0.0059
P2
MQA (All Layers)
1.07
0.60
1.18
-31.69
3.33
44.85%
10.56M
7.10ms
966.50M
0.0071
P3
MQA ( L=1,3 )
1.08
0.60
1.18
-31.81
3.31
44.90%
9.38M
5.69ms
907.41M
0.0058
P4
GLARE ( L=1,3 )
1.08
0.61
1.19
-32.61
3.33
41.36%
9.85M
5.61ms
926.55M
0.0057
Appendix
Table 14: Ablation Study on Different Attentions
Dimension
L ↓
ST ↑
P ↑
SD ↑
N ↑
W ↓
Pr ↓
IT ↓
FLOPs ↓
RTF ↓
P0
CD=128
1.13
0.57
1.15
-34.35
3.09
58.88%
0.91M
5.53ms
59.64M
0.0055
P1
CD=256
1.11
0.59
1.16
-33.09
3.20
52.14%
3.59M
5.59ms
234.030M
0.0056
P2
CD=1024
1.08
0.61
1.19
-32.67
3.34
45.10%
56.82M
15.72ms
3.690G
0.0088
P3
CD=512
1.08
0.61
1.19
-32.61
3.33
41.36%
14.34M
5.61ms
926.55M
0.0057
Appendix
Table 15: Ablation Study on Channel Dimension
Count
L ↓
ST ↑
P ↑
SD ↑
N ↑
W ↓
Enc. ↓
IT ↓
FLOPs ↓
RTF ↓
P0
Lc=2
1.10
0.58
1.16
-32.75
3.20
55.94%
5.67M
4.52ms
721.48M
0.0045
P1
Lc=3
1.08
0.60
1.18
-31.45
3.33
45.18%
8.51M
5.37ms
863.65M
0.0057
P2
Lc=5
1.08
0.61
1.19
-33.06
3.25
41.91%
14.18M
8.50ms
1.148G
0.0091
P3
Lc=4
1.08
0.61
1.19
-32.61
3.33
41.36%
9.85M
5.61ms
926.55M
0.0057
Appendix
Table 16: Ablation Study on Layer Count of GLAME
Component
Ref.
Bottleneck
Without Component
With Proposed Component
Efficiency Gain
GLAME
Tab. 3
Macro
Gaddy and Klein (2021) Encoder
GLAME
-34.18M Params
(Encoder)
Architectural Bloat
(Enc: 44.03M, FLOPs: 2400M)
(Enc: 9.85M, FLOPs: 926.6M)
-1473.4M FLOPs
Chaos Losses
Tab. 4
Acoustic Fidelity,
w/o LER+MSDFA ( P1 )
w/ LER+MSDFA ( P4 )
N/A
(LER + MSDFA)
Phonetic Accuracy
(WER: 51.56%, LSD: 1.10)
(WER: 41.36%, LSD: 1.08)
LyRIC Activation
Tab. 5
Computational
SiLU + SwiGLU ( P0 )
LyRIC + GeLU ( P5 )
-111.4M FLOPs
(in Mamba block)
Cost, Latency
(WER: 51.16%, FLOPs: 1038M)
(WER: 41.36%, FLOPs: 926.6M)
-1.51ms Latency
Appendix
Table 17: Consolidated Summary of Individual Component Contributions. This table highlights the isolated impact of each proposed mechanism compared to its standard counterpart.
Figure 6: Output Density Distributions of ReLU, GELU, SiLU against Log-divergence Equation (Equation 6 )
Function
FLOPs/element ↓
P0
ReLU
1
P1
GeLU
14
P2
SiLU
22
P3
Log-divergence
47
Appendix
Table 18: Comparison of Computational Profiling among ReLU, GeLU, SiLU, and Log-divergence Equation (see Equation 6 )
Figure 7: Topological Sensitivity Analysis of the LyRIC Activation Surface ( f(x,α)=x∣x∣α ).
Figure 8: Geometric Representation of the LyRIC Activation Function at the Optimal Expansion Coefficient α=1.0 .
Figure 9: Gradient Landscape Analysis of ReLU, GeLU, SiLU, and LyRIC ( α=1 ).
Figure 10: Parametric Influence of the Expansion Coefficient ( α ) on the Gradient Landscape of the LyRIC Activation.
Figure 11: The Software Interface Used in Subjective Tests.
Figure 12: The MOS score.
Hyperparameter
Value
Use Case & Rationale
Optimization & Training Dynamics
Max epochs
80
Provide sufficient gradient updates for convergence while preventing catastrophic representational over-fitting.
Ensure robust initial gradient descents across the highly non-convex objective manifold.
Learning rate warmup
500 steps
Strictly prevent early gradient spikes and stabilize initial optimization dynamics.
Weight decay
1×10−7
Subtly regularize network weights to prevent over-fitting on specific real-world EMG physiological artifacts.
Appendix
Table 19: Comprehensive Summary of Critical Hyperparameters, Architectural Configurations, and Their Theoretical Justifications within the Proposed CGS-ETS Framework.
Module
Stage
Layer Type
Input→Output
Kernel
Stride
Padding
Params
CFE Block
Block 1
Conv1d (Main)
8→512
3
2
1
12 800
BatchNorm1d + ReLU
512→512
–
–
–
1 024
Conv1d (Main)
512→512
3
1
1
786 944
BatchNorm1d
512→512
–
–
–
1 024
Conv1d (Residual)
8→512
1
2
0
4 608
BatchNorm1d + ReLU
512→512
–
–
–
1 024
Appendix
Table 20: Comprehensive Layer-wise Parameter Breakdown of the Proposed CGS-ETS Architecture. The summary block demonstrates the strict architectural compression achieved by our GLAME and its selective application of GLARE attention modules.
We propose a chaos-inspired new architecture for EMG-to-Speech (ETS) synthesis called CS-ETS, which combines a Samba-based encoder with two novel chaos-inspired loss functions -- Lyapunov Exponent Regularization (LER) and Multi-Scale Detrended Fluctuation Analysis (MSDFA). LER is designed based on Lyapunov exponents to capture nonlinear fluctuations and sensitivity to initial conditions. MSDFA exploits detrended fluctuation analysis to quantify fractal-like, long-range temporal chaotic correlation. CS-ETS surpasses prior work with a 40.79% lower parameter count (32M vs 54.1M) and introduces a new Post-Vocoder Alignment approach that improves LSD by 2.1x, STOI by 4.7x, and SI-SDR by 1.25x. CS-ETS reduces computation by 13.33% while maintaining improved performance. To the best of our knowledge, for the first time, we show how ETS can be supervised by the subtle non-linear chaotic physics with Samba attention to achieve a significantly smaller model with superior performance.
Sajid Fardin Dipto, Tarikul Islam Tamiti, David Vergano +2
Language models increasingly serve as the backbone of text-to-speech (TTS) systems, yet we understand little about the representations they build when text and generated speech tokens share a single residual stream. We train BatchTopK sparse autoencoders on the LM backbone of CosyVoice3 and introduce a modality-aware auto-interp pipeline that labels each feature from where it fires-text-prefix context, 1-second speech clips, or both. The recovered features are interpretable, spanning phonemes, laughter, accent prompts and speaker gender. Steering through the SAE latent space shows these features are causal rather than merely descriptive: targeted interventions raise laughter probability from 0.02 to 0.79, flip perceived speaker gender, and control speech rate while preserving spoken content. SAE features thus serve both as interpretability objects and as control directions for TTS synthesis.
Nikita Koriagin, Georgii Aparin, Nikita Balagansky +1
The loss of speech limits communication for individuals with paralysis. Direct neural-to-speech synthesis is challenging due to the limited availability of neural data for training speech brain-computer interfaces. Most existing systems rely on cascaded neural-to-text-to-speech pipelines, which increase inference latency and propagate errors across stages. We present Brain2Speech-Net, a single-stage neural-to-speech generation framework without intermediate text decoding. We use a differentiable phoneme bottleneck and a deep-HMM alignment mechanism to map long neural recordings into the latent space of a text-to-speech (TTS) model, enabling high-quality speech synthesis. Brain2Speech-Net is the only system in our comparison that produces intelligible speech while generating faster than real time.
Shreeram Suresh Chandra, Zexin Cai, Yu Tsao +2
Johns Hopkins University, USA · Academia Sinica, Taiwan · University of Edinburgh, UK