Lyapunov-Inspired LyRIC Activation and GLARE Attention in Chaos-Guided State Space Modeling for EMG-To-Speech (ETS) Synthesis
Organizations: Department of Cyber Security Engineering George Mason University Fairfax, VA, USA
Abstract
Electromyography-to-Speech (ETS) synthesis is typically a non-linear, chaotic dynamical system. However, no prior work has studied the chaotic behavior of ETS synthesis to date. Yet, prior works strictly rely on standard reconstruction metrics with parameter-heavy transformers that systematically over-smooth natural acoustic dynamics. To close this gap, for the first time, we propose a chaos-inspired Lyapunov-derived activation function (LyRIC) with two novel chaotic loss functions, Lyapunov Exponent Regularization and Multi-Scale Detrended Fluctuation Analysis, to explicitly capture the deterministic chaos of human phonation. In addition, we introduce a compressed novel encoder, GLAME, which synergizes global Mamba state-space modeling with localized GLARE attention. We comprehensively perform frame-level acoustic evaluation in a multilingual and multi-speaker setup using English and Mandarin datasets. The proposed system outperforms the established baseline with a 4.69x increase in objective intelligibility (STOI: 0.61 vs. 0.13) and a 2.08x improvement in spectral reconstruction (LSD: 1.08 vs. 2.25). Importantly, this improvement is achieved with 73.49% fewer parameters (14.34M vs. 54.10M), establishing a new baseline for ETS synthesis. To the best of our knowledge, this is the first work demonstrating that integrating non-linear chaotic physics into neural networks yields superior yet compact inductive biases for real-time ETS synthesis.
Figures & tables
| Function | FLOPs/element | |
| 11 | ||
| 1 | ||
| 12 | ||
| 2 | ||
| 13 | ||
| 3 |
| Model | Cr | L | ST | P | SD | N | W | C | Pr | |
| Baseline 1 | EN | 3.23 | 0.04 | 1.17 | -52.5 | 3.28 | 68.00 | 38.2 | 60.21 | |
| Baseline 1 | ZH | 3.50 | 0.03 | 1.05 | -58.1 | 1.31 | 59.1 | 60.21 | ||
| Baseline 2 | EN | 2.25 | 0.13 | 1.18 | -42.0 | 3.30 | 42.20 | 23.7 | 54.10 | |
| Baseline 2 | ZH | 2.67 | 0.10 | 1.08 | -39.5 | 1.84 | 39.8 | 54.10 | ||
| Baseline 3 | EN | 46.14 | ||||||||
| Baseline 4 | ZH | 38.0 |
| Model | CFE | Enc | Total | IT | Fs | RTF |
| Gaddy and Klein (2021) | 10.17 | 44.03 | 54.10 | 5.97 | 2400 | 0.0057 |
| Proposed CGS-ETS | 4.49 | 9.85 | 14.34 | 5.61 | 926.6 | 0.0057 |
| Model Config. | L | ST | P | SD | N | W | Pr | |
| w/o LER+MSDFA | 1.14 | 0.60 | 1.12 | -35.5 | 3.25 | 51.56 | 14.34 | |
| w LER w/o MSDFA | 1.10 | 0.60 | 1.18 | -33.7 | 3.29 | 48.34 | 14.34 | |
| w MSDFA w/o LER | 1.09 | 0.60 | 1.19 | -32.6 | 3.31 | 45.65 | 14.34 | |
| w LER+MSDFA | 1.08 | 0.61 | 1.19 | -32.6 | 3.33 | 41.36 | 14.34 |
| CFE | Mam. | FFN | SD | N | W | IT | Fs | RTF | |
| ReLU | SiLU | SwiGLU | -31.5 | 3.25 | 51.16 | 7.12 | 1038 | 0.0075 | |
| ReLU | SiLU | GeLU | -31.3 | 3.28 | 47.06 | 5.92 | 929.8 | 0.0059 | |
| LyRIC | SiLU | GeLU | -33.2 | 2.90 | 71.86 | 5.92 | 929.8 | 0.0059 | |
| ReLU | SiLU | LyRIC | -31.9 | 3.23 | 52.63 | 5.70 | 928.4 | 0.0057 | |
| ReLU | LyRIC | SwiGLU | -32.6 | 3.30 | 45.71 | 6.21 | 1035 | 0.0062 | |
| ReLU | LyRIC | GeLU | -32.6 | 3.33 | 41.36 | 5.61 | 926.6 | 0.0057 |
| Layers | SD | N | W | Pr | IT | Fs | RTF | |
| None | -34.4 | 2.91 | 73.22 | 12.68 | 4.23 | 848.3 | 0.0046 | |
| -32.4 | 3.30 | 46.14 | 14.34 | 5.91 | 926.6 | 0.0059 | ||
| -33.0 | 3.25 | 50.25 | 14.34 | 5.97 | 926.6 | 0.0059 | ||
| -33.5 | 3.30 | 49.01 | 14.34 | 5.91 | 926.6 | 0.0061 | ||
| -33.3 | 3.26 | 46.38 | 15.05 | 6.86 | 966.5 | 0.0068 | ||
| -32.8 | 3.28 | 44.91 | 15.05 | 7.11 | 966.5 | 0.0071 |
| L | ST | P | SD | N | W | IT | Fs | RTF | ||
| 1.09 | 0.60 | 1.18 | -34.1 | 3.30 | 48.22 | 5.98 | 927.0 | 0.0058 | ||
| 1.09 | 0.61 | 1.18 | -30.8 | 3.26 | 44.73 | 5.95 | 929.9 | 0.0060 | ||
| 1.08 | 0.60 | 1.18 | -32.3 | 3.30 | 45.40 | 5.95 | 927.2 | 0.0059 | ||
| 1.09 | 0.60 | 1.18 | -33.0 | 3.30 | 49.26 | 6.00 | 927.3 | 0.0059 | ||
| 1.08 | 0.61 | 1.19 | -32.6 | 3.33 | 41.36 | 5.61 | 926.6 | 0.0057 |
| Attention Type | SD | N | W | Enc | IT | Fs | RTF | |
| MHA, All Layers | -32.4 | 3.30 | 43.32 | 14.49 | 7.42 | 1166 | 0.0074 | |
| MHA, | -32.4 | 3.27 | 44.73 | 12.39 | 5.92 | 1061 | 0.0059 | |
| MQA, All Layers | -31.7 | 3.33 | 44.85 | 10.56 | 7.10 | 966.5 | 0.0071 | |
| MQA, | -31.8 | 3.31 | 44.90 | 9.38 | 5.69 | 907.4 | 0.0058 | |
| GLARE, | -32.6 | 3.33 | 41.36 | 9.85 | 5.61 | 926.6 | 0.0057 |
| Agg. | Model | Mean | SDv | SEM | 95% CI | |
| Utter. | Gaddy 2021 | 99 | 3.301 | 0.593 | 0.060 | [3.193, 3.409] |
| Utter. | CGS-ETS | 99 | 3.329 | 0.560 | 0.056 | [3.217, 3.440] |
| Sess. | Gaddy 2021 | 7 | 3.303 | 0.195 | 0.074 | [3.132, 3.474] |
| Sess. | CGS-ETS | 7 | 3.328 | 0.059 | 0.022 | [3.273, 3.383] |
| Agg. | Median(d) | T | p | ||||
| Utter. | 99 | 0.028 | -0.064 | 2405.0 | 2545.0 | 2405.0 | 0.807 |
| Sess. | 7 | 0.025 | -0.008 | 13.0 | 15.0 | 13.0 | 0.938 |
| Lyapunov (LER) | Fractal (MSDFA) | |||||
| Phase | MACs | FLOPs | Lat. | MACs | FLOPs | Lat. |
| Training | 31.81M | 63.62M | 1.28ms | 0.358M | 0.716M | 0.16ms |
| Inference | 0 | 0 | 0 | 0 | 0 | 0 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| CFE | Mam. | FFN | L | ST | P | SD | N | W | IT | FLOPs | RTF | |
| ReLU | SiLU | SwiGLU | 1.10 | 0.59 | 1.16 | -31.47 | 3.25 | 51.16% | 7.12ms | 1.038G | 0.0075 | |
| ReLU | SiLU | GeLU | 1.09 | 0.60 | 1.18 | -31.30 | 3.28 | 47.06% | 5.92ms | 929.76M | 0.0059 | |
| LyRIC | SiLU | GeLU | 1.16 | 0.54 | 1.13 | -33.22 | 2.90 | 71.86% | 5.92ms | 929.76M | 0.0059 | |
| ReLU | SiLU | LyRIC | 1.10 | 0.59 | 1.17 | -31.85 | 3.23 | 52.63% | 5.70ms | 928.43M | 0.0057 | |
| ReLU | LyRIC | SwiGLU | 1.08 | 0.61 | 1.19 | -32.63 | 3.30 | 45.71% | 6.21ms | 1.035G | 0.0062 | |
| ReLU | LyRIC | GeLU | 1.08 | 0.61 | 1.19 | -32.61 | 3.33 | 41.36% | 5.61ms | 926.55M | 0.0057 |
| Layers | L | ST | P | SD | N | WER | Para. | IT | FLOPs | RTF | |
| No Attention | 1.15 | 0.53 | 1.14 | -34.42 | 2.91 | 73.22% | 12.68M | 4.23ms | 848.33M | 0.0046 | |
| 1.08 | 0.60 | 1.18 | -32.41 | 3.30 | 46.14% | 14.34M | 5.91ms | 926.55M | 0.0059 | ||
| 1.10 | 0.59 | 1.17 | -31.96 | 3.25 | 50.25% | 14.34M | 5.97ms | 926.55M | 0.0059 | ||
| 1.10 | 0.59 | 1.17 | -33.54 | 3.30 | 49.01% | 14.34M | 5.91ms | 926.55M | 0.0061 | ||
| 1.08 | 0.61 | 1.19 | -33.31 | 3.26 | 46.38% | 15.05M | 6.86ms | 966.45M | 0.0068 | ||
| 1.09 | 0.60 | 1.19 | -32.81 | 3.28 | 44.91% | 15.05M | 7.11ms | 966.45M | 0.0071 |
| Attention Type | L | ST | P | SD | N | W | Enc. | IT | FLOPs | RTF | |
| MHA (All Layers) | 1.08 | 0.61 | 1.19 | -32.43 | 3.30 | 43.32% | 14.49M | 7.42ms | 1.166G | 0.0074 | |
| MHA ( ) | 1.08 | 0.61 | 1.18 | -32.36 | 3.27 | 44.73% | 12.39M | 5.92ms | 1.061G | 0.0059 | |
| MQA (All Layers) | 1.07 | 0.60 | 1.18 | -31.69 | 3.33 | 44.85% | 10.56M | 7.10ms | 966.50M | 0.0071 | |
| MQA ( ) | 1.08 | 0.60 | 1.18 | -31.81 | 3.31 | 44.90% | 9.38M | 5.69ms | 907.41M | 0.0058 | |
| GLARE ( ) | 1.08 | 0.61 | 1.19 | -32.61 | 3.33 | 41.36% | 9.85M | 5.61ms | 926.55M | 0.0057 |
| Dimension | L | ST | P | SD | N | W | Pr | IT | FLOPs | RTF | |
| 1.13 | 0.57 | 1.15 | -34.35 | 3.09 | 58.88% | 0.91M | 5.53ms | 59.64M | 0.0055 | ||
| 1.11 | 0.59 | 1.16 | -33.09 | 3.20 | 52.14% | 3.59M | 5.59ms | 234.030M | 0.0056 | ||
| 1.08 | 0.61 | 1.19 | -32.67 | 3.34 | 45.10% | 56.82M | 15.72ms | 3.690G | 0.0088 | ||
| 1.08 | 0.61 | 1.19 | -32.61 | 3.33 | 41.36% | 14.34M | 5.61ms | 926.55M | 0.0057 |
| Count | L | ST | P | SD | N | W | Enc. | IT | FLOPs | RTF | |
| 1.10 | 0.58 | 1.16 | -32.75 | 3.20 | 55.94% | 5.67M | 4.52ms | 721.48M | 0.0045 | ||
| 1.08 | 0.60 | 1.18 | -31.45 | 3.33 | 45.18% | 8.51M | 5.37ms | 863.65M | 0.0057 | ||
| 1.08 | 0.61 | 1.19 | -33.06 | 3.25 | 41.91% | 14.18M | 8.50ms | 1.148G | 0.0091 | ||
| 1.08 | 0.61 | 1.19 | -32.61 | 3.33 | 41.36% | 9.85M | 5.61ms | 926.55M | 0.0057 |
| Component | Ref. | Bottleneck | Without Component | With Proposed Component | Efficiency Gain |
| GLAME | Tab. 3 | Macro | Gaddy and Klein (2021) Encoder | GLAME | -34.18M Params |
| (Encoder) | Architectural Bloat | (Enc: 44.03M, FLOPs: 2400M) | (Enc: 9.85M, FLOPs: 926.6M) | -1473.4M FLOPs | |
| Chaos Losses | Tab. 4 | Acoustic Fidelity, | w/o LER+MSDFA ( ) | w/ LER+MSDFA ( ) | N/A |
| (LER + MSDFA) | Phonetic Accuracy | (WER: 51.56%, LSD: 1.10) | (WER: 41.36%, LSD: 1.08) | ||
| LyRIC Activation | Tab. 5 | Computational | SiLU + SwiGLU ( ) | LyRIC + GeLU ( ) | -111.4M FLOPs |
| (in Mamba block) | Cost, Latency | (WER: 51.16%, FLOPs: 1038M) | (WER: 41.36%, FLOPs: 926.6M) | -1.51ms Latency |
| Function | FLOPs/element | |
| ReLU | 1 | |
| GeLU | 14 | |
| SiLU | 22 | |
| Log-divergence | 47 |
| Hyperparameter | Value | Use Case & Rationale |
| Optimization & Training Dynamics | ||
| Max epochs | 80 | Provide sufficient gradient updates for convergence while preventing catastrophic representational over-fitting. |
| Batch size | 32 | Mathematically balance gradient variance stabilization with strict GPU VRAM (OOM) constraints. |
| Initial learning rate ( ) | Ensure robust initial gradient descents across the highly non-convex objective manifold. | |
| Learning rate warmup | 500 steps | Strictly prevent early gradient spikes and stabilize initial optimization dynamics. |
| Weight decay | Subtly regularize network weights to prevent over-fitting on specific real-world EMG physiological artifacts. | |
| Module | Stage | Layer Type | Input→Output | Kernel | Stride | Padding | Params |
| CFE Block | Block 1 | Conv1d (Main) | 8→512 | 3 | 2 | 1 | 12 800 |
| BatchNorm1d + ReLU | 512→512 | – | – | – | 1 024 | ||
| Conv1d (Main) | 512→512 | 3 | 1 | 1 | 786 944 | ||
| BatchNorm1d | 512→512 | – | – | – | 1 024 | ||
| Conv1d (Residual) | 8→512 | 1 | 2 | 0 | 4 608 | ||
| BatchNorm1d + ReLU | 512→512 | – | – | – | 1 024 |