Jev Matches 7B Language Models for Speech-Neuroprosthesis Rescoring
Abstract
A speech neuroprosthesis decodes attempted speech from brain activity and ends by rescoring the decoder's candidate sentences with a language model of several billion parameters, the only component that needs a GPU. Replacing that model with a cheaper one is hard: general language models asked to pick one sentence from a list answer from where a label sits in the list rather than from the sentence itself. We pose rescoring as a single typed decision, one call that returns a probability for every candidate, served by Jev, a hosted model trained for calibrated decisions, and combine it with the decoder's own score. On 978 held-out sentences from a participant with ALS, where the published decoder alone reaches 8.1% word error, Jev reaches 7.5% against 7.8% for both OPT-6.7b and Qwen2.5-7B; with the decoder's weight re-tuned, 6.9% against 7.2% and 7.4%. Jev is ahead in all four comparisons and at most 0.2 points behind at the 95% bound. It costs 0.07 USD per thousand sentences and needs no GPU; a dedicated GPU running a 7B model is cheaper per sentence only above 43% utilisation, far beyond what one user generates. End-to-end latency over the internet is 262 ms, of which 62 ms is spent at the provider, the same order as a 7B model on a local GPU (27 ms) but not faster.
Figures & tables
| published ( ) | re-tuned | |||||
|---|---|---|---|---|---|---|
| Arm | WER | WER | ||||
| No rescoring | 8.09 | — | — | (0.3, —) | 8.09 | — |
| OPT-6.7b | 7.80 | 0.070 | (0.40, 0.45) | 7.20 | ||
| Qwen2.5-7B | 7.75 | 0.070 | (0.45, 0.45) | 7.42 | ||
| Jev (typed) | 7.46 | 0.018 | (2.0, 0.75) | 6.93 | ||
| Protocol | Comparison | WER | 95% CI | |
|---|---|---|---|---|
| published | Jev OPT-6.7b | 0.082 | ||
| published | Jev Qwen2.5-7B | 0.129 | ||
| re-tuned | Jev OPT-6.7b | 0.112 | ||
| re-tuned | Jev Qwen2.5-7B | 0.019 |
| Typed arm | Labels used | Share on A | Truth picked | Truth picked ( A) | WER |
|---|---|---|---|---|---|
| Jev | 57 / 100 | 72% | 67.9% | 26.6% | 7.46 |
| Laya (421M) | 44 / 100 | 69% | 53.7% | 2.3% | 8.09 |
| Qwen2.5-7B | 31 / 100 | 28% | 29.3% | 2.4% | 8.11 |
| OPT-6.7b | 13 / 100 | 35% | 27.3% | 0.2% | 8.09 |
| Decoder’s top candidate | — | — | 68.0% | — | 8.09 |
| Cost ($) | Latency per list (ms) | |||
|---|---|---|---|---|
| Arm | per 1k sentences | one user, per day | median | p95 |
| Jev, API | 0.069 | 0.07 | 282 | 346 |
| OPT-6.7b, L40S | 0.029 ∗ | 43.68 | 27 | 134 |
| Qwen2.5-7B, L40S | 0.034 ∗ | 43.68 | 32 | 158 |