We find that major reported improvements in decoding words from non-invasive brain recordings are largely reproducible without any brain data. In the influential work of d'Ascoli et al. (2025), time series of brain activity from subjects perceiving continuous speech are segmented into fixed-length windows starting at each word. A neural network then generates predictions for all of the words in a sentence together. Neighbouring windows partially overlap, implicitly revealing the interval between words. Since these intervals indicate the duration of the words spoken, and different words tend to have different durations - for example, "the" is much shorter than "supercalifragilisticexpialidocious" - the neural network can improve its predictions of words without relying on the underlying brain activity. Consistent with this, the method reaches 22.0% balanced accuracy on synthetic signals containing no brain information, compared with 22.3% on real brain recordings. To prevent the network from learning this shortcut, we make a single, simple change. Instead of jointly encoding all windows in a sentence, we process each independently. As a result, the neural network achieves better performance by learning underlying word-specific information from brain recordings. This makes two existing strategies become much more effective than before. Both aggregating predictions from distinct neural responses to the same word and using a pretrained LLM as a linguistic prior now substantially improve results. On our perceived speech benchmark, this simple recipe (SimpleB2T) achieves a word error rate of 36.6% with five observations per word, approaching past invasive speech decoding performance, albeit under different conditions. The results in this work expose an important shortcut in brain-to-text decoding and show that removing it leads to a simple and considerably more effective strategy.
Figures & tables
Figure 1: Overlapping word-aligned windows leak word duration. (a) Overlapping word-aligned windows contain shared samples whose relative offsets reveal the duration of words. (b) To test whether jointly decoding words uses neural information, we compare MEG with synthetic signals having no relation to brain activity. With MEG, isolated means we decode each window independently; with synthetic data, it means we make windows independent. Timing tests what happens when the same model receives an encoding of the interval between words directly. Jointly decoding words reaches nearly the same accuracy on synthetic and real inputs, whereas removing the shortcut substantially reduces performance. Error bars indicate ±1σ across training seeds.
Figure 2: Jointly decoding words is sensitive to word duration. (a) Each point is a word with accuracy plotted against the standard deviation of its duration across training occurrences. (b) We test whether the decoder predicts words whose typical durations match the test occurrence. We measure how far the typical durations of the model’s top-10 predictions are from an occurrence’s duration. We then replace the duration with that of another occurrence of the same true word and recompute the error. The bar shows how much the error increases after this replacement, so positive values mean the predictions have durations more similar to the specific occurrence than expected. Error bars show standard deviation. ∗p<.05 , ∗∗∗p<.001 .
Figure 3: SimpleB2T. (a) Instead of jointly decoding words, we decode each neural response independently. We also test the effect of aggregating responses from distinct observations and using an LLM with beam search for sentence decoding. (b) Combining predictions with an LLM substantially reduces WER. (c) Decoded examples. Results are on our core set of 100 clinical sentences with k=5 observations per word. Error bars indicate ±1 standard deviation across five training seeds.
Core
Expanded
Full
Method
k
WER (%) ↓
SMR (%) ↑
WER (%) ↓
SMR (%) ↑
WER (%) ↓
SMR (%) ↑
LM only
-
73.8
2.0
91.8
0.0
83.4
1.0
d’Ascoli
1
98.9±0.3
0.0±0.0
98.8±0.1
0.0±0.0
98.8±0.2
0.0±0.0
d’Ascoli + LM
1
77.3±1.1
2.4±0.3
87.5±0.8
0.2±0.1
82.7±0.2
1.3±0.2
SimpleB2T
1
65.6±0.8
6.4±0.5
76.3±0.5
4.2±0.4
71.3±0.5
5.3±0.3
d’Ascoli
5
98.9±0.3
0.0±0.0
98.8±0.4
0.0±0.0
98.9±0.1
0.0±0.0
Table 1: Effect of decoding words independently. We compare SimpleB2T and d’Ascoli et al. , with and without the same LLM rescoring. The table tests how using k observations and LLM rescoring behave when words are decoded jointly and when words are decoded independently. Results are means across five seeds and subscripts show sample standard deviations.
Table 2: Ablations. Values selected on a development set of 50 sentences, except for (c) which is retrospective and did not inform q . Results on the core set of test sentences with k=5 . Bold indicates the selected value. Results are means across training seeds, with sample standard deviations.
Figure 4: SimpleB2T results. (a, b) WER and exact sentence match as the number of observations increases. (c) Oracle selection among the top- N beam candidates. (d) Sentence WER distribution. (e) Confusion matrix across the 92-word vocabulary. (f) Word accuracy by part of speech for the LM, neural decoder, and their combination. Panels c–f use k=5 ; all results are on the Core test set. Shading and error bars indicate ±1 standard deviation across five training seeds.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Intervals between word onsets closely track word duration. Relationship between the annotated duration of a word and the interval from its onset to the onset of the following word across adjacent word pairs. The two quantities are strongly correlated ( r=0.90 ). The dashed line indicates equality and the solid line shows a linear fit.
Figure 6: The overlapping window shortcut reproduces across all subjects. Subject-0-trained MEG and shared-synthetic decoders evaluated on the 32 other subjects in LibriBrain100. (a) Mean across subjects. (b) Per-subject results. Error bars show standard deviation across five training seeds.
Figure 7: Replication of Figure 1 b with Armeni et al. (2022) and the listening component of Le Petit Prince MEG ( d’Ascoli et al., 2025 ) .
Figure 8: Joint decoders aligns with duration-based predictions. We measure how much the Jensen–Shannon divergence between the decoder probabilities and a duration-based word predictor’s probabilities increase when their pairing across test occurrences is randomly shuffled. Thus, larger values indicate stronger alignment with the duration-based word predictor. Bars show means across five training seeds and error bars indicate sample standard deviations.
Figure 9: Jointly decoding words with and without overlapping inputs. We evaluate d’Ascoli et al. (2025) , the same architecture trained without overlapping neighbouring windows, and isolated decoding on both natural overlapping sentences and non-overlapping versions of the same word sequences. Error bars indicate ±1 sample standard deviation across training seeds.
Figure 10: The timing shortcut does not explain Brain2Qwerty performance. We repeat our synthetic-input control for Brain2Qwerty ( Lévy et al., 2026 ) . Replacing MEG with synthetic signals that preserve the aligned-window structure reduces balanced character accuracy from 37.0% to 6.5% on validation and from 34.9% to 6.8% on test, close to the 3.45% chance level. Dots show individual training seeds and error bars show standard deviations.
Session ID
Split
Tasks
Recordings
Hours
1
Train
MOCHATIMIT, Sherlock1–9, TIMIT, TheMoth
12
6.96
2
Train
MOCHATIMIT, Sherlock1–9, TIMIT, TheMoth
12
6.85
3
Train
MOCHATIMIT, Sherlock1–9, TIMIT, TheMoth
12
6.45
4
Train
MOCHATIMIT, Sherlock1–9, TIMIT, TheMoth
12
6.47
5
Train
Sherlock1–9, TIMIT, TheMoth
11
5.51
6
Train
Sherlock1–9, TIMIT, TheMoth
11
6.61
Appendix
Table 3: LibriBrain100 subject-0 session split. Session IDs and task names are taken from recording filenames. Task ranges are inclusive: for example, Sherlock1–9 denotes the nine tasks Sherlock1 through Sherlock9. Each listed task contributes one recording for that session ID. All recordings sharing a session label are assigned to the same split, including across tasks. The split contains 121 training recordings (66.29 hours), 10 validation recordings (5.57 hours), and 31 test recordings (14.71 hours). Durations refer to complete recordings before windowing. Core and Expanded use the test partition; development sentences use the validation partition.
Hyperparameter
Setting
Input
MEG channels / sampling rate
306 / 50 Hz
Word window / baseline correction
3 s from onset / subtract channel-wise mean of first 0.5 s
Bandpass filter / clipping
0.1–40 Hz / [−5,5] after scaling
Channel scaling
Recording-wise, channel-wise RobustScaler ( Pedregosa et al., 2011 )
Model
Appendix
Table 4: SimpleB2T hyperparameters. The ablations use λ=0.5 , while other results use λ=(1.5,1,0.5,0.5,0.5) for k=1,…,5 .
Figure 11: Language-model weight decreases as neural evidence improves. (a) Development-selected λ falls as more observations are aggregated. (b) WER is minimised at lower LM weights for larger k .
Overall split
Word-matched split
Method
Low
High
Low
High
LM
26.3
25.8
28.3
28.3
Brain
27.5±2.0
23.8±2.7
24.0±2.1
27.1±2.6
Brain + LM
62.6±1.3
64.0±2.2
65.3±2.0
64.9±1.9
Appendix
Table 5: Word accuracy grouped by the predictability of subsequent context. Subscripts denote sample standard deviations across five training seeds. The LM is deterministic.
Core
Expanded
Full
Method
k
WER (%) ↓
SMR (%) ↑
WER (%) ↓
SMR (%) ↑
WER (%) ↓
SMR (%) ↑
d’Ascoli
1
97.7±0.4
0.0±0.0
97.0±0.7
0.0±0.0
97.3±0.6
0.0±0.0
d’Ascoli + LM
1
79.9±1.5
2.3±0.2
86.2±0.9
0.6±0.0
83.2±0.2
1.4±0.1
d’Ascoli
5
98.0±0.6
0.0±0.0
97.2±1.1
0.0±0.0
97.5±0.8
0.0±0.0
d’Ascoli + LM
5
79.2±1.6
2.3±0.6
85.5±0.7
1.0±0.0
82.6±0.5
1.7±0.3
Appendix
Table 6: Controlling for train–test mismatch in the joint word decoder. We retrain the decoder of d’Ascoli et al. (2025) on non-overlapping sentence constructions and evaluate at k=5 . Jointly decoding words remains much less accurate than SimpleB2T. Values are means across training seeds and subscripts show sample standard deviations.
Figure 12: Word-level decoding analysis. Word accuracy is weakly related to (a) phonetic distinguishability, (b) semantic distinguishability, and (c) mean word duration. It is negatively related to (d) word duration variability and positively related to (e) LM predictability, with little association with (f) training frequency. (g) Words grouped by accuracy bins.
Core
Expanded
Full
Prompt
λ
LM only
Brain + LM
LM only
Brain + LM
LM only
Brain + LM
A
0.25
97.3
55.2±1.2
98.1
53.9±0.8
97.7
54.5±1.0
B
0.5
96.0
51.3±3.6
91.0
52.5±2.6
93.3
51.9±3.1
C
0.25
91.8
51.8±3.3
90.0
49.8±3.1
90.8
50.7±3.2
D
0.25
84.4
49.3±2.0
95.0
48.1±3.2
90.0
48.7±2.6
E (ours)
0.5
73.8
36.6±1.7
91.8
44.8±3.7
83.4
41.0±2.2
Appendix
Table 7: Prompt specificity improves sentence reconstruction. Brain + LM uses k=5 , beam width 50, and frozen Qwen3-8B-Base. Subscripts report sample standard deviations. LM only decoding is deterministic. Bold, shaded cells mark the lowest WER within each condition and split. A line break before Patient: is part of prompts E and F.
Reference p
k
Core
Expanded
Full
SUBTLEX-UK
1
1.40±0.09
1.05±0.03
1.23±0.05
5
6.41±0.48
6.09±0.83
6.25±0.54
Switchboard
1
1.88±0.12
1.41±0.04
1.66±0.06
5
8.53±0.63
8.11±1.09
8.32±0.71
UCV
1
5.68±0.36
4.29±0.13
5.04±0.18
5
24.32±1.71
23.17±2.97
23.72±1.93
Appendix
Table 8: SimpleB2T normalised OVMI scores across reference distributions. OVMI measures the information conveyed per word by a model relative to a reference communication distribution. Values are 100×OVMI/H(p) , where H(p) is the entropy of the full reference distribution. These estimates use balanced word accuracies of 16.04/13.29/14.80% at k=1 and 46.83/45.08/45.93% at k=5 for Core/Expanded/Full, respectively. Results report means across five training seeds, with sample standard deviations as subscripts. See Jayalath et al. (2026) for further details on OVMI.
ID
Sentence
ID
Sentence
1
Can you help me?
51
Can you open the door?
2
I would like some help.
52
Can you open the window?
3
Can you come here?
53
Can you put the light on?
4
Can you come back?
54
Can you take this away?
5
Can you be here with me?
55
Can you put that over me?
6
Can you give me more time?
56
Can you put this by my hand?
Appendix
Table 9: Core communication benchmark sentences.
ID
Sentence
ID
Sentence
1
I would like some water now.
51
Tell her to come back in a little while.
2
Give me a little water now.
52
Tell him to come back before night.
3
I would like water before you go.
53
I would like to see them now.
4
Take the water away now.
54
I would like to have more time with her.
5
I have enough water for now.
55
I would like to have more time with him.
6
I would like the water by my hand.
56
Ask them to be here with me.
Appendix
Table 10: Expanded communication benchmark sentences.
Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine-grained reconstruction of phonemes or individual words difficult. Motivated by neuroscientific evidence that high-level semantic representations are distributed across cortical regions and evolve over slower temporal scales, we hypothesize that semantic content may provide a more suitable target for non-invasive decoding than low-level acoustic or lexical features. We introduce Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space. Our model maps sentence-level MEG responses into a semantic manifold and then inverts the predicted embeddings into natural language. This semantic bottleneck enables recovery of high-level meaning without word-level alignment. We describe the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping. Finally, we compare against prior non-invasive Brain2Text methods and show improved sentence-level results.
Gilad D. Landau, Dulhan Jayalath, Oiwi Parker Jones
The loss of speech limits communication for individuals with paralysis. Direct neural-to-speech synthesis is challenging due to the limited availability of neural data for training speech brain-computer interfaces. Most existing systems rely on cascaded neural-to-text-to-speech pipelines, which increase inference latency and propagate errors across stages. We present Brain2Speech-Net, a single-stage neural-to-speech generation framework without intermediate text decoding. We use a differentiable phoneme bottleneck and a deep-HMM alignment mechanism to map long neural recordings into the latent space of a text-to-speech (TTS) model, enabling high-quality speech synthesis. Brain2Speech-Net is the only system in our comparison that produces intelligible speech while generating faster than real time.
Shreeram Suresh Chandra, Zexin Cai, Yu Tsao +2
Johns Hopkins University, USA · Academia Sinica, Taiwan · University of Edinburgh, UK
Brain-to-language decoding translates neural activity associated with language production, internal speech and perception into linguistic or expressive outputs. It offers a route to restoring communication after speech loss and a means of studying how the brain represents language. Advances in neural recording and representation learning have expanded the field from constrained recognition and acoustic reconstruction to text generation, streaming personalised speech and facial animation. This survey synthesises these developments across invasive and non-invasive measurements, drawing on a search without a lower year limit and source-led updates through September 2026. We connect Articulated, Inner and Perceived tasks to the neural populations they engage, the representations available to decoders and the outputs those representations can support. We examine model development, public resources and the evolution of evaluation, and compare published performance and communication costs within their reported protocols. The synthesis identifies complementary routes to progress: phonetic, acoustic and semantic targets preserve different aspects of a message; shared representations support reuse across recording conditions and tasks; and online communication increasingly depends on calibration, feedback and user control alongside decoding accuracy. Shared benchmarks enable algorithmic comparisons, while longitudinal studies reveal the demands of sustained use. We discuss these developments and their remaining limitations, then outline a prospective five-level trajectory from commands and language to meaning, scenarios and bidirectional cognitive exchange
Yiqian Yang, Yiqun Duan, Chenyu Liu +4
Uploading Inc · Human-centric Artificial Intelligence Centre, University of Technology Sydney · Stanford Institute for Human-Centered Artificial Intelligence (HAI), Stanford University