Context Spanning: A Communication Framework for Full-Duplex Speech Models and External LLM Backends
Organizations: Mindlogic
Abstract
Full-duplex spoken dialogue models can listen and speak simultaneously like the real-time dynamics of human conversation. For natural dialogue, the ability to search for external information in real-time is also an important capability. Many models remain trapped in parametric knowledge, leaving them unable to access real-time information and tool execution. Furthermore, even when Large Language Models (LLM) retrieve information, many duplex speech models process it within a compressed latent space rather than in its raw text form, which can lead to information loss from compression. To address this issue, we propose Context Spanning, a framework for information injection between a full-duplex speech model and an external LLM backend via real-time chunked prefill. The injected frame is encoded in a single forward pass inside the real-time frame budget. It feeds the retrieved information to the speech model as-is, enabling it to reason over the information independently and generate responses. With this approach, our model achieves high performance on Full-Duplex benchmarks and strong results on Question Answering tasks, demonstrating its conversation potential. Context Spanning shows that external information can be injected directly into a duplex speech model, introducing a new simple and powerful mechanism for duplex systems.
Figures & tables
| (frames) | 0 | 16 | 64 | 256 | 600 |
|---|---|---|---|---|---|
| Latency ( ms ) | |||||
| ( ms ) |
| Spoken QA | Mathematical Reasoning | |||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LlamaQ | WebQ | TriviaQA | HaluEval | AddSub | MultiArith | SinglEq | SVAMP | GSM8K | ||||||||||
| Model | ref. | resp. | ref. | resp. | ref. | resp. | ref. | resp. | ref. | resp. | ref. | resp. | ref. | resp. | ref. | resp. | ref. | resp. |
| GLM-4-Voice | 64.7 | 32.2 | 39.1 | 21.2 | 59.4 | 62.0 | 71.0 | 4.0 | 29.0 | |||||||||
| STITCH-S | 73.3 | 50.2 | 50.0 | – | 81.7 | 87.9 | 91.7 | 72.2 | 56.7 | |||||||||
| MoshiRAG | 83.0 | 80.3 | 71.5 | 67.2 | 73.7 | 69.6 | 42.0 | 36.3 | 76.6 | 61.7 | 87.1 | 69.0 | 83.2 | 68.2 | 74.1 | 55.0 | 66.2 | 33.9 |
| MoshiRAG | 87.8 | 80.6 | 77.7 | 68.9 | 86.8 | 78.2 | 61.2 | 51.3 | 87.9 | 64.8 | 87.1 | 76.0 | 89.6 | 72.9 | 80.5 | 61.1 | 70.8 | 43.2 |
| Task | Metric | Ours | PPlex | Moshi | Gemini |
|---|---|---|---|---|---|
| Turn-Taking | TOR ( ) | 0.899 | 0.992 | 0.941 | 0.655 |
| Latency, s ( ) | 0.064 | 0.070 | 0.265 | 1.301 | |
| Pause Handling | TOR, Candor ( ) | 0.824 | 0.662 | 0.980 | 0.310 |
| TOR, synthetic ( ) | 0.861 | 0.584 | 0.985 | 0.255 | |
| Backchannel | TOR ( ) | 0.782 | 0.327 | 1.000 | 0.091 |
| Frequency, per s ( ) | 0.150 | 0.025 | 0.001 | 0.012 |
| Task | Metric | Ours | MoshiRAG | GPT | Gemini |
|---|---|---|---|---|---|
| Tool Use | Tool selection ( ) | 0.855 | 0.738 | 0.876 | 0.817 |
| Argument accuracy ( ) | 0.567 | 0.440 | 0.680 | 0.588 | |
| Response quality ( ) | 0.411 | 0.255 | 0.792 | 0.718 | |
| Pass rate ( ) | 0.470 | 0.280 | 0.600 | 0.540 | |
| Turn-Taking Dynamics | Take-turn rate, % ( ) | 95.0 | 94.0 | 96.0 | 78.0 |
| Latency, s ( ) | 5.83 | 7.69 | 6.89 | 4.25 |