LLMs are General Asynchronous Agents
Organizations: Yandex · Together AI · HSE university · Yandex School of Data Analysis
Abstract
Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform another task. Modern LLMs address this with specialized architectures for voice interaction and video streams, VLAs for robot control, asynchronous tool calling for API usage, and others. In this work, we generalize from different asynchronous tasks to general asynchronous agents that can adapt to different types of concurrency. To achieve this, we develop an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states. We showcase that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.
Figures & tables
| Agent & Model | SoccerNet (streaming) | ProactiveVideoQA PAUC ( ) | Trigger | TimVal | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| AUROC | Trigger | TimVal | WEB | EGO | TV | VAD | ALL | ALL | ALL | |
| Qwen3.5-9B | 0.677 | 62.82 | 39.62 | 0.493 | 0.563 | 0.638 | 0.358 | 0.541 | 52.99 | 17.19 |
| Qwen3.8-27B | 0.608 | 54.88 | 33.27 | 0.501 | 0.482 | 0.616 | 0.366 | 0.504 | 52.45 | 16.82 |
| Q3.6-35B-A3B | 0.651 | 62.55 | 37.46 | 0.476 | 0.578 | 0.624 | 0.344 | 0.539 | 52.30 | 16.37 |
| Mage-VL | 0.555 | 52.79 | 27.87 | 0.323 | 0.538 | 0.390 | 0.284 | 0.428 | 43.27 | 10.03 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Pattern | Training-free implementations | Implementations with additional training |
|---|---|---|
| Probes and monitors | Prompted interruption classification ( Cao et al., 2025 ) ; reasoning-safety monitoring ( Wang et al., 2026b ) | EgoSpeak ( Kim et al., 2025a ) ; StreamMind ( Ding et al., 2025 ) ; LTS-VoiceAgent’s semantic trigger ( Zou et al., 2026 ) |
| Parallel inference streams | Skeleton-of-Thought ( Ning et al., 2024 ) ; Hogwild! ( Rodionov et al., 2025 ) ; Asynchronous Reasoning ( Yakushev et al., 2025 ) | Moshi ( Défossez et al., 2024 ) ; Hume ( Song et al., 2025 ) ; StreamingThinker ( Tong et al., 2025 ) |
| Subroutines and delegation | LLMCompiler ( Kim et al., 2024 ) ; ReDel ( Zhu et al., 2024 ) ; RLM ( Zhang & Khattab, 2025 ) | PASTA ( Jin et al., 2025 ) ; Multiverse ( Yang et al., 2026b ) ; Parallel-R1 ( Zheng et al., 2026a ) |
| Interruptions and incremental inputs | Asynchronous Reasoning ( Yakushev et al., 2025 ) ; AsyncLM’s prompted GPT-4o variant ( Gim et al., 2024 ) | StreamChat ( Liu et al., 2024b ) ; VITA-E ( Liu et al., 2025b ) ; AsyncLM’s fine-tuned Llama variant ( Gim et al., 2024 ) |
| Shared and evolving context | Hogwild! ( Rodionov et al., 2025 ) ; LiveVLM ( Ning et al., 2025 ) | StreamingVLM ( Xu et al., 2026 ) ; VideoStreaming ( Qian et al., 2024 ) |
| Our framework | All five patterns through a common inference interface | No additional training required |
| Qwen3.5-9B | Qwen3.8-27B | Qwen3.6-35B-A3B | ||||
|---|---|---|---|---|---|---|
| Coroutines | CUDA graphs | Eager | CUDA graphs | Eager | CUDA graphs | Eager |
| 1 | 106 | 30 | 43 | 16 | 96 | 17 |
| 2 | 192 | 59 | 80 | 31 | 169 | 34 |
| 4 | 339 | 116 | 144 | 60 | 284 | 67 |
| 8 | 552 | 220 | 239 | 116 | 438 | 131 |
| 16 | 817 | 408 | 358 | 221 | 624 | 246 |
| Operation | Both eager | Decode graphs | Both CUDA graphs |
|---|---|---|---|
| 32-tok context prefill | 63.4 | 64.4 | 25.1 |
| 30-tok probe (3 ctx blocks) | 65.5 | 65.5 | 25.8 |
| 256-tok context prefill | 74.8 | 64.5 | 28.8 |
| 1024-tok context prefill | 88.1 | 66.1 | 65.0 |
| 4096-tok context prefill | 179.6 | 178.3 | 649.7 |
| 4096-tok flat prefill | 136.3 | 135.0 | 302.7 |
| Source | Original split | Pairs |
|---|---|---|
| MathVista ( Lu et al., 2024 ) | testmini | 100 |
| MathVision ( Wang et al., 2024b ) | test | 72 |
| CharXiv ( Wang et al., 2024f ) | validation | 64 |
| ChartQA ( Masry et al., 2022 ) | test | 72 |
| TabMWP ( Lu et al., 2022 ) | test | 113 |
| MapQA-U ( Chang et al., 2022 ) | test | 32 |
| Segment sec. | FPS | #Frames | Codec | TriggerAcc | TimVal | ROC-AUC |
|---|---|---|---|---|---|---|
| 16 | 2 | 32 | default | 78.59 | 8.52 | 56.69 |
| 8 | 2 | 16 | default | 84.82 | 9.62 | 57.75 |
| 8 | 1 | 8 | default | 85.50 | 8.44 | 58.49 |
| 8 | 4 | 32 | default | 85.58 | 8.94 | 57.84 |
| 4 | 2 | 16 | default | 90.63 | 4.31 | 53.73 |
| 8 | 2 | 16 | HEVC | 52.79 | 27.87 | 55.50 |
| Model | Domain | PAUC ( ) | PAUC ( ) | PAUC ( ) |
|---|---|---|---|---|
| Qwen 9b | WEB | 0.4033 | 0.4932 | 0.5832 |
| EGO | 0.4965 | 0.5630 | 0.6295 | |
| TV | 0.5400 | 0.6378 | 0.7355 | |
| VAD | 0.3241 | 0.3583 | 0.3925 | |
| ALL | 0.4638 | 0.5409 | 0.6179 | |
| Qwen27b | WEB | 0.4128 | 0.5009 | 0.5890 |
| Model | Qwen 9b | Qwen 27b | Qwen 35a3 |
|---|---|---|---|
| WEB | 3.8 | 2.3 | 2.0 |
| EGO | 4.2 | 2.9 | 2.3 |
| TV | 1.8 | 1.4 | 1.4 |
| VAD | 5.1 | 3.4 | 3.0 |