Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform another task. Modern LLMs address this with specialized architectures for voice interaction and video streams, VLAs for robot control, asynchronous tool calling for API usage, and others. In this work, we generalize from different asynchronous tasks to general asynchronous agents that can adapt to different types of concurrency. To achieve this, we develop an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states. We showcase that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.
Figures & tables
Figure 1: An example of AsyncLLM agent design with two parallel sub-agents solving a math task with shared memory: (left) the agent is defined as a set of coroutines that write into shared cache blocks; (middle) the cache blocks are arranged in “views” that define how coroutines see each other’s work; (right) the engine groups coroutines into batches for efficient inference. Details in Section 3 .
Figure 3: Accuracy with one asynchronous input; the x-axis denotes when the input arrived (step #). (left) text clarifications on MATH-500-Sharded, (right) image changes on ShardedVQA (Section 4.1 ).
Agent & Model
SoccerNet (streaming)
ProactiveVideoQA PAUC ( ω=0.5 )
Trigger Acc
TimVal
AUROC
Trigger Acc
TimVal
WEB
EGO
TV
VAD
ALL
ALL
ALL
Qwen3.5-9B
0.677
62.82
39.62
0.493
0.563
0.638
0.358
0.541
52.99
17.19
Qwen3.8-27B
0.608
54.88
33.27
0.501
0.482
0.616
0.366
0.504
52.45
16.82
Q3.6-35B-A3B
0.651
62.55
37.46
0.476
0.578
0.624
0.344
0.539
52.30
16.37
Mage-VL
0.555
52.79
27.87
0.323
0.538
0.390
0.284
0.428
43.27
10.03
Table 1: Streaming video understanding evaluation with AsyncLLM Qwen 3.x models and a non-AsyncLLM Mage-VL streaming model: (left) SoccerNet results in streaming video understanding protocol, (right) ProactiveVideoQA evaluation across four subsets and totals, using the metrics from the original evaluation protocol (all higher is better). See details and discussion in Section 4.2 .
Figure 4: AsyncLLM Qwen3.6-35B-A3B evaluation on VizDoom environments (left) DoomHealthGathering and (right) DoomDeadlyCorridor; the y-axis is the average reward and the x-axis is the mean delay (forward passes) from observation to taking an action, averaged over 100 episodes.
Table 5
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Pattern
Training-free implementations
Implementations with additional training
Probes and monitors
Prompted interruption classification ( Cao et al., 2025 ) ; reasoning-safety monitoring ( Wang et al., 2026b )
EgoSpeak ( Kim et al., 2025a ) ; StreamMind ( Ding et al., 2025 ) ; LTS-VoiceAgent’s semantic trigger ( Zou et al., 2026 )
Parallel inference streams
Skeleton-of-Thought ( Ning et al., 2024 ) ; Hogwild! ( Rodionov et al., 2025 ) ; Asynchronous Reasoning ( Yakushev et al., 2025 )
Moshi ( Défossez et al., 2024 ) ; Hume ( Song et al., 2025 ) ; StreamingThinker ( Tong et al., 2025 )
Subroutines and delegation
LLMCompiler ( Kim et al., 2024 ) ; ReDel ( Zhu et al., 2024 ) ; RLM ( Zhang & Khattab, 2025 )
PASTA ( Jin et al., 2025 ) ; Multiverse ( Yang et al., 2026b ) ; Parallel-R1 ( Zheng et al., 2026a )
Interruptions and incremental inputs
Asynchronous Reasoning ( Yakushev et al., 2025 ) ; AsyncLM’s prompted GPT-4o variant ( Gim et al., 2024 )
StreamChat ( Liu et al., 2024b ) ; VITA-E ( Liu et al., 2025b ) ; AsyncLM’s fine-tuned Llama variant ( Gim et al., 2024 )
Shared and evolving context
Hogwild! ( Rodionov et al., 2025 ) ; LiveVLM ( Ning et al., 2025 )
StreamingVLM ( Xu et al., 2026 ) ; VideoStreaming ( Qian et al., 2024 )
Our framework
All five patterns through a common inference interface
No additional training required
Appendix
Table 4: Representative concurrency mechanisms grouped by pattern and training requirements. Patterns are non-exclusive, so a method may appear in multiple rows. Training-free means no additional method-specific parameter updates. AsyncLM appears in both columns because its reported variants differ.
Qwen3.5-9B
Qwen3.8-27B
Qwen3.6-35B-A3B
Coroutines
CUDA graphs
Eager
CUDA graphs
Eager
CUDA graphs
Eager
1
106
30
43
16
96
17
2
192
59
80
31
169
34
4
339
116
144
60
284
67
8
552
220
239
116
438
131
16
817
408
358
221
624
246
Appendix
Table 5: AsyncLLM decoding throughput (total across coroutines) for different models with and without CUDA graphs with different number of active coroutines, 1× H200.
Figure 5: Comparison of GPU inference throughput, decode step latency and prefill latency under load across model sizes. We use 1 × H200 GPU with synthetic decode and prefill requests. Our agents in Section 4 have, on average, 1.5 to 3.5 simultaneous active coroutines.
Operation
Both eager
Decode graphs
Both CUDA graphs
32-tok context prefill
63.4
64.4
25.1
30-tok probe (3 ctx blocks)
65.5
65.5
25.8
256-tok context prefill
74.8
64.5
28.8
1024-tok context prefill
88.1
66.1
65.0
4096-tok context prefill
179.6
178.3
649.7
4096-tok flat prefill
136.3
135.0
302.7
Appendix
Table 6: GPU inference latency for prefill, probe, and decode coroutines in different setups, 1× H200. Comparing eager prefill and decode (left), CUDA graphs on decode but not prefill (middle), and CUDA graphs for both prefill and decode (right).
Source
Original split
Pairs
MathVista ( Lu et al., 2024 )
testmini
100
MathVision ( Wang et al., 2024b )
test
72
CharXiv ( Wang et al., 2024f )
validation
64
ChartQA ( Masry et al., 2022 )
test
72
TabMWP ( Lu et al., 2022 )
test
113
MapQA-U ( Chang et al., 2022 )
test
32
Appendix
Table 7: Dataset composition. Each example contains an initial image and a corrected image.
Figure 6: Examples from the ShardedVQA dataset. (Upper) The question is: “Is the sum of the smallest two values greater than the largest value?” (Lower) The question is: “Ruth runs around the perimeter of the pool while Sarah swims its length. Ruth runs three times as fast as Sarah swims. Sarah swims six lengths in the same time Ruth completes five laps. How wide is the pool?” (Left) Before correction, edited image. (Right) Correct image, from source.
Figure 7: A summary of AsyncLLM agent design for streaming video understanding.
Segment sec.
FPS
#Frames
Codec
TriggerAcc
TimVal
ROC-AUC
16
2
32
default
78.59
8.52
56.69
8
2
16
default
84.82
9.62
57.75
8
1
8
default
85.50
8.44
58.49
8
4
32
default
85.58
8.94
57.84
4
2
16
default
90.63
4.31
53.73
8
2
16
HEVC
52.79
27.87
55.50
Appendix
Table 8: Hyperparameter tuning for SoccerNet-Caption: we use the official Mage-VL streaming inference pipeline (inference_streaming.py) and vary segment length, FPS, num. frames, and codec settings. We chose the row highlighted in the green as our main evaluaiton configuration on the balance of metrics.
Model
Domain
PAUC ( ω=0 )
PAUC ( ω=0.5 )
PAUC ( ω=1 )
Qwen 9b
WEB
0.4033
0.4932
0.5832
EGO
0.4965
0.5630
0.6295
TV
0.5400
0.6378
0.7355
VAD
0.3241
0.3583
0.3925
ALL
0.4638
0.5409
0.6179
Qwen27b
WEB
0.4128
0.5009
0.5890
Appendix
Table 9: Additional evaluations on ProactiveVideoQA sub-domains using the default evaluation protocol with varying ω parameter. Intuitively, ω=1 does not take reply time into account, ω=0 takes reply time into account fully, and ω=0.5 (recommended) takes reply time into account with half weight, see ( Wang et al., 2025b ) for details.
Model
Qwen 9b
Qwen 27b
Qwen 35a3
WEB
3.8
2.3
2.0
EGO
4.2
2.9
2.3
TV
1.8
1.4
1.4
VAD
5.1
3.4
3.0
Appendix
Table 10: AsyncLLM inference speed in seconds per second at 1× H200 on the 4 subset of ProactiveVideoQA. The TV subset has both visual and speech streams, the others are visual-only. Time varies because of how many inference steps it takes to process an average event from a given subset.
Function calling, also known as tool use, is a core capability of modern LLM agents but is typically constrained by synchronous execution semantics. Under these semantics, LLM decoding is blocked until each function call completes, resulting in increasing end-to-end latency. In this work, we introduce AsyncFC, a pure execution-layer framework that decouples LLM decoding from function execution, enabling overlap between model decoding and function execution as well as inter-function parallelism when dependencies permit. AsyncFC layers over existing models and unmodified function implementations, requiring no fine-tuning or changes to the standard synchronous function-calling protocol. Across standard function-calling benchmarks and adapted software engineering benchmarks, AsyncFC significantly reduces end-to-end task completion time while preserving task accuracy. Furthermore, these results reveal that LLMs possess a native capability to reason over symbolic futures that represent unresolved execution results, enabling an asynchronous paradigm for model-tool interaction.
Every major LLM agent framework gives the LLM the role of orchestrator; the model decides what to do next, when to call tools, and when to stop. We argue that token explosion, control-flow hallucination, and unreliable completion are not implementation bugs but architectural consequences of assigning the deterministic work of looping, branching, and sequencing to a probabilistic system. A better prompt or a stronger model cannot guarantee the reliability of the LLM agent. We therefore propose Agentic Programming, in which the program governs all control flow, and the LLM is itself part of it, an adaptive component we call LLM-as-Code and invoke only where a task calls for reasoning or generation. Within each call the model keeps full flexibility, but it cannot alter the program's execution path. With control in the program, the LLM's context is built from the execution history's call tree and forms a directed acyclic graph (DAG). Each call's context length is then determined by its call depth rather than by accumulation over steps. A case study of computer-use agents shows that the design is practical, not just a theoretical stance, substantially improving the stability of long visual operation sequences.
Modern language-model agents are built around the agent loop: the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their output. However, certain capabilities such as long-term memory and self-improvement currently require specialized systems beyond the agent loop itself. We built an LLM agent framework, JAZ, to explore the extent to which a minimal harness that is little more than the agent loop itself can accomplish tasks these specialized systems are built for. JAZ exposes a single LLM-based primitive invoke and provides a set of built-in hooks that allow the programmer to apply constraints and perform monitoring. Generalizing existing code-mode agent loops, invoke is the simplest loop that satisfies two defining properties: (1) the LLM can write arbitrary executable code that can include recursive invoke; (2) everything visible to the LLM - all inputs to invoke as well as its interaction history with the code environment - are variables in the code environment. We motivate our design from first principles, viewing invoke as a language primitive representing a function whose implementation is provided at runtime by an LLM every time it is called. To validate the design of our core invoke primitive, we evaluate invoke - with only prompting, no manually designed tools, harness, or external systems (e.g., memory or the file system) - on workflows traditionally implemented through specialized harnesses. On long-horizon workflows requiring recall far beyond the context window, JAZ invoke outperforms Letta (MemGPT) by 8% at half its cost on the recall-heavy portion of StuLife. On continual self-improvement, JAZ invoke outperforms ACE by 4% at a lower cost on AppWorld.