LLM agents generate intermediate reasoning and actions token by token, making extended interactions slow and computationally expensive. Jev-style models offer fast probabilistic predictions over finite fields, but require those fields to be specified in advance. This requirement limits autonomous task solving, where the available actions must be derived from natural language instructions and adapted through interaction. We introduce JevSpawn, a compositional policy that connects natural language task specifications to finite probabilistic exploration. Parallel action spawning is coupled with feedback driven branch selection, representation revision, and recovery from retained alternatives. Shared action structure and model prefixes reduce repeated generation and context computation without additional training. Evaluations on eight benchmark tasks against seven agent baselines and a TypeSafe Jev variant establish JevSpawn as a promising approach to structured agentic inference, with improved task performance and faster navigation.
Figures & tables
(a) Quality and latency
(b) Task scores over time
Figure 2 : JevSpawn architecture. (a) Compositional actions derived from natural language. (b) Probabilistic spawning, feedback-driven transitions, and recovery. (c) Finite action probabilities. (d) Shared computation in the interaction loop.
Figure 3 : Stability of the quality–latency tradeoff under paired bootstrap resampling. Symbols mark paired bootstrap medians, and error bars show 95% marginal intervals from 2,000 resamples. Horizontal bars show Pareto frontier frequency under resampling of paired instances.
Figure 4 : Instance-level effects of component ablations. Stacked bars show fractions with lower, equal, and higher scores than complete JevSpawn. Boxes summarize paired latency ratios with medians, interquartile ranges, and fifth to ninety-fifth percentile whiskers. Ratios below one indicate shorter execution.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Request scheduling and prefix reuse during JevSpawn execution on PPNL. (a) Request durations and returned observations during asynchronous execution. Teal spans indicate spawn scoring, with the mean batch latency annotated. Text generation can span intervening finite evaluations. (b) Computed and reused input tokens for successive finite evaluation batches.
Figure 6 : GPU activity during finite action evaluation of eight requests and 77 assignments. Teal indicates model computation, purple indicates collective communication, and light gray indicates gaps between kernels. Each timeline begins at the first kernel on the corresponding GPU.
Figure 7 : Paired task times for JevSpawn and AgentPrune. Each point represents the same instance under both methods. Points below the diagonal indicate shorter JevSpawn execution. Filled teal, gray, and open navy marks denote higher, equal, and lower JevSpawn scores. Crosses identify pairs containing a timeout. Marginal histograms show elapsed time distributions. The dark curve and band summarize median and interquartile JevSpawn times within equal-count groups of AgentPrune times, excluding timeout pairs.
Figure 9 : Task performance and E2E latency across expansion widths. Circled labels indicate the maximum number of spawned actions, K . Dotted lines mark the reference setting, K=4 .
Figure 10 : Task performance and timeout incidence across interaction horizons. Shading separates scores with a fixed deadline and unrestricted time. Brackets mark the gap at 108 rounds, with timeout rates below.
Figure 11 : Throughput, latency, and memory under increasing physical batch size. (a) The throughput curve highlights the best measured batch and the gain over batch 8. (b) ITL and TTFT are normalized to batch 8, with absolute values annotated. (c) Memory intervals connect allocated and reserved maxima across four GPUs. (d) Paired finite throughput measurements contrast short and long histories, with ratios and batch latencies shown. Finite throughput includes context processing. The next tested text and finite batches, 128 and 32, exceed memory.