LLM-driven agentic applications automate complex, multi-step tasks, but serving them efficiently remains difficult due to heterogeneous components, dynamic model-driven control flow, long-lived state, and highly variable latencies. Nalar is a serving framework for agent workflows that separates workflow specification from execution while providing the runtime visibility and control needed for robust performance. Nalar preserves ordinary Python interfaces and control flow through lightweight auto-generated stubs that turn agent and tool invocations into futures carrying dependency and execution-context metadata. A two-level control architecture combines global policy computation with local event-driven enforcement to support adaptive routing, scheduling, and resource management across evolving workflows. A workflow-aware KV-cache layer enables the runtime to manage cache placement and lifetime. Together, these mechanisms enable scalable, efficient, policy-driven serving of heterogeneous agentic applications without burdening developers with orchestration logic. Across three agentic workloads, Nalar reduces tail latency by 34-74% and achieves up to 3.38x speedups.
Figures & tables
Figure 1 . An example agentic application: OpenSWE ( LangChain, 2025 ) workflow for software development.
Figure 2 . Challenges serving agentic application: Serving concurrent agent workflows without global orchestration. Independent component scheduling leads to head-of-line blocking and wasted recomputation, demonstrating the need to account for workflow progress as well as cluster-wide resource information.
Figure 3 . Workload heterogeneity in a workflow.
Figure 4 . Nalar Overview: Nalar takes user-specified files and generates stubs (§ 3.1 , § 3.2 ) that replace original function calls with controllable hooks to generate futures (§ 3.3 ). These stubs connect the user program to the runtime controllers. At deployment, Nalar launches and manages the runtime (§ 4 ), where component-level controllers and the global controller coordinate to enforce scheduling, routing, and resource policies while managing KV caches.
Figure 5 . Example Agent : A software developer agent definition. It calls a documentation lookup tool, a shared inference engine and another testing agent. These calls look like calls to local objects.
Figure 6 . Three-agent workflow: The planner agent decomposes a natural-language coding request into subtasks. Each subtask is sent to a Developer agent from Figure 5 , which returns a future indicating test success or failure. The program creates and consumes these futures, retrying failing subtasks. Unless the user explicitly desires(Line 29) they don’t need to interact with futures.
Metadata
Structure
Description
dependencies
list(agentA:ip, …)
List of dependencies required to compute the future’s output
creator
agentName:ip
Component that created the future
executor
agentName:ip
Component instance currently assigned to execute the future
consumers
list(future-id, …)
Components waiting to consume the future’s value
session-id
str:UUID
The user-request associated with the future
Table 1 . Nalar ’s future metadata.
Figure 7 . Future Generation Timeline: For the agent workflow depicted in Figure 6 we depict a timeline for future generation and how their consumers are updated and their values realized in Nalar .
Figure 8 . Nalar ’s two-level control. Each component has an associated controller with it. Each node has a local node store. The global controller communicates with each agent and workflow driver, through the node store.
Hint
Values
Descriptions
stateful
Boolean
True indicates for a session successive calls to the agent will be routed to the same instance
batchable
Boolean
True indicates that module can accept a batch of request
reusable
Boolean
Resource created during a session can be reused by subsequent invocations within the same session.
preemptable
function
True indicates that running invocations of the component may be preempted by the runtime.
max_instances
Integer
Indicates the max number of instances to initialize
min_instances
Integer
Indicates the min number of instances the framework should keep alive
Table 2 . Nalar ’s hint interface
Figure 9 . End-to-End Evaluation The bars represent average latency, the whiskers represent P50, P95 and P99 latencies.
Figure 12
Number of Futures
One-Level Design
Two-level Design
Time(ms)
Time(ms)
1024
1.2
0.1
2048
2.3
0.1
4096
2.8
0.2
8192
3.4
0.4
16384
3.9
0.4
Table 3 . Impact of Two-level Control.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Metric
Description
Future metadata
Session ID
UniqueID associated with a session.
Future ID
Unique ID associated
Component
Identifies the agent or tool invoked by the future.
Creator
Component that created the future.
Executor
Agent or tool instance assigned to execute the future.
Dependencies
Futures whose values must be available before the future becomes ready.
Appendix
Table 4 . Nalar telemetry collected by component-level controllers. Per-future telemetry captures workflow dependencies and timing, while instance-level metrics summarize local load and resource availability.
Figure 13 . Control Interaction in Nalar . The above figure shows the interaction and relevant updates to metadata when a future is being migrated. An important feature is that it’s entirely locally coordinated, i.e. , the global controller only issues the migrate command, the component level controllers coordinate it among themselves.
As LLM applications grow more complex, developers are increasingly adopting multi-agent architectures to decompose workflows into specialized, collaborative components, introducing structure that constrains agent behavior and exposes useful semantic predictability. Unlike traditional LLM serving, which operates under highly dynamic and uncertain conditions, this structured topology enables opportunities to reduce runtime uncertainty\unicodex2015yet existing systems fail to exploit it, treating agentic workloads as generic traffic and incurring significant inefficiencies. Our analysis of production traces from an agent-serving platform and an internal coding assistant reveals key bottlenecks, including low prefix cache hit rates, severe resource contention from long-context requests, and substantial queuing delays due to suboptimal scaling. To address these challenges, we propose Pythia, a multi-agent serving system that captures workflow semantics through a simple interface at the serving layer, unlocking new optimization opportunities and substantially improving throughput and job completion time over state-of-the-art baselines.
Agentic AI applications form an emerging serving workload in which a request creates a workflow: a directed acyclic graph of LLM and tool calls that exposes per-node model choices and optional quality operators such as verifiers. This workload falls between two existing layers. Model-serving engines execute individual calls efficiently but cannot see workflow structure, while agent frameworks fix the workflow but cannot see backend load, so neither jointly chooses each node's model, verifier, and backend under serving-time conditions. We present Dyserve, a workflow-aware serving layer that fills this gap. Dyserve compiles each workflow's per-node model and verifier choices in one integer linear program (ILP) over a heterogeneous backend pool, priced by skill-conditioned offline profiles that transfer across workflows. This couples with hardware entering only through per-model throughput sweeps, and is weighted to concentrate strong models and verification on the nodes whose errors propagate the furthest. Because no single latency-quality preference fits every workload mix, Dyserve pre-solves the program at several pressure levels at admission and shifts a workflow's uncommitted suffix among these strategies under load, keeping the solver off the load-shift path; a failed tool call triggers a one-time residual re-solve that preserves committed work.
Jiayi Qian, Zishen Wan, Hanchen Yang +3
Columbia University New York, NY, USA · Intel Los Angeles, CA, USA
Large Language Model (LLM)-based agents demonstrate strong reasoning and execution capabilities on complex tasks when guided by structured instructions, commonly referred to as workflows. However, existing workflow-assisted agent serving systems typically rely on predefined templates and shallow matching mechanisms, which limit their ability to capture deep semantic relationships and generalize to previously unseen tasks. To address these limitations, we propose a new workflow management paradigm that represents workflows using a unified graph, termed wGraph, where each node corresponds to an atomic operation. wGraph serves as a shared substrate from which task-specific workflows are dynamically instantiated. Building on wGraph primitives, we introduce GraphFlow, a system that efficiently integrates workflows into agent serving through two key designs. First, adaptive workflow generation dynamically constructs workflows from wGraph based on task semantics and constraint requirements. Second, workflow state management exploits wGraph structure to efficiently manage Key-Value (KV) caches, reducing redundant computation during agent serving. Extensive experiments across five benchmark datasets show that GraphFlow consistently outperforms state-of-the-art methods, yielding an average performance improvement of approximately 4.95 percentage points, while achieving an approximately 4× reduction in memory footprint.
Ao Li, Shangpeng Yang, Fahao Chen +3
School of Cyber Science and Engineering, Xi’an Jiaotong University, Xi’an, China · School of Artificial Intelligence, Shandong University, Jinan, China · 3Shanghai Advanced Research Institute, Chinese Academy of Sciences, Shanghai, China.