cs.AISep 28, 2026
SaveReport: Progressive Disclosure of Agent Skills
Organizations: Workday AI Research
Abstract
Users of Workday's deployed LLM-based agents often request features which can be addressed by defining named procedures, also known as skills, in the LLM context, effectively augmenting agents' capabilities. However, as an agent's skills library grows in size, so does the agent's operational cost. Progressive disclosure (lazy-loading) of skills as needed may reduce operational costs, but its impact on overall latency and skill-retrieval quality remains unclear. In this report, we investigate the impact empirically and find that progressive disclosure improves skill-retrieval quality but marginally degrades overall latency.
Figures & tables
Figure 1: (a) Under the eager loading regime, skill-retrieval quality degrades as the skills library size grows (Qwen3-8B). (b) Under the eager loading regime, the larger core LLM (Qwen3-14B) is reliable against easy/medium-difficulty skill distractors, but susceptible to hard distractors. (c) Under the progressive disclosure regime, the LLM core is invoked more often which increases latency, especially for smaller values of the library size .
Figure 2: (a) Prompt token usage grows faster in the eager loading regime. (b) Skill-retrieval quality degrades faster in the eager loading regime, as grows. (c) The rate at which each skill management regime overflows the maximum context size at library size .
| Token usage | Skill retrieval | Reliability | ||||||
|---|---|---|---|---|---|---|---|---|
| LLM core | EL tok | PD tok | Tok | EL succ | PD succ | EL crash | PD crash | |
| Qwen2.5-7B | 5 | 1648 | 1215 | 26.3% | 1.00 | 1.00 | 0.00 | 0.00 |
| 20 | 6057 | 1943 | 67.9% | 0.17 | 1.00 | 0.00 | 0.00 | |
| 50 | 14871 | 3481 | 76.6% | 0.40 | 0.67 | 0.00 | 0.00 | |
| 100 | crash | 6223 | — | 0.00 | 1.00 | 1.00 | 0.00 | |
| Qwen3-8B | 5 | 1652 | 1644 | 0.5% | 1.00 | 1.00 | 0.00 | 0.00 |
Table 1: Eager loading (EL) vs. progressive disclosure (PD) results as a function of library size . Tokens are per-task means over rollouts. crash denotes context overflow. Token reductions at are significant at , based on one-sided Mann–Whitney test on per-rollout totals.