A Systematic Survey of Agentic Skills: Architecture, Lifecycle, and Security
Organizations: Google LLC · Purdue University
Abstract
Autonomous large language model (LLM) agents increasingly face reliability, context consumption, and execution stability bottlenecks when deployed on complex, long-horizon tasks. While monolithic prompt engineering and stateless tool-calling paradigms struggle to scale, the field is rapidly converging toward \emph{agentic skills}: modular procedural abstractions that externalize execution knowledge into reusable, executable, and portable artifacts. This paper establishes a unified systems foundation and reference architecture for the agentic skills ecosystem. We formalize skills as externalized procedural knowledge bridging high-level cognitive planning with deterministic execution environments, and systematically delineate the architecture across a nine-stage lifecycle: autonomous discovery, authoring and representation formats, memory storage, dynamic retrieval and routing, composition and orchestration, execution and repair, lifelong adaptation, empirical evaluation, and security governance. We further examine marketplace dynamics, public registries, and emerging adversarial threat vectors, alongside runtime verification and defense mechanisms. Finally, we categorize system implementations across software engineering, operating system navigation, embodied robotics, and scientific discovery, while highlighting critical open challenges in continual learning and benchmark realism. This work establishes agentic skills as a foundational paradigm for building scalable, robust, and verifiable autonomous language agents.
Figures & tables
| Abstraction | Components Instantiated | Persists across resets | Invoked by | Executes outside | Composable at runtime | Modifiable w/o retrain | Verifiable before use | Cost per invocation | Reference |
|---|---|---|---|---|---|---|---|---|---|
| System prompt | × | always-on | × | × | ✓ | × | full, every turn | Yao et al. (2023b) ; Wei et al. (2022) ; Sumers et al. (2024) | |
| Episodic memory | (traces only) | ✓ | similarity match | × | × | ✓ | × | retrieval only | Packer et al. (2023) ; Park et al. (2023) ; Zhou et al. (2026) |
| Workflow / DAG | ( fixed at authoring) | ✓ | scheduler | partial | × (fixed graph) | × | ✓ | per node | Hong et al. (2024) ; Wu et al. (2024) ; Qian et al. (2024) |
| Tool / API | ✓ | model emits call | ✓ | ✓ (stateless) | × (server-side) | ✓ (schema) | per call | Schick et al. (2023) ; Patil et al. (2024) ; Yuan et al. (2024) | |
| Agentic skill | ✓ | conditional on state | ✓ | ✓ | ✓ | partial | metadata, then body | Wang et al. (2024) ; Yang et al. (2026b) ; Lu et al. (2026b) ; Ling et al. (2026) |
| System | Core Contribution | Trajectory Source | Authoring Modality | Failure Signal Usage | Verified Before Commit | Updates Over Time |
| Autonomous Reinforcement Learning (RL-Induced Procedural Abstractions) | ||||||
| SkillRL Xia et al. (2026a) | RL-induced procedural policy | Task rollout | Autonomous RL | Contrastive pairs | Unit test | Policy, membership |
| Skill-Pro Mi et al. (2026) | Non-parametric PPO over Skill-MDP | Task rollout | Autonomous RL | Contrastive pairs | PPO Gate | Membership, instructions |
| Skill-R1 Vishe et al. (2026) | RL-based skill evolution | Task rollout | Autonomous RL | Contrastive pairs | Unit test | Membership |
| Skill0 Lu et al. (2026a) | In-context RL internalization | Task rollout | Autonomous RL | Contrastive pairs | None | Policy (internalized) |
| AI-Synthesized Skills (LLM-Generated SOPs, Trajectory Consolidation & Textual Evolution) | ||||||
| Format Class | Pre-Execution Verifiability | Portability | Execution Capability | Injection Channel | Load-Time Token Cost | Repair Signal on Failure | Representative Exemplars |
|---|---|---|---|---|---|---|---|
| Natural language | none (no component) | high | none | prose | high | re-prompt only | SKILL.md (AutoSkill) Yang et al. (2026b) , SkillDex Saha and Hemanth (2026) , SkillAnalysis Ling et al. (2026) , PromptSkills Maloyan and Namiot (2026) , ProceduralMining Bi et al. (2026) |
| Structured (JSON / DSL) | (schema) | high | none | schema strings | medium | schema error | EasyTool Yuan et al. (2024) , SkillDroid Chen et al. (2026b) , Toolformer Schick et al. (2023) |
| Executable specification | (static analysis) | medium | arbitrary code | code + prose | high | stack trace | Voyager Wang et al. (2024) , LLMCompiler Kim et al. (2024) , VisProg Gupta and Kembhavi (2023) , SkVM Chen et al. (2026a) , ProgSkills Wang et al. (2025) |
| Contract-based | (pre/postconditions) | low | constrained | code + prose | medium-high | counterexample | ContractSkill Lu et al. (2026b) , SemiA Wen et al. (2026) , SkillAuditSAST Lv et al. (2026) |
| Hybrid & adaptive | schema + (static analysis) | medium | arbitrary code | code + prose | low at load, high on activation | schema error stack trace | GraphSkill Wang et al. (2026b) , PSN Shi et al. (2026) |
| System | Core Contribution | Primary Tier | Index Structure | Eviction / Compression Policy | Failure Mode Addressed |
| Virtual OS Paging & Mutable State (Bounded Memory Management) | |||||
| MemGPT Packer et al. (2023) | OS-level virtual context paging | virtual paging | OS page table | LLM-controlled paging | context overflow |
| SKILL.state Badhe et al. (2026) | Mutable execution state runtime | procedural / state | JSON state schema | Ephemeral reasoning disposal ( ) | context bloat / cost |
| Utility-Aware Curation (Pruning Low-Value, Redundant, or Polluting Skills) | |||||
| SkillOps Pu et al. (2026) | 5-dimension library health maintenance | procedural | topological graph (HSEG) | utility pruning (5-dim score) | skill technical debt |
| MEMP Fang et al. (2026) | Learnable lifelong procedural memory | procedural | dense vector | utility pruning | retrieval pollution |
| System | Core Contribution | Retrieval / Routing Paradigm | Retrieval Granularity | Selection / Ranking Signal | Pollution / Redundancy Control |
| Retrieve-then-Rerank (Dense Embedding Search & Listwise / Contrastive Reranking Pipelines) | |||||
| SkillRet Cho et al. (2026) | Large-scale retrieve-then-rerank IR | Retrieve-then-Rerank | Atomic skill SOP | Dense vector + cross-encoder | Contrastive reranking |
| SkillRouter Zheng et al. (2026) | Full-text retrieve-and-rerank (1.2B pipeline, 80K skills) | Retrieve-then-Rerank | Atomic skill SOP | Semantic sim. + listwise loss | Metadata-only routing blindness (31–44% drop) |
| SkillRAG Su et al. (2026) | Hybrid BM25 skill retrieval | Retrieve-then-Rerank | Atomic skill SOP | BM25 + vector hybrid | None (top- only) |
| Graph-Structured Search (Traversing Structural Dependency & Documentation Graphs) | |||||
| GraphOfSkills Liu et al. (2026b) | Dependency-aware structural retrieval | Graph-Structured Search | Sub-graph workflow | Structural dependency edges | Dependency constraints |
| System | Core Contribution | Orchestration Paradigm | Communication Topology | Task Decomposition Mechanism | Verification & Error Recovery |
|---|---|---|---|---|---|
| SkillGraph Nie et al. (2026) | Multimodal graph topology evolution | Dynamic Graph Topology | Dynamic graph (MMGT) | Query-conditioned graph predictor | Policy gradient on edge log-probs |
| SingleAgentSkills Li (2026) | Compiling multi-agent into single-agent skill selection | Single-Agent Skill Switching | In-context prompt loading | Skill router / selector | Capacity phase transition (sharp drop at critical size) |
| MetaGPT Hong et al. (2024) | SOP-guided multi-agent assembly lines | Multi-Agent SOP Collaboration | Assembly line chain | Role-specific SOP prompts | Role-based code review & testing |
| ChatDev Qian et al. (2024) | Software development agent society | Multi-Agent SOP Collaboration | Assembly line chain | Role-specific SOP prompts | Role-based code review & testing |
| AgentVerse Chen et al. (2024) | Emergent multi-agent peer collaboration | Multi-Agent SOP Collaboration | Dynamic broadcast / peer chat | Conversational dialogue turn | Multi-agent peer critique / voting |
| AutoGen Wu et al. (2024) | Conversable multi-agent programming | Multi-Agent SOP Collaboration | Dynamic broadcast / peer chat | Conversational dialogue turn | Multi-agent peer critique / voting |
| System / Attack | Core Contribution | Threat Vector Class | Target Component | Injection Surface / Channel | Target Lifecycle Stage | Compromise Impact |
|---|---|---|---|---|---|---|
| HiddenCommentInject Wang et al. (2026e) | Invisible HTML comment prompt injection | Indirect Prompt Injection | HTML / Markdown comments | Authoring | Goal hijacking | |
| PromptInject Schmotz et al. (2025) | Markdown instruction prompt injection | Indirect Prompt Injection | Natural language prose | Authoring | Goal hijacking | |
| SeeingIsNot Jia et al. (2026b) | Multimodal hidden instruction attack | Indirect Prompt Injection | Multimodal prose | Authoring | Goal hijacking | |
| PhantomSkill Lin and Yu (2026) | VulMask auxiliary script injection | Indirect Prompt Injection | Code body / auxiliary script | Authoring | Goal hijacking | |
| SkillTrojan Feng et al. (2026) | Split encrypted backdoor composition | Backdoor / Trojan | Code body | Authoring | Sandbox escape | |
| BadSkill Tie et al. (2026) | Model-in-skill weight poisoning | Backdoor / Trojan | Model weights in skill | Authoring | Sandbox escape |
| System / Defense | Core Contribution | Defense Lifecycle Stage | Target Threat Defended | Serving Latency Overhead |
|---|---|---|---|---|
| SemiA Wen et al. (2026) | Constraint-guided formal skill auditing | Pre-Admission Static Audit | Backdoor / Trojan | Zero (offline audit) |
| SkillFortify Bhardwaj () | Formal supply chain security analysis | Pre-Admission Static Audit | Malicious Code | Zero (offline audit) |
| SkillAuditSAST Lv et al. (2026) | Cross-file SAST security auditing | Pre-Admission Static Audit | Malicious Code | Low (SAST filter) |
| SkillSieve Hou and Yang (2026) | Hierarchical triage & static scanning | Pre-Admission Static Audit | Prompt Injection | Low (SAST filter) |
| SkillClone Zhu et al. (2026) | Multi-modal clone detection across 196K skills | Pre-Admission Static Audit | Clone Infringement | Zero (offline audit) |
| FlowGuard Etteib et al. (2026) | Attention & flow graph malware detection | Runtime Guardrail & Filtering | Malicious Code | Low (SAST filter) |
| System / Framework | Core Contribution | Skill Functional Domain | Action Space Granularity | Environment Feedback Loop | Domain Bottleneck Addressed |
|---|---|---|---|---|---|
| ProceduralMining Bi et al. (2026) | Large-scale procedural mining from repos | Instructional SOP (§4.1) | Mined open-source SOP checklists | Static repository mining (no loop) | Manual SOP authoring bottleneck |
| SkillAnalysis Ling et al. (2026) | Structural analysis of natural SOP skills | Instructional SOP (§4.1) | Markdown SKILL.md structure | Empirical structural usage audit | Opaque community skill quality |
| Gorilla Patil et al. (2024) | Massive API calling with REST docs | Tool & API Calling (§4.2) | REST API call w/ doc constraints | AST sub-tree matching (eval) | API doc hallucination & drift |
| Toolformer Schick et al. (2023) | Self-supervised tool-calling token insertion | Tool & API Calling (§4.2) | Special API token [API(...) res] | Self-supervised perplexity filter | Manual tool annotation cost |
| AnyTool Du et al. (2024) | Hierarchical 16,000+ RapidAPI tool calling | Tool & API Calling (§4.2) | Hierarchical API pool selector | Multi-round API execution feedback | Massive API search space saturation |
| EasyTool Yuan et al. (2024) | Concise tool instruction wrappers | Tool & API Calling (§4.2) | Concise unified tool schema | None (standardized prompt format) | Excessive API doc token overhead |
| System / Suite | Core Contribution | Ecosystem / Benchmark Class | Evaluated Scale / Pool Size | Primary Evaluation Metric | Empirical Finding / Bottleneck |
|---|---|---|---|---|---|
| SkillsBench Li et al. (2026b) | First comprehensive agent skill usage benchmark | Lifelong Skill Learning | 87 tasks across 8 domains | Pass@1 task completion rate | Skills consistently improve domain pass rates |
| SkillFlow Zhang et al. (2026e) | 20-workflow lifelong skill evolution suite | Lifelong Skill Learning | 166 tasks / 20 workflow families | Workflow evolution gain ( pass) | Workflow-level skills transfer better than SOPs |
| SkillLearnBench Zhong et al. (2026) | Continual procedural skill learning benchmark | Lifelong Skill Learning | 20 tasks across 15 sub-domains | Skill quality, trajectory & outcome | All CL methods improve over no-skill baseline |
| SkillSec-Eval Badhe and Tiwari (2026) | Lifecycle-aware security & empirical evaluation | Security & Vetting Suite | 327 real-world skills | Attack Success Rate (ASR) | Vulnerabilities span all 5 lifecycle stages |
| HarmfulSkillBench Jiang et al. (2026) | Weaponizing agent skills safety benchmark | Security & Vetting Suite | 200 harmful skills | ASR & guardrail bypass rate | Substantial fraction harmful; bypasses guards |
| SkillSafetyBench Jin et al. (2026) | Comprehensive agent skill safety suite | Security & Vetting Suite | 155 adversarial cases across 47 tasks | Multi-dimensional ASR score | Natural-language skills create stealthy paths |