cs.AIOct 1, 2026

Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills

Authors: Ngoc Phuoc An Vo, Aarya Doshi, Vadim Sheinin

Organizations: IBM · Georgia Institute of Technology · IBM Research

Abstract

Enterprise AI agent skills evolve as tool APIs, models, and specifications change, yet final-output evaluation can miss process-level behavioral drift. We present a continuous evaluation framework combining outcome-level and process-level checks, applied to Revenue and Productivity variants of a Business Value Determination skill in an enterprise Value Aware Resiliency system. The framework independently computes per-run ground truth, materializes reusable template tests, and evaluates tool selection, arguments, execution order, and database integrity through programmatic checks and a narrowly scoped LLM judge. We evaluate 240 trials across two skills, two specification variants, two agent harnesses, and three models. Of 175 trials passing all applicable final numerical checks, 162 (92.6 percent; Wilson 95 percent CI: 87.7-95.6 percent) contained another evaluator-detected deviation. Under a broader seven-check final-state definition, 151 of 164 passing runs (92.1 percent; 95 percent CI: 86.9-95.3 percent) still violated a trajectory check. Dependency attribution reduced a mean of 6.34 failed checks per run to 2.65 roots. Specification sensitivity varied by model and harness, with exploratory bootstrap interaction intervals excluding zero for all three Revenue comparisons and one of three Productivity comparisons. Runtime-resolved templates provided reusable regression coverage across the evaluated configurations; longitudinal validation under actual API evolution remains future work.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing

    Jun 12, 2026Haowen Gao, Haoran Chen, Can Wang +5Skill EvolutionSkills

  2. SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use

    Jul 2, 2026Jiayin Zhu, Kelong Mao, Yudong Guo +4SkillsRubrics

  3. Agent Skill Evaluation and Evolution: Frameworks and Benchmarks

    Jun 9, 2026Kexin Ding, Yang Zhou, Can Jin +3Self-Evolving Skill LibrariesEvolutionary Optimization Methods