EIO-Agents: The Missing Semantic Layer for AI Agent Evaluation
Organizations: The University of Chicago, USA
Abstract
AI agents are entering production in increasingly consequential environments without a shared semantic standard for what their evaluations actually mean. Scores, traces, judge outputs, and multi juror findings are increasingly used to justify readiness and release decisions, yet they often do not specify what evidence supports a claim, what that evidence can establish, or how the claim leads to a decision. We introduce EIO-Agents, an open specification for interoperable AI agent evaluation built on two layers. The Evaluation Intelligence Ontology (EIO) provides the semantic layer through typed evidence, versioned behavioral predicates, evidence contracts, claims, witness rules, proof status, recurrence, and computable derivations for metrics, findings, controls, and PASS, REVIEW, or BLOCK decisions. The Portable Evaluation Record (PER) provides the system of record: a canonical, content addressed representation of one evaluation that preserves the evidence to decision chain and can be re derived, explained, and verified. Scores summarize, juries interpret, and traces record, but none of them define what the evidence means or what it can prove. EIO provides that missing semantic contract, while PER preserves the resulting evaluation as a portable and verifiable system of record. As AI agents assume greater operational responsibility, evaluation must become more than a collection of scores and verdicts; it must become an accountable artifact whose meaning, evidence, limitations, and decisions can be independently checked.
Figures & tables
| Item | Count | Item | Count |
|---|---|---|---|
| Modules, pinned by digest | 41 | Claim states | 7 |
| Predicates (risk, safeguard, obs.) | 58 (49, 7, 2) | Recurrence bands | 4 |
| Evidence types (can prove) | 11 (7) | Explanation templates | 36 |
| Source types (may witness) | 10 (5) | ||
| Resolvers (det., sem., human) | 11 (8, 2, 1) | Domain modules | 11 |
| Framework views | 30 | Coverage obligations | 126 |
| Group | Checks | What is recomputed |
|---|---|---|
| Schema and release | S1–S5, A1–A3 | JSON Schema; claim schema; PER and EIO versions; module digests; every EIO identifier; domain and coverage |
| Evidence | E1–E4 | Cited refs resolve; the witness rule for every ref; cited-refs hash; array orders |
| Claims | C1–C3 | Claim identifiers; conditional parameters and votes; evidence contracts |
| Views | F1, K1, M1, L1 | Findings and the precondition of Proven ; control statuses; caps and drivers; reliability ledgers |
| Explanations | X1, X2 | Re-rendering from registered templates; forbidden wording such as “compliant” |
| Decision | R1, R2 | Gates; the release recommendation and its invariants |
| Command | What it does |
|---|---|
| eio-agents project | Turns an evaluation bundle into a canonical PER record. |
| eio-agents validate | Checks that the record obeys the EIO and PER rules. |
| eio-agents verify | Goes back to the source evidence, rebuilds what can be rebuilt, re-projects the record, and compares digests. |
| eio-agents explain | Traces a metric, finding, control, readiness score, or release decision back to its claims and evidence. |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Finding | Decided by | Recurrence | Proof | Severity |
|---|---|---|---|---|
| t05 untrusted-instruction-execution | code | CONFIRMED 5/5 | Proven | critical |
| t01 claim-contradicts-grounding | jury | UNCONFIRMED | Unproven | high |
| t02 claim-contradicts-grounding | jury | UNCONFIRMED | Unproven | high |
| t03 applicable-policy-abandoned | jury | UNCONFIRMED | Unproven | high |
| t04 applicable-policy-abandoned | jury | UNCONFIRMED | Unproven | high |
| t14 applicable-policy-abandoned | jury | INTERMITTENT | Unproven | high |
| Record | Claims | Findings | Jury findings | Proven | Readiness | State |
|---|---|---|---|---|---|---|
| FIN_3 | 243 | 12 | 8 | 1 | 49 (raw 65.3) | BLOCK |
| EXAM_B | 132 | 12 | 9 | 0 | 49 (raw 65.0) | REVIEW |
| MED_1 | 243 | 34 | 30 | 0 | 68 | REVIEW |
| Total | 618 | 58 | 47 | 1 |