Finding Blind Spots in AppWorld and WorkArena Task Verifiers
Organizations: OpenAdapt.AI (MLDSAI Inc.)
Abstract
Execution-based task verifiers decide whether an agent succeeded. We audit shipped AppWorld and WorkArena verifiers with source-informed mutation tests. The main audit never modifies a shipped checker. In AppWorld, duplicating a non-idempotent write creates an extra record while preserving every checked field value. The verifier accepts all three task variants from two of five eligible generators: 6/15 constructed effects. A cardinality patch applied to checker copies after the census makes all six cells fail while preserving valid controls. In WorkArena, we prospectively rerun 23 extra-field candidates selected for earlier checker-PASS outcomes. Independent Table API readback confirms nondefault persisted values in 21, while all 23 receive PASS. Two requested strings are aliases of stored defaults. The 21 confirmed wrong effects span three form templates. These selected cases confirm wrong effects under the audit's protocol; they do not estimate a population rate. No other construction produces an independently confirmed false accept. Other checker-PASS cases are effect-correct degeneracies. We report zero-PASS families separately because retained evidence differs. In fixed intent-swap grids, the checkers return no PASS on 2,689 off-diagonal executions. This is a rejection census: 57 WorkArena cells use session-scoped evidence; the other 2,632 lack classified rejection causes and independent target ground truth. Each increment is specified before its own cells are scored. A supplement accompanies the OpenReview submission with the construction grammar, evidence, content-bound stage lineage and count reproducer.
Figures & tables
| Suite | Rung | Valid cells | Checker Pass |
| AppWorld (2,294 cells) | Diag (control) | 31/31 pass | — |
| Rung-0: twin intent, 1 parameter | 43 | 0 | |
| Rung-1: sibling, same scenario | 144 | 0 | |
| Rung-2: cross-scenario, shared models | 369 | 0 | |
| Rung-3: cross-scenario, disjoint | 1,707 | 0 | |
| WorkArena (468 cells; 3 errors) | Diag (control) | 39/39 pass | — |
| Checker:family | Wrong effect | Evidence | N | FA |
| appworld: claim-only | claim w/o write | runner no-write | 48 | 0 |
| appworld: extra | duplicate write (changed state) | replay write counters | 15 | 6 |
| workarena: claim-only | claim w/o request | runner no-write | 27 | 0 |
| workarena: no-submit | leave form unsubmitted | runner no-write | 39 | 0 |
| workarena: extra-field * | out-of-scope write | Table API readback | 21 | 21 |
| Split | Budget | Pairs surviving | Top-1 unique? | ||
|---|---|---|---|---|---|
| test_challenge | Observed magnitude | .4000 | .0000 | 0/91 | no |
| Illustrative 95% | .8107 | .2384 | 0/91 | no | |
| test_normal | Observed magnitude | .4000 | .0000 | 6/91 | no |
| Illustrative 95% | .8107 | .2384 | 0/91 | no |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Construction |
|---|---|
| claim-only | Deliver only the claim/completion step; no write occurs. |
| omit | Drop the final write call from the reference program. |
| partial | Keep only the first of a repeated write sequence. |
| wrong-value | Move one written value to a different legal value. |
| extra | Duplicate the final write call. |
| extra-ni | Duplicate the final write call restricted to endpoints that a committed idempotence census marks non-idempotent (a duplicated POST creates an extra payment request and notification; a duplicated action doubles a player record). |
| Checker:family | Sched. | Ref. | Inc. | Deg. | Cand. | Class. | Pass | FA |
|---|---|---|---|---|---|---|---|---|
| appworld: claim-only | 48 | 0 | 0 | 0 | 0 | 48 | 0 | 0 |
| appworld: extra | 48 | 0 | 0 | 33 | 0 | 15 | 39 | 6 |
| appworld: omit | 48 | 0 | 0 | 0 | 0 | 48 | 0 | 0 |
| appworld: partial | 48 | 9 | 0 | 0 | 0 | 39 | 0 | 0 |
| appworld: wrong-value | 48 | 15 | 0 | 0 | 0 | 33 | 0 | 0 |
| workarena: claim-only | 52 | 25 | 0 | 0 | 0 | 27 | 0 | 0 |