Recursive Organization Improvement: A Modeling Specification for Human--Agent Organizations
Organizations: CataX AI
Abstract
Stronger AI agents do not automatically produce better organizations: teams must also learn which work arrangements to retain and when to reconsider them. We propose a modeling specification for recursive organization improvement and evaluate it through an executable checker, a public-record mapping, and controlled simulation. The specification connects actor-visible histories, organizational memory, decision rights, and evidence-carrying change contracts. The mechanism study crosses six decision rules, three memory conditions, and three task environments under fixed resource ceilings. In a stationary environment, cumulative evidence raises balanced evaluation's normalized net value per task from 0.45224 to 0.48007. Repeated reassessment's disadvantage relative to this comparator falls from 0.01702 with reset evidence to 0.00007 with cumulative evidence. A reversal of the best workflow reveals the opposite cost: indefinite retention delays adaptation, while a finite window restores eventual performance at a transition cost. In exploratory controls, matching trial acquisition and label reuse reduces the apparent reassessment gain from 0.00607 to 0.00191. Program replacement adds no stable benefit across the tested reversal times. The study identifies evidence acquisition, reuse, and timely updating as mechanisms that must be separated from evaluator replacement when assessing organizational improvement.
Figures & tables
| Approach | Established modeling object | Required connection for a change comparison |
|---|---|---|
| MOISE+; OperA | Roles, missions, norms, interactions | Target-specific authority, prior version, evidence and comparator for a revision. |
| Process mining | Cases, activities, event order, conformance | Organization/procedure versions, actual observation access, and sampling coverage. |
| Organizational learning | Experience, routines, retention | Link a retained routine to the experiment and criterion used to accept it. |
| Agent/workflow search | Executable candidates and evaluations | Human resources, rights, generated-output costs, and revision of the evaluator. |
| Generic event log | Arbitrary recorded attributes | All the same relations, if explicitly supplied; missing fields remain unknown. |
| Symbol | Meaning | PR example |
|---|---|---|
| Actors, roles, tasks, versioned artifacts | Reviewer, agent, PR, test report. | |
| Task and resource dependencies | Approval depends on tests and reviewer availability. | |
| Decision rights | Maintainer may change routing; review board may change auditing. | |
| Observation availability | Which test results or private assessments each actor can see. | |
| Interaction and allocation protocols | Assignment order, review sequence, and workload allocation. | |
| Retained evidence and versions | Audit labels, trial scores, and the adopted routing rule. |
| Relation | Required condition | Failure or missing record |
|---|---|---|
| Actor–evidence | Each decision input was available to its actor beforehand. | Hidden/future input; unknown exposure history. |
| Patch–state | The target version exists and matches the expected version. | Stale patch; unversioned configuration. |
| Actor–target | The actor holds the applicable editing right at that version. | Unauthorized change; unknown historical permission. |
| Evidence–population | Sampling scope and label source support the stated population. | Uncovered stratum; unknown inclusion mechanism. |
| Outcome–criterion | Both arms use the declared criterion and horizon. | Changed success definition; censored follow-up. |
| Cost–resource | Search, outputs, labels, and transitions enter one ledger. | Omitted counterfactual output; duplicate charge. |
| Arm | Program decision | Role in the study |
|---|---|---|
| Biased / Balanced / Neyman | Program specified before the experiment | Fixed designs with different coverage/allocation rules. |
| Successive rejection | Fixed elimination rule across templates | A stronger selection-oriented comparator. |
| Discover once | Compare programs once; retain the winner | Discovery followed by reuse and budget reallocation. |
| Repeated discovery | Reassess every eight rounds | Incremental value of continued program assessment. |
| Trial-matched Balanced | Run the same trials; keep Balanced | Exploratory control for evidence acquisition and timing. |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| PR | Commit artifacts | Reviews | Head checks | Events |
|---|---|---|---|---|
| #6133 | 1 | 0 | 14 | 16 |
| #6096 | 2 | 5 | 14 | 21 |
| #6095 | 1 | 3 | 14 | 19 |
| Environment | Memory | B | Bal | N | SR | O | R | Mix |
|---|---|---|---|---|---|---|---|---|
| Stationary harm | Reset | 0.3967 | 0.4522 | 0.4539 | 0.4540 | 0.4366 | 0.4352 | 0.4343 |
| Stationary harm | Cumulative | 0.4692 | 0.4801 | 0.4805 | 0.4806 | 0.4793 | 0.4800 | 0.4766 |
| Stationary harm | Window 8 | 0.4594 | 0.4792 | 0.4796 | 0.4800 | 0.4754 | 0.4779 | 0.4727 |
| Workflow reversal | Reset | 0.5260 | 0.5828 | 0.5830 | 0.5844 | 0.5644 | 0.5628 | 0.5639 |
| Workflow reversal | Cumulative | 0.4721 | 0.4803 | 0.4805 | 0.4829 | 0.4796 | 0.4835 | 0.4776 |
| Workflow reversal | Window 8 | 0.5609 | 0.5758 | 0.5760 | 0.5786 | 0.5723 | 0.5821 | 0.5709 |
| Program share | Trials/program | Stationary, cumulative | Reversal, Window 8 |
|---|---|---|---|
| 0.2 | 1 | +0.0012 0.0012 | +0.0071 0.0024 |
| 0.2 | 3 | +0.0019 0.0020 | +0.0053 0.0023 |
| 0.2 | 6 | +0.0006 0.0008 | +0.0062 0.0020 |
| 0.4 | 1 | +0.0002 0.0004 | +0.0105 0.0025 |
| 0.4 | 3 | +0.0002 0.0001 | +0.0098 0.0026 |
| 0.4 | 6 | +0.0002 0.0001 | +0.0080 0.0019 |
| Environment | Review | B (%) | Bal (%) | N (%) | Net given Bal | Harm (%) | |
|---|---|---|---|---|---|---|---|
| Stationary | 1 | 27.3 | 37.5 | 35.2 | 48 | 0.4747 | 0.00 |
| Stationary | 9 | 26.6 | 35.9 | 37.5 | 46 | 0.4819 | 0.00 |
| Stationary | 17 | 33.6 | 34.4 | 32.0 | 44 | 0.4819 | 0.00 |
| Stationary | 25 | 33.6 | 35.9 | 30.5 | 46 | 0.4819 | 0.00 |
| Stationary | 33 | 35.9 | 32.0 | 32.0 | 41 | 0.4819 | 0.00 |
| Stationary | 41 | 38.3 | 31.2 | 30.5 | 40 | 0.4819 | 0.00 |
| Stationary | Slow shifts | ||||||
|---|---|---|---|---|---|---|---|
| Prior | Threshold | F | A | E | F | A | E |
| 1:1 | 0.14 | 0.8479 | 0.8530 | 0.8338 | 0.6616 | 0.6753 | 0.6758 |
| 1:1 | 0.16 | 0.8523 | 0.8571 | 0.8342 | 0.6597 | 0.6729 | 0.6757 |
| 1:9 | 0.14 | 0.8581 | 0.8649 | 0.8656 | 0.6620 | 0.6706 | 0.6599 |
| 1:9 | 0.16 | 0.8592 | 0.8686 | 0.8660 | 0.6515 | 0.6642 | 0.6526 |
| 1:19 | 0.14 | 0.8595 | 0.8720 | 0.8795 | 0.6468 | 0.6571 | 0.5673 |