Comprehension Audits to Mitigate Risks from Automated AI Research
Organizations: Independent · Institute for Law and AI · Johns Hopkins University
Abstract
AI is already writing a majority of code for frontier AI labs. This creates a safety risk if there is insufficient human oversight. Existing work proposes minimum comprehension thresholds and unaided checks to mitigate this. To our knowledge, however, there is currently no published frontier-AI assurance regime that requires demonstrated evidence that the responsible humans understand what they are building as a precommitted condition for continuing development or usage. We propose comprehension audits, a novel development-process assurance mechanism in which the responsible people explain R&D contributions to auditors to demonstrate understanding. With independent administration and graded reports, they provide a gate: development of a contribution stops based on a failure to demonstrate human understanding until remediated, with escalating consequences for repeated failures. Our analysis of leading open-source AI projects finds increased output of code with reduced human review commentary rates per line of code, with far lower rates for automated fleet accounts. We advocate for labs to conduct them with embedded independent auditors.
Figures & tables
| Tier | Required depth | Aspects covered |
|---|---|---|
| Deep understanding | A thorough understanding of these aspects of the contribution | Architecture and design; implementation behavior; process |
| Familiarity | Good familiarity and appropriate diligence and understanding | Key findings; analysis of results |
| Limited | Contributors are not expected to explain root causes of model behavior or otherwise exceed the level of understanding of deep learning that prevails in the industry | Model internals; root causes of model behavior |
| Stage | Description |
|---|---|
| Auditor onboarding | Selecting and training auditors to be prepared to conduct audits and integrating audits into the R&D processes and monitoring of the organization. |
| Auditor monitoring | Auditors track R&D progress in the organization as well as tracking proxy metrics for comprehension. |
| Audit selection | Audited contributions can be selected based on coverage sampling and/or evidence of comprehension issues, and are subject to a budget of how many may be conducted. |
| Conduct audits | The core process of auditing a specific contribution and producing a report. |
| Audit response | Audit findings may be appealed. Those that find issues impose consequences until remediated. |
| Strategic updates | Provides overall analysis of audits with recommendations for improvement of R&D and auditing. |
| Failure rate | 50% detection | 90% detection | 99% detection |
|---|---|---|---|
| 5% | 14 | 45 | 90 |
| 10% | 7 | 22 | 44 |
| 25% | 3 | 9 | 17 |
| Topical area | Understanding assessed |
|---|---|
| Architecture and design | The overall software components, how they work together, what key decisions were made in terms of how they would work and cooperate, optimization approach and what datasets were used. |
| Software implementation behavior | What choices were made in implementing key parts of the contribution, including algorithms, security, scalability and tests. |
| Process | What process was followed, research decisions, why were changes made, what were key tradeoffs considered in design and during experiments that motivated changes or updates. |
| Key findings | The overall results of experiments or implementations and how they were derived. |
| Analysis of results | Results of evaluations performed, any metrics computed, any failures encountered and reasons for them, and any anomalies identified and updates in response to them. |
| Grade | Criteria |
|---|---|
| Pass | Demonstrating minimal acceptable understanding in all components and aspects examined. There may be acceptable but suboptimal knowledge gaps or other areas that could be improved that are addressed in optional recommendations. Acceptable understanding should be found for a component if both of the following are true: (1) for any questions about all deep understanding tier aspects (architecture and design, software implementation, and process) at least one contributor is able to adequately answer them; and (2) for both familiarity tier aspects (key findings and analysis of results) for any question, at least one contributor is able to demonstrate due consideration. |
| Conditional pass | A failure in a noncritical component or area may be assigned a conditional pass with a requirement to improve understanding and/or adoption of changed processes to improve understanding. If there is a lack of knowledge due to staff turnover in an area during a transition period, a conditional pass would require a current team member to obtain sufficient knowledge in the area. |
| Failure | Any significant lack of understanding of important concepts or a general pattern of insufficient understanding in an important component or central aspect of a contribution will lead to a failure. |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Period | Repos | AI share of submitted PRs | AI share of merged PRs | AI-PR share of submitted LOC | AI-PR share of merged LOC |
|---|---|---|---|---|---|
| Q1 2024 | 465 | 0.0% [0.0, 0.0] | 0.0% [0.0, 0.0] | 0.0% [0.0, 0.0] | 0.0% [0.0, 0.0] |
| Q4 2025 | 825 | 3.0% [2.3, 3.7] | 2.2% [1.6, 2.9] | 4.3% [3.4, 5.4] | 2.7% [1.9, 3.6] |
| Q1 2026 | 1009 | 10.3% [9.0, 11.7] | 8.8% [7.5, 10.4] | 13.6% [11.7, 15.4] | 11.1% [9.5, 12.9] |
| Q2 2026 | 1152 | 16.8% [14.8, 18.7] | 16.4% [14.3, 18.4] | 22.5% [19.6, 25.0] | 20.6% [17.9, 23.2] |
| Quarter | Repos with 5+ PRs | PRs (all authors) | Merged (8-week mature) | Human-authored PRs | Human authors | Active insider authors (non-fleet) | Active fleet accounts |
|---|---|---|---|---|---|---|---|
| Q1 2024 | 465 | 72,290 | 54,355 | 66,095 | 9,571 | 1,432 | 13 |
| Q4 2025 | 825 | 134,088 | 96,516 | 121,775 | 16,996 | 3,002 | 27 |
| Q1 2026 | 1,009 | 197,860 | 126,696 | 177,275 | 25,035 | 3,325 | 41 |
| Q2 2026 | 1,152 | 285,088 | 161,565 | 255,580 | 37,897 | 3,526 | 40 |
| Rule (min every quarter / mean over four) | n | Ever-AI / never-AI | Submitted kLOC Q1 2024 to Q2 2026 | Submitted multiple | Merged multiple | Ever-AI / never-AI submitted multiple |
|---|---|---|---|---|---|---|
| 3 / 5 | 438 | 250 / 188 | 4.69 to 13.00 | 2.77 | 2.75 | 3.54 / 2.32 |
| 5 / 10 (paper) | 337 | 202 / 135 | 6.26 to 17.72 | 2.83 | 2.73 | 3.02 / 2.32 |
| 10 / 20 | 211 | 140 / 71 | 8.89 to 26.05 | 2.93 | 2.28 | 3.09 / 2.40 |
| 5 / 5 | 348 | 204 / 144 | 6.12 to 16.80 | 2.75 | 2.67 | 3.13 / 2.40 |
| 10 / 10 | 233 | 150 / 83 | 8.66 to 23.58 | 2.72 | 2.40 | 3.22 / 2.27 |
| Month | Slice | Merged PRs | Pooled % | Wilson 95% | Contributor-weighted repo mean % | Repos |
|---|---|---|---|---|---|---|
| Jan 2024 | Non-AI | 15,451 | 77.2 | [76.6, 77.9] | 70.0 [61.7, 77.5] | 366 |
| Jan 2024 | Fleet | 745 | 44.2 | [40.6, 47.7] | 44.9 [21.0, 80.3] | 17 |
| Feb 2024 | Non-AI | 15,171 | 77.3 | [76.6, 77.9] | 72.0 [63.4, 79.8] | 378 |
| Feb 2024 | Fleet | 877 | 54.6 | [51.3, 57.9] | 37.1 [11.2, 79.5] | 24 |
| Mar 2024 | Non-AI | 17,417 | 78.0 | [77.4, 78.6] | 74.1 [66.4, 80.9] | 390 |
| Mar 2024 | Fleet | 745 | 48.2 | [44.6, 51.8] | 42.7 [20.2, 75.4] | 19 |
| Month | Slice | Reviewed PRs | kLOC | Comments per kLOC | Bootstrap 95% | % of Q1 2024 non-AI |
|---|---|---|---|---|---|---|
| Jan 2024 | Non-AI | 11,906 | 4,137 | 4.13 | [3.81, 4.49] | 104% |
| Jan 2024 | Fleet | 328 | 106 | 1.92 | [1.24, 2.98] | 48% |
| Feb 2024 | Non-AI | 11,702 | 4,436 | 3.88 | [3.56, 4.20] | 98% |
| Feb 2024 | Fleet | 479 | 78 | 2.23 | [1.26, 3.95] | 56% |
| Mar 2024 | Non-AI | 13,557 | 5,030 | 3.90 | [3.61, 4.23] | 98% |
| Mar 2024 | Fleet | 359 | 82 | 1.35 | [0.79, 2.33] | 34% |
| Month | Slice | Merged PRs | Review % [Wilson 95%] | Contributor-weighted repo mean % | Comments per kLOC [bootstrap 95%] | % of Q1 2024 non-AI |
|---|---|---|---|---|---|---|
| Jan 2024 | Non-AI | 14,243 | 77.7 [77.0, 78.4] | 70.1 | 4.26 [3.92, 4.63] | 105% |
| Jan 2024 | Fleet | 745 | 44.2 [40.6, 47.7] | 44.9 | 1.92 [1.24, 2.98] | 47% |
| Feb 2024 | Non-AI | 13,947 | 77.4 [76.7, 78.1] | 72.1 | 3.95 [3.64, 4.29] | 98% |
| Feb 2024 | Fleet | 872 | 54.6 [51.3, 57.9] | 36.8 | 2.21 [1.24, 4.02] | 55% |
| Mar 2024 | Non-AI | 16,112 | 78.2 [77.6, 78.8] | 74.3 | 3.95 [3.65, 4.28] | 98% |
| Mar 2024 | Fleet | 731 | 48.6 [45.0, 52.2] | 42.7 | 1.30 [0.73, 2.21] | 32% |