Correct Code, Broken Contributions? SWE-CC: Benchmarking Repository Policy Compliance for Coding Agents
Organizations: Singapore Management University · Singapore University of Technology and Design
Abstract
Autonomous coding agents now resolve a substantial share of real-world GitHub issues. However, passing functional tests differs fundamentally from producing a high-quality contribution acceptable for merging. Mature open-source projects publish repository-specific contribution policies, spanning style, git, testing workflows, to ensure code quality and long-term maintainability. Because existing benchmarks evaluate patches solely on unit tests, agent compliance with repository governance remains unknown. In this paper, we introduce SWE-CC, a benchmark evaluating code and process compliance in autonomous software engineering. We develop a semi-automated pipeline that converts developer documentation across 12 open-source repositories into 823 machine-checkable atomic policies. SWE-CC introduces two features: 1) lightweight, deterministic checker functions that represent each policy, 2) a comprehensive auditing mechanism that inspects both agent runtime behaviors and final deliverables. We evaluate the compliance of agent workflows in 500 end-to-end software contribution tasks extended from SWE-bench Verified. Our evaluation of four LLMs under two agent scaffolds shows that modern agents suffer from coding compliance issues: although agents produce functionally correct patches, they still violate 43.1 percent of applicable project policies, with nearly half of all violations occurring during intermediate execution steps. These results show that functional correctness does not guarantee real-world readiness, highlighting that future software engineering agents must reliably conform to repository governance to enable safe and trustworthy deployment.
Figures & tables
| Benchmark | Evaluation scope | Policy source | # Policy | Agent behavior | Policy retrieval | Policy evaluator |
|---|---|---|---|---|---|---|
| SWE-NFI | Patch quality | Literature, common practice | 92 | Executable | ||
| SWE-Gate | Patch quality | Code review history | 303 | Executable | ||
| RepoComplianceBench | AI declaration | Contribution file | 455 | Executable + LLM | ||
| SWE-SHIELD | Patch quality | Code review history | 1,787 | LLM-as-a-judge | ||
| SWE-CC (ours) | Coding workflow | Project documentation | 823 | Executable |
| mini-SWE-agent | OpenHands | |||||
|---|---|---|---|---|---|---|
| Agent models | Triggering rate (%) | Compliance rate (%) | Resolve rate (%) | Triggering rate (%) | Compliance rate (%) | Resolve rate (%) |
| Native Setting | ||||||
| GPT-5.6 Luna | 23.5 | 51.8 | 78.6 | 24.4 | 55.3 | 86.6 |
| Gemini 3.7 Flash | 23.8 | 53.9 | 80.2 | 24.3 | 56.3 | 78.4 |
| DeepSeek V4 Flash | 25.6 | 61.3 | 93.4 | 26.0 | 63.8 | 94.4 |
| Kimi K2.5 | 24.6 | 55.9 | 74.6 | 24.8 | 57.2 | 76.2 |
| mini-SWE-agent | OpenHands | ||||
|---|---|---|---|---|---|
| Category | Native (%) | Consolidated (%) | Native (%) | Consolidated (%) | |
| Git and commit conventions | Triggering rate | 52.3 | 59.7 | 55.7 | 60.4 |
| Compliance rate | 59.9 | 87.0 | 63.4 | 90.1 | |
| PR and release metadata | Triggering rate | 36.3 | 45.6 | 37.5 | 51.8 |
| Compliance rate | 36.4 | 51.6 | 39.6 | 58.9 | |
| Code and quality | Triggering rate | 83.7 | 83.8 | 82.5 | 82.8 |
| Native (autonomously discover) | Consolidated | ||||
|---|---|---|---|---|---|
| Model | Scaffold | Attempt (%) | Retrieve (%) | Opened (%) | Delivered (%) |
| GPT-5.6 Luna | mini-SWE-agent | 50.0 | 13.2 | 100.0 | 100.0 |
| OpenHands | 75.6 | 43.2 | 100.0 | 99.9 | |
| Gemini 3.7 Flash | mini-SWE-agent | 10.0 | 0.4 | 100.0 | 100.0 |
| OpenHands | 2.8 | 1.0 | 98.6 | 100.0 | |
| DeepSeek V4 Flash | mini-SWE-agent | 34.8 | 10.6 | 100.0 | 99.7 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| rules | evidence type | |||||||
|---|---|---|---|---|---|---|---|---|
| project | instances | extracted | in scope | binding | output | diff. | traj. | approx. |
| astropy | 22 | 255 | 129 | 111 | 103 | 5 | 3 | 85% |
| django | 231 | 143 | 89 | 78 | 67 | 8 | 3 | 55% |
| matplotlib | 34 | 293 | 199 | 167 | 148 | 14 | 5 | 82% |
| scikit-learn | 32 | 274 | 169 | 146 | 131 | 8 | 7 | 83% |
| sympy | 75 | 279 | 188 | 142 | 112 | 11 | 19 | 24% |
| category | # | example policy |
|---|---|---|
| Documentation and docstrings | 224 | Give every public class, method, and function a docstring. (ASTROPY-C085) |
| Language and framework style | 163 | Do not raise the bare Exception class. (ASTROPY-C100) |
| Tests and test style | 160 | Name test modules test_*.py or *_test.py . (ASTROPY-C002) |
| Specialized changes | 107 | Keep in-repository data files under about 100 kB and host anything larger off the repository. (ASTROPY-C092) |
| PR and release metadata | 57 | Add a changelog fragment under docs/changes/<sub-package>/ describing your change. (ASTROPY-C061) |
| Git and commit conventions | 40 | Phrase commit subject lines in past tense and end them with a period. (DJANGO-C046) |
| Native | Consolidated | ||||||
|---|---|---|---|---|---|---|---|
| Scaffold | Model | All | Exact | All | Exact | All | Exact |
| mini-SWE-agent | GPT-5.6 Luna | 51.8 | 52.0 | 67.6 | 69.9 | 15.8 | 17.9 |
| Gemini 3.7 Flash | 53.9 | 50.7 | 65.3 | 69.0 | 11.4 | 18.3 | |
| DeepSeek V4 Flash | 61.3 | 56.6 | 67.7 | 66.1 | 6.4 | 9.5 | |
| Kimi K2.5 | 55.9 | 51.8 | 56.9 | 55.2 | 1.0 | 3.4 | |
| OpenHands | GPT-5.6 Luna | 55.3 | 55.8 | 67.5 | 69.2 | 12.2 | 13.4 |
| Commit subject | Outcome | Reason |
|---|---|---|
| Fixed #123 -- Added truncation to Truncator. | satisfied | |
| Made Truncator keep HTML entities. | satisfied | irregular verb |
| Add truncation to Truncator. | violated | not in the past tense |
| Added truncation to Truncator | violated | no final period |
| (no commit) | inapplicable |
| Technique | Policy | What the check tests | Case it gets wrong |
|---|---|---|---|
| Word list | DJANGO-C046: commit subjects in the past tense | the leading verb ends in -ed or is a listed irregular form | misses past forms outside the list |
| Text pattern | ASTROPY-C153: astropy for the package, Astropy for the Project | each mention matches one of the two sanctioned spellings | cannot tell which meaning a sentence intends |
| Code structure | SCIKIT-LEARN-C158: estimators inherit from BaseEstimator | a new class that defines fit is an estimator | selects other classes with a fit method |
| File location | MATPLOTLIB-C262: imported code carries a compatible licence | a new file under extern/ is imported code | misses imported code placed elsewhere |
| Co-change | SCIKIT-LEARN-C118: deprecations are listed in the API reference | a @deprecated name comes with an edit to the reference list | does not check that the right name was added |
| project | checks | drawn | diff. | traj. |
|---|---|---|---|---|
| matplotlib | 167 | 24 | 2 | 1 |
| scikit-learn | 146 | 21 | 1 | 1 |
| SymPy | 142 | 21 | 1 | 3 |
| astropy | 111 | 18 | 1 | 1 |
| Django | 78 | 14 | 2 | 1 |
| pylint | 49 | 11 | 1 | 1 |
| raw agr. (%) | Cohen’s | |
|---|---|---|
| verdict (accept / reject) | 94.0 | 0.72 |
| C1 encodes the rule | 99.3 | 0.89 |
| C2 precondition | 96.0 | 0.73 |
| C3 post-condition | 98.0 | 0.72 |
| C4 exact or approximate | 96.7 | 0.93 |
| Checkpoint | Checks |
|---|---|
| 1. Manifest (author) | every URL is at the pinned version; no maintainer or triage page admitted whole; partial pages name their in-scope sections; configuration files are context-only; nothing is claimed about a file that was not retrieved |
| 2. Extraction (author) | row count plausible against the other repositories; section context and notes populated on every row; every rule records its source sentence and URL; identifiers contiguous; every category from the closed list; every rule cross-referenced with its source sentence on the live page; mid-sentence quotations checked against the preceding clause |
| 3. Acceptance (automated) | rules-to-source-length ratio within band; judgment and in scope never co-occur; modal-verb and rubric levels differ on some rules, with direction recorded; reasoning distinct across rows; not-observable counted separately from other out-of-scope reasons; schema conformance |
| Symptom | Cause | Mitigation |
|---|---|---|
| Version drift mid-run | stable documentation retrieved | pin the developer documentation |
| False version mismatch | pinned at build rather than release | halt only on a different release |
| Template rules missing or inverted | rendered view strips comments | retrieve raw text in the pre-pass |
| Configuration files never retrieved | not reachable from navigation | resolve by path in the pre-pass |
| Agent-directed policy missing | agent files outside navigation | admit them in traversal and the pre-pass |
| Rules asserted unread | model answers from prior familiarity | require retrieval or an explicit non-retrieval |
| Model | Ours (Native) | Vals AI | |
|---|---|---|---|
| GPT-5.6 Luna | 78.6 | 93.0 | |
| Gemini 3.7 Flash | 80.2 | 80.8 | |
| DeepSeek V4 Flash 0731 | 93.4 | 88.8 | |
| Kimi K2.5 | 74.6 | 70.0 |
| Category | Policies | Never triggered | Share |
|---|---|---|---|
| Specialized changes | 107 | 81 | 76% |
| Documentation and docstrings | 224 | 93 | 42% |
| Language and framework style | 163 | 59 | 36% |
| Tests and test style | 160 | 50 | 31% |
| Git and commit conventions | 40 | 10 | 25% |
| Code and quality | 35 | 5 | 14% |
| Native | Consolidated | ||||
| Category | Scaffold | Triggered | Graded | Triggered | Graded |
| Git and commit conventions | mini-SWE-agent | 4,100 | 4,092 | 4,676 | 4,671 |
| OpenHands | 4,361 | 4,353 | 4,730 | 4,722 | |
| PR and release metadata | mini-SWE-agent | 4,619 | 4,619 | 5,793 | 5,793 |
| OpenHands | 4,764 | 4,764 | 6,586 | 6,586 | |
| Code and quality | mini-SWE-agent | 5,385 | 1,453 | 5,388 | 1,484 |
| Value | What the check opens |
|---|---|
| output | the final files or the commit message, read once, with no comparison |
| differential | two states, compared: pre-patch against post-patch |
| trajectory | a record of what the agent did, not only what it produced |
| judgment | nothing; the standard is not written down and a human decides |
| Condition | What the annotator does | |
|---|---|---|
| C1 | observability | names the file, diff, or output that the check opens |
| C2 | decidability | writes down the failure condition |
| Exists | Does not exist |
|---|---|
| repository working tree, the agent’s diff | pull request object, PR template, review thread |
| one commit and its message | multiple commits, branch history, rebase or squash |
| a local test suite run | CI service, coverage bot, issue tracker |
| the files the agent writes | built or rendered documentation, browser, screenshots |
| a second human, contributor identity |
| Route | The rule… |
|---|---|
| M1 | says so outright ( must , required ) |
| M2 | forbids something ( never , do not , must not ) |
| M3 | is a condition of acceptance: a pre-merge or review checklist, or a named CI job |
| M4 | names an exact thing, limit, form, or ordering |
| M5 | is stated non-mandatorily, but admits only one satisfying state |
| M6 | states a consequence in the source |
| Route | The rule… |
|---|---|
| D1 | uses preference wording |
| D2 | permits deviation in its documentation |
| D3 | has a fixed form, but whether it applies is a judgment call |
| D4 | instructs on how , not what |