CompOrca: Corpus-Scale Compliance Labelling of Instruction-Tuning Data
Organizations: Department of Computer Science, Brunel University of London
Abstract
Studying how fine-tuning shapes refusal and noncompliance behaviour requires identifying training examples that refuse, evade or otherwise fail to fulfil the requested task. But existing annotation covers evaluation sets of a few thousand prompts at most. We present CompOrca, a compliance labelling over the entirety of the 4,233,923-example OpenOrca corpus. Every example was classified as compliant or noncompliant by five independent passes of an open-weight LLM judge (LongCat-2.0, 1.6T parameters), and the corpus is released as unanimous compliance (94.75%), unanimous noncompliance (1.28%), and nonunanimous rows (3.97%) along with the raw vote counts. A single pass flags 2.7-3.2% of the corpus as noncompliant, while only 1.28% is flagged by all five, allowing for filtering the most ambiguous samples. Against 450 human-annotated examples, 150 of them annotated twice (human-human ), the unanimous compliance and noncompliance labels are 97.3% and 86.7% precise, the latter a high-precision subset, not a complete enumeration, of noncompliance. Published refusal-detection methods recall only between 0.4% and 94.1% of the noncompliance class. We release the full corpus with its per-row labels and vote counts at https://huggingface.co/datasets/cemiu/CompOrca
Figures & tables
| bucket | rows | share |
|---|---|---|
| unanimous compliance (0/5) | 4,011,762 | 94.75% |
| nonunanimous (1–4 of 5) | 168,157 | 3.97% |
| unanimous noncompliance (5/5) | 54,004 | 1.28% |
| compliance-clean subset | 3,769,249 | 89.02% |
| total | 4,233,923 | 100% |
| sample | corpus-rw. | |||
|---|---|---|---|---|
| judge label | agr. | agr. | ||
| unanimous (released) | 300 | 92.0 | 0.840 | 97.2 0.439 |
| majority ( 3/5) | 450 | 82.4 | 0.649 | 95.7 0.504 |
| single pass (greedy) | 450 | 83.3 | 0.667 | 95.8 0.524 |
| vs. CompOrca | vs. human labels | |||||
|---|---|---|---|---|---|---|
| method | flagged | prec. | rec. | F1 | rec. | F1 |
| substring: Zou et al. ’s prefix list ( 2023 ) | 29,163 | 12.6% | 6.3% | 8.4% | 4.1% | 7.9% |
| substring: XSTest prefix list ( Röttger et al., 2024 ) | 146,450 | 7.1% | 15.9% | 9.8% | 5.4% | 9.1% |
| classifier: DistilRoBERTa-rejection ( ProtectAI.com, 2024 ) | 15,640 | 14.8% | 3.8% | 6.1% | 1.7% | 3.3% |
| classifier: Do-Not-Answer Longformer ( Wang et al., 2024b ) | 3,918 | 26.0% | 1.8% | 3.3% | 0.4% | 0.8% |
| classifier: Minos-v1 ( Suphavadeeprasit et al., 2025 ) | 60,495 | 21.5% | 19.7% | 20.5% | 17.0% | 28.4% |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| judge | LongCat-2.0 (open-weight) |
|---|---|
| serving | OpenRouter (as owl-alpha ), via LiteLLM |
| passes | 5 (1 ; 4 ) |
| output | forced JSON; reasoning disabled |
| input | question + response (no system prompt) |
| truncation | none |
| max tokens | 256 |