Verifiable, Articulable, and Tacit Components of Preference
Organizations: Stanford University · University of Toronto
Abstract
What makes a short story gripping; a news article newsworthy; or a math proof elegant? These constructs resist articulation or verification; their meaning is at least partially tacit. However, modern AI models are improved primarily via articulated constitutions, rubrics and verifiers (i.e. in RLAIF and RLVR); tacit components of preferences are typically understudied. We introduce a large, labeled preference dataset CreativePreferences, containing 2.8M texts labeled by 317M human preference judgments across 7 creative domains, with 42 benchmark tasks. We model these labels with executable programs, rubric banks and densely trained models (V, A and VAT, respectively). We observe robust articulability gaps, VAT-VA; and verifiability gaps, VAT-V; we estimate upper and lower bounds for each gap with a novel measurement approach that discovers articulable and verifiable metrics, identifies spurious variables and estimates the value of undiscovered metrics using capture-recapture. These gaps occur across all domains, even in domains traditionally treated as fully verifiable: correctness-centered domains (i.e. mathematics and software engineering) and claim- and novelty-centric domains (i.e. news, patents, peer review). The size of the gap varies based on domain (e.g. peer review and creative writing have the largest articulability gaps) and widens as more people take part in the judgment, consistent with Collins' collective tacit knowledge. We show two consequences: (1) on human generations, the full model more closely matches human preferences, often in disagreement with articulated criteria, and (2) in an analogy to Goodhart's law, articulating preference shifts it away from the tacit dimension. Articulability and verifiability gaps are consequential; we give recommendations on when tasks can be prompted; how learning mechanisms might improve; and when to leave judgments with humans.
Figures & tables
| Top verifiable checks | Top articulated criteria | |
|---|---|---|
| software code : PR merged | no pass fail regressions tests flip fail pass patch applies; suite executes | message says what changed, why structured subject/body, issue links single-purpose, reviewable change |
| creative writing : story upvotes | plot-arc shape proxy: paragraphing, rising action density form/medium-fit check worldbuilding-specificity check | deliberate contrasts of image and tone no stumbles read aloud; rhythm holds micro-tension keeps pages turning |
| math : answer accepted | lint checks machine-checkable-proof check formula-correctness check | every inference warranted, no jumps brief, inevitable, clear argument airtight steps, no hidden assumptions |
| task | regret (full model articulated) | 95% CI | disagreement odds |
|---|---|---|---|
| peer review: citations | |||
| peer review: accept/reject | |||
| math.SE: accepted answers | |||
| peer review: spotlight | |||
| creative writing: upvotes |
Appendix figures & tables57 assets
Supplementary material from the paper’s appendix.
Appendix
| Field | Dataset | Status | Notes (1s/0s; sub-corpora) | |
|---|---|---|---|---|
| Coding | GitHub PR merged (enriched) | inst. | 44,751 | 594 repositories; within-repository balance |
| Stack Overflow accepted answer | inst. | 16,001 | 5,972 / 10,029 | |
| Math | mathlib PR merged | inst. | 8,494 | 7,509 / 985; canonical clean slice 7,956 |
| Math.SE accepted answer | inst. | 13,001 | 4,960 / 8,041; 4,960 questions | |
| Law & policy | Patent application granted | inst. | 579,084 | 290,198 / 288,886; nested claim-level corpus 59,937 |
| Comment received a response | inst. | 9,521 | 7,439 / 2,082; 1,814 dockets |
| Field | Dataset | Status | Notes (1s/0s; sub-corpora) | |
|---|---|---|---|---|
| Coding | Competition editorial-approach match | inst. | 6,353 | AtCoder 2,495 LeetCode 1,995 CodeChef 995 CF 868 |
| Stack Overflow bounty award | inst. | 23,498 | 7,103 / 16,395; within-question | |
| Release-highlight picks | coll. | 4,654 | Rust 4,031-PR pool mathlib 623 picks | |
| Math | AoPS editorial approach | inst. | 25,454 | 17,560 / 7,894; 5,202-row scored task |
| Math.SE bounty award | inst. | 23,972 | 8,503 / 15,469; within-question | |
| Law & policy | Comment: change made | inst. | 7,084 | 2,824 / 4,260; second label agency-agreement 5,046 |
| Field | Dataset | Status | Notes (1s/0s; sub-corpora) | |
|---|---|---|---|---|
| Coding | Stack Overflow answer votes | inst. | 12,202 | 6,391 / 5,811; 137,994 ranked pairs collected |
| PR linked-issue reactions | coll. | 154,633 | 13 repositories | |
| Math | Math.SE answer votes | inst. | 100,000 | 50,000 / 50,000; 11,629-row scored task |
| Law & policy | Comment co-signing | inst. | 9,520 | 342 positive |
| Law SE answer votes | coll. | 4,192 | balanced; 72,369-post pool | |
| Academia | Citation percentile | inst. | 2,387 | 1,176 / 1,211; from 27,233-paper collected corpus |
| task | decision-maker | candidate set | information at decision time | prov. | positives | groups | gap 95% CI | label uncertainty / selection effects | maturity | |
|---|---|---|---|---|---|---|---|---|---|---|
| GitHub PR merged | repository maintainers | pull requests opened on 594 repositories | diff, description, CI results, author identity, review thread | O | 11,452 | 9,205 | 255 | [+.020, +.073] | merge reflects maintainer priorities and timing as well as the patch | mature |
| Stack Overflow accepted answer | the question’s asker | answers to the same question | all answers, votes at the time, asker’s own problem | O | 16,001 | 5,972 | 5,972 | [+.070, +.085] | one person’s choice; asker may accept early or never | mature |
| Competition editorial-approach | contest editorial (official solution) | accepted solutions per problem | problem statement, editorial | A | 999 | n/r | n/r | n/r | label = similarity of the solution to the editorial approach; model-assisted, not a gatekeeper verdict on the item | mature |
| Stack Overflow bounty award | the bounty poster | answers to the bountied question | all answers, votes, poster’s own problem | O | 23,498 | 7,103 | 7,036 | [+.036, +.044] | one person’s choice under a deadline | mature |
| Stack Overflow answer votes | site voters (aggregate) | answers to the same question | answers in vote order, accepted mark | O | 12,202 | 6,391 | 5,972 | n/r | vote order and acceptance bias exposure; threshold on score | mature |
| mathlib PR merged | mathlib maintainers | pull requests to mathlib | diff, CI build, reviewer comments | O | 7,956 | 7,502 | 31 | n/r | 94% positives; full model unstable at this balance (u/q) | mature |
| task | executable check | share of items the check applies to | AUC, executable columns alone | AUC, existing bank | + executable |
|---|---|---|---|---|---|
| Code (GitHub PR merged) | tests pass after patch (17 cols); patch well-formed; added Python/JS/Go code parses (“compiles”); lint errors | 3.7% (tests); 99% (patch); 94% (compiles) | .553 /.548 within-repository (tests);.528 (compiles, univariate) | – | no change |
| Code (GitHub PR merged; deep) | added dependencies exist on PyPI/npm; bandit findings on added Python | 5.3% (deps); 3% high findings | .518 /.494 | – | – |
| Math (AoPS editorial approach) | final answer matches the reference key | 46.2% | .572 (univariate.649) | .706 | .706 |
| Law (N&C, three labels) | cited CFR/USC/FR sections exist; cited part matches the rule; docket id matches | 11–13% cite anything | .44–.50 | .59–.60 | .59–.60 |
| Journalism (editorial pickup) | internal arithmetic identities; URLs resolve; dates and weekdays agree | 0.1% / 20% / 50% | .462 | .611 | .607 |
| Journalism (editorial pickup) | stated revenue, income and EPS match the filer’s SEC XBRL facts | 0.6% | .458 | .611 | .610 |
| campaign | proposing rounds | proposers per round | families |
|---|---|---|---|
| joke upvotes | 5 | 8 (rounds 1–5); 5 in round 6, Claude slots missing | Claude, Codex, GLM |
| N&C received a response | 4 | 4 (Opus, Sonnet, two Codex); GLM in round 2 only | Claude, Codex (GLM once) |
| editorial pickup (press) | 2 | 6 | Claude, Codex, GLM |
| BBC most-read | 6 | 8 (four Codex, four GLM) | Codex, GLM |
| spotlight/oral pick | 5 | 4 in rounds 1–2; 6 from round 3 | Claude, Codex, GLM |
| citation percentile | 6 | 4 in rounds 1–2; 6 from round 3 | Claude, Codex, GLM |
| task | articulated judge (Gemma-4-31B) | full model (Llama-3.1-8B LoRA) |
|---|---|---|
| creative writing, story upvotes (and other story tasks) | head tail characters ( tokens) | first tokens ( characters) |
| code, GitHub PR merged (v3, enriched record) | full record (median tokens) | first tokens |
| peer review (abstract tasks) | full abstract and metadata | full abstract and metadata |
| press releases, editorial pickup | first characters ( tokens; was under the retired Llama scoring) | first tokens ( characters) |
| math.SE bounty | SO bounty | |||
|---|---|---|---|---|
| Judge | ||||
| Gemma-4-31B (record) | .696 | .723 | .753 | .779 |
| Llama-3.1-70B | .686 | .711 | .745 | .772 |
| Qwen2.5-14B | .680 | .708 | .736 | .766 |
| phi-4 (14B) | .672 | .708 | .726 | .766 |
| gemma-2-9b | .678 | .708 | .738 | .765 |
| Judge | family | anchors (pos vs. neg / coherent vs. scrambled) | ||
|---|---|---|---|---|
| Gemma-4-31B (record rerun) | Gemma | .648 | .660 | .57 / 1.00 |
| phi-4 (14B) | Phi | .634 | .647 | .66 /.98 |
| Qwen-2.5-14B-Instruct | Qwen | .632 | .640 | .71 / 1.00 |
| Mistral-7B-Instruct-v0.3 | Mistral | .592 | .630 | .66 /.97 |
| Llama-3.1-8B-Instruct | Llama | .560 | .619 | .57 /.99 |
| joke upvotes ( ) | N&C responded ( ) | ||||
|---|---|---|---|---|---|
| Judge | family | ||||
| Gemma-4-31B (record) | Gemma | .748 | .752 | .672 | .733 |
| Gemma-4-31B (re-run, control) | Gemma | .749 | .752 | .670 | .736 |
| Qwen-2.5-14B-Instruct | Qwen | .687 | .699 | .631 | .718 |
| phi-4 (14B) | Phi | .671 | .687 | .644 | .719 |
| Llama-3.1-8B-Instruct | Llama | .662 | .679 | .669 | .732 |
| campaign | truncated after | bound | realized after | used | holds |
|---|---|---|---|---|---|
| BBC most-read | round 1 | .007 | .018 | 255% | no |
| BBC most-read | round 2 | .010 | .011 | 104% | borderline |
| BBC most-read | round 3 | .020 | .005 | 24% | yes |
| BBC most-read | round 4 | .029 | .002 | 0% | yes |
| r/Jokes upvotes | round 1 | .012 | .010 | 86% | yes |
| r/Jokes upvotes | round 2 | .039 | .001 | 2% | yes |
| Task | mined gain | whisker | + whisker | status | whisker c | ||
|---|---|---|---|---|---|---|---|
| Editorial pickup (press) | .589 | .091 | .879 | bound | .022 | .032 | |
| Citation percentile | .456 | .069 | .782 | bound | .048 | .040 | |
| Math.SE votes | .283 | .029 | .651 | bound | .016 | .006 | |
| Joke upvotes | .400 | flat 3 | .743 | bound+tick | .028 | .019 | |
| BBC most-read | .608 | flat 2 | .843 | bound+tick | .041 | .064 | |
| Spotlight/oral pick | .589 | .060 | .675 | bound | .003 | .000 |
| campaign | bound | realized | used | holds | ||
|---|---|---|---|---|---|---|
| BBC most-read | 1/6 | 0.45 | +.007 | +.020 | 297% | no |
| BBC most-read | 2/6 | 0.41 ∗ | +.010 | +.014 | 130% | no |
| BBC most-read | 3/6 | 0.49 ∗ | +.020 | +.008 | 39% | yes |
| BBC most-read | 4/6 | 0.51 ∗ | +.029 | +.001 | 4% | yes |
| BBC most-read | 5/6 | 0.49 ∗ | +.025 | +.003 | 12% | yes |
| r/Jokes upvotes | 1/6 | 0.42 | +.012 | +.011 | 94% | yes |
| campaign / round | similarity-only | merged | strict two-judge | leave-one-proposer-out range |
|---|---|---|---|---|
| HashtagWars r4 | 0.73 | 0.73 | 0.37 | [0.73, 0.91] |
| Spotlight r5 | 0.78 | 0.78 | 0.59 | [0.75, 0.84] |
| Citations r5 | 0.57 | 0.57 | 0.60 | [0.52, 0.60] |
| AoPS r1 | 0.50 | 0.56 | 0.49 | [0.53, 0.64] |
| BBC r1 | 0.55 | 0.45 | 0.45 | [0.38, 0.53] |
| BBC r6 | 0.54 | 0.61 | 0.61 | [0.58, 0.67] |
| task | 95% CI | se | whisker | 95% CI | ||
|---|---|---|---|---|---|---|
| Editorial pickup (press) | .589 | [.405,.772] | .022 | .015 | .032 | [ .000, .108] |
| Citation percentile | .456 | [.288,.623] | .048 | .015 | .040 | [ .012, .093] |
| Math.SE votes | .283 | [.166,.401] | .016 | .007 | .006 | [ .001, .015] |
| Joke upvotes | .400 | [.219,.581] | .028 | .003 | .019 | [ .008, .040] |
| BBC most-read | .608 | [.458,.759] | .041 | .003 | .064 | [ .034, .133] |
| Spotlight/oral pick | .589 | [.411,.767] | .003 | .013 | .000 | [ .000, .037] |
| task | ours | HypotheSAEs | HypoGeniC | HypotheSAEs pkg | ExperiGen | union | ours+union | whisker | ceiling | full | gap | ||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Joke upvotes | 3,163 | .747 | .689 | .604 | .686 | .556 | .691 | .707 | .716 | .019 | .734 | .747 | .013 |
| Math.SE accepted | 2,600 | .644 | .574 | .532 | .554 | .526 | .582 | .579 | .575 | .001 | .576 | .644 | .068 |
| Math.SE votes | 2,326 | .654 | .603 | .535 | .611 | .569 | .628 | .632 | .631 | .006 | .637 | .660 | .023 |
| N&C rule change | 1,417 | .624 | .612 | .545 | .607 | .487 | .632 | .644 | .633 | 0 | .633 | .634 | .001 |
| N&C responded | 1,904 | .817 | .791 | .501 | .675 | .470 | .723 | .736 | .802 | .013 | .815 | .841 | .026 |
| ICLR spotlight/oral | 1,571 | .594 | .529 | .517 | .557 | .455 | .514 | .556 | .573 | 0 | .573 | .594 | .021 |
| case | truth | scramble AUC | surface | fingerprint | AUC pooled within stratum | stage 5 | routed |
|---|---|---|---|---|---|---|---|
| planted token, disguised criterion | spurious | .51 | .89 | .60 | .70 .50 | incidental | yes |
| planted token, literal criterion | spurious | .51 | 1.00 | .62 | .71 .50 | incidental | yes |
| control: reveal is present and locatable | quality | .49 | .05 | .22 | .59 .59 | quality | no |
| control: point survives as a quotable formulation | quality | .49 | .03 | .13 | .62 .63 | quality | no |
| 112 record quality criteria (47 base, 65 mined) | quality | 0 of 112 flagged | no | ||||
| 68 record nuisance criteria | spurious | 6 of 68 flagged by statistics | incidental | yes (stage 5) | |||
| Task | Grouping unit | Declared | Largest declared spurious features (by label AUC on their own) |
|---|---|---|---|
| Creative writing (crowd) | story prompt | 70 | Fragmented staccato paragraphing / line-break repetition; Blank-line and short-paragraph density; Dialogue punctuation mark count |
| N&C responded | docket | 56 | First-hand local exposure; First-person credential or stake self-disclosure; Identified accountable author with stated standing |
| Peer review (citation) | paper | 57 | Extent of currently fashionable subfield vocabulary; Trend-Aligned Vocabulary; Count of enumerated evaluation settings and resource markers |
| Press releases (pickup) | company | 23 | Extent of departure from templated corporate register; Extent of newswire dateline, ticker and boilerplate formatti; Formal wire-service dateline / distribution convention |
| Jokes (crowd) | topic cluster | 58 | Extent of overlap with stock circulated joke material; Pre-owned chestnut: the pun arrives already familiar; Canonical-chestnut retelling fingerprint (position in retell |
| Math SE (accepted) | question | 36 | number of answers; answer position; is the first answer |
| Task | Full-model AUC | Residual AUC (all ) | Asymptote | placebo drop (share) | |
|---|---|---|---|---|---|
| Creative writing (crowd) | 70 | .792 | .737 | .738 | +.008 (0.15) |
| N&C responded | 56 | .817 | .733 | .733 | +.035 (.42) |
| Peer review (citation) | 57 | .884 | .773 | .777 | +.029 (0.26) |
| Press releases (pickup) | 23 | .774 | .729 | .726 | +.034 (0.73) |
| Jokes (crowd) | 58 | .747 | .641 | .641 | +.004 (0.04) |
| Math SE (accepted) | 36 | .644 | .579 | .560 e | +.016 (0.24) |
| Field | task | Channel | mined | ||||
|---|---|---|---|---|---|---|---|
| Coding | GitHub PR merged | verdict | .550 | .631 | .638 | u/q m | .692 |
| Stack Overflow accepted answer | verdict | .639 f | .733 | n/c | .743 | .764 [.736, .791] | |
| Competition editorial-approach | curated | .743 | .734 | n/c | .724 m | .737 m | |
| Stack Overflow bounty award | curated | .739 f | .787 | n/c | .804 | .811 [.791, .829] | |
| Stack Overflow answer votes | community | .638 | .710 | n/c | .723 | .728 | |
| Math | Math.SE accepted answer | verdict | .591 | .632 | .618 f | .644 | .630 |
| task | groups | AUC | boot. 95% CI | LOGO | half sd | pairwise | |
|---|---|---|---|---|---|---|---|
| GitHub PR merged (transition frame) | 7,563 | 184 | .630 | [.587,.677] | .012 | .035 | .646 |
| Homepage placement | 12,998 | 1,229 | .734 | [.722,.746] | .001 | .006 | .708 |
| Kindle Scout | 726 | 3 | .724 | [.714,.790] | .059 | .028 | .715 |
| Math.SE bounty | 23,972 | 8,278 | .723 | [.716,.729] | n/r | .003 | .711 |
| RoyalRoad followers | 3,604 | 149 | .630 | [.611,.648] | .003 | .011 | .655 |
| RoyalRoad market pickup | 1,274 | 1,274 | .562 | [.530,.593] | .001 | .021 | n/r |
| Task | groups | ||||
|---|---|---|---|---|---|
| math.SE bounty | 23,972 | 8,278 | .691 [.684,.699] | .723 [.716,.730] | [+.028,+.035] |
| SO bounty | 23,498 | 7,036 | .739 [.732,.746] | .779 [.772,.785] | [+.036,+.044] |
| SO accepted | 16,001 | 5,972 | .639 [.631,.648] | .717 [.708,.725] | [+.070,+.085] |
| r/Jokes removal | 18,879 | 55 | .603 [.542,.857] | .701 [.651,.903] | [+.043,+.111] |
| Kindle Scout | 726 | 3 | .612 [.608,.685] | .724 [.714,.790] | [+.105,+.142] |
| Macro-norm humor (pooled) | 73,268 | 5 | .561 [.548,.572] | .714 [.690,.742] | [+.135,+.185] |
| Source | Side | Measured bound | Where |
|---|---|---|---|
| Judge family (7 judges, 6 families) | App. C.5 | ||
| Judge panel (7-family ensemble) | App. C.5 | ||
| Per-judge template optimization | App. C.5 | ||
| Per-judge prompt optimization | (peer review) | App. H.3 | |
| Prompt/implementation headroom | parallel forms ( Spangher, 2026 ) | ||
| Criteria-mining exhaustion | asymptote + whisker | § 4 , App. D.2 |
| field | tacit share | jargon fraction | distinctiveness | most discipline-specific words |
|---|---|---|---|---|
| peer review | llms, neural, adversarial, outperforms | |||
| creative writing | wp, didn, sighed, nodded, stared | |||
| humor | contest, prompt, entry, hashtag | |||
| code | import, com, self, github, def | |||
| notice and comment | docket, FAA, NPRM, airworthiness | |||
| mathematics | frac, math, sqrt, mathbb, infty |
| property | all ( ) | community (6) | verdict (6) | curated (7) | with the object control |
|---|---|---|---|---|---|
| rare-word share | ( ) | ||||
| out-of-vocabulary share | ( ) | ||||
| mean word length | ( ) | ||||
| lexical diversity | ( ), sign flips | ||||
| compressibility | ( ), length proxy | ||||
| reading grade | ( ) |
| task | tacit | who decides | deciders per item | producers | institutions |
|---|---|---|---|---|---|
| story upvotes | anonymous crowd | mean / median upvotes ℓ | not in data | none by construction | |
| citation percentile | the citing community | / citing works ℓ | not in data | not in data | |
| ICLR accept/reject | reviewers + area chair | / reviews | authors, /paper | not in data | |
| spotlight/oral | programme committee | committee absent from data | authors, /paper | not in data | |
| RoyalRoad followers | anonymous readers | / followers ℓ | not in data | none by construction | |
| pull request merged | repository insiders | / review comments | not in data | orgs, HHI |
| task | full model | articulated | memory alone | articulated + memory | gap | share closed | |
|---|---|---|---|---|---|---|---|
| citation percentile | |||||||
| #HashtagWars | |||||||
| story upvotes | |||||||
| ICLR accept/reject | |||||||
| Math.SE accepted | |||||||
| Math.SE votes |
| Task | Community | ||||
|---|---|---|---|---|---|
| Humor | pooled (5 subs) | 73,268 | .561 | .712 | .714 |
| r/tifu | 12,650 | .605 | .798 | .802 | |
| r/Showerthoughts | 17,942 | .616 | .758 | .769 | |
| r/me_irl | 7,324 | .582 | .749 | .752 | |
| r/nottheonion | 12,332 | .570 | .727 | .737 | |
| r/funny | 23,020 | .595 | .715 | .727 |
| Task | Channel | 95% CI | |||
|---|---|---|---|---|---|
| r/Jokes moderator removal | verdict | .904 | .922 | .923 | [.908,.939] |
| Kindle Scout publisher acceptance p | verdict | .744 | .744 | .792 | [.627,.935] |
| StackOverflow accepted answer | verdict | .733 | .743 | .764 | [.736,.791] |
| StackOverflow bounty award | curated | .787 | .804 | .811 | [.791,.829] |
| math.SE bounty award | curated | .714 | .733 | .738 | [.717,.759] |
| News-article tweet engagement | community | .600 | .643 | .646 | [.627,.665] |
| field | density ( count) | task | tacit share |
|---|---|---|---|
| math Q&A | 4.79 | accepted answers; answer votes | .083;.255 |
| creative writing | 5.62 | story upvotes | .435 |
| humor | 5.85 | joke upvotes; HashtagWars | .150;.125 |
| notice and comment | 6.32 | responded; outcome | .133;.145 |
| peer review | 6.57 | verdict; curation; revealed | .394;.360;.417 |
| press releases | 7.10 | pickup | .109 |
| pooled AUC | within-subfield AUC | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| task | subfield (levels) | rate range | artic. | +subfield | full | artic. | +subfield | separate | full | gap pooled / within / separate |
| creative writing: upvotes | genre tag (9) | .07–.57 | .668 | .684 | .792 | .653 | .656 | .645 | .779 | .124 /.126 /.135 |
| N&C: received response | agency (14) | .39–.89 | .796 | .827 | .817 | .772 | .768 | .691 | .804 | .021 /.032 /.113 |
| N&C: rule change | agency (13) | .22–.56 | .618 | .619 | .624 | .605 | .594 | .545 | .615 | .006 /.009 /.070 |
| N&C: cosigner agreement | agency (12) | .43–.87 | .562 | .556 | .603 | .538 | .506 | .523 | .593 | .041 /.055 /.070 |
| peer review: accept/reject | subject area (6) | .32–.55 | .647 | .653 | .777 | .649 | .654 | .643 | .776 | .130 /.127 /.133 |
| agency | rows | own criteria | pooled bank | + own | + others’ | + all | full model |
|---|---|---|---|---|---|---|---|
| DOT | 476 | 7 | 0.879 | 0.865 | 0.860 | 0.856 | 0.604 |
| FWS | 416 | 8 | 0.677 | 0.671 | 0.716 | 0.720 | 0.868 |
| FDA | 415 | 8 | 0.653 | 0.655 | 0.673 | 0.658 | 0.655 |
| CMS | 413 | 8 | 0.693 | 0.694 | 0.694 | 0.699 | 0.836 |
| NHTSA | 408 | 8 | 0.722 | 0.707 | 0.714 | 0.701 | 0.780 |
| ED | 406 | 8 | 0.755 | 0.771 | 0.788 | 0.778 | 0.884 |
| base | instruct | |||||||
|---|---|---|---|---|---|---|---|---|
| pair | readout | task | anchor | bank | anchor | bank | bank | bank of record |
| Llama-3.1-8B | free generation | peer review: accept/reject | .99 | .524 | .996 | .582 | +.058 | .654 |
| N&C: received response | .78 | .545 | .94 | .677 | +.132 | .796 | ||
| creative writing: upvotes | .94 | .558 | .99 | .543 | .015 | .668 | ||
| Gemma-3-12B | first-token log-prob | peer review: accept/reject | .75 | .559 | .999 | .663 | +.104 | .654 |
| N&C: received response | .47 | .602 | .81 | .650 | +.048 | .796 | ||
| peer review: accept/reject | N&C: received response | creative writing: upvotes | |||||
|---|---|---|---|---|---|---|---|
| judge | training | record | optimized | record | optimized | record | optimized |
| Gemma-3-12B | base | .559 | .568 | .601 | .596 | .564 | .573 |
| Gemma-3-12B | instruct | .663 | .668 | .651 | .653 | .642 | .646 |
| Llama-3.1-8B | base | .546 | .570 | .601 | .638 | .603 | .601 |
| Llama-3.1-8B | instruct | .625 | .632 | .674 | .691 | .599 | .592 |
| Gemma-4-31B (judge of record) | instruct | .642 | .694 | .681 | .672 | .640 | .652 |
| task | cutoff | late rows | paper | early | drift | paper | early | drift | gap early |
|---|---|---|---|---|---|---|---|---|---|
| press pickup | 2021 | 203 | .734 | .668 | [ , ] | .691 | .718 | ||
| joke upvotes | 2018 | 781 | .783 | .741 | [ , ] | .724 | .734 | ||
| ICLR accept/reject | 2024 | 413 | .796 | .785 | [ , ] | .640 | .641 | ||
| BBC most-read | 2022 | 1,535 | .779 | .772 | [ , ] | .730 | .727 | [ , ] | |
| spotlight/oral | 2024 | 397 | .595 | .587 | [ , ] | .489 | .505 | [ , ] | |
| story upvotes | 2021 | 1,385 | .762 | – | – | .635 | .617 | [ , ] | – |
| Task | Channel (how computed) | Items with it | Gap, raw | Gap, within strata |
|---|---|---|---|---|
| Press pickup | issuing company (group id; read within company) | – | ||
| Creative writing (crowd) | non-[WP] feature-thread tag in the prompt line (pattern) | 9.8% | ||
| Homepage placement | section kicker or format label glued to the headline (pattern) | 5.6% | ||
| Math SE (vote) | answer says it supplements earlier answers (pattern) | 4.4% | ||
| Math SE (accepted) | same | 5.0% | ||
| GitHub PR merged | bot dependency update; status token ([WIP], do not merge); maintainer share of review comments (pattern) | 7.6%, 1.1% |
| Rival account | Tasks | Key result | Verdict |
|---|---|---|---|
| One decider’s taste (decider holdout, offerer-disjoint retrain) | Math.SE bounty, SO bounty, SO accepted | seen vs unseen deciders , ; retrain | not supported |
| Timing and exposure | jokes, stories, SO votes, RoyalRoad, tweets | jokes ; stories , within-quintile shrink ; others | mixed (stories) |
| Coarse criterion readout | jokes, stories, ICLR, citations | , , , [CI crosses 0] | not tripped; instrument |
| Judge noise | same four | paraphrase ICC ; six-readout average | not tripped; instrument |
| One holistic judgment | citations, jokes, stories, ICLR | +holistic vs full ; stories below bank | instrument only (reading retracted) |
| Vocabulary limit (description bottleneck) | stories | description vs raw at 6k rows (three seeds), vs at 24k, vs on all rows; neutral rewrite | description holds 77% of the gap at full size |
| Task | Property | Measured by | Share of items | High-gap / matched (%): preferred not | Label AUC | Scope | Flag |
|---|---|---|---|---|---|---|---|
| Story upvotes | Dialogue-heavy: 25% or more of the story’s characters in quotes | pattern | 35.4% | 16/30 56/28 | text | ||
| Payoff on a reference the reader must already know | judge | 16.8% | 14/20 13/11 | outside | N | ||
| Written as a document or account (report, chronicle, log, letter) | judge | 14.9% | 19/21 9/6 | text | |||
| Scenes of dialogue that do not change the situation | judge | 11.2% | 7/5 13/10 | text | |||
| HashtagWars pick | Minimal-edit substitution into a known title or lyric | judge | 50.5% | – 73/61 | outside | N | |
| BBC most-read | Curiosity, oddity or explainer | judge | 41.6% | 65/12 15/68 | text | C |
| descriptions in the bank | new spurious channels | ||||||
|---|---|---|---|---|---|---|---|
| Task | added | gap: raw with them | declared | gap within strata | both | : raw with | : raw with |
| Story upvotes | 9 | 3 | |||||
| Joke upvotes | 7 | 3 | |||||
| HashtagWars pick | 4 | 3 | |||||
| BBC most-read | 11 | 2 | |||||
| Press pickup | 3 | 7 | |||||
| Task | Excerpt (preferred item; the full model rates it well, the bank poorly) | ||
|---|---|---|---|
| Story upvotes | “The creatures were all smiles. A bunch of tall, pale, ghostly things with these dark smiles.” (prompt: humans are a thousand times more sensitive to the smell of petrichor than sharks are to blood) | ||
| Story upvotes | “Before we first discovered alien life, our best astronomers believed that we were the only ones in our galaxy.” (a future history in essay form, no scene and no dialogue) | ||
| Joke upvotes | “I like to imagine the guy who invented the umbrella was going to call it the ‘brella’… But he hesitated…” | ||
| Joke upvotes | “These no nut November memes / They’re really getting out of hand” | ||
| Math.SE answer votes | “Yes. Since is free, the sequence splits and you get what you want.” | ||
| Math.SE accepted answer | “This may actually be a matter of ‘just practice more.’ If you’ve done enough factoring, you’ll recognize the coefficients…” |
| Task | Description | Verdict | Reason |
|---|---|---|---|
| Story upvotes | Uses the prompt as a springboard and goes elsewhere | invalid | fires on 73% of stories; most flagged stories carry out the prompt |
| Single narrating voice (under 3% of characters in quotes) | noisy | dialogue marked with dashes passes; the cut splits mostly narrated stories | |
| Withholds a key fact about its situation or narrator | noisy | fires on 49%; misses stories that do withhold one | |
| Serialised part or excerpt | noisy, process | records how the story was posted | |
| Joke upvotes | Pun by sound or spelling, including one left unspelled | noisy | flags non-puns; misses puns whose answer is not written out |
| Deadpan literal statement that refuses to be a joke | noisy | reads title-only fragments as deadpan |
| task | readout | units | hit (artic.) | hit (full) | hit (random) | regret | 95% CI | swap rate | wins full:artic. |
|---|---|---|---|---|---|---|---|---|---|
| peer review: citations ∗ | pairwise | 56,901 pairs | .714 | .884 | .500 | +.171 | [+.125, +.216] | 0.27 | 12565:2852 |
| humor: #HashtagWars picks | grouped | 8 groups | .125 | .250 | .099 | +.125 | [ .250, +.500] | 1.00 | 2:1 |
| peer review: accept/reject ∗ | pairwise | 383,426 pairs | .668 | .777 | .500 | +.109 | [+.078, +.140] | 0.31 | 80645:38878 |
| N&C: rule change | grouped | 69 groups | .420 | .522 | .453 | +.101 | [ .043, +.246] | 0.65 | 16:9 |
| humor: joke upvotes | grouped | 10 groups | .900 | 1.000 | .492 | +.100 | [+.000, +.300] | 1.00 | 1:0 |
| N&C: received response | grouped | 20 groups | .750 | .850 | .684 | +.100 | [ .100, +.300] | 0.65 | 3:1 |
| task | criteria | decidable groups | induced gap | regret | [95% CI] | |
|---|---|---|---|---|---|---|
| math.SE: answer votes | 92 | 1,140 | .17–.77 | .021–.119 | +.87 [.57,.97] | .001 |
| creative writing: upvotes | 144 | 523 | .43–.80 | .061–.174 | +.77 [.40,.95] | .001 |
| math.SE: accepted answers | 80 | 992 | .41–.77 | .033–.107 | +.71 [.28,.93] | .002 |
| humor: joke upvotes | 124 | 10 | .10–.59 | .000–.600 | +.65 [.22,.87] | .007 |
| humor: #HashtagWars picks | 106 | 8 | .56–1.18 | .125–.250 | +.64 [.16,.91] | .007 |
| journalism: editorial pickup | 155 | 8 | .25–1.06 | .000–.500 | +.54 [.03,.88] | .031 |
| task | removed ( ) | bank AUC | induced gap | regret | swap rate |
|---|---|---|---|---|---|
| creative writing: upvotes | none (full bank) | .665 | .435 | .061 | .46 |
| 32 most predictive | .652 | .478 | .121 | ||
| 32 least predictive | .662 | .446 | .096 | ||
| 32 random (mean of 3) | .662 | .446 | .091 | ||
| math.SE: accepted answers | none (full bank) | .569 | .518 | .067 | .45 |
| 32 most predictive | .550 | .652 | .080 |
| top pick above the floor | real vs. generated | picks a real item | ||||||
|---|---|---|---|---|---|---|---|---|
| task | pools | real base rate | agree | full on artic. | artic. on full | [95% CI] | artic. / full | artic. / full |
| creative writing (150 prompts) | 150 | .11 | .26 | +.210 | +.176 | .74 /.54 | .39 /.19 | |
| creative writing (402 prompts) | 402 | .25 | .25 | +.264 | +.216 | .71 /.53 | .74 /.55 | |
| joke upvotes | 300 | .08 | .26 | +.234 | +.183 | .80 /.58 | .44 /.23 | |
| Math.SE answer votes | 150 | .15 | .13 | +.121 | +.133 | .46 /.48 | .26 /.27 | |