The Argument and the Letterhead: Source-Position Coherence in AI Evaluation
Abstract
An argument can be surprising coming from a particular speaker without being a bad argument. Do AI evaluators keep these judgments apart? Two preregistered descriptive studies and a later Jev supplement collected 2,976 usable evaluations of six fixed texts about US AI policy, Germany's debt brake and Swiss nuclear energy. Each text was presented under several source attributions. The key comparison asks whether the gap between two sources changes when the argument changes. On Sol, for example, a national-security argument received mean ratings of 0.359 under CODEPINK and 0.639 under College Republicans; a civil-rights argument received 0.742 and 0.721. A constant preference for one source cannot explain that pattern. Related interactions appeared across topics and recent model configurations, including those with reasoning enabled, while several comparisons yielded small effects. The later European Jev supplement yielded five interactions below the adopted absolute reference of 0.05; its distinct rubric and interrupted collection limit comparison with the chat systems. Some written evaluations explicitly invoked a mismatch between a source and its attributed position. Taken together, the numerical and verbal evidence supports source-position coherence as a plausible explanation, alongside competing accounts involving credibility, authenticity and interpretation of the task. The paper develops this inference through controlled comparisons, reports conditional post hoc p-values in an appendix, and documents the human decisions and delegated checks behind an AI-conducted study.
Figures & tables
| Attributed source | A: reform the debt brake | B: retain the debt brake |
|---|---|---|
| Grüne Jugend | 0.620 | 0.350 |
| Junge Liberale | 0.622 | 0.448 |
| AfD | 0.484 | 0.424 |
| Pair | Gap A | Gap B | I |
|---|---|---|---|
| DE GJ - JL | +0.0027 | -0.0075 | +0.0102 |
| DE DIW - IFO | +0.0019 | +0.0005 | +0.0014 |
| DE AFD - JL | -0.0005 | -0.0084 | +0.0080 |
| CH JG - JF | -0.0042 | +0.0019 | -0.0061 |
| CH SES - AV | -0.0181 | +0.0133 | -0.0314 |
| Configuration | Recorded model identifier | Generation limit | Relevant settings |
|---|---|---|---|
| GPT-4o | gpt-4o-2024-08-06 | 1,024 | temperature 1 |
| Sonnet 4.5 | claude-sonnet-4-5-20250929 | 1,024 | temperature 1; thinking disabled |
| Sol, reasoning off | gpt-6-sol | 1,024 | reasoning_effort none; temperature 1 |
| Sol, reasoning on | gpt-6-sol | 8,192 | reasoning_effort medium; temperature not supplied |
| Sonnet 5, reasoning off | claude-sonnet-5 | 1,024 | thinking disabled; temperature not supplied |
| Sonnet 5, reasoning on | claude-sonnet-5 | 8,192 | adaptive thinking; effort high; temperature not supplied |
| Collection | Release published, UTC | First request, UTC | Finished, UTC | Usable / planned | Client attempts |
|---|---|---|---|---|---|
| Study 1, 24 September | 16:33:48 | 16:39:02 | 17:34:55 | 1,024 / 1,024 | 1,038 |
| Extension, 25 September | 11:04:59 | 11:06:50 | 13:17:44 | 1,664 / 1,664 | 1,669 |
| Jev Europe, 26-27 September | 26 Sep 10:24:02 | 26 Sep 10:26:38 | 27 Sep 17:49:25 | 288 / 288 | 347 |
| Amendment | Public release, UTC | Operational scope |
|---|---|---|
| Repair | 26 Sep 10:53:14 | Recover 18 named exhausted slots; preserve prior attempts |
| Pacing | 26 Sep 15:02:38 | At least 60 seconds after calls; explicit historical hold releases |
| Routing | 26 Sep 16:07:45 | Two eligible providers; first retrospective rating review |
| Bounded failover | 27 Sep 08:57:21 | Up to two internal attempts; second retrospective rating review |
| Account resumption | 27 Sep 12:50:14 | Review first exact 403 access failure; no cap reset |
| Verified paid access | 27 Sep 13:07:28 | Neutral access verified; review second exact 403 |
| Pair | Blocks 1–8: I | Blocks 9–16: I |
|---|---|---|
| DE GJ - JL | +0.0109 | +0.0094 |
| DE DIW - IFO | +0.0012 | +0.0016 |
| DE AFD - JL | +0.0072 | +0.0087 |
| CH JG - JF | -0.0066 | -0.0056 |
| CH SES - AV | -0.0322 | -0.0306 |
| US1 configuration | Pair | First 16: I / p0 / p | Last 16: I / p0 / p |
|---|---|---|---|
| Sonnet 4.5 | CP - CR | -0.3200 / 1.6e-18 / 2e-17 | -0.3188 / 7.8e-19 / 9.9e-18 |
| Sonnet 4.5 | CE - AEI | -0.1000 / 2.2e-07 / 0.00046 | -0.0887 / 5.4e-07 / 0.0025 |
| Gemini Flash | CP - CR | -0.4375 / 3.7e-14 / 2.2e-13 | -0.4612 / 2.1e-14 / 1.1e-13 |
| Gemini Flash | CE - AEI | -0.1206 / 1e-11 / 1.8e-08 | -0.1225 / 7.9e-11 / 1e-07 |
| GPT-4o | CP - CR | -0.0375 / 0.023 / 1 | -0.0500 / 0.00045 / 1 |
| GPT-4o | CE - AEI | +0.0081 / 0.32 / 1 | -0.0137 / 0.054 / 1 |
| Episode | Documented human intervention | Evidence of human control |
|---|---|---|
| 1. Meaning of the contrast | Challenged averaging source-pair interactions; retained pair-specific comparisons (E0056). | Direct conceptual scrutiny of the estimand. |
| 2. Status of the boundary | Asked what 0.05 meant and why it was justified; adopted it as a contestable assumption (E0057). | Scrutiny of measurement and interpretative convention. |
| 3. Inferential ambition | Challenged doubtful independence and the cost of pursuing statistical relevance; requested and adopted the descriptive alternative (E0075-E0076). | Reasoned control over the intended scope of the conclusions. |
| 4. Adversarial checking | Requested adversarial review, asked for counterarguments, and required distinctions between AI proposals and human adoption (conversation; E0005, E0065, E0075). | An expressed disposition to seek counterarguments. |
| 5. Treatment of imperfect responses | Adopted up to two additional attempts to obtain a rating and analysis of obtained ratings (recorded design dialogue). | Control over inclusion policy. |
| 6. Methodological continuity | Rejected adding anonymous baselines and required the extension to preserve the existing comparison method (E0095). | Awareness of how additional conditions change the scientific comparison. |
| Stage | Recorded checks or corrective actions | Evidential contribution and boundary |
|---|---|---|
| Feasibility pilot | Integrity checks on 384 pilot requests; reporting defects corrected and three regression checks recorded (E0052). | Technical feasibility and evidence of correction; pilot outputs do not enter the study results. |
| Study 1 preparation | 37 offline checks and a synthetic end-to-end rehearsal with 1,098 simulated attempts (E0085). | Exercises collection, failure and reporting paths; synthetic responses are not empirical observations. |
| Study 1 completion | Audit of 1,024 selected ratings, 1,038 client attempts, saved identities, hashes and collection chronology (E0090). | Integrity and selection checks; no certification of argument quality or causal interpretation. |
| Extension preparation | Four neutral API checks, 22 offline tests and a 1,664-slot synthetic rehearsal (E0113). | Checks model access, settings and workflow; neutral probes are excluded from scientific results. |
| Extension completion | 1,664 first-usable selections; 5,004 raw-file hash comparisons; 104 cell summaries and 84 full/half interactions recalculated (E0115 and report verification). | Numerical and provenance consistency; repeated checks are neither independent reviewers nor substantive validation of every explanation. |
| Interpretation and manuscript | Focused reading of five paired examples (ten responses), chosen after collection from illustrations obtained under the preregistered first-usable-response rule; all 36 full-sample interactions and 52 reported full/half t calculations cross-checked during manuscript assembly. | Bounded qualitative scrutiny and numerical transfer checks; the five-pair focus was post hoc. No full-corpus qualitative coding or external peer review. |
| Study / configuration | Source | n / cell | Text A | Text B |
|---|---|---|---|---|
| US1 / Gemini Flash | AEI | 32 | 0.7797 | 0.7500 |
| US1 / Gemini Flash | CE | 32 | 0.7516 | 0.8434 |
| US1 / Gemini Flash | CP | 32 | 0.4438 | 0.7750 |
| US1 / Gemini Flash | CR | 32 | 0.7681 | 0.6500 |
| US1 / GPT-4o | AEI | 32 | 0.7475 | 0.8453 |
| US1 / GPT-4o | CE | 32 | 0.7478 | 0.8484 |