Two scientific findings can disagree without contradicting each other. Determining whether they conflict requires knowing whether they describe comparable measurements. We study how language models behave at this decision point. In a controlled task, we generate an unsatisfiable XOR constraint system and translate its constraints into scientific reports from different laboratories. One assignment satisfies more constraints, while another satisfies fewer but better matches expected biology. This creates a simple dilemma: does the model choose the assignment that best fits the constraints, or the one that better matches biological expectations? When the constraints are stated directly, GPT-5.6 Sol and Claude Opus 5 recover the best-supported assignment in 90% and 96% of cases, respectively. In scientific prose, however, the models behave differently. Claude Opus 5 often prefers the biologically expected assignment. Removing that biological preference increases recovery of the better-supported assignment from 27% to 79% (p<.001); recovery reaches 92% when the same Biology-favored record is accompanied by a formalization request and an explicit paired-design cue (p<.001). GPT-5.6 Sol is less sensitive, with neither corresponding change reaching statistical significance. These results suggest that reliable scientific verification depends not only on formal reasoning, but also on how models decide which findings should be compared and what relations they imply.
Figures & tables
Figure 1: Scientific-framing manipulation. Reversing both reports in a paired experiment changes its biological interpretation while preserving the same XOR relation. In this example, the two laboratories still disagree after both reports are reversed, so the formal constraint remains xA⊕xC=1 . We apply this transformation to selected pairs to change the biological preference between the two leading assignments without changing their formal support.
Condition
What the model receives
Formal
24 constraints directly, without scientific prose.
Scientific framing
Biology-favored
Same constraints in prose; the 22/24 alternative better matches expected biology.
Biology-neutral
Same constraints in prose; that biological preference is removed.
Structural guidance
Formalize only
Biology-favored record with a general request to formalize.
Table 1: Five versions of the same underlying formal problem.
Figure 2: The two primary models respond differently to scientific framing and structural guidance. Formal is the direct-constraint reference. Panel (a) compares Biology-favored and Biology-neutral prose; panel (b) compares the same Biology-favored record with and without the paired-design statement. Labels report the paired change in exact 23/24 recovery ( Δ ) with 95% case-bootstrap confidence intervals. Four additional model variants are reported in Appendix F . Full counts for the 23/24, 22/24, and Other outcomes are reported in Appendix Table 5 .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Check
Verified result
Prospective balance
48 cases in two 24-case tranches: 3 graph families × 2 prose families × 8 repetitions overall
Formal landscape
One 23/24 assignment; one 22/24 assignment; all others ≤21/24
Exhaustive verification
All 212=4,096 assignments scored for every case
Biology-favored fit
22/24 minus 23/24 biological-fit range: 15 to 21 reports
Biology-neutral fit
23/24 minus 22/24 biological-fit range: −1 to 1 report
Prompt integrity
Unique prompts, one response schema, and frozen file hashes before evaluation
Appendix
Table 2: Pre-call construction and manipulation checks.
ABC
A=B
B=C
A=C
Satisfied
000
✓
✓
–
2/3
001
✓
–
✓
2/3
010
–
–
–
0/3
011
–
✓
✓
2/3
100
–
✓
✓
2/3
101
–
–
–
0/3
Appendix
Table 3: All eight assignments in the three-laboratory toy example. Here, 0 denotes unchanged records and 1 denotes flipped records.
Figure 3: Case families used in the robustness analysis. a , The defining qualitative structure of each graph family; drawings are schematic rather than particular sampled cases. b , The two narrative organizations used to present the scientific records.
GPT-5.6 Sol
Claude Opus 5
Case family
Favored
Neutral
Favored
Neutral
Graph topology
Double hub
12/16
14/16
3/16
12/16
Modular
10/16
15/16
5/16
13/16
Ring with chords
9/16
8/16
5/16
13/16
Prose family
Appendix
Table 4: Exact 23/24 recovery in the Biology-favored and Biology-neutral conditions by prespecified graph topology and prose family. Cells show exact recovery divided by the number of cases in each subgroup.
Condition
n
23/24
22/24
Other
GPT-5.6 Sol
Formal
48
43
3
2
Biology-favored
48
31
15
2
Biology-neutral
48
37
9
2
Formalize only
48
42
5
1
Formalize + paired design
48
41
4
3
Appendix
Table 5: Exact outcome counts for the two primary models in every condition. “Other” includes any different assignment, invalid output, or timeout. Each condition contains 48 scheduled calls.
Model
Comparison
Matched n
Target recovery
Change, pp [95% CI]
p
GPT-5.6 Sol
Favored → Neutral
48
31/48→37/48
+12.5[−4.2,29.2]
.210
GPT-5.6 Sol
Formalize only → Formalize + paired design
48
42/48→41/48
−2.1[−14.6,10.4]
1.000
Claude Opus 5
Favored → Neutral
48
13/48→38/48
+52.1[33.3,68.8]
<.001
Claude Opus 5
Formalize only → Formalize + paired design
48
21/48→44/48
+47.9[31.2,64.6]
<.001
Appendix
Table 6: Paired comparisons of exact 23/24 recovery. Changes are the second condition minus the first. Intervals are 95% case-bootstrap intervals; p is the exact two-sided McNemar test over discordant pairs. All comparisons contain 48 matched cases.
Figure 4: Direct comparison of manipulation effects across the tested model variants. Panel (a) shows the paired change from Biology-favored to Biology-neutral; panel (b) shows the paired change from Formalize only to Formalize + paired design. Small points are the three model-specific changes within each family. Large symbols show the mean across the three tested variants in each family, and horizontal bars show ±1 standard deviation across those model-level changes. GPT includes GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.5; Claude includes Claude Opus 5, Claude Opus 4.8, and Claude Sonnet 5. Values use all 48 scheduled calls per condition, with invalid outputs and timeouts counted as failures.
Model
Condition
n
23/24
22/24
Other
GPT-5.6 Sol
Formal
48
43
3
2
Biology-favored
48
31
15
2
Biology-neutral
48
37
9
2
Formalize only
48
42
5
1
Formalize + paired design
48
41
4
3
GPT-5.6 Terra
Formal
48
35
8
5
Appendix
Table 7: Exact outcome counts across all evaluated models and conditions. The 23/24 and 22/24 columns report exact recovery of the constraint-maximizing assignment and the prespecified nearby alternative, respectively. “Other” includes every different assignment, invalid output, or timeout. Each condition contains 48 scheduled calls.
Figure 5: Descriptive efficiency comparison for the primary models. Each arrow goes from Formalize only to Formalize + paired design. Tokens and latency are medians. GPT cost is an API-equivalent estimate from logged usage; Claude cost is provider-reported. Reasoning-token scales are model-specific and should not be compared across panels.