Language-model agents increasingly answer questions over customer-relationship management (CRM) records, such as whether to qualify a sales lead. We identify a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism - here the sales representative, a witness recorded in the CRM - the model treats the assertion as evidence and clears deals the company's own records deem unacceptable. Across 100 lead-qualification tasks from CRMArena-Pro, the representative asserts an acceptable timeline in every call and an acceptable budget in 76; on the 31 tasks where such an assertion contradicts the price list and installation policy, a model reading only the transcript clears the deal in 29 of 31 cases. The signature is consistent across seven models from four providers (misled on 87-97%); scale and explicit reasoning confer no resistance. Only 3 of 35 genuine failures involve no assertion: the failure is persuasion, not missing information. We contribute a diagnostic method rather than an architecture: (i) a bucket analysis that separates persuasion from information gaps, (ii) a same-information control showing that supplying the records to the model lowers strict accuracy from 41 to 18 while raising recall - precision collapses - and (iii) a compute-step control that holds extraction fixed and varies only who computes Budget and Timeline. The margin ranges from 42 points on an inexpensive model to 2-5 points on models that already compute correctly; on the strongest models the arms are within confidence intervals, so the pattern is a consistent direction and a soundness property, not a proved performance floor. We pre-specify a generalization test that returns a negative result, characterize the precondition (a policy exactly specified in the inputs), and release all evaluation artifacts.
Figures & tables
Dimension
Claims acceptable
Genuinely fails
… and claimed acceptable
Model misled
Timeline
100/100
12
12
12/12
Budget
76/100
23
19
17/19
Contradiction cases (pooled)
—
35
31
29/31 (94%)
TABLE I: What the representative claims versus what the records say, across 100 lead-qualification calls. The claim that the deal is acceptable is asserted almost always; on the deals where it is false, the model accepts it.
Model
Provider
Misled % [95% CI]
Caught % [95% CI]
Current-generation
GPT-5
OpenAI
90.3 [75–97]
80.6 [64–91]
gpt-5.6-sol
OpenAI
90.3 [75–97]
90.3 [75–97]
Claude Sonnet 5
Anthropic
96.8 [84–99]
83.9 [67–93]
Claude Fable 5.1
Anthropic
93.5 [79–98]
87.1 [71–95]
Kimi K2.6
Moonshot
87.1 [71–95]
83.9 [67–93]
TABLE II: Contradiction persists across model families. On the 31 contradiction cases, the fraction of deals each model clears despite the records (“misled by the pitch”) and the fraction the grounded extract-then-compute procedure recovers, with 95% Wilson confidence intervals ( n=31 ). Every model is misled on at least 87%. Kimi K2.6 and Qwen 3.8 Max permit only temperature 1, so those rows are one sample of a non-deterministic run; all others are temperature 0. Figure 4 carries two older OpenAI models (GPT-4o-mini and o3-mini) that we keep for the compute-step chart because they show what code does when the model is worst at the arithmetic.
Rung — method
Reader
Tr.
Te.
Full
Direct from transcript
efficient
44
40
41
+ chain-of-thought
standard
56
—
—
+ records, model does arithmetic
standard
28
—
—
Extract → compute in code
standard
88
90
89
same, GPT-4o-mini
efficient
—
—
84
TABLE III: The method ladder. Train and test are the CRMArena 80/20 split; “full” is all 100 tasks.
Condition
Records go to
Strict % [95% CI]
Contain
Transcript only
—
41 [32–51]
58
+ records, model reasons to verdict
model
18 [12–27]
67
same, conservative prompt
model
11 [6–19]
62
Extract → compute
code
85 [77–91]
86
Compute-step control: one extraction, same four fields, catalog in context; only who computes Budget/Timeline differs
GPT-4o-mini, model computes
model
43 [34–53]
60
TABLE IV: Same-information control (all 100 tasks; strict shown with 95% Wilson intervals, n=100 ). Upper block: an GPT-4o-mini throughout. Strict = exact set match; contain = all gold factors present, extras tolerated. Handing the model the records raises contain and lowers strict: recall up, precision down. Lower block: the compute step isolated on five models from three providers—identical extraction, only who computes Budget/Timeline differs. Code never lowers either metric; the margin is largest on the inexpensive model, 14–19 points pooled over two runs on the frontier and reasoning models, shrinks to 2–4 points on Claude Sonnet 5 (three runs) and 5 on Kimi K2.6. † 4 Claude Sonnet 5 responses were not parseable JSON and 5 Kimi K2.6 calls timed out; each is graded wrong in both arms. Over parsed rows only, the pairs are 79/83 (Claude) and 79/84 (Kimi). Kimi runs at temperature 1. Confidence intervals are marginal Wilson intervals on each arm. Because both arms score the same 100 cases, the correct uncertainty for the arm-to-arm difference is a paired McNemar-style interval; the marginal intervals shown here are an upper bound on it. Per-case outcomes are released so the paired test can be computed.
Factor
Correct
Answerable from…
Authority
27/27
transcript content
None (qualifies)
20/20
transcript content
Budget
18/19
catalog prices + computation
Need
12/15
transcript (subjective)
Timeline
9/11
install policy + computation
TABLE V: Per-factor accuracy at the top of the ladder (GPT-4o). Budget and Timeline are computed by code from the extracted quantities; Authority, Need, and None are the model’s own narrow booleans, relayed by code. The computed factors become answerable; the residual Need misses are model judgments.
Extract-then-compute requires a policy that is exactly specified and identifiable from the inputs, not merely documented in prose. Lead qualification satisfies this; invalid_config does not, which bounds what any deterministic method can achieve on it.