DelegationBench: Measuring When AI Agents Should Ask Before Acting
Abstract
AI agents that send emails, edit files, and make purchases must decide when to act on their own and when to check with the user first. This decision is usually evaluated by showing a model a proposed action, asking whether it should proceed, and scoring agreement with human labels. We introduce DelegationBench to test whether such scores can be trusted. It has 156 scenarios with four possible responses (act, ask for permission, ask for missing information, refuse), and most scenarios come in matched pairs that change a single feature: whether the action was requested, what is at stake, whether it can be undone, or who will see it. Across ten models from five families, agreement scores mislead in three ways. A simple keyword rule, which we wrote after seeing the benchmark, agrees with our annotators more often than eight of the models, yet its decision changes in only 9 of 48 matched pairs. Equivalent ways of asking the same question change how often a model acts by up to 52.5 percentage points. And every model stops to ask the user less often when it must carry out the task with tools than when it judges a proposed action. When rules are stated explicitly, the same models follow them almost perfectly, so the gaps are not explained by a general inability to follow rules. We release the benchmark and evaluation tools and recommend reporting these properties separately rather than as one score.
Figures & tables
| Model | 4-way agree | Act/non-act agree | Under-intervention | Over-intervention |
|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | 73.4 | 75.7 | 20.2 | 29.0 |
| Gemini 3.6 Flash | 71.9 | 75.6 | 25.3 | 23.5 |
| Claude Sonnet 5 | 68.7 | 71.1 | 36.5 | 20.0 |
| Claude Opus 4.8 | 65.5 | 68.7 | 40.3 | 20.9 |
| GPT-5.6 Terra | 60.4 | 66.3 | 43.8 | 22.0 |
| GPT-5.6 Luna | 60.0 | 67.4 | 41.5 | 22.3 |
| Model | Natural policy | Structured policy | Decision given | -bench |
|---|---|---|---|---|
| GPT-5.6 Luna | 100.0 | 100.0 | 100.0 | 100.0 |
| Claude Sonnet 5 | 97.6 | 100.0 | 99.4 | 100.0 |
| Gemini 3.5 Flash-Lite | 97.3 | 98.5 | 100.0 | 100.0 |
Appendix figures & tables36 assets
Supplementary material from the paper’s appendix.
Appendix
| System | Judges a proposed action | Separates permission from information | Tests format sensitivity | Tests acting with tools | Continues after the agent asks |
|---|---|---|---|---|---|
| DelegationBench | Yes | Yes | Yes (5 formats) | Yes (Tool-Loop) | Yes (Post-Approval; denial checked against state) |
| When2Call ( Ross et al., 2025 ) | Partial (whether to call a tool) | Not reported | Yes (multiple choice vs. free form) | Not reported | Not reported |
| Suri et al. ( Suri et al., 2026 ) | Not reported | Partial (information side only) | Not reported | Not reported | Not reported |
| Janus ( Brigham et al., 2026 ) | Yes | No | Not reported | Not reported | Partial |
| AuthBench ( Yan et al., 2026 ) | Yes | Not reported | Not reported | Not reported | Not reported |
| AgentAbstain ( Liu et al., 2026 ) | No (agent behavior only) | No | Not reported | Yes | Not reported |
| Provider | Model identifier | Setting |
|---|---|---|
| OpenAI | gpt-5.6-sol | temperature 1; medium reasoning |
| OpenAI | gpt-5.6-terra | temperature 1; medium reasoning |
| OpenAI | gpt-5.6-luna | temperature 1; medium reasoning |
| Groq | openai/gpt-oss-120b | temperature 1; medium reasoning |
| Groq | openai/gpt-oss-20b | temperature 1; medium reasoning |
| Anthropic | claude-sonnet-5 | adaptive thinking; medium effort |
| Model | ACT | ASK | REQ | REF | INV | ASK ACT | REF ACT | Entropy |
|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol | 65.6 | 22.2 | 10.1 | 2.1 | 0.0 | 64.6 | 6.0 | 0.049 |
| GPT-OSS-120B (Groq) | 64.0 | 16.9 | 11.7 | 7.4 | 0.0 | 47.0 | 20.6 | 0.137 |
| GPT-OSS-20B (Groq) | 65.3 | 13.1 | 12.4 | 9.2 | 0.0 | 37.6 | 26.6 | 0.157 |
| GPT-5.6 Terra | 58.5 | 22.9 | 13.7 | 4.9 | 0.0 | 55.2 | 11.7 | 0.067 |
| GPT-5.6 Luna | 57.6 | 23.6 | 14.1 | 4.7 | 0.0 | 55.6 | 11.2 | 0.087 |
| Claude Sonnet 5 | 55.8 | 31.2 | 9.1 | 4.0 | 0.0 | 70.4 | 9.0 | 0.037 |
| Model | 5/5 | 4/5 | Pairwise | Norm. entropy |
|---|---|---|---|---|
| GPT-5.6 Sol | 87.8 | 96.2 | 0.944 | 0.049 |
| GPT-OSS-120B (Groq) | 68.6 | 84.6 | 0.840 | 0.137 |
| GPT-OSS-20B (Groq) | 66.0 | 81.4 | 0.821 | 0.157 |
| GPT-5.6 Terra | 84.6 | 92.9 | 0.922 | 0.067 |
| GPT-5.6 Luna | 80.1 | 92.3 | 0.901 | 0.087 |
| Claude Sonnet 5 | 91.7 | 96.2 | 0.958 | 0.037 |
| Domain | ACT median [range] | ASK median [range] | Majority agreement median [range] | Stability median [range] |
|---|---|---|---|---|
| Calendar | 65.2 [56.6, 72.4] | 13.8 [8.3, 24.8] | 68.6 [59.3, 73.8] | 82.8 [48.3, 96.6] |
| Code/system | 77.1 [46.7, 86.7] | 9.6 [5.0, 45.8] | 55.2 [50.0, 73.6] | 87.5 [79.2, 95.8] |
| Communication | 62.2 [51.1, 63.7] | 25.4 [16.3, 30.4] | 69.2 [63.1, 73.1] | 90.7 [74.1, 100] |
| Files | 30.3 [18.9, 42.3] | 42.6 [26.3, 59.4] | 58.2 [37.6, 70.9] | 75.7 [57.1, 97.1] |
| Publishing/social | 92.1 [85.7, 100] | 4.3 [0, 11.8] | 81.4 [72.9, 87.1] | 85.7 [57.1, 100] |
| Purchasing | 50.4 [12.6, 77.0] | 35.7 [7.4, 76.3] | 49.2 [20.0, 90.4] | 79.6 [48.1, 96.3] |
| Family | Lexical cue | Consequential context | Interaction |
|---|---|---|---|
| Calendar change | 5.0 | 10.0 | |
| Code change | 6.0 | 12.0 | |
| Email send | 1.0 | 2.0 | |
| File cleanup | 2.0 | ||
| Purchase commitment | 10.0 | 20.0 | |
| Social publish | 7.0 | 14.0 |
| Model | Consistency | Large shifts | Model | Consistency | Large shifts |
|---|---|---|---|---|---|
| Sol | 93.8 | 1 | Sonnet | 87.5 | 2 |
| GPT-OSS-120B | 81.3 | 2 | Opus | 97.5 | 0 |
| GPT-OSS-20B | 82.5 | 1 | Gemini Flash | 100.0 | 0 |
| Terra | 90.0 | 2 | Flash-Lite | 92.5 | 1 |
| Luna | 82.5 | 2 | Qwen | 83.8 | 0 |
| Probe | ACT | ASK | REQUEST | Stable models | Majority label |
|---|---|---|---|---|---|
| Inbox cleanup 1 | 96 | 2 | 2 | 8 | ACT |
| Inbox cleanup 2 | 36 | 54 | 8 | 5 | ACT |
| Storage cleanup 1 | 54 | 46 | 0 | 6 | ASK |
| Storage cleanup 2 | 12 | 86 | 0 | 7 | ASK |
| Calendar management 1 | 94 | 4 | 0 | 8 | ASK |
| Calendar management 2 | 44 | 44 | 10 | 5 | ASK |
| Model | Offline ASK | Tool-Loop ASK | Difference |
|---|---|---|---|
| Sol | 19.2 | 15.8 | |
| GPT-OSS-120B | 15.8 | 4.6 | |
| GPT-OSS-20B | 10.0 | 0.4 | |
| Terra | 17.5 | 12.5 | |
| Luna | 24.6 | 14.6 | |
| Sonnet | 32.5 | 20.0 |
| Model | Factor | Mean | 95% CI | 0 | Sign-flip | ||
|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol | scope | -0.917 | [-1.000, -0.750] | 11 | 1 | 0 | 0.0010 |
| GPT-5.6 Sol | stakes | -0.417 | [-0.667, -0.183] | 7 | 5 | 0 | 0.0156 |
| GPT-5.6 Sol | reversibility | -0.467 | [-0.733, -0.200] | 7 | 5 | 0 | 0.0156 |
| GPT-5.6 Sol | external_visibility | +0.000 | [-0.150, +0.183] | 2 | 9 | 1 | 1.0000 |
| GPT-OSS-120B (Groq) | scope | -0.883 | [-1.000, -0.700] | 11 | 1 | 0 | 0.0010 |
| GPT-OSS-120B (Groq) | stakes | -0.600 | [-0.800, -0.400] | 11 | 1 | 0 | 0.0010 |
| Source | Scheme | ACT | ASK | REQUEST | REFUSE | |
|---|---|---|---|---|---|---|
| gpt-5.6-sol | emitted | 780 | 65.6 | 22.2 | 10.1 | 2.1 |
| gpt-5.6-sol | joint_consensus | 775 | 66.1 | 22.2 | 9.7 | 2.1 |
| gpt-5.6-sol | claude | 779 | 65.7 | 22.3 | 9.6 | 2.3 |
| gpt-5.6-sol | gemini | 780 | 65.6 | 22.1 | 10.0 | 2.3 |
| openai/gpt-oss-120b | emitted | 780 | 64.0 | 16.9 | 11.7 | 7.4 |
| openai/gpt-oss-120b | joint_consensus | 766 | 65.1 | 19.8 | 10.4 | 4.6 |
| Held out | Four-way [95% CI] | Binary [95% CI] | Reason [95% CI] | |||
|---|---|---|---|---|---|---|
| A | 101 | .782 [.703,.861] | 115 | .896 [.835,.948] | 48 | .729 [.604,.854] |
| B | 107 | .738 [.654,.822] | 125 | .824 [.752,.888] | 41 | .854 [.732,.951] |
| C | 99 | .798 [.717,.879] | 122 | .844 [.779,.910] | 37 | .946 [.865,1.000] |
| Model | Annotator A | Annotator B | Annotator C |
|---|---|---|---|
| Gemini 3.5 Flash-Lite | 65.4 | 57.3 | 78.8 |
| Gemini 3.6 Flash | 64.4 | 55.6 | 77.9 |
| Claude Sonnet 5 | 60.4 | 52.7 | 77.9 |
| Claude Opus 4.8 | 56.2 | 51.2 | 72.8 |
| GPT-5.6 Terra | 56.8 | 49.1 | 68.6 |
| GPT-5.6 Luna | 56.4 | 47.8 | 67.4 |
| Model | Multiclass Brier | Binary Brier | ACT bias |
|---|---|---|---|
| Gemini 3.5 Flash-Lite | .336 | .292 | |
| Gemini 3.6 Flash | .415 | .337 | |
| Qwen 3.6 27B | .455 | .405 | |
| Claude Sonnet 5 | .463 | .417 | |
| GPT-5.6 Terra | .545 | .463 | |
| Claude Opus 4.8 | .546 | .484 |
| Policy | 4-way | All pairs | Scope | Unchanged | Paraphrase |
|---|---|---|---|---|---|
| Ten-model range | 50.6–73.4% | 45.8–70.8% | 11–12/12 | 25.0–50.0% | 81.3–100% |
| Keyword rule | 69.1% | 9/48 | 3/12 | 39/48 | 15/16 |
| Annotator ACT vote share | – | 20/48 | 9/12 | 27/48 | 83.3% ∗ |
| Model | Majority agreement | Brier | Responsiveness | Under | Over | Elicit. shift | TL tool/strict | Approve act |
|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol | 58.1 | 0.599 | 56.3 | 54.8 | 19.4 | 8.3 | 64.8/32.4 | 76.3 (29/38) |
| GPT-OSS-120B | 54.5 | 0.599 | 64.6 | 55.3 | 23.5 | 47.5 | 66.9/36.6 | 72.7 (8/11) |
| GPT-OSS-20B | 50.6 | 0.648 | 56.3 | 58.8 | 25.8 | 52.5 | 77.1/37.5 | – † |
| GPT-5.6 Terra | 60.4 | 0.545 | 62.5 | 43.8 | 22.0 | 30.0 | 62.8/45.5 | 90.0 (27/30) |
| GPT-5.6 Luna | 60.0 | 0.547 | 64.6 | 41.5 | 22.3 | 8.3 | 60.0/34.5 | 77.1 (27/35) |
| Claude Sonnet 5 | 68.7 | 0.463 | 56.3 | 36.5 | 20.0 | 11.7 | 60.7/38.6 | 83.3 (40/48) |
| Stage | Question | Recommended measurement |
|---|---|---|
| Candidate judgment | Is this action authorized? | ACT/ASK/REQUEST/REFUSE against matched controls, not aggregate agreement alone |
| Elicitation | Does presentation change the operating point? | Counterbalanced/alternative formats; report the shift, not one number |
| Permission resolution | What does an ASK actually require? | Classify permission vs. information vs. mixed before scoring resolution |
| Constructed action | What tool/action does the agent propose? | Candidate-tool dispatch plus argument checks, not judgment alone |
| Semantic correspondence | Does the call implement the candidate? | Audited correspondence; strict matching is a conservative floor |
| State update | Did the authorized/denied outcome occur? | Deterministic state transition check when an environment exists |
| Model | Valid / scheduled | Offline ACT | Strict match | Cand.-tool | Any-tool | ||
|---|---|---|---|---|---|---|---|
| Sol | 145/145 | 72.4 | 32.4 | 64.8 | 64.8 | .467 | .671 |
| GPT-OSS-120B | 145/145 | 70.3 | 36.6 | 66.9 | 66.9 | .376 | .583 |
| GPT-OSS-20B | 144/145 | 68.8 | 37.5 | 77.1 | 80.6 | .490 | .610 |
| Terra | 145/145 | 69.7 | 45.5 | 62.8 | 62.8 | .396 | .440 |
| Luna | 145/145 | 64.8 | 34.5 | 60.0 | 60.0 | .488 | .595 |
| Sonnet | 145/145 | 64.1 | 38.6 | 60.7 | 60.7 | .387 | .437 |
| Model | Parent | ASK/eligible | Valid | Act | Cand. tool | Meas. eligible | Meas. tool | Strict | Ask again | Request | Refuse | Other | Tech. |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Sol | 240 | 38/38 | 38 | 29 | 29 | 23 | 19 | 11 | 1 | 4 | 0 | 4 | 0 |
| GPT-OSS-120B | 240 | 11/11 | 11 | 8 | 8 | 8 | 6 | 0 | 1 | 0 | 1 | 1 | 0 |
| Sonnet | 240 | 48/48 | 48 | 40 | 40 | 25 | 22 | 9 | 3 | 4 | 0 | 1 | 0 |
| Gemini Flash | 240 | 16/16 | 16 | 15 | 15 | 10 | 10 | 5 | 0 | 1 | 0 | 0 | 0 |
| Terra | 240 | 30/30 | 30 | 27 | 27 | 19 | 17 | 9 | 3 | 0 | 0 | 0 | 0 |
| Luna | 240 | 35/35 | 35 | 27 | 27 | 26 | 23 | 10 | 3 | 3 | 1 | 1 | 0 |
| Rule stratum | Audit | R1 equiv. | R1 non-equiv. | R1 uncertain | R2 equiv. | R2 non-equiv. |
|---|---|---|---|---|---|---|
| Strict ACCEPT | 100 | 100 | 0 | 0 | 100 | 0 |
| Strict REJECT | 100 | 88 | 9 | 3 | 92 | 8 |
| Probability-weighted quantity | Reviewer 1 | 95% interval | Reviewer 2 | 95% interval |
|---|---|---|---|---|
| Strict-ACCEPT precision | 100.0 | 96.3–100.0 | 100.0 | 96.3–100.0 |
| Strict-REJECT false rejection | 87.1 | 80.7–92.5 | 91.4 | 86.1–95.8 |
| Overall semantic correspondence | 94.6 | 91.9–96.9 | 96.4 | 94.1–98.2 |
| Tool-Loop correspondence candidate tool | 93.9 | 90.9–96.5 | 96.0 | 93.4–98.0 |
| Overall uncertain mass | 1.3 | 0.4–2.6 | 0.0 | 0.0–0.0 |
| Reviewer 1 | Reviewer 2 | ||||||
|---|---|---|---|---|---|---|---|
| Automated category | Eq. | Non-eq. | Unc. | Eq. | Non-eq. | Unc. | |
| Representation candidate | 79 | 78 | 0 | 1 | 78 | 1 | 0 |
| Changed value | 12 | 7 | 3 | 2 | 11 | 1 | 0 |
| Narrower or safer | 6 | 0 | 6 | 0 | 0 | 6 | 0 |
| Changed target | 3 | 3 | 0 | 0 | 3 | 0 | 0 |
| Family | Overall acc. | ACT acc. | ASK acc. | Balanced acc. | Cell consistency |
|---|---|---|---|---|---|
| GPT-5.6 | 89.2% | 78.3% | 100% | 89.2% | 97.9% |
| Claude | 87.9% | 75.8% | 100% | 87.9% | 97.9% |
| Gemini | 85.8% | 71.7% | 100% | 85.8% | 97.9% |
| GPT-OSS | 83.8% | 69.2% | 98.3% | 83.7% | 91.7% |
| Qwen | 86.2% | 72.3% | 100% | 86.1% | 97.9% |
| Pooled (1,199 valid) | 86.6% | 73.5% | 99.7% | 86.6% | – |
| Condition | N valid | Accuracy (lenient / strict) |
|---|---|---|
| Permission present | 250 | 90.4% / 89.6% |
| Permission absent | 250 | 79.6% / 64.8% |
| Information complete | 250 | 91.6% / 91.6% |
| Information missing | 250 | 78.4% / 62.8% |
| Condition | Family | Compliance rate | Fabricated | N valid/scheduled |
|---|---|---|---|---|
| APPROVAL | GPT-5.6 | 100% | 0 | 15/15 |
| Claude | 66.7% | 0 | 15/15 | |
| Gemini | 100% | 0 | 15/15 | |
| GPT-OSS | 40.0% | 0 | 15/15 | |
| Qwen | 100% | 0 | 13/15 | |
| MODIFICATION | GPT-5.6 | 100% | 0 | 15/15 |
| Model | Condition | Valid | Balanced acc. | Unauthorized | Completion | |
|---|---|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | Oracle decision | 192 | 192 | 100.0 | 0.0 | 100.0 |
| Gemini 3.5 Flash-Lite | Natural language | 192 | 189 | 97.3 | 0.0 | 96.4 |
| Gemini 3.5 Flash-Lite | Structured policy | 192 | 189 | 98.5 | 0.0 | 98.8 |
| GPT-5.6 Luna | Oracle decision | 192 | 192 | 100.0 | 0.0 | 100.0 |
| GPT-5.6 Luna | Natural language | 192 | 192 | 100.0 | 0.0 | 100.0 |
| GPT-5.6 Luna | Structured policy | 192 | 192 | 100.0 | 0.0 | 100.0 |
| Model | Condition | Valid | Balanced acc. | Unauthorized | Completion | |
|---|---|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | Eligibility rule | 192 | 192 | 100.0 | – | – |
| GPT-5.6 Luna | Eligibility rule | 192 | 192 | 97.7 | – | – |
| Claude Sonnet 5 | Eligibility rule | 192 | 192 | 98.1 | – | – |
| Model | Condition | Valid | Balanced acc. | Unauthorized | Completion | |
|---|---|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | Hidden / QUESTION | 144 | 141 | 97.9 | 0.0 | 100.0 |
| Gemini 3.5 Flash-Lite | Hidden / ask_user | 144 | 143 | 99.3 | 0.0 | 100.0 |
| Gemini 3.5 Flash-Lite | Shown / QUESTION | 144 | 144 | 100.0 | 0.0 | 100.0 |
| Gemini 3.5 Flash-Lite | Shown / ask_user | 144 | 140 | 97.2 | 0.0 | 97.2 |
| GPT-5.6 Luna | Hidden / QUESTION | 144 | 144 | 100.0 | 0.0 | 100.0 |
| GPT-5.6 Luna | Hidden / ask_user | 144 | 144 | 100.0 | 0.0 | 100.0 |
| Model | Condition | Valid | Balanced acc. | Unauthorized | Completion | |
|---|---|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | Natural language | 68 | 68 | 100.0 | 0.0 | 100.0 |
| Gemini 3.5 Flash-Lite | Structured policy | 68 | 68 | 100.0 | 0.0 | 100.0 |
| GPT-5.6 Luna | Natural language | 68 | 68 | 100.0 | 0.0 | 100.0 |
| GPT-5.6 Luna | Structured policy | 68 | 68 | 100.0 | 0.0 | 100.0 |
| Claude Sonnet 5 | Natural language | 68 | 68 | 100.0 | 0.0 | 100.0 |
| Claude Sonnet 5 | Structured policy | 68 | 68 | 100.0 | 0.0 | 100.0 |
| Model | Condition | Valid | Balanced acc. | Unauthorized | Completion | |
|---|---|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | Oracle decision | 32 | 32 | 100.0 | 0.0 | 100.0 |
| Gemini 3.5 Flash-Lite | Natural language | 32 | 32 | 100.0 | 0.0 | 100.0 |
| Gemini 3.5 Flash-Lite | Structured policy | 32 | 31 | 97.2 | 0.0 | 100.0 |
| GPT-5.6 Luna | Oracle decision | 32 | 32 | 100.0 | 0.0 | 100.0 |
| GPT-5.6 Luna | Natural language | 32 | 32 | 100.0 | 0.0 | 100.0 |
| GPT-5.6 Luna | Structured policy | 32 | 32 | 100.0 | 0.0 | 100.0 |
| Model | Study / contrast | Cluster | Change (pp) | 95% interval |
|---|---|---|---|---|
| GPT-5.6 Luna | Structured minus natural | Family | 0.0 | [0.0, 0.0] |
| GPT-5.6 Luna | Structured minus natural | Domain | 0.0 | [0.0, 0.0] |
| GPT-5.6 Luna | Candidate visibility | Family | 0.0 | [0.0, 0.0] |
| GPT-5.6 Luna | Question tool | Family | 0.0 | [0.0, 0.0] |
| GPT-5.6 Luna | Interaction | Family | 0.0 | [0.0, 0.0] |
| GPT-5.6 Luna | Structured minus natural | Task | 0.0 | [0.0, 0.0] |
| Model | Fact wording | Valid | Balanced acc. | Completion | |
|---|---|---|---|---|---|
| GPT-5.6 Luna | Natural | 192 | 192 | 99.4 | 98.8 |
| GPT-5.6 Luna | Paraphrased | 192 | 192 | 98.8 | 97.6 |
| Claude Sonnet 5 | Natural | 192 | 192 | 97.0 | 94.0 |
| Claude Sonnet 5 | Paraphrased | 192 | 192 | 97.0 | 94.0 |
| Gemini 3.5 Flash-Lite | Natural | 192 | 191 | 98.9 | 98.8 |
| Gemini 3.5 Flash-Lite | Paraphrased | 192 | 192 | 98.8 | 97.6 |
| Model | Cluster | Accuracy change (pp) | 95% interval |
|---|---|---|---|
| GPT-5.6 Luna | Family | -0.3 | [-1.0, 0.0] |
| GPT-5.6 Luna | Domain | -0.6 | [-1.8, 0.0] |
| Claude Sonnet 5 | Family | 0.0 | [-1.6, 1.6] |
| Claude Sonnet 5 | Domain | 0.0 | [-1.8, 1.8] |
| Gemini 3.5 Flash-Lite | Family | -0.2 | [-1.6, 1.0] |
| Gemini 3.5 Flash-Lite | Domain | -0.1 | [-1.8, 1.4] |