From Token-Max to Outcome-Max: How You Use AI Determines Its Productivity
Organizations: Carnegie Mellon University · Amazon
Abstract
Generative artificial intelligence (AI) models can perform increasingly complex tasks, yet greater AI usage does not necessarily translate into proportional productivity gains. We identify token-max as one source of this inefficiency: when token consumption is treated as productive effort, agents are encouraged to over-exert and expend computation beyond what is necessary. We instead propose outcome-max, which rewards independently verified task completion per unit cost and induces a principled stopping rule. Then, to study these objectives, we develop a three-level simulation framework spanning immediate interaction, long-run behavioral adaptation, and organizational collaboration. Across all three levels, outcome-max improves the efficiency of AI-assisted production while largely preserving verified task performance. To further align these incentives with outcome-max, we introduce OutcomeShare, an incentive mechanism. Theory and simulation show that OutcomeShare can induce participation while generating shared gains for employees, firms, and LLM providers. Together, our results suggest that AI productivity not only depends on model capability, but also on how to construct the objectives governing AI use.
Figures & tables
| Success (%) | Tokens (k) | Efficiency gain (%) | |||||
| Setting | Objective | Overall | Easy | Medium | Hard | ||
| L1: single-round human–AI interaction | |||||||
| SWE-bench | Token-max | 48.8 | 62.2 | 43.9 | 39.8 | 21.1 | 0.0 |
| Outcome-max | 48.2 | 66.7 | 48.9 | 25.0 | 9.4 | +120.7 | |
| HumanEval+ | Token-max | 87.2 | 94.3 | 89.3 | 73.8 | 5.4 | 0.0 |
| Outcome-max | 87.5 | 95.9 | 91.4 | 69.6 | 4.7 | +15.3 | |
| Objective | Employee behavior after the first patch | Verified result; total tokens |
| Case A: both agents’ first patches already pass the task tests. | ||
| Token-max | Requests extra tests, import/changelog checks, then edge-case and documentation work. | Four passing patches; 33,602. |
| Outcome-max | Accepts the first patch and ends the interaction. | One passing patch; 7,696. |
| Why wasteful: repeated checks and edits increase tokens without improving verified task correctness. | ||
| Case B: both agents’ first patches fail the task tests. | ||
| Token-max | Questions backward compatibility of existing dim= calls, prompting a repair. | The next patch passes; 30,901. |
| Deliveries / week | Weekly accounts (USD) | ||||
| Production / contract | Correct | Wrong | Provider | Firm | Employee |
| Token-max / token billing | 9.8 | 15.6 | 2 | 5,429 | 2,400 |
| Outcome-max / token billing | 38.3 | 61.3 | 3 | 14,746 | 2,400 |
| Outcome-max / OutcomeShare | 19.4 | 2.9 | 535 | 275 | 3,443 |
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
| Dimension | Low propensity | High propensity |
| Iterates and refines | Accepts initial output | Requests revisions |
| Clarifies goals | Leaves goals implicit | States goals and constraints |
| Provides examples | Gives no examples | Supplies examples |
| Specifies format | Leaves format open | Specifies output structure |
| Sets interaction mode | Uses no explicit role | Assigns a collaboration role |
| Communicates tone | Leaves style open | States style preferences |
| Parameter | Scenario value |
| Annual backlog | 2,400 tasks |
| Task-value tiers | 40, 640 |
| Loaded wage | $80 per hour |
| Employee-time anchor | 15 minutes per attempt |
| Input, output, cache prices | 1.10, $0.07 per million tokens |
| Employees and periods | 12 employees; five selection periods |
| Setting | Token (%) | Outcome (%) | Gap [95% CI] (pp) |
| (a) Primary SWE benchmark | |||
| Overall | 51.3 | 50.3 | -1.0 [-8.2, +5.2] |
| Easy | 62.0 | 67.6 | +5.6 [+0.0, +11.4] |
| Medium | 46.8 | 50.0 | +3.2 [+0.0, +7.4] |
| Hard | 43.6 | 26.9 | -16.7 [-24.5, +0.0] |
| (b) Model and persona robustness | |||
| Setting | Objective | Success (%) | Tokens (k) | |
| Haiku 4.5 | Token-max | 51.9 | 15.5 | 33.57 |
| Outcome-max | 61.6 | 8.3 | 73.86 | |
| DeepSeek v4 Pro | Token-max | 58.1 | 281.3 | 2.06 |
| Outcome-max | 61.0 | 78.5 | 7.78 | |
| Detailed guidance | Token-max | 45.8 | 30.2 | 15.16 |
| Outcome-max | 50.0 | 7.4 | 67.33 |
| Different backbones and human persona | |||||||
| SWE-bench; human: Opus 4.6 | |||||||
| Agent | Human behavior | Overall | Easy | Medium | Hard | Tokens (k) | Eff. gain (%) |
| Haiku 4.5 | Minimal input | 26.7 | 20.0 | 40.0 | 20.0 | 9.8 | 0.0 |
| Detailed guidance | 26.7 | 40.0 | 20.0 | 20.0 | 11.6 | -15.5 | |
| Active checking | 46.2 | 40.0 | 66.7 | 40.0 | 23.1 | -26.1 | |
| Guidance + checking | 28.6 | 40.0 | 25.0 | 20.0 | 25.0 | -57.8 | |
| Different backbones and human persona | |||||||
| SWE-bench; human: Haiku 4.5 | |||||||
| Agent | Human behavior | Overall | Easy | Medium | Hard | Tokens (k) | Eff. gain (%) |
| Haiku 4.5 | Minimal input | 40.0 | 60.0 | 40.0 | 20.0 | 8.6 | 0.0 |
| Detailed guidance | 54.5 | 60.0 | 66.7 | 33.3 | 43.2 | -72.8 | |
| Active checking | 25.0 | 40.0 | 33.3 | 0.0 | 38.4 | -86.0 | |
| Guidance + checking | 50.0 | 60.0 | 66.7 | 25.0 | 49.9 | -78.4 | |
| Policy | Objective | Success (%) | Tokens (k) | Eff. gain (%) |
| Haiku 4.5 agent; Sonnet 4.6 employee | ||||
| Haiku 4.5 | Token-max | 51.6 | 15.5 | 0.0 |
| Outcome-max | 60.5 | 8.4 | +117.4 | |
| DeepSeek v4 Pro agent; Haiku 4.5 employee | ||||
| DeepSeek v4 Pro | Token-max | 55.8 | 262.7 | 0.0 |
| Outcome-max | 63.0 | 91.5 | +224.2 | |
| Persona | Objective | Success (%) | Tokens (k) |
| Guidance and checking | Token-max | 50.0 | 34.8 |
| Outcome-max | 54.2 | 11.3 | |
| More guidance | Token-max | 42.3 | 29.7 |
| Outcome-max | 50.0 | 7.4 | |
| More checking | Token-max | 54.5 | 17.7 |
| Outcome-max | 52.2 | 11.4 |
| Token-max | Outcome-max | |||
| Ending | Runs | Solved | Runs | Solved |
| Verified 50-turn cap | 18 | 5 | 46 | 6 |
| Employee acceptance | 0 | 0 | 32 | 14 |
| Confidence threshold | 2 | 2 | 4 | 1 |
| Other agent ending | 73 | 30 | 10 | 2 |
| Success (%) | Tokens (k) | Efficiency improvement (%) | |||||
| Adaptation | Objective | Overall | Easy | Medium | Hard | ||
| Employee adapts | Token-max | 58.1 | 73.4 | 57.3 | 25.0 | 281.3 | 0.0 |
| Outcome-max | 61.0 | 80.6 | 56.7 | 28.3 | 78.5 | +276.5 | |
| Agent adapts | Token-max | 62.0 | 78.1 | 59.7 | 30.4 | 207.3 | 0.0 |
| Outcome-max | 61.2 | 79.6 | 58.0 | 26.8 | 94.6 | +116.3 | |
| Both adapt | Token-max | 56.2 | 72.5 | 55.0 | 20.4 | 261.7 | 0.0 |
| (a) Adapting role | ||||
| Role | Objective | Success (%) | Tokens (k) | Efficiency (solutions per M tokens) |
| Neutral (zero profile) | — | 57.0 | 54.2 | 10.51 |
| Employee | Token-max | 57.6 | 281.0 | 2.05 |
| Outcome-max | 60.9 | 78.9 | 7.72 | |
| Agent | Token-max | 62.4 | 206.8 | 3.02 |
| Outcome-max | 61.0 | 96.0 | 6.36 | |
| Objective | Success (%) | Tokens (k) | Efficiency | Continued (%) | Later tokens (%) |
| Token-max | 51.3 | 21.5 | 23.8 | 19.7 | 5.6 |
| Outcome-max | 50.3 | 9.4 | 53.6 | 3.5 | 1.3 |
| Initial patch | Objective | Employee decision and result | Patches | Tokens |
| Correct | Token-max | Requests tests and further checks; all saved patches pass. | P, P, P, P | 33,602 |
| Outcome-max | Accepts the first correct fix. | P | 7,696 | |
| Incorrect | Token-max | Questions backward compatibility; the next patch passes. | F, P, P, P | 30,901 |
| Outcome-max | Accepts a patch that breaks existing dim= calls. | F | 4,456 |
| Aspect | Evidence or scope |
| Efficient stopping | High-guidance/low-checking profile, repetition 0. Both first patches pass; outcome-max stops, while token-max continues without improving correctness. |
| Premature stopping | Revision-only profile, repetition 0. Token-max asks about backward compatibility and obtains a passing patch; outcome-max accepts a patch failing both task-specific integration tests. |
| Selection | The efficient example is the first task–profile repetition within this task with complete replays, two correct first patches, and earlier, cheaper outcome-max acceptance. The failure example is a selected counterexample. |
| Verification | All final patches match the tested artifacts, with complete logs and no identified disk-space installation failure. Shared background failures are distinguished from task-specific resolution. |
| Interpretation | Approval is not verification. The examples show possible stopping mechanisms, not their frequency or whether extra interaction would repair the same outcome-max trajectory. |
| Decision | Token-max | Outcome-max |
| Scope | Search related methods and compare alternative fixes. | Request a minimal diff with the relevant context upfront. |
| Verification | Enumerate edge cases after every change; challenge confident answers. | Check a specific expected outcome; target the exact failure on retry. |
| Stopping | Revisit assumptions or request explanations after a solution is accepted. | Accept a plausible complete fix; cap failed attempts. |
| Metric | Token-max | Outcome-max |
| Initial-period net profit (USD) | -6,917 | +2,141 |
| Final-period net profit (USD) | -15,778 | +3,826 |
| KPI–efficiency rank correlation | -0.46 | +0.44 |
| Component | Result and interpretation |
| Production | Lower token use per resolution supports more attempted work under the capacity assumption. |
| Provider | Outcome fees raise revenue by sharing verified production value; revenue is before compute costs. |
| Employee | A bonus on verified output raises modeled income. |
| Firm | The residual value yields a positive profit estimate relative to the token-regime loss; its uncertainty interval still includes losses. |
| Scope | These accounts illustrate feasible value sharing, not a field deployment of the full payment-and-liability contract. |
| Production / payment | Verified successes/wk | Provider revenue ($) | Firm profit ($) | Employee income index |
| Token / token fees | ||||
| Outcome / token fees | ||||
| Outcome / outcome fee | ||||
| Outcome / equal surplus | ||||
| Outcome / provider–firm | ||||
| Outcome / employee–firm |
| Production / payment | Verified successes/wk | Provider revenue | Firm profit | Employee income index |
| Token / token fees | ||||
| Outcome / token fees | ||||
| Outcome / outcome fee | ||||
| Outcome / equal surplus | ||||
| Outcome / provider–firm | ||||
| Outcome / employee–firm |
| Deliveries / week | Weekly accounts (USD) | ||||
| Policy | Correct | Wrong | Provider | Firm net profit | Employee |
| Production baselines: submit every output | |||||
| Token-max / token billing | 9.8 | 15.6 | 2 | 5,429 | 2,400 |
| Outcome-max / token billing | 38.3 | 61.3 | 3 | 14,746 | 2,400 |
| Alternative submission rules on outcome-max outputs | |||||
| Random screen | 3.2 | 5.2 | 179 | 1,774 | 549 |
| Task value / reference hour | Provider | Firm net profit | Firm 95% interval | Employee |
| $80 | 178 | -1,508 | 2,748 | |
| $160 | 357 | -617 | 3,095 | |
| $240 (central) | 535 | 275 | 3,443 | |
| $320 | 713 | 1,166 | 3,790 |
| Fine level | Success (%) | Tokens (k) | Efficiency | Acceptance (%) | Bad delivery (%) |
| Matched tasks (same cohort as the main figure) | |||||
| 0 | 40.0 | 54.9 | 7.3 | 90.0 | 50.0 |
| 0.25 | 40.0 | 58.1 | 6.9 | 90.0 | 50.0 |
| 0.5 | 30.0 | 60.6 | 5.0 | 90.0 | 60.0 |
| 1 | 40.0 | 72.5 | 5.5 | 70.0 | 40.0 |
| 2 | 40.0 | 61.6 | 6.5 | 70.0 | 30.0 |