Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
Organizations: Stanford University · Together AI
Abstract
Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Demand growth strains this paradigm faster than providers can scale. Two advances create an opportunity to rethink it: small, local LMs (<=20B active parameters) now achieve competitive performance to frontier models on many tasks, and local accelerators (e.g., Apple M4 Max) can host these models at interactive latencies. This raises the question: can local inference viably redistribute demand from centralized infrastructure? This requires measuring both whether local LMs can accurately answer real-world queries and whether they can do so efficiently on power-constrained devices (e.g., laptops). We propose intelligence per watt (IPW), task accuracy per unit of power, as a unified metric for the capability and efficiency of local inference across model-accelerator configurations. We evaluate 20+ state-of-the-art local LMs, 8 hardware accelerators (local and cloud), and 1M real-world single-turn chat and reasoning queries. For each query, we measure accuracy (local LM win rate against frontier models), energy, latency, and power. We find three key results. First, local LMs successfully answer 88.7% of these queries, with accuracy varying by domain. Second, longitudinal analysis from 2023-2025 shows IPW improved 5.3x, driven by both algorithmic and accelerator advances, with locally-serviceable query coverage rising from 23.2% to 71.3%. Third, local accelerators achieve at least 1.4x lower IPW than cloud accelerators running identical models, revealing significant headroom for local accelerator optimization. These findings demonstrate that local inference can meaningfully redistribute demand from centralized infrastructure for a substantial subset of queries, with IPW serving as the critical metric for tracking this transition.
Figures & tables
| Dataset Origin | Category | |
| Wildchat | Chat | 500K |
| NaturalReasoning | Reasoning | 500K |
| MMLU Pro | Knowledge | 12K |
| SuperGPQA | Grad. Reasoning | 26.5K |
| 2023 | 2024 | 2025 | |
| SOTA Local Model | Mixtral-8x7B-v0.1 | Llama-3.1-8B-Instruct | GPT-OSS-120B |
| SOTA Accelerator | NVIDIA Quadro RTX 6000 | NVIDIA RTX 6000 Ada | Apple M4 Max |
| Success Rate | |||
| Intelligence per Watt | |||
| YoY Efficiency Gain | — | 2.27× | 2.32× |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Life, physical, and social science | Computer and mathematical |
| Architecture and engineering | Education instruction and library |
| Installation, maintenance, and repair | Business and financial operations |
| Legal services | Transportation and material moving |
| Arts, design, sports, entertainment, and media | Production services |
| Farming, fishing, and forestry | Healthcare support |
| Food preparation and serving related | Healthcare practitioners and technical |
| Domain | WC Count | WC % | WC Solv % | NR Count | NR % | NR Solv % |
| Computer and mathematical | 90,662 | 18.1 | 99.5 | 174,242 | 34.8 | 67.3 |
| Arts, design, sports, entertainment, and media | 235,658 | 47.1 | 98.7 | 2,648 | 0.5 | 52.9 |
| Life, physical, and social science | 28,079 | 5.6 | 98.8 | 180,065 | 36.0 | 60.5 |
| None | 49,014 | 9.8 | 97.3 | 79,752 | 16.0 | 65.6 |
| Education instruction and library | 23,196 | 4.6 | 97.2 | 13,864 | 2.8 | 80.4 |
| Architecture and engineering | 5,782 | 1.2 | 98.9 | 28,762 | 5.8 | 40.8 |
| Category | Wildchat | MMLU Pro | SuperGPQA | Average |
| Computer and mathematical | 93.4% | 90.6% | 72.8% | 85.6% |
| Life, physical, and social science | 91.1% | 84.7% | 50.4% | 75.4% |
| Sales and related | 86.8% | 74.2% | 64.3% | 75.1% |
| Business and financial operations | 89.3% | 82.9% | 52.5% | 74.9% |
| Production services | 89.7% | 85.7% | 48.8% | 74.7% |
| Office and administrative support | 88.5% | 83.3% | 44.8% | 72.2% |
| Metric | Description |
| flops_per_request | FLOPs per query. |
| macs_per_request | MACs per query; proxy for compute. |
| per_query_joules | Energy per query (J). |
| total_joules | Total energy across queries. |
| per_token_ms | Latency per token (ms). |
| throughput_tokens_per_sec | Token output rate (toks/s). |
| Hardware | Memory | Bandwidth | Power |
| NVIDIA A100 (Ampere) | 40 GB HBM2 | 1,555 GB/s | 400 W TDP |
| NVIDIA H200 (Hopper) | 141 GB HBM3e | 4.8 TB/s | Up to 700 W TDP |
| NVIDIA B200 (Blackwell) | 192 GB HBM3e | 8 TB/s | 1000 W TDP |
| NVIDIA GH200 (Grace Hopper) | 144 GB HBM3e (+624 GB LPDDR5X) | 4.8 TB/s (GPU) | 1000 W TDP |
| NVIDIA Quadro RTX 6000 (Turing) | 24 GB GDDR6 | 672 GB/s | 295 W TDP |
| NVIDIA RTX 6000 Ada Generation | 48 GB GDDR6 | 960 GB/s | 300 W TDP |
| Cost Savings | Compute Savings | Energy Savings | ||||
| Size Threshold ( ) | Qwen + GPT-OSS | Qwen | Qwen + GPT-OSS | Qwen | Qwen + GPT-OSS | Qwen |
| 4B | 65.2% | 65.2% | 65.1% | 65.1% | 63.5% | 63.5% |
| 8B | 80.8% | 80.8% | 83.1% | 83.1% | 79.6% | 79.6% |
| 14B | 89.0% | 89.0% | 93.0% | 93.0% | 87.0% | 87.0% |
| 20B | 90.5% | — | 97.4% | — | 89.4% | — |
| 32B | 91.3% | 91.9% | 97.4% | 92.8% | 90.4% | 90.5% |
| Cost Savings | Compute Savings | Energy Savings | ||||
| Size Threshold ( ) | Qwen + GPT-OSS | Qwen | Qwen + GPT-OSS | Qwen | Qwen + GPT-OSS | Qwen |
| 4B | 52.9% | 52.9% | 54.5% | 54.5% | 46.3% | 46.3% |
| 8B | 60.5% | 60.5% | 62.5% | 62.5% | 54.0% | 54.0% |
| 14B | 68.7% | 68.7% | 70.1% | 70.1% | 62.5% | 62.5% |
| 20B | 73.3% | — | 72.2% | — | 67.8% | — |
| 32B | 76.9% | 75.9% | 75.1% | 75.1% | 72.4% | 71.6% |
| Model | Input Cost (USD / 1M tokens) | Output Cost (USD / 1M tokens) |
| Qwen3-4B | 0.000 | 0.000 |
| Qwen3-8B | 0.035 | 0.138 |
| Qwen3-14B | 0.060 | 0.124 |
| Qwen3-32B | 0.100 | 0.450 |
| Qwen3-235B | 0.220 | 0.880 |
| GPT-OSS-20B | 0.03 | 0.14 |
| Model | Type | WildChat | NaturalReasoning | MMLU Pro | SuperGPQA | Average |
| gpt-5-2025-08-07 | Closed | 81.9% | 82.9% | 86.5% | 64.4% | 78.9% |
| gemini-2.5-pro | Closed | 89.5% | 77.9% | 87.4% | 66.5% | 80.3% |
| claude-sonnet-4-5 | Closed | 88.1% | 76.9% | 86.4% | 60.1% | 77.9% |
| Qwen3-235B-A22B (Best OSS) | Open | N/A ∗ | 70.0% | 82.3% | 63.1% | 71.8% |
| Qwen3-32B | Open | 76.1% | 69.7% | 77.9% | 56.5% | 70.1% |
| gpt-oss-120b | Open | 89.2% | 65.0% | 78.3% | 55.3% | 72.0% |
| Metric | WildChat | NaturalReasoning | MMLU Pro | SuperGPQA |
| Closed Best | 89.5% | 82.9% | 87.4% | 66.5% |
| Best Closed Model | gemini-2.5-pro | gpt-5 | gemini-2.5-pro | gemini-2.5-pro |
| Open Best | 89.2% * | 70.0% | 82.3% | 63.1% |
| Best Open Model | gpt-oss-120b | Qwen3-235B-A22B | Qwen3-235B-A22B | Qwen3-235B-A22B |
| Gap | ||||
| Local Best ( Active) | 89.2% | 67.3% | 80.3% | 50.5% |
| Qwen3-4B | Qwen3-8B | Qwen3-14B | Qwen3-32B | |
| Success Rate | ||||
| Apple M4 Max | ||||
| Intelligence per Watt | ||||
| NVIDIA B200 | ||||
| Intelligence per Watt | ||||
| SambaNova SN40L | ||||
| Qwen3-8B | Qwen3-32B | GPT-OSS-20B | GPT-OSS-120B | |
| Apple M4 Max | ||||
| Intelligence per Joule | ||||
| NVIDIA B200 | ||||
| Intelligence per Joule | ||||
| SambaNova SN40L | ||||
| Intelligence per Joule | — | — | ||
| Model | bs=1 IPJ | bs=64 IPJ | IPJ Gain | Architecture |
| ( ) | ( ) | |||
| Qwen3-8B | 8.92 | 104.93 | Dense | |
| Qwen3-14B | 7.22 | 80.61 | Dense | |
| GPT-OSS-120B | 6.72 | 132.21 | MoE ( B active) |
| Model | vLLM | SGLang | llama.cpp |
| IPW ( ) | IPW ( ) | IPW ( ) | |
| Qwen3-4B | 1.40 | 1.35 | 1.52 |
| Qwen3-8B | 1.63 | 1.58 | 1.71 |
| Qwen3-14B | 1.69 | 1.62 | 1.78 |
| GPT-OSS-120B | 4.18 | 4.05 | 4.31 |
| Model | Precision | Hardware | Acc. | Power | Lat. | IPW | IPJ |
| (%) | (W) | (s/q) | ( ) | ( ) | |||
| Smartphone-class accelerator (Apple A18 Pro, iPhone 16 Pro) | |||||||
| Qwen3-4B | FP16 | A18 Pro | 42.5 2.0 | 12.0 0.4 | 92.5 9.4 | 11.8 0.7 | 38.3 3.9 |
| Qwen3-4B | FP8 | A18 Pro | 40.5 1.8 | 11.0 0.3 | 55.2 5.6 | 12.4 0.7 | 66.7 6.8 |
| Qwen3-4B | FP4 | A18 Pro | 38.0 1.6 | 9.5 0.3 | 36.8 3.7 | 13.3 0.8 | 108.7 11.0 |
| Gemma3-4B | FP4 | A18 Pro | 32.0 1.5 | 8.8 0.3 | 33.5 3.5 | 11.6 0.7 | 108.5 11.2 |
| Benchmark | Model | Hardware | Acc. | Power | Lat. | Energy | IPW | IPJ |
| (%) | (W) | (s/q) | (kJ/q) | ( ) | ( ) | |||
| GAIA (165 multi-turn general-assistant queries with tool use) | ||||||||
| GAIA | MiniMax-M2.5 | 8 H100 | 16.4 2.9 | 1558 95 | 4.69 0.41 | 7.31 0.62 | 0.105 0.020 | 2.24 0.39 |
| GAIA | Qwen3-235B | 8 H100 | 5.5 1.8 | 1595 88 | 0.74 0.09 | 1.18 0.13 | 0.034 0.012 | 4.66 1.55 |
| GAIA | Qwen3-30B | 8 H100 | 8.6 2.2 | 822 64 | 1.22 0.14 | 1.00 0.11 | 0.105 0.027 | 8.60 2.20 |
| GAIA | MiniMax-M2.5 | M4 Max | 14.2 2.7 | 305 24 | 51.6 5.4 | 15.74 1.69 | 0.466 0.092 | 0.90 0.18 |