Same Pieces, Different Servers: A Tetris Benchmark for AI Agents as Served
Organizations: Massachusetts Institute of Technology
Abstract
An agent meets a model as served: through an endpoint with a price card, a shared cache and other tenants, or on whatever hardware a self-hosted model runs. Benchmarks rank the weights. We introduce a Tetris benchmark that measures what agents get from models as served: every move is scored against an oracle, and every agent receives the same pieces. In five pre-specified experiments with nine open-weight models on one serverless provider, plus an open decision model self-hosted on a CPU, we find that price and size do not predict decision quality; that resending history costs almost nothing when cached input is free, would cost eleven to twelve times more if it were not, and makes play worse; that deployments keep between 4 and more than 64 agent contexts warm, in line with their throughput rather than the model's KV-cache size; that an agent's own long requests slow its slowest short decisions more than tenfold, which a simple admission rule cuts by 59% without hurting play; and that the decision model plays mid-pack but takes 27 s per move on a CPU, about 650 times longer than reported on a GPU. Architecture predicts some of what an agent sees; the deployment sets the rest.
Figures & tables
| Layer | What the agent varies (experiment) | What it observes | What stays hidden |
|---|---|---|---|
| Task: Tetris | nothing: every agent gets the same seeded pieces | board, current and next piece, legal placements; oracle regret per move | nothing (fully observed) |
| Agent harness | model (E1), history resent (E2), concurrent agents (E3), request mix and admission (E4) | its own queue and timing | nothing |
| Serverless API | output limits, JSON schema, priority hints (E4) | latency; reported prompt, cached and output tokens; refusals (HTTP 429); the public price card | billing records, routing, replica count, other tenants |
| Inference engine | nothing | only indirectly: cache hits, prefill throughput | batching, scheduler, KV memory pool, eviction and offloading |
| Weights | which served model id (E1–E4) | model card and config: parameters, attention and KV design | the exact checkpoint and quantization behind an id |
| Decision model | hardware (E5) | the probability of every option; exact compute time | nothing (self-hosted) |
| Model | Total / active | Attention and KV design | KV/ctx | $/M in, out | Source |
|---|---|---|---|---|---|
| gpt-oss-20b | 21B / 3.6B | alternating full and 128-token window | 132 MB | 0.07, 0.28 | OpenAI (2025) |
| gpt-oss-120b | 117B / 5.1B | same pattern, 36 layers | 198 MB | 0.15, 0.60 | OpenAI (2025) |
| gemma-4-31B | 31B dense | 5 sliding (1,024) : 1 global layer | 1,015 MB | 0.14, 0.56 | Gemma Team (2026) |
| DeepSeek-V4-Flash | 284B / 13B | compressed sparse attention; FP8 cache | 21 MB | 0.14, 0.28 | DeepSeek-AI (2026) |
| MiniMax-M2.5 | 229B / 10B a | full attention in all 62 layers | 1,335 MB | 0.30, 1.20 | MiniMax (2025) |
| Qwen3.5-397B | 397B / 17B | 3 Gated DeltaNet : 1 full, 60 layers | 342 MB | 0.60, 3.60 | Qwen Team (2026) |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Experiment | Design | Main measures |
|---|---|---|
| E1 model ladder | 9 configurations 10 seeds, up to 100 pieces; no history; output constrained to legal placements | score relative to the bot, oracle regret, price and latency per decision |
| E2 history policy | 4 models 3 history policies (Qwen3.8: 2) 9 seeds | uncached tokens and cost per move, oracle regret |
| E3 cache capacity | 9 models concurrent agents 1–5 repetitions | share of contexts still cached after 30 s idle, refusals |
| E4 interference | 2 models 10 time blocks 4 client policies; held-out request corpus | 95th-percentile time to a valid move, useful decisions per second |
| E5 decision model | Intelif on a 4-core CPU, 3 games (seeds 1000–1002) | score relative to the bot, oracle regret, latency |