cs.LGMar 2, 2026

Same Pieces, Different Servers: A Tetris Benchmark for AI Agents as Served

Authors: Haochuan Wang

Organizations: Massachusetts Institute of Technology

Abstract

An agent meets a model as served: through an endpoint with a price card, a shared cache and other tenants, or on whatever hardware a self-hosted model runs. Benchmarks rank the weights. We introduce a Tetris benchmark that measures what agents get from models as served: every move is scored against an oracle, and every agent receives the same pieces. In five pre-specified experiments with nine open-weight models on one serverless provider, plus an open decision model self-hosted on a CPU, we find that price and size do not predict decision quality; that resending history costs almost nothing when cached input is free, would cost eleven to twelve times more if it were not, and makes play worse; that deployments keep between 4 and more than 64 agent contexts warm, in line with their throughput rather than the model's KV-cache size; that an agent's own long requests slow its slowest short decisions more than tenfold, which a simple admission rule cuts by 59% without hurting play; and that the decision model plays mid-pack but takes 27 s per move on a CPU, about 650 times longer than reported on a GPU. Architecture predicts some of what an agent sees; the deployment sets the rest.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Switchcraft: AI Model Router for Agentic Tool Calling

    May 8, 2026Sharad Agarwal, Pooria Namyar, Alec Wolman +3Agentic InferenceLarge Language Model Routing

  2. RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions

    Jun 2, 2026Zongwei Lv, Zhewen Tan, Yaoming Li +7Agentic BenchmarksOpenclaw

  3. Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads

    Oct 5, 2026Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona ZahidAgentic Inference