cs.LGSep 23, 2026

Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools

Authors: Shunya Nagashima

Abstract

Time-series foundation models (TSFMs) provide forecasts for operational decisions, but accuracy alone does not determine their value. Evaluating agents that use these models requires measuring decision quality and forecast cost. FWBench evaluates this capability on 1,251 electricity and cycle-hire cases using fixed forecast tools and simulated capacity contracts. Agents select models, histories and horizons, then submit capacities to minimize a stated loss-cost objective. We evaluated two hosted and eight local configurations, including small language models, and tested local models with and without TSFMs. GPT-6 Astra bought inexpensive short-horizon forecasts selectively, using 2.5% of the budget, and outperformed fixed policies when the saved decisions were scored with three loss-cost weightings. FWBench enables reproducible evaluation of how language models select and use time-series forecasts to make decisions under cost constraints.

Explore similar work

CardsList
  1. t0t_0: A Time-Series Foundation Model for Forecasting with Context

    Sep 21, 2026Lucas Meyer, Claudio Sole, Huikan Xiang +6