Terminal agents are commonly trained by imitating long teacher trajectories, yet how much of each trajectory to supervise remains unexplored. We study the \emph{supervision horizon}, the number of trajectory tokens retained for training, and show that it is a key design axis for reliability and cost. Reliability improves with longer horizons but saturates: on Terminal-Bench, a 12K-token horizon solves more tasks than 16K (29±0.7 vs.\ 26±0.8) while requiring 30% less training time. The horizon also shapes agent behavior: short horizons cause premature termination, intermediate horizons yield productive error recovery, and long horizons induce over-persistence. We analyze this saturation through a bias--complexity bound, in which longer supervision reduces temporal supervision bias but increases finite-sample estimation error from more heterogeneous late-stage histories. Guided by this analysis, we propose \emph{selective long-horizon refinement}, which first trains on short prefixes and then refines only on continuations that are most likely under the warm-start model. It consistently outperforms full long-horizon training. At 16K, it raises successful attempts from 110±2.7 to 126±2.1 and tasks solved in at least six of eight attempts from 9±0.7 to 14±0.6; with half of the long-horizon data, it still reaches 122±2.4 while cutting training time by 23%. The gains transfer across benchmarks, from 64±2.6 to 73±2.1 on Terminal-Bench v2.0 and from 137±2.7 to 155±2.2 on OpenThoughts-TBLite. For long-horizon supervision, selecting the right trajectories matters more than training on all of them.