StreamDecisionBench: Evaluating Decisions in Force on Evolving Language Streams
Organizations: National Yang Ming Chiao Tung University
Abstract
Language models increasingly make real-time decisions in applications that apply the latest answer until a newer one arrives. A late answer can prolong an outdated decision, such as a call recorder still running while a customer reads out card details, an error offline accuracy misses. We make three contributions. First, we release StreamDecisionBench (SDB), a dataset of eight streaming scenarios in four application families, with executable reference decisions derived from public rules. Second, we propose an evaluation protocol and a metric, in-force accuracy: the share of time the applied decision is correct across update intervals of 0.5-8 s. It reflects accuracy and latency jointly, attributing each error to judgment, latency or both. Third, we evaluate thirteen single-model settings, and this attribution separates speed-limited from judgment-limited models: slower, more accurate models lose 42-51% of the time to outdated answers, a fast model 34% to wrong ones. We therefore test hybrids in which a slow model corrects a fast one; with the right pairing and configuration, a hybrid outperforms every single model. However, even the best evaluated system keeps a correct decision in force only about two-thirds of the time, leaving a substantial gap for real-time use.
Figures & tables
| Normalized log-AUC over 0.5–8 s (%) | ||||||
| Setting | IDE | Assembly | Support | Presenter | Macro | Untimed |
| Luna low | 46.7 | 43.3 | 45.7 | 45.4 | 45.3 | 88.8 |
| Luna none | 14.3 | 19.5 | 43.3 | 45.5 | 30.7 | 43.8 |
| Terra low | 44.6 | 46.6 | 50.3 | 50.6 | 48.0 | 95.4 |
| Terra none | 48.8 | 49.9 | 56.2 | 61.5 | 54.1 | 82.1 |
| Astra low | 45.6 | 47.2 | 39.2 | 48.7 | 45.2 | 99.8 |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Sc. | Qs | Routes | Options | Trans. |
| IDE debugging | A | 7 | 6 | 2–6 | 20 |
| B | 7 | 6 | 2–6 | 21 | |
| Assembly station | A | 6 | 7 | 5–7 | 21 |
| B | 6 | 7 | 5–9 | 23 | |
| Support call | A | 7 | 4 | 3–7 | 22 |
| B | 7 | 6 | 3–8 | 24 |
| Scenarios | Route options | Used in every branch | Branch questions |
|---|---|---|---|
| IDE debugging (A, B) | wait, rerun, inspect, control, delegate, ready | process state, target result | rerun: scope; inspect: file; control: continue or stop; delegate: owning team; wait, ready: none |
| Assembly station (A, B) | wait, advance, repair, escalate, handoff, hold, release | stage | advance: next step; repair: target, method; escalate: target, destination; handoff, hold, release: destination; wait: none |
| Support call A | payment, service, hold, wrap-up | recorder command | payment: payment stage, card; service: target, action; hold: hold action; wrap-up: none |
| Presenter control (A, B) | talk, questions, clip, closed | slide, captions | talk: chair cue; questions: question card, chair cue; clip: clip state; closed: none |
| Support call B | delivery, cancel, repair, hold (wait), hold (return), closed | follow-up contact channel | delivery: parcel, delivery action; cancel: parcel, cancellation action; repair: device, repair action; hold and closed: none |
| Luna low | Luna none | Terra low | Terra none | Astra low | Jev | |
| Valid responses | 480 | 480 | 480 | 480 | 480 | 480 |
| Attempts | 480 | 480 | 480 | 480 | 480 | 480 |
| Failed attempts | 0 | 0 | 0 | 0 | 0 | 0 |
| Retried requests | 0 | 0 | 0 | 0 | 0 | 0 |
| Accepted updates | 464 | 478 | 464 | 478 | 466 | 480 |
| Discarded: after horizon | 9 | 0 | 8 | 0 | 8 | 0 |
| Model | Net. | Range | Pre. | Dec. | p50 | p95 |
|---|---|---|---|---|---|---|
| (s) | (s) | (ms/1k tok) | (ms/tok) | (s) | (s) | |
| Luna low | 0.58 | (0.31–0.69) | 59.5 | 8.3 | 2.41 | 4.10 |
| Luna none | 0.86 | (0.50–1.09) | 0.0 | 4.8 | 1.32 | 1.88 |
| Terra low | 0.67 | (0.56–0.75) | 0.0 | 12.1 | 2.44 | 3.77 |
| Terra none | 1.26 | (1.05–1.28) | 0.0 | 0.0 | 1.49 | 2.08 |
| Astra low | 0.94 | (0.63–1.30) | 0.0 | 28.3 | 2.67 | 4.45 |
| Luna low | Luna none | Terra low | Terra none | Astra low | Jev | |
| Tokens per request (median) | ||||||
| Input | 3034 | 3034 | 3034 | 3034 | 3034 | 3613 |
| Output | 148 | 46 | 117 | 46 | 48 | 396 |
| Reasoning | 100 | 0 | 69 | 0 | 0 | 0 |
| Tokens per pass (total) | ||||||
| Input | 1,476,489 | 1,476,489 | 1,476,489 | 1,476,489 | 1,476,489 | 1,754,113 |
| Network removed | ||||
| Setting | Untimed | In force | Est. | Range |
| IDE debugging | ||||
| Luna low | 84.2 | 51.6 | 59.3 | (55.5–60.6) |
| Luna none | 19.2 | 15.9 | 17.7 | (16.9–18.1) |
| Terra low | 99.2 | 48.2 | 59.4 | (57.4–60.8) |
| Terra none | 72.5 | 55.6 | 67.7 | (65.6–67.8) |
| Normalized AUC | In force at | |||||||
|---|---|---|---|---|---|---|---|---|
| Log | Linear | Log | Log | Log | Log | Log | ||
| Setting | 0.5–8 s | 0.5–8 s | 1–4 s | 0.5–4 s | 1–8 s | 0.1–4 s | 0.1–8 s | 2 s |
| Luna low | 45.3 | 61.0 | 46.8 | 36.0 | 55.6 | 22.1 | 30.2 | 47.7 |
| Luna none | 30.7 | 36.2 | 32.5 | 27.6 | 35.0 | 17.9 | 21.3 | 33.3 |
| Terra low | 48.0 | 65.0 | 49.2 | 38.0 | 58.8 | 23.4 | 32.1 | 50.4 |
| Terra none | 54.1 | 66.0 | 58.2 | 47.5 | 63.4 | 29.0 | 36.1 | 59.8 |
| Setting | |||||
|---|---|---|---|---|---|
| Luna low | 48.6 | 90.9 | 1.1 | 43.1 | 45.3 |
| Luna none | 66.1 | 44.4 | 1.3 | 28.8 | 30.7 |
| Terra low | 49.6 | 95.9 | 0.5 | 47.3 | 48.0 |
| Terra none | 62.5 | 84.4 | 1.4 | 51.3 | 54.1 |
| Astra low | 45.2 | 100.0 | 0.0 | 45.1 | 45.2 |
| Jev | 93.0 | 63.8 | 0.3 | 59.3 | 59.6 |
| Luna | Terra | Astra | |||||
|---|---|---|---|---|---|---|---|
| Scale | Interval (s) | low | none | low | none | low | Jev |
| 0.5 | 4.00 | 67.1 | 38.3 | 71.6 | 70.8 | 71.2 | 62.2 |
| 2/3 | 3.00 | 60.2 | 36.6 | 64.0 | 67.1 | 62.0 | 61.7 |
| 1 | 2.00 | 47.7 | 33.3 | 50.4 | 59.8 | 45.8 | 60.7 |
| 1.5 | 1.33 | 33.2 | 28.5 | 33.9 | 49.1 | 27.9 | 59.2 |
| 2 | 1.00 | 23.7 | 24.1 | 23.8 | 39.0 | 18.9 | 57.7 |
| Terra low vs. Luna low | ||||
|---|---|---|---|---|
| Both | Luna only | Terra only | Neither | |
| IDE debugging | 100 | 1 | 19 | 0 |
| Assembly station | 104 | 1 | 7 | 8 |
| Support call | 97 | 7 | 15 | 1 |
| Presenter control | 112 | 4 | 4 | 0 |
| All | 413 | 13 | 45 | 9 |
| Error time (s) | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Scenario | Model | Untimed | Log-AUC | In force | Seg.-bal. | None | Stale | Judg. | Comp. | p50 (s) |
| IDE debugging A | Luna low | 83.3 | 47.0 | 52.3 | 47.7 | 2.3 | 39.9 | 10.7 | 4.3 | 2.18 |
| Luna none | 18.3 | 11.8 | 13.4 | 13.3 | 2.6 | 5.2 | 73.8 | 22.4 | 1.33 | |
| Terra low | 100.0 | 44.2 | 48.1 | 43.7 | 2.4 | 59.9 | 0.0 | 0.0 | 2.50 | |
| Terra none | 76.7 | 51.3 | 59.1 | 54.6 | 2.1 | 20.8 | 16.0 | 10.2 | 1.58 | |
| Astra low | 100.0 | 45.0 | 48.8 | 44.5 | 3.4 | 58.1 | 0.0 | 0.0 | 2.54 | |
| Luna | Luna | Terra | Terra | Astra | ||
| low | none | low | none | low | Jev | |
| Share of observed time (%) | ||||||
| Judgment | 4.0 | 40.7 | 1.7 | 10.7 | 0.0 | 34.4 |
| Stale | 40.2 | 11.1 | 43.4 | 22.9 | 51.6 | 3.0 |
| Compound | 5.8 | 13.6 | 2.6 | 5.1 | 0.2 | 1.6 |
| No decision | 2.2 | 1.3 | 2.0 | 1.5 | 2.4 | 0.4 |
| Reference schedule | Changes | Luna low | Luna none | Terra low | Terra none | Astra low | Jev |
|---|---|---|---|---|---|---|---|
| SDB (eight scenarios) | – | 46.0 | 32.2 | 50.6 | 57.5 | 47.5 | 60.5 |
| sparse | 5 | 77.2 | 40.8 | 83.4 | 75.6 | 85.5 | 62.9 |
| medium | 11 | 65.8 | 37.8 | 71.3 | 69.3 | 71.4 | 62.0 |
| uniform | 22 | 44.6 | 32.2 | 49.4 | 57.6 | 45.7 | 60.5 |
| dense | 40 | 20.6 | 23.4 | 23.9 | 38.9 | 18.8 | 57.9 |
| bursty | 22 | 53.0 | 32.4 | 57.7 | 58.0 | 58.1 | 60.5 |
| Setting | Near error (%) | Far error (%) |
|---|---|---|
| Luna low | 12.8 | 3.8 |
| Luna none | 56.5 | 55.0 |
| Terra low | 5.3 | 1.3 |
| Terra none | 19.5 | 10.0 |
| Astra low | 0.3 | 0.0 |
| Jev | 37.5 | 30.0 |
| Slow setting | Arbitration | 1 s | 2 s | AUC |
|---|---|---|---|---|
| Luna low | Freshest | 57.7 | 61.2 | 63.9 |
| Luna low | Lag | 56.0 | 58.2 | 62.9 |
| Luna low | Override | 47.0 | 57.6 | 59.3 |
| Luna none | Freshest | 57.3 | 52.8 | 52.0 |
| Luna none | Lag | 43.1 | 52.6 | 48.2 |
| Luna none | Override | 43.1 | 52.6 | 47.4 |
| Normalized log-AUC over 0.5–8 s (%) | ||||||
|---|---|---|---|---|---|---|
| Setting | IDE | Assembly | Support | Presenter | Macro | Untimed |
| Laya English | 0.00 | 0.00 | 1.67 | 0.00 | 0.42 | 0.42 |
| Laya typed-decisions | 0.00 | 0.00 | 1.67 | 4.17 | 1.46 | 1.46 |
| Laya multilingual | 0.00 | 0.84 | 0.00 | 0.00 | 0.21 | 0.21 |
| DJev / DiffusionGemma | 22.59 | 6.15 | 28.63 | 26.43 | 20.95 | 21.88 |
| Kev-4B | 15.69 | 12.94 | 36.71 | 18.30 | 20.91 | 21.46 |
| Setting | GPU | p50 (s) | p95 (s) |
|---|---|---|---|
| Laya English | RTX PRO 6000 (96 GB) | 0.126 | 0.240 |
| Laya typed-decisions | RTX PRO 6000 (96 GB) | 0.129 | 0.235 |
| Laya multilingual | RTX PRO 6000 (96 GB) | 0.068 | 0.132 |
| DJev / DiffusionGemma | RTX PRO 6000 (96 GB) | 0.257 | 0.403 |
| Kev-4B | L40S (48 GB) | 0.160 | 0.261 |
| Bespoke Nimble-9B | L40S (48 GB) | 6.162 | 41.782 |
| Provisional setting | Freshest | One-tick lag | Late override |
|---|---|---|---|
| Laya English | 27.27 | 33.65 | 34.94 |
| Laya typed-decisions | 27.77 | 34.07 | 35.33 |
| Laya multilingual | 26.22 | 33.06 | 34.64 |
| DJev / DiffusionGemma | 42.07 | 45.11 | 45.16 |
| Kev-4B | 41.05 | 44.73 | 44.87 |
| Bespoke Nimble-9B | 54.08 | 54.08 | 54.08 |