The Decision Value of Perception Compute
Organizations: Hanoi University of Science and Technology (HUST), Hanoi, Vietnam
Abstract
Adaptive perception spends extra computation on inputs where perception is expected to improve. When perception feeds a downstream decision system, a better perception output need not produce a better decision. We define the decision value of perception compute as the change in downstream loss from escalating an input from a cheap to an expensive perception mode. Because this value can be negative, the allocation of perception compute should be judged against a budget-constrained decision oracle, with uniform full-fidelity inference as a baseline rather than an upper bound. We introduce DEEP (Decision Evaluation for Escalated Perception), a benchmark that scores pre-escalation allocators against this oracle under selection, latency and energy budgets, charging each allocator for its own computation. With deployed monocular geometry on KITTI and nuScenes, we find that 34--54% of the escalations that change downstream loss make it worse; harmful escalations also occur for the published PDM-Closed planner, evaluated open-loop on nuPlan with real detector outcomes. On nuScenes, perception-level gain frequently disagrees in sign with decision value. This mismatch has practical consequences: choosing among fixed deployable signals by missed-object perception gain rather than by decision value reduces realized test decision gain by 7.4% of the all-cheap loss on average. Learned allocators recover part of the oracle's value by finding beneficial escalations but select nearly as much harm as random, and once their own computation is charged at a 20% latency budget, only the lightweight routers, at about 3.5% of a full detector pass, still beat random.
Figures & tables
| Dataset, system, geometry | Transition / loss | Affected | Harmed | All-full | Oracle@20 | |
| nuScenes, , oracle | YOLOv8s 320 640 | 141 | 35.5% | 0.42 | 14.01% | 24.07% |
| nuScenes, , mono | YOLOv8s 320 640 | 273 | 39.6% | 0.41 | 16.47% | 27.77% |
| nuScenes, , oracle † | YOLOv8s 320 640 | 1,179 | 50.9% | 0.70 | 0 1.31% | 0 4.41% |
| nuScenes, , mono † | YOLOv8s 320 640 | 1,612 | 49.3% | 0.71 | 0 1.59% | 0 5.31% |
| KITTI, , mono | YOLOv8s 320 640 | 678 | 48.4% | 0.28 | 12.78% | 17.81% |
| KITTI, , mono | YOLOv8s 384 640 | 458 | 48.7% | 0.46 | 0 5.97% | 11.03% |
| Controls | Heuristic | Gate | Detection-list router | |||||||
| Cell | Rand. | Ego sp. | Unc. | Crit. | ridge | GBM | MLP, | MLP, | GBM, | GBM, |
| nuScenes oracle, | 0.148 | 0.208 | 0.015 | 0.470 | 0.254 | 0.155 | 0.354 | 0.354 | 0.196 | 0.421 |
| nuScenes mono, | 0.132 | 0.148 | 0.088 | 0.407 | 0.293 | 0.482 | 0.366 | 0.255 | 0.383 | 0.369 |
| nuScenes oracle, | 0.031 | |||||||||
| nuScenes mono, | 0.092 | 0.013 | 0.063 | 0.042 | 0.250 | 0.102 | 0.090 | |||
| KITTI oracle, | 0.139 | 0.584 | 0.238 | 0.226 | 0.255 | 0.313 | 0.293 | 0.360 | ||
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Braking | Published learned planner | |||
| Allocation score | class | |||
| Random (16 seeds) | deployable | 0.122 | 0.061 | 0.123 |
| Cheap-detection uncertainty | deployable | 0.005 | 0.134 | 0.096 |
| Pre-escalation gate, GBM | deployable | 0.309 | 0.079 | — |
| Criticality (reference geometry) | diagnostic | 0.182 | 0.016 | 0.219 |
| False-negative gain | diagnostic | 0.171 | 0.228 | 0.328 |
| Downstream system | Pair | Geometry | Affected | Harmed | All-full / Oracle@20 | |
| nuScenes, | Y8 320 640 | oracle | 141 | 35.5% | 0.42 | 14.01 / 24.07% |
| nuScenes, | Y8 320 640 | oracle | 1,179 | 50.9% | 0.70 | 0 1.31 / 0 4.41% |
| nuScenes, | Y8 320 640 | oracle | 1,413 | 40.9% | 0.46 | 0 7.49 / 13.67% |
| KITTI, | Y8 320 640 | mono | 0, 678 | 48.4% | 0.28 | 12.78 / 17.81% |
| KITTI, | Y8 320 640 | oracle | 0, 463 | 0 3.0% | 0.003 | 68.34 / 68.58% |
| KITTI, | Y8 384 640 | mono | 0, 458 | 48.7% | 0.46 | 0 5.97 / 11.03% |
| Harmed | Harmed | |||||||
| Pair | mono | oracle | mono | oracle | mono | oracle | mono | oracle |
| YOLOv8s 320 640 | 34.2 | 24.6 | 0.168 | 0.080 | 48.4 | 3.0 | 0.282 | 0.003 |
| YOLOv8s 384 640 | 39.8 | 35.3 | 0.473 | 0.295 | 48.7 | 5.1 | 0.459 | 0.002 |
| YOLOv8s 512 640 | 42.9 | 39.7 | 0.615 | 0.467 | 53.8 | 17.0 | 1.114 | 0.051 |
| RT-DETR-l 320 640 | 44.2 | 47.5 | 0.701 | 0.691 | 49.7 | 19.9 | 0.762 | 0.023 |
| Cell | S0 (shared) | S1 (equal boxes) | S2 (F1) | S3 (precision) |
| nuScenes oracle, | 35.5% / 0.42 | 38.5% / 0.54 | 34.5% / 0.39 | 34.1% / 0.41 |
| nuScenes oracle, | 50.9% / 0.70 | 50.2% / 0.78 | 50.1% / 0.71 | 50.2% / 0.72 |
| nuScenes mono, | 39.6% / 0.41 | 36.2% / 0.34 | 37.0% / 0.34 | 37.5% / 0.37 |
| nuScenes mono, | 49.3% / 0.71 | 49.6% / 0.83 | 49.3% / 0.73 | 49.2% / 0.71 |
| KITTI Y8 384 640, | 48.7% / 0.46 | 48.1% / 0.48 | 48.2% / 0.47 | 48.6% / 0.47 |
| KITTI Y8 384 640, | 39.8% / 0.47 | 34.9% / 0.44 | 37.1% / 0.45 | 38.8% / 0.47 |
| Detection-list router | |||||
| Cell | MLP, | MLP, | GBM, | GBM, | Pixel router |
| nuScenes oracle, | 0.354 | 0.354 | 0.196 | 0.421 | |
| nuScenes oracle, planner ADE | 0.031 | ||||
| nuScenes oracle, planner FDE | 0.037 | 0.065 | |||
| nuScenes mono, | 0.366 | 0.255 | 0.383 | 0.369 | 0.215 |
| nuScenes mono, planner ADE | 0.250 | 0.102 | 0.090 | 0.082 | |
| Cell | Random | Ego speed | Uncertainty | Criticality | Gate, ridge | Gate, GBM |
| nuScenes oracle, | 0.148 | 0.208 | 0.015 | 0.470 | 0.254 | 0.155 |
| nuScenes oracle, planner ADE | ||||||
| nuScenes oracle, planner FDE | 0.059 | 0.024 | 0.030 | 0.125 | ||
| nuScenes mono, | 0.132 | 0.148 | 0.088 | 0.407 | 0.293 | 0.482 |
| nuScenes mono, planner ADE | 0.092 | 0.013 | 0.063 | 0.042 | ||
| nuScenes mono, planner FDE | 0.080 | 0.007 | 0.079 |
| Cell | Criticality (reference) | PKL | TIP | ||
| nuScenes oracle, | 0.287 | 0.258 | 0.687 | 0.461 | 0.365 |
| nuScenes oracle, planner ADE | 0.123 | ||||
| nuScenes oracle, planner FDE | 0.121 | 0.124 | |||
| nuScenes mono, | 0.323 | 0.210 | 0.330 | 0.215 | 0.110 |
| nuScenes mono, planner ADE | 0.350 | 0.404 | 0.126 | 0.117 | |
| nuScenes mono, planner FDE | 0.264 | 0.246 | 0.077 | 0.168 |
| Cell | Half-width | Cell | Half-width |
| nuScenes oracle, planner ADE | 0.7% | KITTI mono, | 3.7% |
| nuPlan IDM, scalar | 1.0% | nuScenes mono, | 5.8% |
| nuScenes oracle, planner FDE | 1.1% | nuScenes oracle, | 6.0% |
| nuScenes mono, planner ADE | 1.1% | KITTI oracle, | 6.5% |
| nuPlan IDM, safety | 1.1% | nuPlan PDM-Closed, scalar | 14.0% |
| nuScenes mono, planner FDE | 1.7% | nuPlan PDM-Closed, safety | 16.0% |
| Allocator | Overhead per input | Escalated at 20% budget | at 50% budget |
| Random | 0 ms | 20.0 / 20.0% | 50.0 / 50.0% |
| Cheap-detection uncertainty | 0.21–0.29 ms | 18.5 / 18.8% | 48.5 / 48.8% |
| Cheap-side criticality | 1.30–1.47 ms | 12.4 / 12.9% | 42.4 / 42.9% |
| Detection-list router, MLP | 0.65 ms | 16.6 / 16.5% | 46.6 / 46.5% |
| Gate, ridge | 3.93 ms | inf. / inf. | 29.7 / 28.7% |
| Gate, GBM, batched inference | 3.54 ms | 1.7 / 0.8% | 31.7 / 30.8% |
| Planner | Loss | Affected | Harmed | All-full / Oracle@20 | |
| Real detector outcomes | |||||
| PDM-Closed | collision | 24 | 25.0% | 0.33 | 11.54 / 17.31% |
| PDM-Closed | clearance shortfall | 45 | 44.4% | 0.34 | 0 5.22 / 0 7.89% |
| PDM-Closed | logged-trajectory deviation | 104 | 37.5% | 0.55 | 0 1.90 / 0 4.25% |
| PDM-Closed | safety aggregate | 45 | 44.4% | 0.33 | 0 9.81 / 14.74% |
| PDM-Closed | scalar aggregate | 104 | 37.5% | 0.35 | 0 8.26 / 12.68% |
| nuScenes, | nuScenes, | |||||
| helped | harmed | unaffected | helped | harmed | unaffected | |
| false positives | ||||||
| false negatives | ||||||
| detections | ||||||
| frames with an added FP | 61.5% | 78.0% | — | 60.0% | 64.2% | — |
| frames recovering no miss | 25.3% | 56.0% | — | 31.0% | 44.6% | — |
| (141 frames) | (1,179 frames) | |||||
| Perception gain | Corr. | Disagree | Corr. | Disagree | ||
| Exact (false-negative count) | 24.4% | 25.0% | 50.6% | 50.4% | ||
| False negatives + false positives | 15.7% | 31.9% | 50.3% | 49.8% | ||
| Combined error ( ) | 19.2% | 32.6% | 51.0% | 50.6% | ||
| Risk-weighted ( ) | 23.8% | 26.8% | 50.5% | 51.0% | ||
| Configuration | Task | nDG of | Multi-metric diagnostic |
| KITTI/Y8 320 640 mono | longitudinal | 0.162 | 0.672 |
| KITTI/Y8 320 640 mono | lateral | ||
| KITTI/Y8 320 640 oracle | longitudinal | 0.200 | 0.810 |
| KITTI/Y8 320 640 oracle | lateral | ||
| KITTI/Y8 384 640 mono | longitudinal | 0.152 | 0.388 |
| KITTI/Y8 384 640 mono | lateral |
| Braking weights | Braking corridor | Trajectory weights | ||||
| Setting | Harmed | Harmed | Harmed | |||
| nuScenes, oracle | 35.5–36.2 | 0.28–0.62 | 35.5–36.5 | 0.33–0.46 | — | |
| nuScenes, mono | 38.1–42.5 | 0.28–0.55 | 39.2–41.6 | 0.37–0.41 | — | |
| KITTI Y8 320 640, mono | 32.6–35.3 | 0.11–0.28 | 32.7–34.6 | 0.15–0.22 | 46.0–48.8 | 0.27–0.32 |
| KITTI Y8 384 640, mono | 38.8–40.1 | 0.35–0.65 | 38.9–39.8 | 0.41–0.51 | 47.3–51.0 | 0.45–0.51 |
| KITTI Y8 512 640, mono | 42.9–43.4 | 0.51–0.72 | 42.6–42.9 | 0.59–0.71 | 51.6–54.5 | 1.10–1.15 |
| Mode | GPU ms | End-to-end ms | Energy mJ |
| YOLOv8s 320 | 5.09 | 13.18 | 41.1 |
| YOLOv8s 384 | 6.09 | 14.38 | 47.8 |
| YOLOv8s 512 | 7.89 | 16.72 | 62.0 |
| YOLOv8s 640 | 9.28 | 18.47 | 84.4 |
| RT-DETR-l 320 | 16.82 | 25.31 | 397 |
| RT-DETR-l 480 | 21.35 | 31.21 | 535 |