Rethinking Streaming-Perception Evaluation on Heterogeneous Edge Platforms
Authors: Misun Yu, Jinyoung Moon, Jemin Lee
Organizations: Electronics and Telecommunications Research Institute (ETRI), Daejeon, Republic of Korea · Division of Electronics and Information Engineering, Jeonbuk National University, Jeonju, Republic of Korea
Multi-camera streaming perception is increasingly deployed on heterogeneous edge platforms shared with co-resident workloads, yet accelerator placement is often evaluated using isolated single-stream experiments and mean streaming average precision (sAP). Using two end-to-end pipelines on a single GPU--NPU platform, we show that isolated evaluation can mis-rank deployment-time placement. Although the GPU pipeline is preferred in isolation, GPU-localized contention introduces deadline misses that make detections stale and can reverse the preferred placement before full GPU saturation. The NPU pipeline is less accurate than the GPU pipeline on small and medium objects in isolation, but nearly matches it on large objects. The largest absolute sAP losses in our latency and contention experiments occur for large objects. In our four-stream experiments, the preferred placement depends on which path becomes stale, and increasing GPU-side contention shifts the best placement from All-GPU to All-NPU. Under a GPU-saturating vision--language co-tenant, All-NPU achieves 5.2× the worst-stream sAP of All-GPU. Because mean sAP can hide severe single-stream degradation, evaluation should report contention sweeps, deadline-miss rates on both paths, and worst-stream sAP alongside mean sAP.
Figures & tables
Figure 1 : Overview of the evaluated system. (a) GPU–NPU placement with GPU-pinned co-tenants and shared host post-processing. (b) Frame skipping and evaluation with stale detections when processing is delayed by contention.
Metric
GPU
NPU
NPU − GPU
end-to-end mean (ms)
8.3
10.0
—
deadline misses (%)
∼0
∼0
—
sAP overall
19.6
18.6
−1.0 ( −5.0% )
sAP small
1.6
0.8
−0.8 ( −48.4% )
sAP medium
18.4
14.8
−3.6 ( −19.7% )
sAP large
47.7
47.7
+0.05 ( +0.1% )
Table 1 : Single-camera evaluation with negligible deadline misses. Latency reports mean end-to-end per-frame processing time at four post-processing threads. Values are 24 -log means over three replays. Relative gaps in parentheses are computed as (NPU−GPU)/GPU from the unrounded 24 -log means.
Component
Small
Medium
Large
Isolated gap (NPU − GPU)
−0.8 ( −48.4% )
−3.6 ( −19.7% )
+0.05 ( +0.1% )
Latency control ( 24−4 thr.)
−0.1 ( −8.0% )
−3.0 ( −16.1% )
−9.8 ( −20.5% )
Table 2 : Size-specific sAP differences for single-stream YOLO11s. Rows compare the isolated NPU and GPU pipelines and the same NPU pipeline with 24 versus 4 post-processing threads. Differences are in sAP points. Parenthesized percentages divide these differences by the isolated GPU sAP, using 24 -log means.
Figure 2 : Per-size All-GPU sAP loss under ResNet50 contention at N=4 , relative to k=0 . Losses are stream-averaged absolute sAP differences (points). Large-object loss rises sharply with GPU deadline misses, while small-object loss remains nearly flat.
All-GPU
All-NPU
Oracle
GPU/NPU
Co-tenant
worst
mean
worst
mean
worst
mean
DM (%)
L1 CNN (ResNet50)
9.8
13.7
8.3
12.6
10.3
13.7
24/1
L2 LM ( + LLM)
6.2
10.0
4.2
8.8
6.3
9.9
83/76
L3 VLM ( + VLM)
1.6
4.7
8.3
12.6
8.3
12.0
100/1
Table 3 : Worst-stream and mean sAP under GPU co-tenancy at N=4 . GPU/NPU DM denotes the GPU deadline-miss rate under All-GPU and the NPU deadline-miss rate under All-NPU (%).
Figure 3 : Worst-stream sAP versus GPU deadline misses under GPU-localized ResNet50 contention. All-GPU degrades as GPU deadline misses increase, while All-NPU remains stable. The ranking of All-GPU and All-NPU reverses between the adjacent measured points at 45.0% and 51.5% GPU deadline misses ( N=4 ), before full GPU saturation.
Contention point
P1/P2
P3
P4
Oracle
sAP (gap to Oracle)
sAP
L1 CNN (24%)
9.8 (+0.5)
9.8 (+0.5) G
9.8 (+0.5) G
10.3
ResNet (52%)
8.2 (+0.2)
8.2 (+0.2) G
8.3 (+0.0) N
8.4
ResNet (67%)
7.1 (+1.2)
7.1 (+1.2) G
8.4 (+0.0) N
8.4
L2 LM (83%)
6.2 (+0.1)
6.2 (+0.1) G†
4.2 (+2.0) N
6.3
L3 VLM (100%)
1.6 (+6.8)
8.3 (+0.0) N
8.3 (+0.0) N
8.3
Table 4 : Worst-stream sAP of placement-selection policies at N=4 . Parentheses report Oracle minus policy sAP in points. Gaps are computed from unrounded values. Row percentages denote All-GPU deadline-miss rates. Superscripts G and N indicate All-GPU and All-NPU. Policies and Oracle are defined in Sec. 4.1 .
Detector
isolated small
isolated large
VLM cont.
worst NPU/GPU
GPU-DM bracket
YOLO12s
GPU ( −36% )
within 5% ( −1.2% )
NPU
8.0/1.1
0.4 – 61.5%
YOLO11s
GPU ( −48% )
within 5% ( +0.1% )
NPU
8.3/1.6
45.0 – 51.5%
YOLOv10s
GPU ( −45% )
within 5% ( −4.6% )
NPU
8.2/1.6
0.0 – 29.7%
YOLOv8s
GPU ( −44% )
within 5% ( −0.3% )
NPU
8.0/1.5
10.8 – 32.4%
YOLOv8n
GPU ( −36% )
GPU ( −8.1% )
NPU
6.0/0.9
34.3 – 46.9%
Table 5 : Placement reversal across detectors. Isolated columns report pipeline preference and relative sAP differences (NPU−GPU)/GPU computed from 24 -log means. “Within 5%” denotes numerical proximity, not statistical equivalence. Under VLM contention, “VLM cont.” identifies the preferred pipeline, and “worst NPU/GPU” reports worst-stream sAP for All-NPU and All-GPU, respectively. GPU-DM brackets give adjacent measured GPU deadline-miss rates between which the All-GPU versus All-NPU worst-stream sAP ordering reverses. They are not interpolated estimates or confidence intervals.
Figure S1 : Qualitative differences between the two evaluated pipelines (YOLO11s, Argoverse-HD). Green: GPU-pipeline detections. Red: detections missed by the NPU pipeline. Missed detections are predominantly small and distant objects.
Figure S2 : Per-size confidence distributions for detections matched by both the GPU and NPU pipelines. Vertical lines mark the mean score for each pipeline (dashed: GPU, solid: NPU), and the panel title reports the mean NPU-minus-GPU score difference. The NPU pipeline shifts scores downward, especially for small and medium objects. Small objects with no nearby NPU-pipeline detection ( 51% of the GPU-matched small objects) are absent from the plotted distributions. The pattern is consistent with quantization but is not separated from other software-stack effects.
Figure S3 : Temporal box overlap and estimated sAP loss using Argoverse-HD ground truth. (a) Mean same-object IoU versus temporal offset δ . (b) Fraction of objects with same-object IoU below 0.5 . (c) Estimated staleness sAP loss (solid) versus the measured same-NPU latency-control losses from Table 2 of the main paper (dotted). The vertical axis in panel (c) is in absolute sAP points. Box overlap decreases fastest for small objects, but the estimated sAP loss is largest for large objects. The estimated losses show the same size ordering as the measured losses but do not reproduce their magnitudes.
statistic
All-GPU
All-NPU
mean
10.7
12.6
median
8.7
11.5
p10
7.2
9.1
worst
6.6
8.4
Table S1 : Metric sensitivity at the ResNet50 k=8 sweep point ( N=4 , four post-processing threads, GPU deadline-miss rate about 80% , single replay). Mean and median mask the degraded camera, whereas worst-stream and lower-tail statistics expose it. Percentiles use linear interpolation. With only four per-stream values, the 10th percentile is coarse and is closer to the worst-stream value than the mean or median for both placements. Removing the worst camera raises the All-GPU system mean from 10.7 to 12.1 .
weights
small
medium
large
COCO-pretrained
−43.1% ( p<0.001 )
−16.5% ( p<0.001 )
+0.6% (n.s., p=0.33 )
fine-tuned
−34.9% ( p<0.001 )
−16.7% ( p<0.001 )
−4.3% (n.s., p=0.079 )
Table S2 : Per-size relative sAP gap (NPU−GPU)/GPU under a single local compile recipe applied to COCO-pretrained and Argoverse-HD fine-tuned YOLO11s. The magnitude of the relative sAP gap is largest for small objects in both model conditions. The values are an internal paired comparison. n.s. denotes not statistically significant ( p≥0.05 ).
Pipeline
log 2
log 22
log 3
log 21
mean
NPU (vendor INT8)
10.90
19.12
12.14
8.35
12.63
GPU (TensorRT INT8)
8.99
10.30
6.87
6.37
8.13
NPU − GPU (per-log mean)
+1.91
+8.83
+5.28
+1.99
+4.50
Table S3 : Single-stream overall sAP in percent per evaluated log, mean over nine repetitions. NPU − GPU differences are in sAP points and are computed from unrounded means. Deadline misses are ≈0 for both pipelines (mean end-to-end latency 14.3 ms NPU, 6.8 ms GPU TensorRT INT8).
Sep 25, 2024·Daghash K. Alqahtani, Muhammad Aamir Cheema, Maria A. Rodriguez +1Yolo26Mobilenetv2
DisNet Lab, School of Computing and Information Systems, The University of Melbourne, Melbourne, VIC, Australia · Monash University, Melbourne, VIC, Australia