Analytical and Convolutional Neural Network-Based Motion-Vector Propagation for Efficient Video Object Detection
Organizations: Department of Computer and Information Engineering, Khalifa University, Abu Dhabi, UAE
Abstract
Continuous video analytics requires accurate localization at low latency within embedded power budgets. This paper presents a hardware-software design methodology that reuses codec motion vectors (MVs) between detector invocations. Two alternative models support translation and scale changes: analytical motion-vector propagation (Analytical-MV) and learned propagation using a convolutional neural network (CNN) (CNN-MV). The learned model uses convolutional operations and independent object updates suited to parallel execution on an edge graphics processing unit (GPU). Analytical-MV combines a harmonic-mean precision-recall score (F1) of 0.909 with a mean end-to-end latency of 9.03 ms and an energy consumption of 0.177 J per frame, yielding the lowest latency and energy among the evaluated configurations. Relative to detection on every frame, it reduces mean latency by 25.9% and energy per frame by 36.4%. CNN-MV offers a different trade-off: its fastest configuration raises recall from 0.871 for Analytical-MV to 0.890 and lowers mean power from 19.64 to 17.32 W, while achieving a latency of 18.42 ms and an energy consumption of 0.319 J per frame. It is therefore useful when recall or operating power is more important than minimum latency and energy. Execution on a deep learning accelerator (DLA) further reduces time-averaged GPU utilization relative to GPU execution. Host-processing optimization substantially improves both latency and energy, demonstrating the value of jointly designing temporal models and their execution pipelines.
Figures & tables
| ID | Method | Motion execution | Host implementation |
| C0 | YOLO-only | None | Detector on every frame |
| C1 | MVP-YOLO (adapted reference) | CPU analytical | Original MVP grid propagation |
| C2 | Analytical-MV (ours) | CPU analytical | Masked grid plus coupled translationscale fit |
| C3 | CNN-MV (ours) | DLA | Serial original |
| C4 | CNN-MV (ours) | DLA | Serial optimized pre/post |
| C5 | CNN-MV (ours) | DLA | Parallel original |
| ID | Accuracy | Detection counts | Detector routing | MV routing | |||||||
| Prec. | Recall | F1 | IoU | TP | FP | FN | Frames | % | Frames | % | |
| C0 | 0.957 | 0.910 | 0.933 | 0.904 | 197,005 | 8,940 | 19,522 | 28,177 | 100.0 | 0 | 0.0 |
| C1 | 0.957 | 0.874 | 0.914 | 0.901 | 189,343 | 8,594 | 27,184 | 21,048 | 74.7 | 7,129 | 25.3 |
| C2 | 0.951 | 0.871 | 0.909 | 0.880 | 188,543 | 9,708 | 27,984 | 13,515 | 48.0 | 14,662 | 52.0 |
| C3 | 0.934 | 0.890 | 0.911 | 0.874 | 192,795 | 13,730 | 23,732 | 15,647 | 55.5 | 12,530 | 44.5 |
| C4 | 0.934 | 0.890 | 0.911 | 0.874 | 192,795 | 13,730 | 23,732 | 15,647 | 55.5 | 12,530 | 44.5 |
| ID | Throughput (FPS) | End-to-end latency (ms) | Inference time (ms) | Calls/ | ||||
| Pipeline | Engine | Mean | P95 | P99 | Engine/call | Compute/frame | frame | |
| C0 | 82.00 | 249.31 | 12.20 | 13.28 | 13.68 | 4.01 | 4.01 | 1.000 |
| C1 | 85.33 | 247.99 | 11.72 | 15.90 | 18.39 | 4.03 | 4.25 | 0.747 |
| C2 | 110.69 | 249.16 | 9.03 | 16.90 | 18.24 | 4.01 | 3.44 | 0.479 |
| C3 | 11.67 | 193.10 | 85.71 | 124.63 | 138.71 | 0.37 | 5.18 | 6.017 |
| C4 | 48.46 | 194.94 | 20.64 | 36.20 | 41.55 | 0.37 | 5.13 | 6.017 |
| ID | YOLO | Motion compute | Propagation route | YOLO recovery/fallback | Calls/ | |||
| Mean | Mean | P95 | Mean | P95 | Mean | P95 | frame | |
| C1 | 4.03 | 1.24 | 3.74 | 3.93 | 8.25 | 14.43 | 16.34 | 0.747 |
| C2 | 4.01 | 1.52 | 3.32 | 3.68 | 6.65 | 14.99 | 17.66 | 0.479 |
| C3 | 4.00 | 2.96 | 8.05 | 76.83 | 92.85 | 94.19 | 130.32 | 6.017 |
| C4 | 4.01 | 2.91 | 7.97 | 13.40 | 19.79 | 27.07 | 38.30 | 6.017 |
| C5 | 4.00 | 3.60 | 8.69 | 76.80 | 91.31 | 93.34 | 127.11 | 6.213 |
| ID | Total power (W) | Mean rail power (W) | Frame mean | Energy efficiency | ||||
| Mean | P95 | GPU-SoC | CPU-CV | 5V | (W) | J/frame | FPS/W | |
| C0 | 22.88 | 23.54 | 12.98 | 3.19 | 6.70 | 22.90 | 0.279 | 3.584 |
| C1 | 20.60 | 21.73 | 11.28 | 3.00 | 6.33 | 20.50 | 0.241 | 4.142 |
| C2 | 19.64 | 21.24 | 10.30 | 3.15 | 6.19 | 19.50 | 0.177 | 5.637 |
| C3 | 14.14 | 14.73 | 5.92 | 3.19 | 5.04 | 14.23 | 1.212 | 0.825 |
| C4 | 17.45 | 18.63 | 8.18 | 3.45 | 5.82 | 17.33 | 0.360 | 2.777 |
| ID | Mean utilization (%) | CPU temperature ( ∘ C) | GPU temperature ( ∘ C) | |||||
| CPU | GPU | DLA0 | EMC | Mean | Max. | Mean | Max. | |
| C0 | 9.2 | 33.1 | - | 13.9 | 56.3 | 59.0 | 52.8 | 55.0 |
| C1 | 8.9 | 25.8 | - | 10.8 | 57.6 | 58.8 | 54.0 | 55.3 |
| C2 | 9.7 | 21.8 | - | 9.2 | 57.1 | 58.8 | 52.9 | 54.2 |
| C3 | 8.1 | 2.9 | 0.0 | 1.3 | 54.9 | 56.4 | 50.4 | 52.1 |
| C4 | 12.1 | 11.8 | 1.4 | 6.1 | 56.8 | 58.3 | 52.2 | 53.7 |
| ID | Decode | Preprocess | Compute (total and components) | Postprocess | Overhead | ||
| Total | YOLO | Motion | |||||
| C0 | 0.60 | 3.90 | 4.01 | 4.01 | - | 1.69 | 2.00 |
| C1 | 1.64 | 2.84 | 4.25 | 3.01 | 1.24 | 1.32 | 1.67 |
| C2 | 1.61 | 1.83 | 3.44 | 1.92 | 1.52 | 0.87 | 1.27 |
| C3 | 1.67 | 73.54 | 5.18 | 2.22 | 2.96 | 2.29 | 3.03 |
| C4 | 1.65 | 9.09 | 5.13 | 2.22 | 2.91 | 1.80 | 2.97 |
| ID | Decode | Preprocess | Compute (total and components) | Postprocess | Overhead | ||
| Total | YOLO | Motion | |||||
| C0 | 13.6 | 89.3 | 91.8 | 91.8 | - | 38.4 | 45.9 |
| C1 | 33.7 | 58.6 | 87.5 | 62.3 | 25.2 | 27.0 | 34.4 |
| C2 | 31.5 | 36.4 | 67.5 | 38.2 | 29.4 | 17.1 | 25.0 |
| C3 | 23.8 | 1038.7 | 74.1 | 32.2 | 41.9 | 32.6 | 43.2 |
| C4 | 28.6 | 158.1 | 90.0 | 39.5 | 50.5 | 31.3 | 51.9 |