Moving Object Detection from Moving Camera Using Focus of Expansion Likelihood and Segmentation
Authors: Masahiro Ogawa, Qi An, Atsushi Yamashita
Organizations: Department of Precision Engineering, Graduate School of Engineering, The University of Tokyo, Japan · Department of Human and Engineered Environmental Studies, Graduate School of Frontier Sciences, The University of Tokyo, Japan
Separating moving and static objects from a moving camera viewpoint is essential for 3D reconstruction, autonomous navigation, and scene understanding in robotics. Existing approaches often rely primarily on optical flow, which struggle to detect moving objects in complex, structured scenes involving camera motion. To address this limitation, we propose Focus of Expansion Likelihood and Segmentation (FoELS), a method based on the core idea of integrating both optical flow and texture information. FoELS computes the focus of expansion (FoE) from optical flow and derives an initial motion likelihood from the outliers of the FoE computation. This likelihood is then fused with a segmentation-based prior to estimate the final moving probability. The method effectively handles challenges including complex structured scenes, rotational camera motion, and parallel motion. Comprehensive evaluations on the DAVIS 2016 and FBMS-59 datasets, along with real-world traffic videos including parallel, cross-direction, opposite-direction, and crowded scenes, demonstrate its effectiveness and state-of-the-art performance.
Figures & tables
Figure 1 : Sample result of FoELS. It detects moving objects from a moving camera at various distances within the scene.
Figure 2 : Detailed flowchart of the proposed method. The side image illustrates the process outlined in the flowchart.
Method
Stop
Go Forward
Rotate
Go Forward and Rotate
Textureless object
Close object
Close dominant object
Flow Orientation [ 1 ]
✓
×
×
×
×
✓
×
FoE [ 4 ]
✓
✓
×
×
×
✓
×
AdversarialNet [ 3 ]
✓
✓
△
△
×
×
×
FoELS (Ours)
✓
✓
△
△
✓
✓
×
by Orientation
by FoE
by Seg
by FoE
Table 1 : Comparison of tractable scenes. × : Not tractable, △ : Partially tractable, ✓: Tractable The possible reasons for tractability are listed in the bottom row for FoELS.
Method
DAVIS 2016
FBMS 59
Adversarial Net
0.599
0.369
FoELS (Ours)
0.773
0.695
Table 2 : Quantitative evaluation result. The values represent the average IoU scores over the DAVIS 2016 train-val-movobj sequences and FBMS-59 Testset scenes.
Figure 3 : Example visual results of FoELS on the DAVIS 2016 bear scene. First row (left to right): (a) Input frame, (b) segmentation result, and (c) prior moving probability derived from segmentation. Second row (left to right): (d) Optical flow, (e) optical flow with FoE inlier (green arrows) and outliers (red arrows), and (f) the FoE-based moving likelihood. Third row (left to right): (g) Joint moving pixel probability, (h) refined object-level moving mask, and (i) the final moving object result.
Figure 4 : Comparison results with AdversarialNet (left) and FoELS (right) across different scenarios. AdversarialNet exhibits limited generalization to unseen scenes, while FoELS maintains robust performance without scene-specific tuning. The dramatic visual improvement reflects the difference between real-world complexity and standard datasets.
Figure 5 : Example visual results of FoELS on various motion types including parallel, opposite-direction, cross-direction, and crowded scenes. See Fig. 3 for the 9-subimage format.
Method
IoU
InternImageT (Semantic)
0.532
Oneformer (Panoptic) with ObjRefine
0.65
+ FoE sign (=FoELS)
0.773
Table 3 : Ablation study results. The values represent the average IoU scores over the DAVIS 2016 train-val-movobj sequences. "OneFormer with ObjRefine" refers to the OneFormer model for panoptic segmentation with object refinement. "+ FoE sign" indicates the addition of the FoE sign, representing the final FoELS configuration.
Optical Flow Method
IoU
MemFlow
0.736
UniMatch (FoELS)
0.773
Table 4 : Comparison of optical flow methods. The values represent the average IoU scores of FoELS over the DAVIS 2016 train-val-movobj sequences using different optical flow methods.
Formula
Fl
Macro-mean IoU
Std
diff_log (FoELS)
∣log10(1+∣Δℓ∣)∣
0.773
0.183
diff_lin
∣Δℓ∣
0.764
0.195
ratio_log
∣log10(dlratio)∣
0.765
0.186
ratio_lin
∣dlratio−1∣
0.769
0.182
Table 5 : Ablation of the length-factor formula Fl in Eq. ( 4 ) on the DAVIS 2016 train-val-movobj benchmark (47 sequences, 3,508 frames), with α=0.25 and all other parameters fixed. Let Δℓ=∥vP∥−∥vP,static∥ and dlratio=∥vP∥/∥vP,static∥ . diff_log is the formula used in the paper (Eq. ( 4 )).
α
Macro-mean IoU
Std
0.00
0.746
0.184
0.25 (FoELS)
0.773
0.184
0.50
0.767
0.188
Table 6 : Sensitivity sweep over the length-factor weight α in Eq. ( 2 ) on the DAVIS 2016 train-val-movobj benchmark (47 sequences, 3,508 frames). Macro-mean IoU is computed by averaging per-sequence mean IoUs across the 47 sequences; the between-sequence standard deviation is reported alongside. α=0.25 is the value used in the paper.