BatSLAM 2.0: Sequence-Verified Sonar Place Recognition in a Robust Pose Graph
Authors: Jan Steckel
Organizations: Cosys-Lab, Faculty of Applied Engineering, University of Antwerp, 2020 Antwerp, Belgium · Flanders Make Strategic Research Centre, 3920 Lommel, Belgium
Echolocating bats can navigate dark and cluttered spaces using echolocation. Over a decade ago, BatSLAM showed that a robot with a biomimetic binaural sonar can build a topological map of the environment, by recognizing places from the received acoustic signals. Sonar place recognition, however, is ambiguous by nature: corridors produce nearly identical echo trains, and wrong loop closure can collapse the topological map. In this paper, we introduce BatSLAM 2.0, a novel sonar-only SLAM system built from three elements: an updated acoustic front-end, a sequence verifier that tracks and verifies loop closure candidates and a pose graph implemented on a high performance factor graph framework. The system was thoroughly evaluated both in simulated as well as real world recordings. In both cases, the BatSLAM2.0 algorithm shows the capability of robust topological map creation, countering map collapse, and robust scaling of map size.
Figures & tables
Fig. 1: Overview of the processing flow of BatSLAM 2.0. For every pulse, the binaural echo is converted into a local view V with an energy image E and a spectral-shape image S (the direction cue), which is compared with all stored templates. The five best matches are candidates for the sequence verifier, which keeps several recognition hypotheses alive and commits to one only when it is long, strong, unambiguous and plausible given the current pose uncertainty. A committed hypothesis injects weak loop closure links into the pose graph, which is solved incrementally with iSAM2. The links of every hypothesis form a group that the link management can withdraw from the graph or re-admit later. New templates are anchored to pose-graph nodes, so that their poses follow every optimization, and no templates are laid while the robot is locked onto known territory.
Fig. 2: Processing steps of the acoustic front-end, illustrated on pulse 2600 of the development drive on the indoor floor. Panel a) shows the received binaural echo train, normalized per ear, as a function of range. Panel b) shows the output of the matched filter for the left ear in decibels, which reveals weak echoes up to 10m that are invisible in the raw signal. Panel c) shows the fine cochleogram (64 bands between 30kHz and 90kHz , 3cm range bins) as linear amplitude, for both ears (bands low to high within each ear), and panel d) shows the same cochleogram after the time-varying gain. Panel e) shows the energy image E of the descriptor (32 pooled bands, 6cm range bins, cube-root compression), and panel f) the spectral-shape image S (red: bands above the mean level of the range bin, blue: below), which carries the direction-dependent filtering of the ears and the emitter. The speckle in panel f) beyond 5m is noise that passes the level mask.
Fig. 3: Snapshots of five places and their local views. The top row shows the scene within 5m of the robot (circle, with an arrow for its heading), with the frontal half-space of the sonar shaded in blue and the complete route in light grey. The middle row shows the energy image E and the bottom row the spectral-shape image S of the local view (red: above the mean level of the range bin, blue: below), each with the left ear above the right ear. Panels a) and b) show the same place in a corridor of the indoor floor, visited 641m of travel apart; panel c) shows the view of another place, 7.7m away, that is most similar to b) among all earlier pulses of the drive; panel d) shows the pillar hall of the indoor floor, and panel e) a street in the city. The similarity s is 0.89 between a) and b), and 0.69 between b) and c).
Fig. 4: Directivity of the simulated sonar (Lambert azimuthal equal-area projection of the frontal hemisphere, seen from behind the head; grid lines every 30\SIUnitSymbolDegree ). Rows: emitter, left ear, right ear and the ERTFs of both ears, at five frequencies between 30kHz and 90kHz , each normalized to its maximum.
Fig. 5: The simulated worlds and exemplary results of BatSLAM 2.0. Top row: the ground-truth drive (colored from start to end) in a) the indoor floor ( 38m by 18m , development drive on route 3, 964m ), b) the city ( 46m by 26m ) and c) the small development hall ( 16m by 10m ). Black crosses: the two copies of the aliasing pattern. Bottom row, d)–f): the final map estimated by BatSLAM 2.0 for the same drives with calibrated odometry (blue), with its trajectory error (ATE); wrong links in the final graph are drawn in red.
Fig. 6: Single-view place recognition for four sensors on the city drive: a) top-1 accuracy, b) AUC. Grey: descriptor of the original BatSLAM; orange: the same with TVG; blue: the final descriptor.
Fig. 7: Capture region of local view templates on the indoor floor, averaged over 16 template poses: fraction of recognized offset views for combined a) along-track and heading offsets, and b) lateral and heading offsets.
Fig. 8: Loop closure matrix of the development drive (route 3, calibrated odometry). Every dot is a link in the final graph between a query pulse and the anchor of the matched template. Grey: true same-direction revisits; blue: correct links; red: wrong links.
Fig. 9: Robustness against systematic odometry errors on the development drive. a) Trajectory error of odometry (orange), BatSLAM 2.0 (blue) and BatSLAM 2.0 after alignment (green) versus the bias factor b ; b) wrong hypotheses in the final graph and percentage of collapsed pairs.
Fig. 10: The real-world recording. a) Occupancy grid of the hall from the lidar scans, with the ground-truth trajectory (colored by time); the left loop was driven twice in the same direction, the right loop once in each direction. b) Camera images at places A, B and C. c) Binaural cochleograms of the closest sonar pulse, formed with the virtual ears of equation 16 .
Fig. 11: Real-world results. a) Map of one run (refitted evidence curve) with odometry, ground truth, BatSLAM 2.0 estimate and links. b) Similarity of correct and wrong candidates, with the simulated (dashed) and refitted (solid) evidence curves.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Parameter
Value
Front-end
Filterbank
64 Gaussian bands, 30kHz90kHz , log-spaced, width of two band spacings
Band pooling
Pairs of bands (32 bands)
Range bins
6cm , from 0.25m to 10m (162 bins)
Time-varying gain
Amplitude ×r
Energy image
Cube root of the amplitude, relative to the peak of the view
Spectral-shape image
±15dB scaled to ±1 , only above −40dB
Appendix
TABLE I: Parameters of BatSLAM 2.0. The same values were used for every experiment in this paper.
Indoor floor
City
Variant
Top-1
AUC
Top-1
AUC
Cues
Energy
0.875
0.915
0.878
0.918
Energy + shape (ours)
0.913
0.945
0.888
0.938
Energy + shape + ILD
0.892
0.968
0.878
0.960
Compression of the energy image (energy + shape)
Appendix
TABLE II: Single-view place recognition for the design choices of the descriptor, with the design sensor. All variants use 32 channels, 6cm range bins and the TVG. Our descriptor uses a cube-root compression. Best value per column in bold. ILD: interaural level difference.
ATE [m]
Condition (runs)
Odometry
Ours
Aligned
Wrong
Coverage
Calibrated (6)
4.2–11.1
0.47–1.91
0.18–0.46
0/230
0.79–0.84
Residual bias (6)
11.0–19.0
0.62–1.65
0.23–0.53
0/227
0.79–0.84
Degraded sensors (3)
7.5
0.52–1.79
0.25–0.81
1/106
0.67–0.74
Uncalibrated (4)
19.1–22.3
1.38–12.68
0.45–2.96
0/149
0.62–0.81
Appendix
TABLE III: Results of BatSLAM 2.0 per condition (ranges over the runs). ATE: trajectory error of odometry and of BatSLAM 2.0, without and after rigid alignment; wrong: wrong committed hypotheses in the final graph, out of all committed hypotheses; coverage: fraction of true revisits with a correct link.
Variant
ATE [m]
Wrong
Coverage
Full system (ours)
0.87
0
0.81
Naive
18.90
364
0.72
No sequence verification
4.85
10
0.80
Original front-end (8 bands)
1.54
0
0.73
No time-varying gain
2.53
3
0.72
Energy image only
1.38
8
0.74
Appendix
TABLE IV: Ablations over four standard cases (routes 3 and 5 with calibrated odometry, route 7 with a residual bias, and the city). ATE: mean over the cases; wrong: wrong committed hypotheses in the final graphs, summed over the cases. Naive: no sequence verification and no link management.
Evidence curve
ATE
Aligned
Links
Precision
Wrong
Revisits
Collapsed
( m )
( m )
Hyp.
Linked
Pairs
Odometry only
3.62 (0.46)
Simulation evidence curve
0.77 (0.07)
0.60 (0.05)
108
1.000 (0.000)
0
0.50 (0.01)
0.000
Refitted, even stretches
0.67 (0.13)
0.52 (0.11)
141
0.980 (0.018)
0
0.64 (0.10)
0.000
Refitted, odd stretches
0.66 (0.12)
0.52 (0.11)
142
0.975 (0.021)
0
0.64 (0.10)
0.000
Appendix
TABLE V: Real-world results, mean (standard deviation) over ten odometry seeds (distance ×0.95 , rotation ×1.10 ).
Fig. 12: Trajectories of all runs of the main experiment (table VI ), sorted by the odometry condition. Grey: ground truth; orange: odometry; blue: BatSLAM 2.0. Run names: r3 , r5 , r7 : routes 3, 5 and 7 on the indoor floor; city : the city world; b0 , b025 , b1 : bias factor b=0 , 0.25 and 1 ; s1 – s3 : odometry noise realization; noise5e-6 , noise1.5e-5 , oldsensor : degraded sensors.
Fig. 13: Trajectories of all runs of the robustness experiment (table VII ). bias b _s k : bias factor b with odometry realization k ; odomnoise_x3 : three times the random odometry noise; the last three runs use degraded sensors with a residual bias. Colors as in figure 12 .
Case
Odometry
ATE odo
ATE
Aligned
Precision
Wrong (final)
Wrong (all)
Coverage
Collapsed
[m] ↓
[m] ↓
[m] ↓
↑
↓
↓
↑
[%] ↓
Indoor, route 3, seed 1
Calibrated
7.5
0.78
0.25
0.999
0/39
0
0.84
0.0
Indoor, route 3, seed 2
Calibrated
4.2
0.89
0.24
0.999
0/40
0
0.84
0.0
Indoor, route 3, seed 3
Calibrated
7.9
1.02
0.19
1.000
0/36
1
0.84
0.0
Indoor, route 5
Calibrated
10.1
0.47
0.23
0.998
0/45
0
0.82
0.0
Indoor, route 7
Calibrated
7.6
1.91
0.18
1.000
0/46
0
0.80
0.0
Appendix
TABLE VI: Results of BatSLAM 2.0 on all runs of the main experiment (drives of about 1km ). ATE odo: dead reckoning; ATE and aligned: BatSLAM 2.0 without and after rigid alignment; precision: fraction of correct links; wrong (final / all): wrong committed hypotheses in the final graph and over all commits; coverage: fraction of true revisits with a correct link; collapsed: fraction of collapsed pairs.
Case
Bias
ATE odo
ATE
Aligned
Precision
Wrong (final)
Wrong (all)
Coverage
Collapsed
[m] ↓
[m] ↓
[m] ↓
↑
↓
↓
↑
[%] ↓
Indoor, route 3, seed 4
b=0
8.0
1.22
0.30
1.000
0/38
1
0.82
0.0
Indoor, route 3, seed 5
b=0
7.6
0.34
0.20
0.999
0/40
1
0.83
0.0
Indoor, route 3, seed 4
b=0.25
15.9
1.98
0.33
0.999
0/38
1
0.83
0.0
Indoor, route 3, seed 5
b=0.25
20.7
0.45
0.24
0.999
0/37
2
0.83
0.0
Indoor, route 3, sensor −20 dB
b=0.25
11.0
1.98
1.43
0.988
1/33
3
0.65
0.2
Appendix
TABLE VII: Robustness runs: odometry bias factor b , three times the random odometry noise, and degraded sensors with a residual bias. Columns as in table VI .
Variant
−10 dB sensor
−20 dB sensor
Original sensor
Route 3, seed 3
Wrong (final)
Wrong (all)
No risk-scaled commit
0.26 (0)
0.24 (0)
1.06 (1)
0.19 (0)
1
10
Risk-scaled commit, length only (24 pairs)
0.26 (0)
0.24 (0)
0.25 (0)
0.19 (0)
0
5
Plausibility gate on the absolute marginal
0.26 (0)
0.23 (0)
0.69 (1)
0.19 (0)
1
12
No plausibility gate
0.26 (0)
1.37 (2)
0.60 (1)
0.19 (0)
3
19
No length rule
0.26 (0)
0.81 (2)
1.07 (3)
0.20 (2)
7
7
No link management
0.26 (0)
0.81 (2)
1.07 (3)
0.20 (2)
7
7
Appendix
TABLE VIII: Ablation of the safeguards on the four hardest cases (three degraded sensors and route 3 with odometry seed 3). Every cell lists the aligned trajectory error (m) and, in parentheses, the number of wrong hypotheses in the final graph; the last two columns sum the wrong hypotheses in the final graphs and over all commits.
Variant
Route 3
Route 5
Route 7, residual bias
City
Naive (single views, no gates, no link mgmt.)
8.00 (63)
26.70 (68)
8.22 (72)
32.68 (161)
No sequence verification (single views)
5.08 (1)
1.95 (1)
8.17 (6)
4.21 (2)
Original front-end (8 wide bands)
1.97 (0)
0.44 (0)
3.15 (0)
0.59 (0)
No time-varying gain
0.85 (1)
6.82 (2)
1.84 (0)
0.62 (0)
Energy image only (no shape)
0.63 (0)
1.70 (3)
1.26 (4)
1.92 (1)
No range-shift search
0.45 (0)
1.57 (0)
3.28 (0)
0.56 (0)
Appendix
TABLE IX: Ablations per case: trajectory error (ATE, in m) and the number of wrong committed hypotheses in the final graph (in parentheses), for route 3 and route 5 with calibrated odometry, route 7 with a residual bias, and the city with calibrated odometry.