Object permanence, keeping track of an object's identity and position while it is occluded, is central to video representations that track, predict and plan. Trackers that achieve it learn from boxes, track identities and visibility labels. On the other hand, self-supervised object-centric methods discover objects without labels: through slot attention, it represents a video as slots that bind to objects and follow them across frames. However, these slots are lost under occlusion, making the desired permanence impossible. Reasoning permanence is a hard problem because it requires to detect when an object becomes occluded, re-identify when object reappears, and keep the object's hidden position continuous, using reapperance as the only learning cue. To address this, we propose iSEE, a novel framework that offers all three aforementioned requirements, without any labels whatsoever. We built iSEE using the following three proposed components: (i) Object evidence modelling: a slot's attention, compared with its own past, reveals when its object is hidden. (ii) Appearance-position separation: two slot streams let the appearance be held for re-identification while the position keeps changing. (iii) Permanence from reappearance: a walker follows the hidden object's position, trained only on where the object reappears. On LA-CATER static, iSEE returns a reappearing object to its own slot after 86% of occlusions, against 32% for SlotContrast, and localises it while hidden within 4.1 mAP of the label-trained SoTA RAM. The two streams also allow downstream planning, with the position stream as the action of a world model. Project page: https://insait-institute.github.io/iSEE/
Figures & tables
Figure 1: iSEE achieves self-supervised object permanence, close to a supervised tracker. (a) Slot k (green) holds the ball; dashed, its mask predicted with iSEE during the occlusion. A slot’s evidence is the attention it receives from the frame’s patches: iSEE’s clearly signals the occlusion and the reappearance. iSEE’s slot is separated into appearance and position; during the occlusion the appearance is held and the position keeps walking, learnt only from where it reappears (red ring). (b) On LA-CATER (static), each arrow adds one of our contributions to the existing baseline, SlotContrast ( Manasyan et al., 2025 ) .
Figure 2: (a) The two-stream slot in training. (b) TEN measures a slot’s evidence on frame t against the largest evidence the slot had before t . (a) An appearance predictor and a position predictor advance the two streams a and p of every slot; the position predictor reads a through a stop-gradient (sg). The two-stream slot attention updates a from the appearance grid ht and p from the position grid c (Eq. 3 ). The decoder places each object from p alone, and Lssc acts on a only (Sec. 3.1 ). (b) For slot k , etk is the mean of its largest attention values on frame t (Eq. 4 ), and e^tk=etk/rt−1k (Eq. 5 ). The slot switches to held , its appearance kept and its position walked (Sec. 3.2 ), when e^tk falls below τ , and back to updated when it rises above ρ (Eq. 6 ).
Table 3
Figure 4: Object permanence in LA-CATER: a cone is put over the target and carries it across the frame. The target is visible only in the first and last columns. White: its amodal mask. Red: the followed slot, solid for its own mask, dashed where moved onto the predicted position; no red: the slot owns no pixel there. Only iSEE tracks the target while it is hidden.
LA-CATER, static camera
LA-CATER, moving camera
KITTI
requires
0.2–1.0 s
1.0–2.1 s
2.1–4.2 s
≥ 4.2 s
0.2–1.0 s
1.0–2.1 s
2.1–4.2 s
≥ 4.2 s
0.1–2.2 s
RAM
boxes, identities, visibility
90.2
95.5
85.7
88.7
95.0
90.0
77.3
69.5
87.1
Loci-Looped
per-frame background
53.1
33.3
24.5
25.4
39.5
31.1
27.3
21.1
–
RandSF.Q
none
56.1
22.7
16.3
21.1
38.1
23.3
16.7
15.8
–
SlotContrast
none
63.4
27.3
20.4
23.9
30.2
11.7
12.1
13.7
29.0
iSEE
none
95.1
86.4
91.8
77.5
82.0
81.7
75.8
64.2
48.4
Table 2: Re-identification after an occlusion, LA-CATER & KITTI. SAME ↑ (%) by the length of the occlusion in seconds. requires : what the row is given beyond the video, at training or at test time. With no label at any stage iSEE is ahead of every row but the supervised RAM in every bin, and ahead of RAM as well on the 0.2–1.0 s and 2.1–4.2 s occlusions under the static camera.
Figure 5: Occlusions in KITTI: a car passes behind another road user and reappears. Unevenly spaced frames of one tile, cropped to the car and its occluder. White: the car’s visible mask. Green: the slot on the car before the occlusion; cross: the walker’s predicted centre while the car is hidden. Both models’ evidence falls while the car is hidden; only iSEE’s recovers, and its slot returns.
MOVi-C
LA-CATER static
LA-CATER moving
KITTI
method
features
FG-ARI ↑
mBO ↑
FG-ARI ↑
mBO ↑
FG-ARI ↑
mBO ↑
FG-ARI ↑
mBO ↑
SlotContrast
DINOv2
66.9
32.0
94.2
21.5
83.1
13.7
42.4
12.4
SlotContrast + TEN
DINOv2
68.1
31.4
94.5
22.3
86.7
14.7
39.4
12.3
iSEE
DINOv2 NoPE
71.6
36.0
96.6
23.9
89.4
17.8
48.9
22.9
Table 3: Object discovery on MOVi-C, LA-CATER and KITTI. Video FG-ARI ↑ / video mBO ↑ ( ×100 ) on decoder masks over whole clips. Splitting the slot costs no discovery.
Table 8
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
stage
what is trained
signal
what is frozen
1. encoder
gψ , corrector, predictors, decoder
Eq. 2
backbone f
2. TEN
nothing
none
everything
3. walker
the walker of Eq. 7
Eq. 8 , on the positions and holds of stages 1–2
encoder, P∗
Appendix
Table S1: The three stages of iSEE. At inference, TEN and the walker run together, and the walker’s position is written into the held slots. No box, mask, visibility or occlusion label is used at any stage.
static camera
moving camera
test clips read
1,200
1,200
scorable occlusions
695
546
in clips
526
434
target moves ≥ 10 px while hidden, all scored
145
338
target moves less, sampled
50 of 550
50 of 208
occlusions scored
195
388
Appendix
Table S2: From test clips to scored occlusions, LA-CATER. The first 1,200 test clips of each camera. Displacement is that of the target’s amodal centre, between the first hidden frame and the furthest point of the occlusion.
Figure S1: Permanence against the IoU threshold, LA-CATER static camera. mAP@ t↑ ( ×100 ). Grey band : the thresholds 0.10 to 0.30 that Table 1 averages. Open markers : rows given more than the video. ceiling : the share of occlusions on which iSEE’s box overlaps the ground-truth box with IoU >t on the last frame before the occlusion, where the box is on the target by construction. The ceiling falls above t=0.3 because the box is cut from a 37×37 patch-grid mask, so beyond it the metric reads the resolution of the grid, not the predicted position.
mAP@ [0.1,0.3]↑ , occlusion length (frames)
5–25
25–50
50–100
≥ 100
n 50, median 7
n 24, median 38
n 50, median 77
n 71, median 150
RAM
86.0
92.1
82.6
83.5
frozen
63.9
63.3
53.1
48.1
Loci-Looped
69.3
55.7
49.6
37.1
SlotContrast
10.8
9.4
3.2
4.1
Appendix
Table S3: Permanence against the length of the occlusion, LA-CATER static camera. mAP@ [0.1,0.3]↑ ( ×100 ) over all hidden frames of the 195 occlusions, split by length into four disjoint bins, with the rows and boxes of Table 1 . frozen : the last visible position; RAM is given boxes, track identities and visibility labels, Loci-Looped a per-frame background. iSEE keeps its permanence as the occlusion lengthens, while frozen and Loci-Looped fall. Bold: best row that reads the video alone, which excludes frozen , RAM and Loci-Looped; n and median length in frames per bin in the header.
occluded objects
all objects
HOTA ↑
DetA ↑
AssA ↑
HOTA ↑
DetA ↑
AssA ↑
static camera (150 clips; 224 occluded objects in 127 clips; 1,088 objects)
SlotContrast
71.5
83.5
61.3
87.8
92.8
83.2
+ TEN
81.3
82.6
80.0
89.5
92.0
87.1
RandSF.Q
42.3
61.8
29.0
35.1
48.1
25.9
Loci-Looped
54.8
52.7
57.2
70.3
68.8
72.2
Appendix
Table S4: Identity over whole clips (TrackEval). HOTA, DetA and AssA ↑ ( ×100 ) on the first 150 test clips of each camera, protocol above. RandSF.Q: its decoder’s masks (last-layer cross-attention, head-averaged, softmax over slots). Loci-Looped: the argmax over its decoded masks with its background channel as one more track, at every second frame, at the checkpoints of Sec. B.3 .
code
dim
within-object R2
DINOv2 ( Oquab et al., 2024 )
384
0.833
DINOv2 NoPE ( Pawlowsky et al., 2026 )
384
0.552
SlotContrast slot
64
0.597
iSEE appearance a
60
0.525
iSEE position p
4
0.727
Appendix
Table S5: Position is carried by the position stream, MOVi-C. Within-object R2 of an object’s modal centre. Feature rows are each model’s own frozen backbone pooled over the object’s cells, and are the floor for the code that reads them. Position is carried by the 4-dimensional position stream, and the appearance stream reads it below the floor.
contrastive loss on
SAME ↑
FG-ARI ↑
mBO ↑
mAP all ↑
mAP hidden ↑
R2 of p↑
appearance stream (60)
82.0
96.6
23.9
69.0
30.1
0.937
whole slot (64)
83.1
94.7
28.5
25.6
13.0
0.693
Appendix
Table S6: Contrastive loss on the appearance stream or on the whole slot. SAME, discovery, permanence mAP and the within-object R2 of the position stream for iSEE trained with the contrastive loss on the appearance stream only (60 dimensions) or on the whole slot (64 dimensions), LA-CATER, static camera.
position predictor
mAP all ↑
with sg(a)
69.0
without sg(a)
66.4
Appendix
Table S7: Appearance input of the position predictor, LA-CATER, static camera. Permanence mAP@ [0.1,0.3] over all frames of the 195 occlusions, with TEN and without the walker.
Figure S2: The action-conditioned dynamics model. Frozen slots [ak;pk] enter OCVP-Par ( Villar-Corrales et al., 2023 ) ; per slot, the position change pt+1k−ptk passes a two-dimensional bottleneck and is injected into every block by adaptive layer normalisation. The action reaches the predictor only through the position stream.
Figure S3: Object discovery on MOVi-C.
Figure S4: Object discovery on KITTI.
Figure S5: Re-identification: one occlusion in a LA-CATER scene, the target hidden behind another object and reappearing. Columns are frames, the first the last one on which the target is visible and the last the frame Table 2 scores, in one crop of the frame that is the same throughout. GT : the target’s true (amodal) mask, in white on every row, dashed on the frames where none of it is visible. Green: the mask of the slot that owns the target in the first column, empty where that slot claims nothing in view; iSEE’s claims nothing anywhere while the target is hidden, because TEN holds it. Only iSEE has that slot back on the target when it reappears.
Figure S6: A qualitative example on KITTI. Columns are frames of one tile, unevenly spaced; the whole 2× tile the model sees. GT : the car’s visible mask in white, its KITTI box dashed where none of it is visible. Green: the slot on the car before the occlusion. iSEE’s evidence falls while the car is hidden and its slot is back on the car when it reappears; SlotContrast’s slot moves onto the occluding cyclist.
Figure S7: Qualitative comparative example on LA-CATER, static camera , against SlotContrast, RandSF.Q and Loci-Looped.
Figure S8: Qualitative comparative example on LA-CATER, moving camera , against SlotContrast, RandSF.Q and Loci-Looped.
Figure S9: TEN releases the slot while the target is still hidden. Columns are frames of one clip, unevenly spaced, in one crop held fixed through the sequence; the first is the last frame on which the target is at least half visible. GT : the target’s visible mask in white, its amodal box dashed where none of it is visible. Green: the slot on the target before the occlusion; orange: the slot holding the target after.
Figure S10: Two ways iSEE loses a car through a KITTI occlusion. Columns are frames of one tile, unevenly spaced. GT : the car’s visible mask in white, its KITTI box dashed where none of it is visible. Green: the slot on the car before the occlusion.
Figure S11: Five ways iSEE loses the target through a LA-CATER occlusion. GT : the target’s true (amodal) mask, in white. Green: the slot iSEE follows, solid where it is that slot’s own mask and dashed where that outline is moved onto the predicted position; no green where the slot owns no pixel.