Sparse-View Interpretable 3D Animal Behavior Representations for Neural Encoding and Decoding
Authors: Xinming Dai, Qihang Jin, Tianshu Tan, Baiyuan Chen, Hanrui Lyu, Lenny Aharon, Kyle Daruwalla, Xun Helen Hou, +3 more
Organizations: Columbia University · University of Science and Technology of China · Harvard University · University of Cambridge · Northwestern University · Cold Spring Harbor Laboratory
A deeper understanding of brain function requires a precise, structured characterization of behavior. Yet, extracting behavioral representations from video in a form suitable for scientific analysis remains a fundamental challenge. Many prior studies represent behavior via pose estimation or nonlinear video embeddings. However, pose tracking discards rich information beyond predefined keypoints, while nonlinear video embeddings lack interpretability. We address this limitation with SABLE (Sparse-view Animal Behavior Latent Embeddings), a self-supervised framework that leverages a geometric inductive bias to learn behavior representations. By augmenting a multi-view transformer with priors from monocular depth and pose estimation, SABLE reconstructs 3D animal behavior from extremely sparse views while learning explicit 3D latent structure. Without ground-truth 3D labels, it reliably recovers 3D behavior from two-view videos, whereas state-of-the-art (SOTA) methods fail or yield degenerate solutions. Across the International Brain Lab and Cheese3D datasets, we demonstrate that SABLE learns 3D representations that match or exceed prior SOTA performance in neural encoding and decoding. Once pretrained across animals, SABLE serves as an off-the-shelf model that generalizes zero-shot to unseen animals without animal-specific calibration or retraining. Our method establishes 3D-aware video embeddings that capture complex behavior, opening new avenues for studying brain-behavior relationships.
Figures & tables
Figure 1: Geometric inductive bias enables sable to reliably reconstruct 3D animal behavior from sparse-view IBL data, where existing 3D foundation models fail to recover coherent 3D structure. Existing approaches struggle to recover coherent 3D structure in this challenging two-camera setting. Depth Anything 3 (DA3) ( Lin et al., 2025 ) produces plausible monocular point clouds from depth estimation but fails to infer camera poses for cross-view fusion. Visual Geometry Grounded Transformer (VGGT) ( Wang et al., 2025 ) , applied zero-shot, collapses to a degenerate single-view 2D shortcut solution, and E-RayZer ( Zhao et al., 2026 ) exhibits a similar failure mode; even after finetuning, it produces overlapping, flattened point clouds. In contrast, sable learns accurate camera parameters, avoids degenerate shortcut solutions, and reconstructs coherent 3D mouse structure by incorporating geometric inductive bias. Without the corresponding geometric loss (“w/o geom. loss”), reconstruction degenerates.
Figure 2: sable framework. A multi-view transformer predicts camera poses from two input views, which are converted into Plücker ray maps. These rays are concatenated with image tokens from a DINOv3 encoder that capture semantic information, masked with learnable mask tokens, and processed by a second multi-view transformer to produce 3D-structured embeddings that parameterize a Gaussian splatting decoder for 3D reconstruction. To prevent shortcut solutions, we impose geometric regularization via pseudo-point clouds derived from monocular depth and aligned across views using animal pose keypoints as correspondences. The predicted Gaussian point cloud is projected using the estimated camera poses and rendered into 2D images. Training is guided by a masked view reconstruction loss Lrecon for pixel-level accuracy, a perceptual loss Lpercep for perceptual similarity, and a geometric regularization term Lgeom that constrains the 3D locations of the predicted point clouds.
Figure 3: Benchmarking neural decoding of 2D video frames and 3D point clouds from neural activity. (A) For each video frame, we train a TCN decoder to predict image patch embeddings from neural activity, which are then passed to a Gaussian decoder to generate 3D point clouds and render 2D images. (B) Decoding performance is evaluated across baseline models that learn video-derived behavioral embeddings. ResNet AE compresses image features extracted with a ResNet-18 backbone into latent vectors, while BEAST and sable use image patch embeddings produced by a ViT-MAE and a multi-view transformer, respectively. sable outperforms all baselines in video frame decoding quality, as measured by PSNR and SSIM, which evaluate pixel-level accuracy and perceptual quality. Error bars indicate the standard error of the mean across 14,400 frame pairs from held-out sessions. (C) Comparison of decoded video frames from ResNet AE, BEAST, and sable against the ground truth. (D) The Gaussian point clouds decoded from neural activity capture fine-grained paw movements.
Figure 4: Quantitative and qualitative comparison of neural activity prediction using behavioral features extracted from IBL video data. (A) At each time point, we concatenate the DINOv3 [CLS] tokens from the two camera views, then use a TCN over time to predict neural activity. (B) Encoding performance is evaluated across baseline models that learn video-derived behavioral embeddings. Behavioral keypoints provide interpretable summaries derived from sensors or animal pose tracking. PCA learns linear latent representations from video, while ResNet AE learns nonlinear latents. BEAST uses [CLS] tokens from the ViT-B backbone of a ViT-MAE. sable outperforms all baselines in neural prediction quality measured by bits per spike (BPS). We also include a random-embedding baseline as a control, confirming it yields near-zero BPS. Error bars indicate the standard error of the mean. (C) Scatterplot comparison of sable vs BEAST performance in two example sessions. Each dot corresponds to an individual neuron. The values in the bottom-right corner represent the session-averaged BPS. Although the improvement in BPS is modest, sable outperforms BEAST in prediction quality for most neurons. (D) Comparison of predicted firing rates averaged over randomly sampled 1s recording segments for eight example neurons using sable and BEAST. The values in the top-left corner represent per-neuron BPS. (E) Comparison of residual firing-rate variability (heatmaps), obtained by subtracting each neuron’s mean firing rate across recording segments, for two example neurons (columns) using sable and BEAST.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Cheese3D video frames and 3D reconstruction comparison. (A) Six camera views from the Cheese3D dataset. (B) 3D reconstruction from the selected sparse view pair (top-right and top-left) of Cheese3D. sable reliably reconstructs 3D animal behavior, while existing 3D foundation models exhibit distinct failure modes, consistent with the results on IBL (Fig. 1 ). DA3 ( Lin et al., 2025 ) and VGGT ( Wang et al., 2025 ) produce plausible per-view point clouds; however, DA3 fails to consistently align the two point clouds, with noticeable misregistration around the nose, and VGGT fails to disentangle foreground and background in depth, leaving background geometry intermingled with the mouse. Both zero-shot and fine-tuned E-RayZer ( Zhao et al., 2026 ) collapse to degenerate, flattened 3D structures. In contrast, sable recovers consistent cross-view geometry, separates foreground from background in depth, and avoids degenerate solutions, producing a coherent 3D reconstruction of the mouse.
Figure 6: Novel-view synthesis on Cheese3D. Both models take only the top-left and top-right views as input and are evaluated on the held-out top-center view. (A) Ground truth. (B) sable rendering. (C) E-RayZer rendering. sable better matches the ground truth, indicating more accurate 3D reconstruction.
PSNR
View
All Black Image
E-RayZer
sable
Top-Left
11.31
23.69 ± 0.33
29.74 ± 0.06
Top-Right
11.88
28.10 ± 0.36
28.78 ± 0.09
Top-Center (Held-Out)
7.35
10.74 ± 0.10
12.74 ± 0.06
SSIM
View
All Black Image
E-RayZer
sable
Appendix
Table 1: Novel-view synthesis results on Cheese3D, comparing E-RayZer and sable on input views and a held-out top-center view.
ResNet AE
BEAST
sable
Encoding (bps)
0.603 ± 0.144
0.452 ± 0.088
0.655 ± 0.146
Appendix
Table 2: Neural encoding results on Cheese3D.
Figure 7: SABLE with varying numbers of input views. (A) On IBL, SABLE is fine-tuned using a cross-view prediction objective, where one view is used as input to predict the other view and itself. At inference time, SABLE reconstructs both views from a single input view. (B) On Cheese3D, we restrict the input to 3 of the available 6 views. For both settings, the second row visualizes the reconstructed 3D point clouds from different viewing angles.
Method
Decoding (PSNR) ↑
Decoding (SSIM) ↑
Encoding (bps) ↑
ResNet-AE-18
21.013 ± 0.015
0.853 ± 0.000
0.166 ± 0.005
ResNet-AE-152
19.520 ± 0.017
0.840 ± 0.000
0.152 ± 0.005
BEAST (ViT Base)
21.659 ± 0.009
0.858 ± 0.000
0.165 ± 0.005
BEAST (ViT Large)
21.210 ± 0.008
0.852 ± 0.000
0.102 ± 0.003
sable
22.355 ± 0.014
0.893 ± 0.000
0.168 ± 0.006
Appendix
Table 3: Neural encoding (bits per spike) and decoding (PSNR, SSIM) across model sizes on the IBL dataset.
Variant
Lrecon
Lpercep
Lgeom
PSNR ↑
Ours
1.0
0.3
1.0
26.0
No Lgeom
1.0
0.3
0.0
26.9
High Lrecon
2.0
0.3
1.0
26.1
Low Lpercep
1.0
0.1
1.0
25.9
No Lpercep
1.0
0.0
1.0
25.7
High Lpercep
1.0
1.0
1.0
24.7
Appendix
Table 4: Ablation on loss weighting.
Hyperparameter
Value
Input image size
320×320
Patch Size
16
Number of Image Tokens
400
Embedding Dimension
768
Number of Attention Heads
12
Image Encoder Depth
16
Appendix
Table 5: sable model hyperparameters.
Hyperparameter
Value
Input image size
224×224
Image preprocessing
ImageNet normalization
Backbone
ResNet-18
Residual block configuration
(2,2,2,2)
Embedding Dimension
768
Bottleneck mapping
512×7×7→768→512×7×7
Appendix
Table 6: ResNet autoencoder model hyperparameters.
Hyperparameter
Value
Input image size
224×224
Patch size
16
Embedding Dimension
768
Encoder layers
12
Attention heads
12
MLP intermediate size
3072
Appendix
Table 7: BEAST model hyperparameters.
Hyperparameter
Value
Input dimension
Neuron count (decoding) or behavioral feature dimension (encoding)