Sparse-View Interpretable 3D Animal Behavior Representations for Neural Encoding and Decoding
Authors: Xinming Dai, Qihang Jin, Tianshu Tan, Baiyuan Chen, Hanrui Lyu, Lenny Aharon, Kyle Daruwalla, Xun Helen Hou, +3 more
Organizations: Columbia University · University of Science and Technology of China · Harvard University · University of Cambridge · Northwestern University · Cold Spring Harbor Laboratory
A deeper understanding of brain function requires a precise, structured characterization of behavior. Yet, extracting behavioral representations from video in a form suitable for scientific analysis remains a fundamental challenge. Many prior studies represent behavior via pose estimation or nonlinear video embeddings. However, pose tracking discards rich information beyond predefined keypoints, while nonlinear video embeddings lack interpretability. We address this limitation with SABLE (Sparse-view Animal Behavior Latent Embeddings), a self-supervised framework that leverages a geometric inductive bias to learn behavior representations. By augmenting a multi-view transformer with priors from monocular depth and pose estimation, SABLE reconstructs 3D animal behavior from extremely sparse views while learning explicit 3D latent structure. Without ground-truth 3D labels, it reliably recovers 3D behavior from two-view videos, whereas state-of-the-art (SOTA) methods fail or yield degenerate solutions. Across the International Brain Lab and Cheese3D datasets, we demonstrate that SABLE learns 3D representations that match or exceed prior SOTA performance in neural encoding and decoding. Once pretrained across animals, SABLE serves as an off-the-shelf model that generalizes zero-shot to unseen animals without animal-specific calibration or retraining. Our method establishes 3D-aware video embeddings that capture complex behavior, opening new avenues for studying brain-behavior relationships.
Figures & tables
Figure 1: Geometric inductive bias enables sable to reliably reconstruct 3D animal behavior from sparse-view IBL data, where existing 3D foundation models fail to recover coherent 3D structure. Existing approaches struggle to recover coherent 3D structure in this challenging two-camera setting. Depth Anything 3 (DA3) ( Lin et al., 2025 ) produces plausible monocular point clouds from depth estimation but fails to infer camera poses for cross-view fusion. Visual Geometry Grounded Transformer (VGGT) ( Wang et al., 2025 ) , applied zero-shot, collapses to a degenerate single-view 2D shortcut solution, and E-RayZer ( Zhao et al., 2026 ) exhibits a similar failure mode; even after finetuning, it produces overlapping, flattened point clouds. In contrast, sable learns accurate camera parameters, avoids degenerate shortcut solutions, and reconstructs coherent 3D mouse structure by incorporating geometric inductive bias. Without the corresponding geometric loss (“w/o geom. loss”), reconstruction degenerates.
Figure 2: sable framework. A multi-view transformer predicts camera poses from two input views, which are converted into Plücker ray maps. These rays are concatenated with image tokens from a DINOv3 encoder that capture semantic information, masked with learnable mask tokens, and processed by a second multi-view transformer to produce 3D-structured embeddings that parameterize a Gaussian splatting decoder for 3D reconstruction. To prevent shortcut solutions, we impose geometric regularization via pseudo-point clouds derived from monocular depth and aligned across views using animal pose keypoints as correspondences. The predicted Gaussian point cloud is projected using the estimated camera poses and rendered into 2D images. Training is guided by a masked view reconstruction loss Lrecon for pixel-level accuracy, a perceptual loss Lpercep for perceptual similarity, and a geometric regularization term Lgeom that constrains the 3D locations of the predicted point clouds.
Figure 3: Benchmarking neural decoding of 2D video frames and 3D point clouds from neural activity. (A) For each video frame, we train a TCN decoder to predict image patch embeddings from neural activity, which are then passed to a Gaussian decoder to generate 3D point clouds and render 2D images. (B) Decoding performance is evaluated across baseline models that learn video-derived behavioral embeddings. ResNet AE compresses image features extracted with a ResNet-18 backbone into latent vectors, while BEAST and sable use image patch embeddings produced by a ViT-MAE and a multi-view transformer, respectively. sable outperforms all baselines in video frame decoding quality, as measured by PSNR and SSIM, which evaluate pixel-level accuracy and perceptual quality. Error bars indicate the standard error of the mean across 14,400 frame pairs from held-out sessions. (C) Comparison of decoded video frames from ResNet AE, BEAST, and sable against the ground truth. (D) The Gaussian point clouds decoded from neural activity capture fine-grained paw movements.
Figure 4: Quantitative and qualitative comparison of neural activity prediction using behavioral features extracted from IBL video data. (A) At each time point, we concatenate the DINOv3 [CLS] tokens from the two camera views, then use a TCN over time to predict neural activity. (B) Encoding performance is evaluated across baseline models that learn video-derived behavioral embeddings. Behavioral keypoints provide interpretable summaries derived from sensors or animal pose tracking. PCA learns linear latent representations from video, while ResNet AE learns nonlinear latents. BEAST uses [CLS] tokens from the ViT-B backbone of a ViT-MAE. sable outperforms all baselines in neural prediction quality measured by bits per spike (BPS). We also include a random-embedding baseline as a control, confirming it yields near-zero BPS. Error bars indicate the standard error of the mean. (C) Scatterplot comparison of sable vs BEAST performance in two example sessions. Each dot corresponds to an individual neuron. The values in the bottom-right corner represent the session-averaged BPS. Although the improvement in BPS is modest, sable outperforms BEAST in prediction quality for most neurons. (D) Comparison of predicted firing rates averaged over randomly sampled 1s recording segments for eight example neurons using sable and BEAST. The values in the top-left corner represent per-neuron BPS. (E) Comparison of residual firing-rate variability (heatmaps), obtained by subtracting each neuron’s mean firing rate across recording segments, for two example neurons (columns) using sable and BEAST.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Cheese3D video frames and 3D reconstruction comparison. (A) Six camera views from the Cheese3D dataset. (B) 3D reconstruction from the selected sparse view pair (top-right and top-left) of Cheese3D. sable reliably reconstructs 3D animal behavior, while existing 3D foundation models exhibit distinct failure modes, consistent with the results on IBL (Fig. 1 ). DA3 ( Lin et al., 2025 ) and VGGT ( Wang et al., 2025 ) produce plausible per-view point clouds; however, DA3 fails to consistently align the two point clouds, with noticeable misregistration around the nose, and VGGT fails to disentangle foreground and background in depth, leaving background geometry intermingled with the mouse. Both zero-shot and fine-tuned E-RayZer ( Zhao et al., 2026 ) collapse to degenerate, flattened 3D structures. In contrast, sable recovers consistent cross-view geometry, separates foreground from background in depth, and avoids degenerate solutions, producing a coherent 3D reconstruction of the mouse.
Figure 6: Novel-view synthesis on Cheese3D. Both models take only the top-left and top-right views as input and are evaluated on the held-out top-center view. (A) Ground truth. (B) sable rendering. (C) E-RayZer rendering. sable better matches the ground truth, indicating more accurate 3D reconstruction.
PSNR
View
All Black Image
E-RayZer
sable
Top-Left
11.31
23.69 ± 0.33
29.74 ± 0.06
Top-Right
11.88
28.10 ± 0.36
28.78 ± 0.09
Top-Center (Held-Out)
7.35
10.74 ± 0.10
12.74 ± 0.06
SSIM
View
All Black Image
E-RayZer
sable
Appendix
Table 1: Novel-view synthesis results on Cheese3D, comparing E-RayZer and sable on input views and a held-out top-center view.
ResNet AE
BEAST
sable
Encoding (bps)
0.603 ± 0.144
0.452 ± 0.088
0.655 ± 0.146
Appendix
Table 2: Neural encoding results on Cheese3D.
Figure 7: SABLE with varying numbers of input views. (A) On IBL, SABLE is fine-tuned using a cross-view prediction objective, where one view is used as input to predict the other view and itself. At inference time, SABLE reconstructs both views from a single input view. (B) On Cheese3D, we restrict the input to 3 of the available 6 views. For both settings, the second row visualizes the reconstructed 3D point clouds from different viewing angles.
Method
Decoding (PSNR) ↑
Decoding (SSIM) ↑
Encoding (bps) ↑
ResNet-AE-18
21.013 ± 0.015
0.853 ± 0.000
0.166 ± 0.005
ResNet-AE-152
19.520 ± 0.017
0.840 ± 0.000
0.152 ± 0.005
BEAST (ViT Base)
21.659 ± 0.009
0.858 ± 0.000
0.165 ± 0.005
BEAST (ViT Large)
21.210 ± 0.008
0.852 ± 0.000
0.102 ± 0.003
sable
22.355 ± 0.014
0.893 ± 0.000
0.168 ± 0.006
Appendix
Table 3: Neural encoding (bits per spike) and decoding (PSNR, SSIM) across model sizes on the IBL dataset.
Variant
Lrecon
Lpercep
Lgeom
PSNR ↑
Ours
1.0
0.3
1.0
26.0
No Lgeom
1.0
0.3
0.0
26.9
High Lrecon
2.0
0.3
1.0
26.1
Low Lpercep
1.0
0.1
1.0
25.9
No Lpercep
1.0
0.0
1.0
25.7
High Lpercep
1.0
1.0
1.0
24.7
Appendix
Table 4: Ablation on loss weighting.
Hyperparameter
Value
Input image size
320×320
Patch Size
16
Number of Image Tokens
400
Embedding Dimension
768
Number of Attention Heads
12
Image Encoder Depth
16
Appendix
Table 5: sable model hyperparameters.
Hyperparameter
Value
Input image size
224×224
Image preprocessing
ImageNet normalization
Backbone
ResNet-18
Residual block configuration
(2,2,2,2)
Embedding Dimension
768
Bottleneck mapping
512×7×7→768→512×7×7
Appendix
Table 6: ResNet autoencoder model hyperparameters.
Hyperparameter
Value
Input image size
224×224
Patch size
16
Embedding Dimension
768
Encoder layers
12
Attention heads
12
MLP intermediate size
3072
Appendix
Table 7: BEAST model hyperparameters.
Hyperparameter
Value
Input dimension
Neuron count (decoding) or behavioral feature dimension (encoding)
Multi-view video recordings are increasingly used to capture the 3D movements of animals in experimental settings, yet extracting rich 3D representations from these recordings remains challenging. Supervised pose estimation requires extensive manual annotation, while general-purpose 3D reconstruction models trained on generic scene datasets fail on the specialized imagery and sparse-view setting of laboratory experiments. We address these limitations with BEAST3D, a self-supervised pretraining framework that learns 3D visual representations from unlabeled, calibrated multi-view video. BEAST3D uses a vision transformer to predict 3D Gaussian splats that reconstruct held-out views through differentiable rendering, while simultaneously segmenting the animal from the background. BEAST3D reconstructs 3D structure with as few as four views by conditioning directly on known camera parameters--unlike general-purpose models, which must estimate camera geometry from dense overlapping viewpoints that are seldom available in lab settings. Through comprehensive evaluation across four species, we demonstrate that BEAST3D produces rich, viewpoint-invariant features that transfer effectively to three downstream tasks: novel view synthesis, which validates the quality of the learned 3D representations; multi-view pose estimation, which provides the sparse keypoint trajectories widely used in behavioral analysis; and neural encoding, which relates 3D behavioral features to simultaneously recorded neural activity. BEAST3D thus establishes a versatile framework for behavioral analysis that leverages 3D structure in modern multi-view laboratory recordings.
Yanchen Wang, Lenny Aharon, Wangshu Zhu +7
Columbia University · Cold Spring Harbor · Stanford University
We present Casper3D, a lightweight probabilistic framework for converting noisy multi-view 2D foundation-model embeddings into a latent 3D semantic representation. We model view-level semantic features as noisy observations of an underlying 3D semantic state and infer this state with a set-based variational model that incorporates relative pose during multi-view reasoning. Casper3D is trained by predicting held-out semantic observations from novel viewpoints, while remaining aligned with visual and text semantic spaces for open-vocabulary 3D understanding. The framework is backbone-agnostic and applies to both language-aligned and self-supervised embeddings. Experiments show that Casper3D produces more stable 3D semantics than simple multi-view pooling, especially in ambiguous and noisy settings.
Marwane Hariat, Gianni Franchi, David Filliat +1
U2IS, ENSTA – Institut Polytechnique de Paris, Palaiseau, France · Pôle Recherche, Agence Ministérielle pour l’IA de Défense, Palaiseau, France
3D animal reconstruction in the wild remains challenging due to large species variation, frequent occlusions, and the prevalence of multi-animal scenes, while existing methods predominantly focus on single-animal settings. We present SAM 3D Animal, the first promptable framework for multi-animal 3D reconstruction from a single image. Built on the SMAL+ parametric animal model, our method jointly reconstructs multiple instances and supports flexible prompts in the form of keypoints and masks which enable more reliable disambiguation in crowded and occluded scenes. To train such a model, we further introduce Herd3D, a multi-animal 3D dataset containing over 5K images, designed to increase diversity in species, interactions, and occlusion patterns. Experiments on the Animal3D, APTv2, and Animal Kingdom datasets show that our framework achieves state-of-the-art results over both existing model-based and model-free methods, demonstrating a scalable and effective solution for prompt-driven animal 3D reconstruction in the wild.
Xuyi Hu, Jin Lyu, Jiuming Liu +4
University of Cambridge · 2Southern University of Science and Technology · 3Tsinghua University +1