The rapid growth of 3D Gaussian Splatting (3DGS) research demands significant effort to reimplement papers before building on them. We introduce SPLATIFY, a multi-agent framework that converts 3DGS papers into trainable gsplat-based implementations, where generic paper-to-code methods and frontier models fail. SPLATIFY achieves this through five innovations: (1) A context-free grammar for gsplat over a modular method template with extension points for losses, densification, rendering, and optimization, constraining synthesis so generated code satisfies gsplat's architectural invariants by construction. (2) Architectural elements for faithful reproduction: fork-aware citation recovery retrieving component-level code at function-level granularity, Graph-of-Thought synthesis in topological dependency order, RAG-guided in-context example selection from over 20 verified implementations, and visual feedback combining PSNR-guided regeneration, Gaussian-level structural checks, and VLM-driven patching. (3) Knowledge-driven compositional improvement that autonomously finds weaknesses and composes complementary regularizers, losses, and densification strategies to improve upon original results. (4) Interdisciplinary method discovery where agents retrieve physical priors from outside the 3DGS literature and compose them with rendering knowledge to produce methods for previously unaddressed scene types. (5) SPLATIFY-Bench, an evaluation framework across 30 diverse 3DGS papers. On papers without public code, SPLATIFY matches expert implementations while reducing development time from weeks to minutes, and through compositional discovery further improves PSNR by up to 2.4 dB. We additionally demonstrate novel methods for volumetric nebula rendering and other scientific domains, synthesized entirely by SPLATIFY.
Figures & tables
Figure 2 : Simplified SPLATIFY pipeline. A high-level view of reproduction, refinement, and interdisciplinary method discovery. Fig. 3 shows the detailed architecture.
Figure 3 : Splatify overview. Reproduction (left): a parsing agent structures papers into a CFG-grounded knowledge base; fork-aware dependency resolution recovers components via citation graphs and repository diffs; GoT synthesis generates code in topological order with PSNR-guided and VLM-driven refinement. Knowledge discovery and invention (right): the populated knowledge base enables compositional improvement, where agents compose techniques to surpass original results, and interdisciplinary method creation, where physical priors from outside 3DGS literature are retrieved to address novel scene types.
Paper
Scene
Reported
Expert Impl.
Splatify (Ours)
PSNR ↑
SSIM ↑
LPIPS ↓
#GS ( 106 )
PSNR ↑
SSIM ↑
LPIPS ↓
#GS ( 106 )
PSNR ↑
SSIM ↑
LPIPS ↓
#GS ( 106 )
DWTGS [ 39 ]
mipnerf360
19.75
0.590
0.343
–
17.92
0.508
0.471
1.21
18.68
0.510
0.374
1.641
Dec. Densification [ 24 ]
mipnerf360
27.79
0.826
0.214
1.469
26.85
0.792
0.286
1.44
27.17
0.820
0.261
0.650
BOGausS [ 43 ]
mipnerf360
27.54
0.804
0.251
0.32
27.25
0.801
0.276
0.81
27.34
0.809
0.263
0.440
Opti3DGS [ 17 ]
mipnerf360
27.09
0.809
0.237
1.327
27.51
0.816
0.221
6.18
27.71
0.841
0.209
1.406
CLoD-GS [ 11 ]
Truck
25.29
0.880
0.148
1.46
24.83
0.854
0.227
0.40
25.48
0.873
0.128
0.384
Table 1 : Splatify vs. reported results and human implementations on Set 1. All other baselines failed to produce trainable code for any paper. Full results in the supplementary.
Figure 4 : Qualitative comparison of Splatify against baselines on Set 1 papers. Rendered novel views from Splatify and expert human implementation. All baselines fail to produce trainable code; Splatify closely matches expert quality.
Metric
Paper2Code
AutoP2C
GPT-5.2
Gemini 3.1 Pro
DeepSeek 3.2
NERFIFY
SPLATIFY (Ours)
Imports Resolve
✓
×
✓
✓
✓
✓
✓
Compiles / Trainable
×
×
∼
∼
×
✓
✓
Training Stability
×
×
×
×
×
×
✓
Gaussian Invariants
×
×
×
×
×
×
✓
Converges to Paper Results
×
×
×
×
×
∼
✓
Table 2 : Executability comparison. All baselines fail to produce trainable code. ∼ denotes occasional success that never reaches training stability or paper-level results.
Paper
Scene
Original Repository
Splatify (Ours)
PSNR ↑
SSIM ↑
LPIPS ↓
#GS ( 106 )
PSNR ↑
SSIM ↑
LPIPS ↓
#GS ( 106 )
3DGS [ 28 ]
Bicycle
25.246
0.771
0.205
–
26.162
0.765
0.359
0.964
LP3DGS [ 66 ]
Bicycle
24.906
0.737
0.264
2.51
24.639
0.740
0.189
1.310
GNS [ 13 ]
Truck
26.520
0.899
0.119
0.60
25.970
0.891
0.095
0.490
Micro-Splatting [ 31 ]
Kitchen
34.100
0.960
0.011
0.67
33.145
0.935
0.099
1.620
CompGS [ 38 ]
Counter
28.467
0.895
0.222
0.385
28.283
0.894
0.153
0.665
Table 3 : Set 2: Splatify vs. original author repositories. Metrics from original papers compared against Splatify ’s automated gsplat implementations trained under identical conditions. Full results in the supplementary.
Paper2Code
Gemini 3.1 Pro
DeepSeek 3.2
GPT-5.2 Thinking
Splatify
Paper
C↑
I↓
M↓
W↑
Score ↑
C↑
I↓
M↓
W↑
Score ↑
C↑
I↓
M↓
W↑
Score ↑
C↑
I↓
M↓
W↑
Score ↑
C↑
I↓
M↓
W↑
Score ↑
TrickGS [ 1 ]
0.35
0.30
0.35
0.40
0.41
0.50
0.20
0.30
0.50
0.51
0.65
0.15
0.20
0.60
0.64
0.60
0.15
0.25
0.65
0.62
0.90
0.05
0.05
0.92
0.94
2DGS [ 23 ]
0.85
0.08
0.07
0.50
0.55
0.55
0.15
0.30
0.40
0.40
0.80
0.10
0.10
0.55
0.58
0.75
0.10
0.15
0.60
0.61
1.00
0.00
0.00
0.95
0.95
Scaffold-GS [ 34 ]
0.35
0.20
0.45
0.60
0.34
0.45
0.10
0.45
0.45
0.44
0.40
0.35
0.25
0.40
0.42
0.60
0.25
0.15
0.55
0.59
1.00
0.00
0.00
1.00
1.00
GaussianShader [ 26 ]
0.45
0.30
0.25
0.40
0.40
0.55
0.28
0.17
0.45
0.50
0.60
0.22
0.18
0.48
0.58
0.72
0.15
0.13
0.62
0.73
1.00
0.00
0.00
0.95
1.00
Pixel-GS [ 67 ]
0.25
0.45
0.30
0.60
0.40
0.40
0.10
0.50
0.80
0.46
0.55
0.15
0.30
0.80
0.67
0.65
0.10
0.25
0.85
0.72
1.00
0.00
0.00
1.00
1.00
Table 4 : Novelty coverage on Set 3. C : correct, I : incorrect/partial, M : missing, W : hyperparameter fidelity, Score: weighted semantic score (0–1). Full results in the supplementary.
Before
After
Paper
PSNR ↑
SSIM ↑
LPIPS ↓
#GS
PSNR ↑
SSIM ↑
LPIPS ↓
#GS
Things Added
Opti3DGS [ 17 ]
27.71
.841
.209
1.40
28.05
.853
.125
4.95
Adaptive densif. threshold + opacity entropy reg.
MSAA 3DGS [ 55 ]
25.35
.758
.369
2.73
26.19
.791
.293
2.96
Bilateral grid + losses
GNS [ 13 ]
25.97
.891
.095
0.49
26.15
.902
.097
0.61
AA rasterization + denser init + post-densif. losses
DWTGS [ 39 ]
18.68
.510
.374
1.64
21.09
.725
.281
2.43
Patch structure loss + opacity/normal/depth reg.
CompGS [ 38 ]
28.28
.894
.153
0.67
28.50
.899
.146
0.63
EMA centroid updates + grad-based codebook refine
Table 5 : Compositional discovery on Set 1, Set 2 papers. “Before” is the reproduction-only baseline; “After” is the best validated configuration found by the discovery pipeline. All numbers on same scene as Tables 1 and 3 .
Figure 5 : Cross-domain comparison. Ground truth, 3DGS, and agent-generated method for Nebula and Bubble. NebulaGS resolves opaque shell artifacts; BubbleGS captures iridescent color variation. Additional scenes in the supplementary.
Configuration
Score
Train.
C
(%)
Splatify (Full)
0.97
100
1.00
Knowledge Sources:
w/o CFG Constraints (Stage 1)
0.91
70
0.85
w/o In-context Examples (Stage 1)
0.73
100
1.00
w/o Fork-aware Recovery (Stage 2)
0.64
100
0.70
Table 6 : Component ablation. Impact of each component on synthesis quality, averaged over 10 papers from SplatifyBench .