Organizations: Department of Mechanical and Industrial Engineering, University of Toronto, Canada · Department of Mechanical, Industrial and Mechatronics Engineering, Toronto Metropolitan University, Canada · Department of Computer Science, University College London, UK
In robotics, scene representation plays a pivotal role in understanding and interacting with the environment. The advent of Neural Radiance Fields (NeRF) and its variants, as a novel representation, has opened a new frontier of research. In applications such as semantic mapping and simulation, roboticists aim to build scenes using multiple NeRF models, each representing an object. While extensive datasets of 3D mesh models already exist, there is an urgent need to develop tools to convert these assets to NeRF models for rapid algorithm development and testing. This paper presents a new pipeline for converting existing mesh models to NeRF representations by artificially generating a ground truth point-based radiance field through sampling mesh geometry and texture. This approach alleviates the need for camera-based sampling or rendering multi-view images of the original mesh to train the NeRF model. Extensive benchmarking demonstrates that our method yields comparable rendering quality to the baselines. Additionally, the application of this representation is shown by constructing unified NeRF scenes and performing collision simulations with extracted geometry.
Figures & tables
Figure 1 : We present NeRFifyMesh, an architecture and pipeline for converting textured meshes into Neural Radiance Fields (NeRFs). We generate a ground-truth radiance field from the mesh, encompassing color, opacity, point normals, and their respective weighting terms. NeRFifyMesh supports scene rendering, normal prediction, surface recovery, and scene composition.
Figure 2 : An overview of the NeRFifyMesh network architecture. Input points x are queried across multi-resolution surface voxel grids with hash-encoded features. Interpolated features are decoded with network fψ and combined with the original input to predict opacity through network fζ . Positional encoding of the input is provided to networks fθ and fϕ to predict colors and normals, respectively.
Method
GSO [ 13 ]
Poly Haven [ 17 ]
Objaverse [ 12 ]
PSNR ↑
LPIPS ↓
SSIM ↑
PSNR
LPIPS
SSIM
PSNR
LPIPS
SSIM
Ours
31.98
0.024
0.977
35.11
0.034
0.965
31.70
0.037
0.958
I-NGP
37.00
0.027
0.982
38.52
0.052
0.973
35.62
0.043
0.965
NeRF
34.12
0.054
0.965
36.86
0.083
0.956
32.20
0.083
0.927
Table I : Quantitative comparison of rendering performance metrics across different datasets using our method, Instant-NGP (I-NGP) [ 24 ] , and NeRF [ 23 ] .
Figure 3 : Visual comparison of mesh scene rendering results for our method, Instant-NGP [ 24 ] , and NeRF [ 23 ] across the GSO [ 13 ] , Poly Haven [ 17 ] , and Objaverse [ 12 ] datasets. Our method represents fine textural and geometric details with high fidelity.
Model
GSO [ 13 ]
Poly Haven [ 17 ]
Objaverse [ 12 ]
Size (MB)
Comp.
Size (MB)
Comp.
Size (MB)
Comp.
Mesh
12.4 ± 2.0
15.2 ± 4.9
67.3 ± 60.6
NeRF [ 23 ]
5.0
2.5x
5.0
3.0x
5.0
13.5x
I-NGP [ 24 ]
26.4
0.5x
26.4
0.6x
26.4
2.5x
Ours
4.4 ± 0.7
2.8x
4.9 ± 0.9
3.2x
4.7 ± 0.2
14.4x
Table II : Comparison of average model sizes and compression ratios across datasets.
Figure 4 : Relighting comparison on the Objaverse horse-conch scene [ 12 ] . Since our representation predicts surface normals, it supports relighting via a Blinn-Phong illumination model. Top row : renders from our representation. Bottom row : ground-truth mesh rendered under identical lighting. Columns show progressively composed lighting terms: (a) original scene without altered lighting, (b) ambient only, (c) ambient + diffuse, (d) ambient + diffuse + specular at low intensity, and (e) the same with high specular intensity.
Figure 6 : Scene composition results, RGB and Depth renderings. (left) A scene formed from 8 NeRFifyMesh models of objects from GSO [ 13 ] and Polyhaven [ 17 ] . (right) Intercomposition of a NeRF Garden scene (Mip-NeRF 360 dataset) [ 3 ] with a NeRFifyMesh model of a croissant from the Poly Haven dataset [ 17 ] .
Figure 7 : Visual comparison of collision simulations between the VTech Roll and Learn Turtle and Thomas the Train from GSO [ 13 ] . Both objects are initialized with a velocity of 1.4 m/s on a ground plane with a friction coefficient μk = 0.8. Simulations are run for 2 seconds at 60 FPS with a physics step of 200 Hz. We compare results from our method to a simulation formed using Instant-NGP [ 24 ] .
Figure 8 : Collision simulation in a composite scene combining our representation of a football from Poly Haven [ 17 ] and a static NeRF Garden scene (Mip-NeRF 360 dataset) [ 3 ] . The football was dropped from a height of 0.78m above the table in the scene.
Object
Method
Pos. Error (m)
Rot. Error (deg)
Turtle [ 13 ]
Instant-NGP [ 24 ]
1.73e-3
1.04
Ours
7.31e-4
0.54
Thomas [ 13 ]
Instant-NGP [ 24 ]
1.10e-2
20.34
Ours
1.06e-2
5.99
Table V : Collision simulation accuracy comparison between Instant-NGP [ 24 ] and our method. Position error is reported as mean Euclidean error, rotation error as mean quaternion angular difference.
Figure 9 : Limitation of our method. For non-watertight meshes, inaccuracies in the SDF calculation leads to the presence of artifacts and results in poor-quality renders.
In this paper, we present NEO, a unified framework providing language-guided NeRF editing for robotic manipulation. Our paper introduces (i) a language-guided object removal that combines neural field resampling with multiview-consistent progressive inpainting, (ii) a direct NeRF weight editing method utilizing knowledge distillation, composing original and edited NeRFs via a teacher-student model, enabling coherent modeling of future scene states before a robot executes an action, and (iii) the first benchmark (NEO-Dataset) for quantitatively evaluating NeRF scene editing methods suitable for robot manipulation. We show that our approach outperforms state-of-the-art baselines in scene editing tasks, including object removal and pick-and-place robotic experiments, yielding visually coherent and geometrically consistent edits that reduce artifacts commonly introduced by prior methods.
Mikołaj Zieliński, David Hall, Dominik Belter +1
Institute of Robotics and Machine Intelligence, Poznan University of Technology, 61-131 Pozna´n, Poland · CSIRO Robotics, CSIRO, Australia
Recent advances in neural rendering have introduced numerous 3D scene representations. Although standard computer vision metrics evaluate the visual quality of generated images, they often overlook the fidelity of surface geometry. This limitation is particularly critical in robotics, where accurate geometry is essential for tasks such as grasping and object manipulation. In this paper, we present an evaluation pipeline for neural rendering methods that focuses on geometric accuracy, along with a benchmark comprising 19 diverse scenes. Our approach enables a systematic assessment of reconstruction methods in terms of surface and shape fidelity, complementing traditional visual metrics.
Reconstructing dynamic surgical scenes is crucial for robot-assisted minimally invasive surgery; however, it continues to be difficult because of tissue deformation, occlusions, specular reflections, and restricted viewpoints. In this study, we introduce Endo-NeRF++, a neural rendering framework that accounts for uncertainty in the reconstruction of dynamic surgical scenes. Expanding on EndoNeRF, the suggested approach incorporates multi-resolution hash-grid encoding, temporal feature merging, and uncertainty-informed adaptive sampling to enhance reconstruction accuracy and temporal coherence in deformable endoscopic scenes.The multi-resolution hash-grid representation within the framework effectively captures both coarse and fine anatomical details, while temporal feature blending ensures stable reconstruction during tissue deformation and surgical tool occlusions. Additionally, uncertainty-driven adaptive sampling assigns more samples to uncertain areas to enhance rendering quality and geometric coherence. Experiments on robotic surgical video sequences demonstrate that the proposed uncertainty-guided adaptive sampling improves PSNR by up to 1.22dB (4.3%), increases SSIM by up to 5.3%, and reduces LPIPS by up to 55.1% compared with the EndoNeRF baseline.
Gousia Habib, Laura Ruotsalainen
Department of Computer Science, University of Helsinki, Finland