Organizations: Department of Mechanical Engineering, Iowa State University, Ames, IA 50011, USA · Department of Architecture, Iowa State University, Ames, IA 50011, USA
Computational Fluid Dynamics (CFD) is widely used to evaluate ventilation and contaminant transport in occupied buildings, but deployment at scale is limited by three bottlenecks: acquiring room geometry without costly scanning hardware or manual CAD modeling, decomposing the scene into individually manipulable objects, and reconfiguring those objects for alternative layouts without re-capturing the room. We present a semi-automated workflow that converts a single 360-degree video of a room into individually editable, simulation-ready geometry assets. A dense point cloud is reconstructed using Neural Radiance Fields (NeRF), and 2D instance masks from text-prompted SAM 3 segmentation are lifted to 3D using multi-view consensus and depth-band filtering. Points are separated into object instances with an octree, and occlusion gaps are healed with a connectivity graph. Chair templates are fitted by Iterative Closest Point (ICP) alignment, and table geometry is generated procedurally. A browser-based editor supports quality assurance and rapid construction of alternative layout configurations. A steady Reynolds-averaged OpenFOAM solution then drives transient passive-scalar transport; the setup is verified using a mesh-sensitivity study and validated against an IEA Annex 20 benchmark. We apply the workflow to two university classrooms and a tiered lecture-hall auditorium. The capture-to-geometry pass takes two to five hours per room on a consumer workstation. In a controlled obstruction sequence in one classroom, the modeled half-clearance time varies non-monotonically as furniture is added, and a cross-room comparison indicates that clearance behavior cannot be reliably extrapolated between rooms, motivating per-room geometry acquisition. By making that acquisition low-cost, the workflow makes geometry-resolved comparative ventilation studies practical for spaces such as classrooms.
Figures & tables
Figure 1 : End-to-end workflow from 360-degree video capture to CFD simulation. The figure groups the five stages of Section 3 into three phases: (1) reconstruction (camera pose estimation and NeRF-based scene capture), (2) semantic processing (concept-prompted segmentation, multi-view filtering, and STL generation, with human-in-the-loop scene editing through a browser-based GUI), and (3) simulation (CFD analysis across multiple room configurations).
Environment
Duration (s)
Sampling
Equirect. frames
Pinhole images
COLMAP poses
Mixed-Furniture Classroom
61
1 per 22
84
672
664
Chair-Dominant Classroom
53
1 per 20
80
640
620
Auditorium
280
1 per 42
201
1,608
1,472
Table 1 : Video capture and frame processing statistics. Pinhole images = equirectangular frames × 8 views; COLMAP poses are the subset of pinhole images successfully registered.
Figure 2 : Raw dense point cloud extracted from the trained nerfacto model. (a, b) Interior views showing the reconstructed furniture layout with visible noise and floating artifacts. (c) Exterior view revealing the overall room envelope with scattered outlier points before semantic processing.
Figure 3 : Visual fidelity comparison. (a) Original photograph captured from the 360-degree camera. (b) NeRF point cloud rendered from the same camera viewpoint, illustrating the geometric and color fidelity of the reconstruction.
Figure 4 : SAM 3 text-prompted segmentation. (a) Pinhole-projected training image from the 360° video. (b) SAM 3 segmentation output with per-instance colored masks overlaid on the image, prompted with the text query chair . (c) Four individual binary masks extracted for separate chair instances; each isolates a single object for the subsequent 3D projection step.
Figure 5 : Category point cloud extraction via mask-based projection. (a) For each camera frame, 3D points are projected onto the image plane and checked against the SAM 3 binary mask; points falling inside the mask (chair points) are retained while surrounding background points are discarded. (b) Full NeRF-reconstructed point cloud of the classroom. (c) Resulting raw category point cloud after projecting across all frames and retaining only points labeled chair , with significant noise from background leakage and projection errors that motivates the multi-view consensus and depth-band filtering steps.
Figure 6 : Projection-based filtering logic. For each frame, points falling within the semantic mask receive a vote. Votes are accumulated across all N frames, and only points meeting the multi-view consensus threshold τmv proceed to depth-band filtering, which removes occluded background geometry.
Figure 7 : Multi-view consensus for point selection. (Left) 3D points ( A,B,C,D ) are projected into multiple camera frames; each frame’s semantic mask determines per-frame votes. (Center) Accumulated view counts for each point, with the consensus threshold τmv (labeled τ in the figure) indicated. (Right) Points meeting the threshold are retained as the object point cloud; points below the threshold are discarded as noise or background.
Figure 8 : Depth-band filtering. (Left) 3D scene with object points and background wall noise viewed from camera C . (Center) Image plane showing the semantic mask M and a selected patch q . (Right) Depth analysis for patch q : the median depth μq and spread σq define a depth band [μq−λσq,μq+λσq] . Points within the band (green) are retained; distant background points at zbg (red) that leak through the 2D mask are rejected. The patch statistics μq , σq , and multiplier λ shown in the figure correspond to dmed , MAD , and σdepth in the text.
Ground truth
Final reconstruction
Environment
Chairs
Tables
Chairs
Tables
Mixed-Furniture Classroom
32
17
32
16
Chair-Dominant Classroom
47
3
45
2
Auditorium
242
2
273
2
Table 2 : Ground-truth and final reconstruction furniture counts per environment. Ground truth was verified manually; final reconstruction counts are after the full pipeline including human QA correction.
Figure 9 : Individual object instances extracted via octree-based connected components from the category cloud. Quality of extracted instances depends on the underlying point cloud: (a, b) well-separated chairs yield clean single-object instances, (c) an isolated table is correctly extracted, and (d) two adjacent tables with insufficient spatial gap are merged into a single connected component, with a nearby chair fragment also included.
Figure 10 : Instance-processing pipeline: instance separation, noise filtering, geometric healing, size classification, and category-specific alignment producing the final CFD-ready assets.
Figure 11 : Effect of graph-based healing on a fragmented chair instance. (a) A front leg is disconnected from the main body due to occlusion gaps in the NeRF reconstruction, appearing as a separate connected component within its own bounding box. (b) After bridging gaps smaller than δgap , the leg is merged back into a single unified point cloud.
Figure 12 : Final geometric alignment results. Each column shows the healed point cloud (top) and the corresponding STL template overlay in green (bottom). (a) Failure case where multi-start ICP converges to a local minimum, placing the chair template approximately 90∘ from the correct orientation. (b, c) Successful chair alignments where the template closely matches the point cloud geometry. (d) Procedurally generated table mesh fitted via PCA-based orientation and z -axis ICP refinement, in close agreement with the scanned point cloud.
Figure 13 : Human-in-the-loop scene editing. (a) The browser-based GUI displays the reconstructed classroom with automatically placed furniture STLs, allowing the operator to inspect and modify placements; the editor also supports batch placement of seated mannequins on detected chairs through an adjustable occupancy slider. (b) Raw automated pipeline output with furniture only. (c) Scene after human editing: standing mannequins have been added and misaligned furniture corrected, producing an alternative simulation configuration without re-scanning the physical room.
Figure 14 : OpenFOAM CFD execution workflow, from reconstructed STL geometry to steady RANS airflow and transient turbulent-diffusivity passive-scalar flushing simulations, executed on the Frontera system at the Texas Advanced Computing Center (TACC).
Patch group
Type
Description
Inlet patches
Velocity inlet
Four ceiling supply diffusers (classroom cases); six ceiling supply diffusers (auditorium case)
Table 5 : Flow and passive-scalar transport models used in the OpenFOAM simulations.
Item
Setting
Time discretization
Euler
Gradients
Gauss linear
Scalar convection
Gauss upwind
Laplacian
Gauss linear corrected
p solver
PCG + DIC; tol 10−6 ; relTol 0.01
U,k,ε solver
PBiCGStab + DILU; tol 10−7 ; relTol 0.1
Table 6 : Numerical schemes and solver settings.
Parameter
Symbol
Chairs
Tables
Semantic segmentation and projection ( Section 3.3 )
SAM 3 segmentation confidence threshold
τconf
0.5
0.5
Depth patch size (pixels)
Spatch
32
48
Depth-band tolerance (MAD multiplier)
σdepth
0.75
1.25
Multi-view consensus threshold (votes)
τmv
10
5
Instance extraction and healing ( Section 3.4.2 )
Table 7 : Configurable pipeline parameters and the values used for the mixed-furniture classroom demonstration of Section 4 . Chair and table values are listed separately where category-specific tuning is applied; a small subset of thresholds is additionally adjusted per environment ( Section 5.2 ). Length-valued parameters are specified in the scale-normalized reconstruction frame, in which the room height is approximately one unit; physical lengths follow from the uniform metric scaling of Section 3.5.5 .
Environment
PSNR (dB) ↑
SSIM ↑
LPIPS ↓
Mixed-Furniture Classroom
24.51±2.53
0.819±0.066
0.370±0.129
Chair-Dominant Classroom
23.09±1.68
0.750±0.063
0.418±0.125
Auditorium
24.55±3.13
0.836±0.066
0.233±0.080
Table 8 : NeRF reconstruction quality metrics (mean ± std over held-out test views), computed with Nerfstudio’s ns-eval on the default nerfacto evaluation split of each capture. Higher PSNR and SSIM indicate better fidelity; lower LPIPS indicates better perceptual similarity.
Figure 15 : What-if reconstruction sequence for the mixed-furniture classroom. Rows show four progressively modified configurations: (i) the empty room enclosure, (ii) the room with 32 chairs, (iii) the room with 32 chairs and 16 tables, and (iv) the full furniture configuration with 25 mannequins (22 seated, 3 standing). Columns show, from left to right, the reconstructed STL assets overlaid on the NeRF point cloud, the STL assets with the point cloud removed, and the final simulation-ready geometry including the room enclosure. For the empty-room case, only the enclosure geometry is shown because no furniture or occupant assets are present.
Figure 16 : Steady-state velocity-magnitude fields for the four Mixed-Furniture Classroom configurations shown in three-dimensional views. Panels (a)–(d) correspond to configurations (i)–(iv): the empty room, +chairs, +chairs+tables, and +chairs+tables+mannequins. The shared colour bar indicates velocity magnitude in m/s. The panels show the steady velocity fields obtained for each configuration.
Figure 17 : Steady-state vertical mid-plane velocity-magnitude slices for the four Mixed-Furniture Classroom configurations. Panels (a)–(d) correspond to configurations (i)–(iv): the empty room, +chairs, +chairs+tables, and +chairs+tables+mannequins. The shared colour bar indicates velocity magnitude in m/s. The panels show the mid-plane velocity fields obtained for each configuration.
Figure 18 : Passive-scalar concentration evolution for the Mixed-Furniture Classroom what-if sequence. Rows correspond to the four configurations shown in Figure 15 : (i) room shell only, (ii) room shell with 32 chairs, (iii) room shell with 32 chairs and 16 tables, and (iv) room shell with 32 chairs, 16 tables, and 25 mannequins (22 seated, 3 standing). Columns show matched scalar-field snapshots at t=5 s, 60 s, and 300 s after the start of the scalar-decay window. All panels are rendered from the same viewpoint and use the same normalized concentration scale, C/C0∈[0,1] , shown by the shared color bar.
Figure 19 : Mixed-furniture classroom what-if sequence: normalized passive-scalar concentration histories for the four configurations. The horizontal reference line marks C/C0=0.5 , corresponding to the t50 half-clearance threshold. The curves show non-monotonic variation of the monitor-based clearance across the four configurations; additional metrics are reported in Table 10 .
Figure 20 : Chair-dominant classroom reconstruction and CFD-analysis summary. The top row shows (a) an input 360° frame, (b) the reconstructed NeRF point cloud, and (c) the simulation-ready room geometry. Panels (d)–(f) show three views of the steady velocity-magnitude field, all sharing the velocity-magnitude scale (m/s) indicated beside each panel, and panel (g) shows a representative passive-scalar concentration field with its corresponding normalized concentration scale. Panel (h) shows the normalized scalar concentration decay history.
Figure 21 : Tiered lecture-hall auditorium reconstruction and CFD-analysis summary. The top row shows (a) an input 360° frame, (b) the reconstructed NeRF point cloud, and (c) the simulation-ready auditorium geometry with reconstructed furniture assets. Panels (d) and (e) show two steady velocity-magnitude views through the room interior, both sharing the velocity-magnitude scale (m/s) indicated beside each panel. Panel (f) shows the passive-scalar concentration field with its normalized concentration scale, and panel (g) shows normalized scalar concentration decay histories at front, middle, and back monitor locations, illustrating spatial variation in auditorium clearance.
Table 9 : Final production CFD cases. Mesh counts were extracted from checkMesh run in parallel on the final OpenFOAM case folders. Airflow endpoint denotes the final iteration of the steady simpleFoam solve used to initialize scalar transport; scalar window denotes elapsed scalar-decay time after initialization from the steady airflow field.
Table 10 : Monitor-based scalar-clearance metrics computed from normalized concentration histories. t50 , t80 , and t90 denote the first elapsed times at which C/C0 falls below 0.5, 0.2, and 0.1, respectively. AUC is the time integral of C/C0 over the simulated decay window of duration T , and C/C0(T) is the normalized concentration at the end of that window. A dash indicates that the threshold was not reached within the simulated window.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Mesh level
Cells
Points
∫Cds (m)
Coarse
2.15M
2.65M
0.03322
Medium
5.99M
7.36M
0.03473
Fine
15.91M
18.82M
0.04648
Extra-fine
55.37M
62.85M
0.04654
Appendix
Table 1 : Mesh-sensitivity summary for the fully obstructed mixed-furniture classroom configuration using Sct=0.7 . The scalar integral was computed over the common valid portion of the vertical occupied-zone sampling line, from approximately z=0.05 m to z=2.00 m above the floor.
Figure 1 : Mesh-sensitivity assessment for the mixed-furniture classroom (fully obstructed) using the turbulent scalar-transport framework with Sct=0.7 . (a) Vertical profile of normalized scalar concentration, C/C0 , along the occupied-zone sampling line for the coarse, medium, fine, and extra-fine meshes. Because the extra-fine mesh produced valid samples over a shorter portion of the requested vertical line, all four profiles and line integrals are compared over the common valid interval, approximately z=0.05 – 2.00 m above the floor. (b) Integrated scalar concentration, ∫Cds , computed over the same interval and plotted against mesh size. The meshes contain approximately 2.15 M, 5.99 M, 15.91 M, and 55.37 M cells, respectively.
Case
Lx (m)
Ly (m)
H model (m)
H reference (m)
Difference
Mixed-furniture: empty
5.71
7.22
2.90
2.87
+0.03 m ( +1.0% )
Mixed-furniture: +chairs
5.71
7.22
2.90
2.87
+0.03 m ( +1.0% )
Mixed-furniture: +tables
5.71
7.22
2.90
2.87
+0.03 m ( +1.0% )
Mixed-furniture: full
5.71
7.22
2.90
2.87
+0.03 m ( +1.0% )
Chair-dominant: complex
6.04
6.29
2.90
2.87
+0.03 m ( +1.0% )
Auditorium
30.01
28.98
12.87
12.88
−0.01 m ( −0.08% )
Appendix
Table 2 : Metric dimensions of the final OpenFOAM room geometries and physical room-height references used for global scale assignment. The in-plane dimensions Lx and Ly are extracted from the scaled computational geometry. The AR Ruler height measurements establish the metric scale and are therefore not treated as independent validation of reconstruction accuracy.
Figure 2 : Room-height reference measurements using the AR Ruler mobile-phone application. These heights establish the metric scale of the reconstructed OpenFOAM domains and are therefore not an independent check of reconstruction accuracy; they are not survey-grade measurements.
Figure 3 : Comparison of normalized streamwise velocity profiles predicted using coarse, medium, and fine meshes with digitized measurement data from Nielsen [57] for the IEA Annex 20 benchmark room at x/H=1.0 (left) and x/H=2.0 (right). The experimental points were digitized from the published figure and therefore represent approximate measurement locations rather than original tabulated data.
School of Electrical Engineering & Robotics, Queensland University of Technology, Australia · Imaging and Computer Vision Group, CSIRO Data61, Australia
Computer Science, School of Digital and Physical Sciences, University of Hull, HU6 7RX, UK · The Institute of Artificial Intelligence, Xiamen University, Xiamen, China · Department of Bioengineering, University of Washington, Seattle, WA, USA +1