Recovering three-dimensional human geometry from electromagnetic measure?ments in a complex static environment is difficult because strong multipath responses from walls, floors, and other objects obscure the weak target per?turbation. We propose R1A-PC, a physics-guided method that reconstructs a 2048-point human cloud from paired complex fields measured with and without the target. Complex background subtraction emphasizes target-induced ampli?tude and phase changes, while the background field remains available as an environmental condition. A frequency-balanced discrete Born adjoint produces a three-dimensional spatial knowledge map. At each of two bounded deformation stages, the decoder combines complex measurement features, background fea?tures, and multiscale physical features queried at the current point coordinates; the second stage queries again after the first coordinate update. We analyze the residual of paired subtraction, the weighted normal-operator structure of the raw adjoint, and the feasible set of the predicted cloud. In a held-out background generated by full-wave simulation under a fixed acquisition geometry, R1A-PC obtains a squared Chamfer distance of 0.001434 m2 and an F-score of 0.963080 at 0.05 m. Compared with TopNet, the Chamfer distance decreases by 70.28%. Removing physical guidance or background subtraction increases the Chamfer distance by 242.06% or 241.18%, respectively. Experiments across background layouts and poses support the complementary roles of paired subtraction and position-dependent adjoint features.
Figures & tables
Figure 1: Three-dimensional electromagnetic acquisition geometry in a complex static environment. The env_004 example shows the DOI, four transmitters, twelve receiver locations, walls, static objects, and the human mesh; coordinates are in metres
Figure 2: R1A-PC architecture. Paired complex responses yield measurement, background, and adjoint spatial conditions. Multiscale queries at current point locations drive two bounded deformation stages to output a 2048-point human cloud
Figure 3: Two-stage point-cloud evolution for four fixed env_003 samples. Columns show standing, raised arm, bent torso, and side-facing poses; rows show stage P1, stage P2, and ground truth. All panels use the same coordinate limits, orthographic camera, point size, and height color scale
Environment
Static structure
Samples
Role
env_000
Floor
640
Training
env_001
Floor, two walls
640
Training
env_002
Floor, three objects
640
Training
env_003
Floor, two walls, four objects
570
Validation
env_004
Floor, two walls, four objects
500
Main test
env_005
Floor, two walls, four objects
70
Extended test
Table 1: Environment structure and data split. Counts are the samples included in each set.
Overall geometry
Threshold coverage
Location and tail error
Method
CD-L2 ↓
Root-CD ↓
F@0.02 ↑
F@0.05 ↑
Centroid ↓
H95 ↓
Direct Born adjoint
0.544516
0.737076
0.040135
0.162922
0.264834
0.975135
Data-only TopNet
0.004825
0.069099
0.188703
0.755998
0.155409
0.092107
R1A-PC
0.001434
0.037753
0.449782
0.963080
0.025276
0.045617
Table 2: Method comparison on the principal test environment env_004. Trainable methods report means over three random seeds. The best mean for each metric is bold and shaded. CD-L2 is in m 2 ; Root-CD, centroid error, and H95 are in m.
Figure 4: Qualitative comparison for four env_003 poses. Rows show standing, raised arm, bent torso, and side-facing poses; columns show ground truth, Born adjoint, data-only TopNet, and R1A-PC. Colour encodes height z . All panels use metric coordinates, common DOI limits, an orthographic camera, and the same point size
Figure 5: Seven extended poses in env_005. The upper row shows ground truth and the lower row R1A-PC. Columns show both arms raised, one arm extended forward, slight forward lean, standing lunge, natural sitting, forward-leaning sitting, and one-knee kneeling
Figure 6: Reconstruction-error distributions for 500 test targets in env_004. (a) Target-level empirical cumulative distributions of CD-L2 for the Born adjoint, data-only TopNet, and R1A-PC. (b) Target-level differences in CD-L2 between R1A-PC and TopNet; the zero line denotes equal errors
Reconstruction quality
Relative change
Method
CD-L2 ↓
F@0.05 ↑
H95 ↓
Δ CD-L2
Without physical guidance
0.004907
0.763553
0.092685
242.06
Without background subtraction
0.004894
0.763412
0.092544
241.18
Without cross-environment consistency
0.001459
0.962987
0.045780
1.7
R1A-PC
0.001434
0.963080
0.045617
0.0
Table 3: Component ablations. Values are means of three independent trainings; relative CD-L2 changes use complete R1A-PC as the reference. CD-L2 is in m 2 and H95 in m.
Overall accuracy
Location and tail error
Environment
CD-L2 ↓
F@0.05 ↑
Centroid ↓
H95 ↓
env_002
0.001210
0.976085
0.023693
0.041576
env_003
0.001592
0.950150
0.029864
0.049623
env_004
0.001434
0.963080
0.025276
0.045617
Table 4: R1A-PC reconstruction results on three background levels; means of three independent trainings. CD-L2 is in m 2 and centroid error and H95 are in m.
New-environment pose reconstruction
Method
CD-L2 ↓
F@0.05 ↑
H95 ↓
Direct Born adjoint
0.608902
0.143379
1.018
Data-only TopNet
0.008265
0.621074
0.114149
R1A-PC
0.004845
0.863498
0.092827
Table 5: Mean performance over seven pose groups in the new environment env_005. Trainable methods report means across three random seeds. CD-L2 is in m 2 and H95 in m.
Model size
Inference cost
Method
Parameters
Output points
Time ↓
Memory ↓
Direct Born adjoint
0
2048
29.803
140.3
Data-only TopNet
1029645
2048
2.164
16.8
R1A-PC
1114023
2048
6.460
68.3
Table 6: Computational efficiency on an NVIDIA GeForce RTX 4060 Laptop GPU. Times are mean milliseconds per sample after input construction; peak memory is in MiB.
Three-dimensional ultrasound (US) is a safe, radiation-free complementary modality to CT and X-rays for longitudinal monitoring, yet its segmentation-derived partial point clouds are extremely artifact-laden. Consequently, it is challenging to recover a clean and complete anatomical structure from such US point clouds. In this paper, we present UBone3D, a novel framework based on physics-rectified conditional flow matching (CFM) that performs point cloud completion directly from partial US observations. UBone3D models deterministic physics artifacts (e.g., surface thickening, streaking, dropouts) via a simulated physics proxy, and introduces test-time physics rectification to steer the shape completion. At inference, the completion is jointly steered by two decoupled forces: (1) anatomical plausibility enforced by a CT-trained generative shape prior, BoneFM, and (2) physics consistency enforced by USimNet in the ultrasound formation space. Extensive experiments on simulated and in-vivo data demonstrate significant improvements in reconstruction accuracy and anatomical fidelity over existing baselines.
GeRaF is the first method to use neural implicit learning for near-range 3D geometry reconstruction from radio frequency (RF) signals. Unlike RGB or LiDAR-based methods, RF sensing can see through occlusion but suffers from low resolution and noise due to its lensless imaging nature. While lenses in RGB imaging constrain sampling to 1D rays, RF signals propagate through the entire space, introducing significant noise and leading to cubic complexity in volumetric rendering. Moreover, RF signals interact with surfaces via specular reflections, requiring fundamentally different modeling. To address these challenges, GeRaF (1) introduces filter-based rendering to suppress irrelevant signals, (2) implements a physics-based RF volumetric rendering pipeline, and (3) proposes a novel lensless sampling and lensless alpha blending strategy that makes full-space sampling feasible during training. By learning signed distance functions, reflectiveness, and signal power through MLPs and trainable parameters, GeRaF takes the first step towards reconstructing millimeter-level geometry from RF signals in real-world settings.
We present an image-conditioned point cloud completion approach that treats images as the primary geometric source rather than a secondary guide. To this end, we introduce an Image-to-Point (I2P) module that can reconstruct complete point clouds directly from a single RGB image, with no need for 3D inputs. Additionally, we introduce a transformer-based Point-to-Point (P2P) refinement module that uses self- and cross-attention between point tokens and image features to iteratively refine the coarse I2P output. The I2P module enables the image encoder to learn rich geometric representations, while the P2P module progressively recovers fine-grained details. Unlike existing multimodal methods that rely on auxiliary losses or fusion modules, our explicit I2P task provides a strong, geometry-aware prior based on images alone. Extensive experiments on ShapeNet-ViPC demonstrate state-of-the-art completion performance with a 12.3% relative Chamfer Distance improvement over prior methods. Code is available at: https://github.com/AzharSindhi/I2PRef.git