Impact sound rendering synthesizes the sound produced when a 3D object is struck, but practical renderers often rely on fixed material presets such as wood, plastic, or steel. These presets limit the range of impact sounds a renderer can express, while manually adjusting the underlying material parameters remains difficult without expertise in material acoustics. We therefore study inverse impact sound rendering: predicting material parameters from a reference impact sound so that a simulator can recreate a similar material response. To support this task, we introduce ImpactMat, a dataset and benchmark of single and blended material impact sounds paired with ground-truth material parameters. We further propose a feed-forward model that predicts these parameters from one or more recordings, using blended materials to learn smooth transitions between material types. Experiments show that our method outperforms competitive baselines and enables re-rendering from real recordings without manual parameter tuning. The project page is available https://material-from-impact.github.io/material-from-impact/.
Figures & tables
Figure 1 : Inverse impact sound rendering from reference audio. Given a reference impact sound from an object with unknown material, our continuous material estimator predicts five renderer-compatible parameters, which are then used with a target mesh for physics-based sound synthesis. This enables re-rendering with a similar material response on a new object.
Figure 2 : Overview of the proposed continuous material estimator. Multiple impact recordings are encoded, pooled as an unordered set, and mapped to continuous material parameters. Auxiliary classification and blend-consistency losses guide training, while inference uses the direct regression head.
Split
Objects
Single
Blend
Total
Train
880
7,040
63,360
70,400
Val
110
880
7,920
8,800
Test
110
880
7,920
8,800
All
1,100
8,800
79,200
88,000
Table 1 : ImpactMat dataset and split statistics. Single-material data are synthesized for each object using eight base materials. Blended-material data cover 24 selected material pairs at three mixing ratios, yielding 72 blend groups per object. Each impact group contains eight recordings.
Method
Setting
Test subset
ρ
E
ν
α
β
Avg. NMAE
Infer. time
VRAM
DSP + Ridge Reg.
Conventional
Single
0.190
0.443
0.047
9.93
3.26×10−7
1.405
16.1 (ms)
CPU
Blend
0.131
0.321
0.032
5.19
1.83×10−7
0.961
Qwen2.5-Omni-7B [ 21 ]
ICL
Single
0.338
0.614
0.071
13.38
1.61×10−1
1.270
5.6 (s)
16.1 GB
Blend
0.257
0.495
0.067
8.23
1.73×10−1
1.162
Audio Flamingo 3 [ 8 ]
ICL
Single
14.1
2.14
0.130
21.7
1.8×10−2
12.42
4.4 (s)
18.3 GB
Blend
13.4
2.33
0.123
11.3
4.4×10−2
10.92
Table 2 : Comparison of physical-parameter estimation methods. Results on ImpactMat. ρ and E are reported in the log-space target scale, ν is linear, and α and β in physical units. Avg. NMAE is computed over all five normalized target dimensions and is omitted for DiffSound, which estimates only E and ν in our evaluation. Evaluated with K=1 , lower is better.
Variant
Reg.
Cls.
Blend
Single
Blend
All
-
✓
0.0520
0.1053
0.0999
-
✓
✓
0.0648
0.1045
0.1005
-
✓
✓
0.0527
0.1032
0.0981
Ours
✓
✓
✓
0.0600
0.0989
0.0950
Table 3 : Ablation of training objectives for continuous estimation. Checkmarks indicate the loss terms used during training; Single, Blend, and All report NMAE on the corresponding test subsets. Evaluated with K=4 , lower is better.
K
Single
Blend
All
1
0.0784
0.1163
0.1124
2
0.0672
0.1043
0.1006
4
0.0600
0.0989
0.0950
8
0.0582
0.0965
0.0926
Table 4 : Effect of multi-impact aggregation at inference. Aggregating multiple reference impacts improves robustness to contact variation. Values report NMAE over the five material parameters. Lower is better.
Figure 3 : Spectrogram comparison for re-rendering validation. Each column compares an input reference with the sound synthesized after feeding the estimated parameters back into the simulator.
Domain
Condition
Sim.%
Unsure
Diff.
Synthetic
Re-render (ours)
100%
0%
0%
Real
G.T. variant
59.0%
17.0%
24.0%
Re-render (ours)
63.3%
20.8%
15.8%
Table 5 : Listening study for perceptual material similarity. Comparing reference and re-rendered impact sounds.
Figure 4 : Material transfer from reference audio to new geometry. Estimated material parameters are transferred to the same target mesh to synthesize ceramic-, plastic-, steel-, and glass-like impact responses.
While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical interaction. Object impact sounds convey material, stiffness, and structural properties that complement vision, yet existing impact sound modeling approaches either rely on expensive physics-based simulation or require large datasets to generalize in a purely data-driven manner. We introduce Audio-Visual Modal Sound Field (AV-MSF), a novel object-level acoustic representation reconstructed from multi-view images and only a few impact sound recordings. AV-MSF builds on 3D Gaussian Splatting integrated with dense 3D visual feature to provide a strong geometry-aware prior, and represents the impact sound field using compact, physically meaningful modal parameters, enabling robust few-shot reconstruction. Experiments on two real-world datasets show that AV-MSF achieves state-of-the-art impact sound rendering, outperforming both physics-based and data-driven baselines. Furthermore, we demonstrate downstream applications enabled by our representation, including contact localization and object sound editing.
Rings like gold, thuds like wood! The sound we hear in a scene is shaped not only by the spatial layout of the environment but also by the materials of the objects and surfaces within it. For instance, a room with wooden walls will produce a different acoustic experience from a room with the same spatial layout but concrete walls. Accurately modeling these effects is essential for applications such as virtual reality, robotics, architectural design, and audio engineering. Yet, existing methods for acoustic modeling often entangle spatial and material influences in correlated representations, which limits user control and reduces the realism of the generated acoustics. In this work, we present a novel approach for material-controlled Room Impulse Response (RIR) generation that explicitly disentangles the effects of spatial and material cues in a scene. Our approach models the RIR using two modules: a spatial module that captures the influence of the spatial layout of the scene, and a material module that modulates this spatial RIR according to a user-specified material configuration. This explicitly disentangled design allows users to easily modify the material configuration of a scene and observe its impact on acoustics without altering the spatial structure or scene content. Our model provides significant improvements over prior approaches on both acoustic-based metrics (up to +16% on RTE) and material-based metrics (up to +70%). Furthermore, through a human perceptual study, we demonstrate the improved realism and material sensitivity of our model compared to the strongest baselines.
Mahnoor Fatima Saad, Sagnik Majumder, Kristen Grauman +1
Realistic visual simulation of food manipulation requires accurate material parameters, yet these are difficult to measure directly and vary across the heterogeneous regions of a single food item. We address the inverse problem of estimating material parameters from a target description of fracture behavior in a non-differentiable continuum damage mechanics simulator. Using orange peeling as a test case, we train a neural surrogate on 2,000 forward simulations and compare Covariance Matrix Adaptation Evolution Strategy (CMA-ES, a gradient-free evolutionary optimizer) with Proximal Policy Optimization (PPO, a reinforcement learning algorithm) across the original 9-dimensional parameter space and two learned 4-dimensional latent representations. Since different oranges have different material properties, a practical inverse system must handle arbitrary targets without retraining. We train a goal-conditioned PPO policy that learns a general inverse mapping: given any target description of peeling behavior, the policy produces a material parameter estimate in a single forward pass (8 surrogate evaluations, approximately 10ms). Operating in a normalizing flow latent space with a shared surrogate evaluator, the goal-conditioned policy achieves 0.642 actual recovery when validated through the simulator, outperforming the original parameter space by 23%. A warm-start extension that initializes CMA-ES refinement from the policy's output further improves recovery to 0.828 with 540 evaluations. These findings provide a practical framework for inverse food physics and lay groundwork for vision-driven material identification from video observations of food manipulation.