We present DynaMesh, a dynamic texture generation method for 3D meshes. Given a textureless shape and a text prompt describing an effect, our method produces an appearance that evolves while the object's geometry remains unchanged. Previous works on dynamic 3D content generation have focused on motion, where an object's geometry and position change while keeping its appearance the same. Methods on texture generation sit on the other side of the problem, painting appearance onto a shape as a fixed surface property and not as an evolving process. Neither addresses a visual effect that propagates on a 3D object. A natural route consists of two generators: a video model that shows the effect from a single view, and an image-to-3D generator that lifts each frame to 3D. However, the latter has no notion of time, so running it per video frame produces a sequence that flickers, loses effect details, and yields a different mesh at every video frame. Our method addresses these failures by conditioning a video model on a render of the mesh and the prompt to obtain a reference video, then running a frozen image-to-3D generator on the video with two changes. The conditioning of each frame is blended over a temporal window, and low-rank adapters are fit per shape to restore the lost details. The mesh is encoded once for the whole sequence, so geometry is constant by construction, and the output is a single mesh with a texture per frame. Applied to various objects and effects, DynaMesh substantially improves over recent video-to-4D and texturing methods, and can generalize its temporal effect to different shapes never seen during training. Our project page is at https://threedle.github.io/dynamesh/.
Figures & tables
Figure 2 : Gallery of results. Four objects driven by four different effects. Each panel shows the driving video as a film strip and, below it, our output rendered from two novel viewpoints. The texture remains coherent even at viewpoints unseen in training.
Figure 3 : System overview. A render of the input mesh, together with a text prompt describing the effect, is passed to a video generation model, which returns the reference video for our pipeline. The 3D generator then produces the textured object at every frame. Our temporal cross-attention (the new module) conditions each frame’s generation on a window of neighboring video frames. Low-rank adapters on the attention modules are fit per scene to recover detail. Geometry is shared across frames, so the output is one mesh whose texture evolves with the video.
Figure 4 : Method overview. The token maps of an eleven-frame window are blended by a temporal attention layer. The blend conditions the generator’s spatial cross-attention, which writes the frame’s appearance onto the voxel tokens of the input mesh, and the 3D self-attention then propagates that appearance across the surface. Every weight except the rank-4 adapters on those two attentions is frozen.
Figure 5 : Generalization. Per-scene adapters fit on a reference video are applied, without any retraining, to meshes never seen during fitting. Both effects clearly propagate on the new shapes according to the video. The adapters change only how appearance is generated, and each object keeps its own geometry, so the same spreading blots or lava cracks work even on different shapes.
Figure 6 : Temporal coherence. A chair with moss spreading over its surface is shown at four time frames. The insets magnify the same patch of the backrest. Frozen TRELLIS.2 changes the texture abruptly from one moment to the next (flickers), while our method keeps the moss growing smoothly, as in the video.
Figure 7 : Making an existing static texture dynamic. Our method can operate on objects with a given texture and make it change over time, as shown for the moving spots over the duck.
Unseen views
Fidelity vs . GT
Method
Flicker ↓
Accel. ↓
PSNR ↑
SSIM ↑
SV4D 2.0 † [ 53 ]
0.023
0.035
34.8
0.963
DG4D † [ 35 ]
0.074
0.142
13.3
0.393
L4GM [ 36 ]
0.007
0.011
21.6
0.728
MeshNCA [ 31 ]
0.018
0.033
7.7
0.105
TRELLIS.2 [ 47 ]
0.018
0.027
12.1
0.303
Table 1: Quantitative comparison over 42 objects and effects and all 150 frames. Flicker and acceleration are per-step temporal errors, and PSNR/SSIM, fidelity metrics to the ground truth video, are computed in the supervised view against the video. We achieve better results than state-of-the-art methods. † SV4D 2.0 and DG4D are grayed out as their outputs have fewer frames and cannot be measured on the same frame grid, discussed more below.
Figure 8 : Qualitative comparison. We compare our method against the baselines on one reference video (top strip), with each method shown at four timesteps from a single unseen view. The baselines either flicker or degrade the geometry at that view. Only our method tracks the effect consistently and preserves the shape’s geometry.
Figure 9 : Failure case. DynaMesh relies on effects where the property of the pattern is local. Here, it replicates the local patterns (flowers, dots, etc.), but it does not recover them from unseen views, since the flowers and lines follow a global arrangement that a single view does not capture.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10 : Additional gallery. DynaMesh generates a different temporal effect for each object, and the effect holds at an unseen viewpoint despite being fit to a single camera.
Figure 11 : Qualitative comparison. Given the same reference video, the baselines either lose the effect or distort the geometry at the unseen viewpoint, whereas DynaMesh preserves both.
Figure 12 : Comparison at the second novel view. All methods are shown from a camera behind the object, which none of them was supervised on, at the same four frames as Fig. 11 . The baselines that construct their own geometry or operate in video space lose the effect. In contrast, our method presents the same texture it shows at the supervised view, since the effect is written onto the input mesh rather than onto a camera.
Figure 13 : Flicker over the clip. We plot flicker at each time step for frozen TRELLIS.2 and for DynaMesh, over the 42 objects of Tab. 1 . The gap holds across the whole clip and is largest at the start, before the frozen texture settles.
Attention
Unseen views
Fidelity vs . GT
Cross
Self
Temp.
Flicker ↓
Accel. ↓
PSNR ↑
SSIM ↑
✗
✗
✗
0.01818
0.02720
12.14
0.3035
✓
✗
✗
0.00727
0.01039
23.80
0.7542
✓
✓
✗
0.00702
0.00998
24.86
0.7970
✓
✓
✓
0.00499
0.00531
24.89
0.7980
Appendix
Table 2: Ablation study. We turn on the components of our method one at a time and evaluate on the same 42 objects as Tab. 1 , with the same metrics. The first row is the frozen generator and the last row is our full method, which is best on every metric.
Figure 14 : Continuing an effect across clips. A second reference video, generated from the last frame of the first, continues the same effect, so a progression that does not fit in one clip can be carried across several.
Figure 15 : MeshNCA conditioning. Conditioned on a single frame, MeshNCA reproduces it and does not advance through time, so its result reflects the conditioning image rather than the progression of the effect.
Q1 effect fidelity (%) ↑
Q2 temporal coherence ↑
Q3 geometry ↑
Method
Supervised
Unseen
Supervised
Unseen
Supervised
Unseen
Frozen TRELLIS.2
85.9
85.6
3.66
3.71
4.90
4.86
MeshNCA
31.4
32.9
2.62
2.71
4.70
4.75
SV4D 2.0
98.4
72.3
4.55
4.04
4.98
3.62
DreamGaussian4D
45.4
27.4
2.66
2.42
4.69
3.65
L4GM
98.9
60.3
4.64
3.75
4.90
3.47
Appendix
Table 3: LLM as a judge. We ask a vision-language model three questions about each rendered clip, at the supervised view and at the three unseen views, over all 42 objects with three repeats per clip. Effect fidelity is the percentage of five yes/no questions passed, while temporal coherence and geometry are rated from 1 (worst) to 5 (best). Our method scores best at the unseen views on all three questions, and the gap to the video-based baselines widens as the camera moves away from the supervised one.
Effect
Shapes
Flick. ↓
Accel. ↓
Q1 (%) ↑
Q2 ↑
Rorschach
source
0.004
0.005
100.0
5.00
unseen
0.006
0.009
100.0
4.73
Lava
source
0.007
0.005
86.7
5.00
unseen
0.007
0.006
85.5
4.64
Appendix
Table 4: LLM as a Judge for Generalization. Flicker, acceleration, and judge scores (Q1 effect fidelity, Q2 temporal coherence) for each adapter on the shape it was fit on and averaged over eleven unseen shapes. Effect fidelity is the percentage of five shape-neutral questions passed per effect, generated from the reference video, so the same questions score every shape. The transferred effects keep the temporal quality and the fidelity of the source result.
Question
Reference
Score
Q1 effect fidelity
reference video
% of 5 questions passed
Q2 temporal coherence
none
1 to 5
Q3 geometry correctness
gray render of the input mesh
1 to 5
Appendix
Table 5: Judge protocol. We ask three questions about every rendered clip, each scored against the reference its claim needs.
Figure 16 : Temporal window width. Widening the conditioning window reduces flicker, but fidelity peaks at a width of five and declines beyond it, so we select eleven to balance the two.
Temporal
Reconstruction
Temporal operator
Flicker ↓
Accel. ↓
PSNR ↑
SSIM ↑
Temporal attention
0.00680
0.00921
24.89
0.7171
Spatio-temporal attn.
0.00683
0.00924
24.86
0.7165
Parameter-free (ours)
0.00506
0.00474
24.82
0.7149
Appendix
Table 6: Temporal operator ablation. We compare two learned temporal operators against the parameter-free blend, at the same eleven-frame window with everything else held fixed. Metrics in the table are computed on the texture rather than on the renders, so they differ from Tab. 2 . The blend has lower flicker and acceleration on all 42 objects, at a cost of 0.07 dB in PSNR and 0.002 in SSIM.
Stage
Setting
Value
Reference video
model
Kling 3.0 [ 24 ]
clip length
121 to 150 frames
training res.
9602
Generator
backbone
TRELLIS.2 [ 47 ]
image encoder
DINOv3 ViT-L/16 [ 39 ]
flow blocks
30
Appendix
Table 7: Implementation details. We list the generator, the blending window, the adapter, the loss, and the hardware used to produce every result in the paper. All objects are fit with these same values.
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan, China · DAMO Academy, Alibaba Group, Hangzhou, China · Hupan Lab, Hangzhou, China