cs.CVJun 24, 2026

EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis

Authors: Huaqiu LiJiahao WangSijia CaiHualian ShengBing DengJieping YeWenhan Luo

Organizations: The Hong Kong University of Science and Technology · Alibaba Group · Xi’an Jiaotong University

Abstract

While image stylization has been studied extensively, video stylization remains a critical and largely unsolved challenge in the field of intelligent content creation. Existing methods, usually utilizing a reference image as the style prior, suffer from content leakage, data scarcity and limited adaptability to long videos, leading to suboptimal results with severe style drift and motion distortion. For these issues, we present EchoStyle, a scalable text-driven framework to achieve high-quality stylization of videos with arbitrary lengths. To start with, we construct a video-to-video architecture to appropriately re-fuse the video content and the text style. To address data scarcity, we pioneer an automatic reverse-synthesis pipeline to establish V-Style20k, a large-scale stylization dataset of 20k high-quality video pairs. To facilitate long video stylization, we devise an init-follow-mode mechanism along with a sliding-window inference strategy. Extensive experiments demonstrate EchoStyle's excellent performance across a wide range of artistic styles, even comparable to leading closed-source solutions.

Explore similar work

CardsList
  1. VISTA: Video-Injected Stylized Text-to-Animation

    Sep 20, 2026Monseej Purkayastha, Anindita Ghosh, Philipp SlusallekControllable Video Generation