Paper ID: 2310.11650

VKIE: The Application of Key Information Extraction on Video Text

Siyu An, Ye Liu, Haoyuan Peng, Di Yin

Extracting structured information from videos is critical for numerous downstream applications in the industry. In this paper, we define a significant task of extracting hierarchical key information from visual texts on videos. To fulfill this task, we decouple it into four subtasks and introduce two implementation solutions called PipVKIE and UniVKIE. PipVKIE sequentially completes the four subtasks in continuous stages, while UniVKIE is improved by unifying all the subtasks into one backbone. Both PipVKIE and UniVKIE leverage multimodal information from vision, text, and coordinates for feature representation. Extensive experiments on one well-defined dataset demonstrate that our solutions can achieve remarkable performance and efficient inference speed.

Submitted: Oct 18, 2023

Topics

Application Proficiency
Multimodal Information
Feature Representation
Video Text
Key Information Extraction
Structured Information
Visual Text
Hierarchical Information

Links

arXiv PDF