cs.CVSep 15, 2026

tcnerv:dual-domain temporal context modeling for implicit neural video compression

Authors: Xuezhi Xiang, Yixin Zhao, Heqi Xiang, Jiayao Liu, Shanjun Zhang

Organizations: Information and Communication Engineering, Harbin Engineering University, Harbin, 150001, China · Department of Computer Science, University of Toronto, Toronto, ON M5S 2E4, Canada · The Department of Computer Science, Kanagawa University, Kanagawa, 221-8686, Japan

Abstract

Video compression aims to minimize reconstruction distor tion under a constrained bit rate. Existing video implicit neural representations (INRs) often decode frames independently, leaving intermediate features unconditioned on previous reconstructions and content embeddings without explicit temporal prediction. We propose TCNeRV, which exploits reconstructed context in both feature and embedding domains. Its multi-scale temporal-context fusion (MTCF) module injects gated historical features at multiple decoder scales, while temporal embedding-residual coding (TERC) predicts each content embedding and codes only its residual. With approximately 3M parameters, TCNeRV achieves an average PSNR of 36.08 dB on the UVG dataset, outperforming HNeRV-Boost by 2.20 dB. It reduces BD-rate by 22.06%, 66.73%, and 29.85% relative to HM, DCVC, and HiNeRV, respectively, demonstrating competitive rate-distortion performance with limited model capacity.

Explore similar work

CardsList