Boundary-sensitive Pre-training for Temporal Localization in Videos
KAUST DepartmentElectrical and Computer Engineering
Computer, Electrical and Mathematical Science and Engineering (CEMSE) Division
Electrical and Computer Engineering Program
Permanent link to this recordhttp://hdl.handle.net/10754/666111
MetadataShow full item record
AbstractMany video analysis tasks require temporal localization for the detection of content changes. However, most existing models developed for these tasks are pre-trained on general video action classification tasks. This is due to large scale annotation of temporal boundaries in untrimmed videos being expensive. Therefore, no suitable datasets exist that enable pre-training in a manner sensitive to temporal boundaries. In this paper for the first time, we investigate model pre-training for temporal localization by introducing a novel boundary-sensitive pretext (BSP) task. Instead of relying on costly manual annotations of temporal boundaries, we propose to synthesize temporal boundaries in existing video action classification datasets. By defining different ways of synthesizing boundaries, BSP can then be simply conducted in a self-supervised manner via the classification of the boundary types. This enables the learning of video representations that are much more transferable to downstream temporal localization tasks. Extensive experiments show that the proposed BSP is superior and complementary to the existing action classification-based pre-training counterpart, and achieves new state-of-the-art performance on several temporal localization tasks. Please visit our website for more details https://frostinassiky.github.io/bsp.
CitationXu, M., Perez-Rua, J.-M., Escorcia, V., Martinez, B., Zhu, X., Zhang, L., Ghanem, B., & Xiang, T. (2021). Boundary-sensitive Pre-training for Temporal Localization in Videos. 2021 IEEE/CVF International Conference on Computer Vision (ICCV). https://doi.org/10.1109/iccv48922.2021.00713
SponsorsThis work was supported by the King Abdullah University of Science and Technology (KAUST) Office of Sponsored Research through the Visual Computing Center funding.
Conference/Event nameCVF International Conference on Computer Vision (ICCV)
RelationsIs Supplemented By: