Multi-Directional Convolution Networks with Spatial-Temporal Feature Pyramid Module for Action Recognition

Authors :: Yi-Ping Phoebe Chen
Hong Lu
Bohong Yang
Wu Ran
Wang Zijian
Source :: ICASSP
Publication Year :: 2021
Publisher :: IEEE, 2021.
Abstract: Recent attempts show that factorizing 3D convolutional filters into separate spatial and temporal components brings impressive improvement in action recognition. However, traditional temporal convolution operating along the temporal dimension will aggregate unrelated features, since the feature maps of fast-moving objects have shifted spatial positions. In this paper, we propose a novel and effective Multi-Directional Convolution (MDConv), which extracts features along different spatial-temporal orientations. Especially, MDConv has the same FLOPs and parameters as the traditional 1D temporal convolution. Also, we propose the Spatial-Temporal Feature Pyramid Module (STFPM) to fuse spatial semantics in different scales in a light-weight way. Our extensive experiments show that the models which integrate with MDConv achieve better accuracy on several large-scale action recognition benchmarks such as Kinetics, AVA and Something-Something V1&V2 datasets.

Subjects :: Dimension (vector space)
business.industry
Computer science
Aggregate (data warehouse)
Fuse (electrical)
Feature (machine learning)
Pattern recognition
Artificial intelligence
Pyramid (image processing)
Semantics
business
FLOPS
Convolution

Database :: OpenAIRE
Journal :: ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Accession number :: edsair.doi...........3aa3c9312b9706bf31a50d3bda5ff4fa

Tools