0

Spatial-temporal Concept based Explanation of 3D ConvNets

A 3D ACE framework provides interpretable high-level supervoxel representations for 3D video recognition ConvNets, uncovering spatial-temporal concepts important for tasks like action classification.

Year
2022
Venue
CVPR 2023 1
Authors
4
Hosting
Abstract onlyARXIV-DEFAULT

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2206.05275ARXIV-DEFAULT
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Recent studies have achieved outstanding success in explaining 2D image recognition ConvNets. On the other hand, due to the computation cost and complexity of video data, the explanation of 3D video recognition ConvNets is relatively less studied. In this paper, we present a 3D ACE (Automatic Concept-based Explanation) framework for interpreting 3D ConvNets. In our approach: (1) videos are represented using high-level supervoxels, which is straightforward for human to understand; and (2) the interpreting framework estimates a score for each voxel, which reflects its importance in the decision procedure. Experiments show that our method can discover spatial-temporal concepts of different importance-levels, and thus can explore the influence of the concepts on a target task, such as action classification, in-depth. The codes are publicly available.

Authors

4