Actions ~ Transformations

Open AccessProceedings Article

Actions ~ Transformations

- 01 Jun 2016

305

TL;DR: A novel representation for actions is proposed by modeling an action as a transformation which changes the state of the environment before the action happens (precondition) to the state after the action (effect).

Chat with Paper

AI Agents for this Paper

Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps

Citations

•Proceedings Article•10.1109/CVPR.2017.502

Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

Joao Carreira, +1 more

- 21 Jul 2017

TL;DR: In this article, a Two-Stream Inflated 3D ConvNet (I3D) is proposed to learn seamless spatio-temporal feature extractors from video while leveraging successful ImageNet architecture designs and their parameters.

...read moreread less

8.4K

•Posted Content

The Kinetics Human Action Video Dataset

Andrew Zisserman, +11 more

- 19 May 2017

- arXiv: Computer Vision and Pattern Recog...

TL;DR: The dataset is described, the statistics are described, how it was collected, and some baseline performance figures for neural network architectures trained and tested for human action classification on this dataset are given.

...read moreread less

4.4K

•Proceedings Article•10.1109/CVPR.2018.00675

A Closer Look at Spatiotemporal Convolutions for Action Recognition

Du Tran, +6 more

- 12 Apr 2018

TL;DR: In this article, a new spatio-temporal convolutional block "R(2+1)D" was proposed, which achieved state-of-the-art performance on Sports-1M, Kinetics, UCF101, and HMDB51.

...read moreread less

3.1K

•Posted Content

A Closer Look at Spatiotemporal Convolutions for Action Recognition

Du Tran, +6 more

- 30 Nov 2017

- arXiv: Computer Vision and Pattern Recog...

TL;DR: A new spatiotemporal convolutional block "R(2+1)D" is designed which produces CNNs that achieve results comparable or superior to the state-of-the-art on Sports-1M, Kinetics, UCF101, and HMDB51.

...read moreread less

1.9K

•Book Chapter•10.1007/978-3-030-01267-0_19

Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification

Saining Xie, +4 more

- 08 Sep 2018

TL;DR: In this article, it was shown that it is possible to replace many of the expensive 3D convolutions by low-cost 2D convolution, and the best result was achieved when replacing the 3D CNNs at the bottom of the network, suggesting that temporal representation learning on high-level semantic features is more useful.

...read moreread less

1.4K

...

Expand

References

•Proceedings Article

Very Deep Convolutional Networks for Large-Scale Image Recognition

Karen Simonyan, +1 more

- 04 Sep 2014

TL;DR: This work investigates the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting using an architecture with very small convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers.

...read moreread less

102.6K

•Proceedings Article

Very Deep Convolutional Networks for Large-Scale Image Recognition

Karen Simonyan, +1 more

- 01 Jan 2015

TL;DR: In this paper, the authors investigated the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting and showed that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 layers.

...read moreread less

51.9K

•Journal Article•10.1007/S11263-015-0816-Y

ImageNet Large Scale Visual Recognition Challenge

Olga Russakovsky, +11 more

- 01 Dec 2015

- International Journal of Computer Vision

TL;DR: The ImageNet Large Scale Visual Recognition Challenge (ILSVRC) as mentioned in this paper is a benchmark in object category classification and detection on hundreds of object categories and millions of images, which has been run annually from 2010 to present, attracting participation from more than fifty institutions.

...read moreread less

41.6K

•Journal Article

ImageNet Large Scale Visual Recognition Challenge

Olga Russakovsky, +11 more

- 01 Apr 2015

- Springer US

TL;DR: The ImageNet Large Scale Visual Recognition Challenge (ILSVRC) has been running annually for five years (since 2010) and has become the standard benchmark for large-scale object recognition.

...read moreread less

23.9K

•Proceedings Article•10.1109/ICCV.2015.510

Learning Spatiotemporal Features with 3D Convolutional Networks

Du Tran, +5 more

- 07 Dec 2015

TL;DR: The learned features, namely C3D (Convolutional 3D), with a simple linear classifier outperform state-of-the-art methods on 4 different benchmarks and are comparable with current best methods on the other 2 benchmarks.

...read moreread less

10.6K

...

Expand

Actions ~ Transformations

Chat with Paper

AI Agents for this Paper

Citations

Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

The Kinetics Human Action Video Dataset

A Closer Look at Spatiotemporal Convolutions for Action Recognition

A Closer Look at Spatiotemporal Convolutions for Action Recognition

Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification

References

Very Deep Convolutional Networks for Large-Scale Image Recognition

Very Deep Convolutional Networks for Large-Scale Image Recognition

ImageNet Large Scale Visual Recognition Challenge

ImageNet Large Scale Visual Recognition Challenge

Learning Spatiotemporal Features with 3D Convolutional Networks

Related Papers (5)

Learning Spatiotemporal Features with 3D Convolutional Networks

UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Two-Stream Convolutional Networks for Action Recognition in Videos

Large-Scale Video Classification with Convolutional Neural Networks

Deep Residual Learning for Image Recognition