Learning Unsupervised Visual Representations using 3D Convolutional Autoencoder with Temporal Contrastive Modeling for Video Retrieval
TL;DR: A new unsupervised video feature learning method based on joint learning of past and future prediction using 3D-CAE with temporal contrastive learning is proposed, which achieves better retrieval performance than state-of-the-art and confirms the superiority of the method in learning underlying features.
read more
Abstract: The rapid growth of tag-free user-generated videos (on the Internet), surgical recorded videos, and surveillance videos has necessitated the need for effective content-based video retrieval systems. Earlier methods for video representations are based on hand-crafted, which hardly performed well on the video retrieval tasks. Subsequently, deep learning methods have successfully demonstrated their effectiveness in both image and video-related tasks, but at the cost of creating massively labeled datasets. Thus, the economic solution is to use freely available unlabeled web videos for representation learning. In this regard, most of the recently developed methods are based on solving a single pretext task using 2D or 3D convolutional network. However, this paper designs and studies a 3D convolutional autoencoder (3D-CAE) for video representation learning (since it does not require labels). Further, this paper proposes a new unsupervised video feature learning method based on joint learning of past and future prediction using 3D-CAE with temporal contrastive learning. The experiments are conducted on UCF-101 and HMDB-51 datasets, where the proposed approach achieves better retrieval performance than state-of-the-art. In the ablation study, the action recognition task is performed by fine-tuning the unsupervised pre-trained model where it outperforms other methods, which further confirms the superiority of our method in learning underlying features. Such an unsupervised representation learning approach could also benefit the medical domain, where it is expensive to create large label datasets.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Monkeypox Disease Diagnosis using Machine Learning Approach
01 Dec 2022
TL;DR: In this paper , a machine learning-based method for analyzing RGB images for signs of monkey pox was presented, which employed 3 convolutional neural network (CNN) models and 6 machine learning classifiers (MLCs).
14
Monkeypox Disease Diagnosis using Machine Learning Approach
Ajay Krishan Gairola,Vidit Kumar +1 more
- 01 Dec 2022
TL;DR: In this paper , a machine learning-based method for analyzing RGB images for signs of monkey pox was presented, which employed 3 convolutional neural network (CNN) models and 6 machine learning classifiers (MLCs).
7
Exploring the strengths of Pre-trained CNN Models with Machine Learning Techniques for Skin Cancer Diagnosis
Ajay Krishan Gairola,Vidit Kumar,Ashok Kumar Sahoo +2 more
- 16 Oct 2022
TL;DR: In this article , the strengths of the CNN features for skin cancer diagnosis were investigated and the results showed that the fusion of multi-CNN features further enhances the accuracy of melanoma detection.
6
Exploring the strengths of Pre-trained CNN Models with Machine Learning Techniques for Skin Cancer Diagnosis
16 Oct 2022
TL;DR: In this paper , the strengths of the CNN features for skin cancer diagnosis were investigated and the results showed that the fusion of multi-CNN features further enhances the accuracy of melanoma detection.
6
Unsupervised Learning of Spatio-Temporal Representation with Multi-Task Learning for Video Retrieval
Vidit Kumar
- 24 May 2022
TL;DR: The C3D network is jointly optimized by using multiple pretext tasks such as: rotation prediction, speed prediction, time direction prediction and instance discrimination, which enhances video representation learning and generalizability and fine-tuned action recognition accuracy, which is better than state-of-the-arts.
4
References
ImageNet classification with deep convolutional neural networks
TL;DR: A large, deep convolutional neural network was trained to classify the 1.2 million high-resolution images in the ImageNet LSVRC-2010 contest into the 1000 different classes and employed a recently developed regularization method called "dropout" that proved to be very effective.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
TL;DR: This work introduces a Region Proposal Network (RPN) that shares full-image convolutional features with the detection network, thus enabling nearly cost-free region proposals and further merge RPN and Fast R-CNN into a single network by sharing their convolutionAL features.
Fully convolutional networks for semantic segmentation
Jonathan Long,Evan Shelhamer,Trevor Darrell +2 more
- 07 Jun 2015
TL;DR: The key insight is to build “fully convolutional” networks that take input of arbitrary size and produce correspondingly-sized output with efficient inference and learning.
Learning internal representations by error propagation
David E. Rumelhart,Geoffrey E. Hinton,Ronald J. Williams +2 more
- 01 Jan 1988
TL;DR: This chapter contains sections titled: The Problem, The Generalized Delta Rule, Simulation Results, Some Further Generalizations, Conclusion.
A fast learning algorithm for deep belief nets
TL;DR: A fast, greedy algorithm is derived that can learn deep, directed belief networks one layer at a time, provided the top two layers form an undirected associative memory.