Open AccessPosted Content
Training Deep Nets with Sublinear Memory Cost.
TL;DR: This work designs an algorithm that costs O( √ n) memory to train a n layer network, with only the computational cost of an extra forward pass per mini-batch, and shows that it is possible to trade computation for memory giving a more memory efficient training algorithm with a little extra computation cost.
read more
Abstract: We propose a systematic approach to reduce the memory consumption of deep neural network training. Specifically, we design an algorithm that costs O(sqrt(n)) memory to train a n layer network, with only the computational cost of an extra forward pass per mini-batch. As many of the state-of-the-art models hit the upper bound of the GPU memory, our algorithm allows deeper and more complex models to be explored, and helps advance the innovations in deep learning research. We focus on reducing the memory cost to store the intermediate feature maps and gradients during training. Computation graph analysis is used for automatic in-place operation and memory sharing optimizations. We show that it is possible to trade computation for memory - giving a more memory efficient training algorithm with a little extra computation cost. In the extreme case, our analysis also shows that the memory consumption can be reduced to O(log n) with as little as O(n log n) extra cost for forward computation. Our experiments show that we can reduce the memory cost of a 1,000-layer deep residual network from 48G to 7G with only 30 percent additional running time cost on ImageNet problems. Similarly, significant memory cost reduction is observed in training complex recurrent neural networks on very long sequences.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
•Proceedings Article
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhenzhong Lan,Mingda Chen,Sebastian Goodman,Kevin Gimpel,Piyush Sharma,Radu Soricut +5 more
- 30 Apr 2020
TL;DR: This work presents two parameter-reduction techniques to lower memory consumption and increase the training speed of BERT, and uses a self-supervised loss that focuses on modeling inter-sentence coherence.
•Posted Content
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
TL;DR: The authors proposed a self-supervised loss that focuses on modeling inter-sentence coherence, and showed it consistently helps downstream tasks with multientence inputs, achieving state-of-the-art results on the GLUE, RACE, and \squad benchmarks.
4.1K
•Posted Content
Longformer: The Long-Document Transformer
TL;DR: Following prior work on long-sequence transformers, the Longformer is evaluated on character-level language modeling and achieves state-of-the-art results on text8 and enwik8 and pretrain Longformer and finetune it on a variety of downstream tasks.
3.9K
•Posted Content
Theano: A Python framework for fast computation of mathematical expressions
Rami Al-Rfou,Guillaume Alain,Amjad Almahairi,Christof Angermueller,Dzmitry Bahdanau,Nicolas Ballas,Frédéric Bastien,Justin Bayer,Anatoly Belikov,Alexander Belopolsky,Yoshua Bengio,Arnaud Bergeron,James Bergstra,Valentin Bisson,Josh Bleecher Snyder,Nicolas Bouchard,Nicolas Boulanger-Lewandowski,Xavier Bouthillier,Alexandre de Brébisson,Olivier Breuleux,Pierre Luc Carrier,Kyunghyun Cho,Jan Chorowski,Paul F. Christiano,Tim Cooijmans,Marc-Alexandre Côté,Myriam Côté,Aaron Courville,Yann N. Dauphin,Olivier Delalleau,Julien Demouth,Guillaume Desjardins,Sander Dieleman,Laurent Dinh,Mélanie Ducoffe,Vincent Dumoulin,Samira Ebrahimi Kahou,Dumitru Erhan,Ziye Fan,Orhan Firat,Mathieu Germain,Xavier Glorot,Ian Goodfellow,Matthew M. Graham,Caglar Gulcehre,Philippe Hamel,Iban Harlouchet,Jean-Philippe Heng,Balázs Hidasi,Sina Honari,Arjun Jain,Sébastien Jean,Kai Jia,Mikhail Korobov,Vivek Kulkarni,Alex Lamb,Pascal Lamblin,Eric Larsen,César Laurent,Sean Lee,Simon Lefrancois,Simon Lemieux,Nicholas Léonard,Zhouhan Lin,Jesse A. Livezey,Cory Lorenz,Jeremiah Lowin,Qianli Ma,Pierre-Antoine Manzagol,Olivier Mastropietro,Robert T. McGibbon,Roland Memisevic,Bart van Merriënboer,Vincent Michalski,Mehdi Mirza,Alberto Orlandi,Chris Pal,Razvan Pascanu,Mohammad Pezeshki,Colin Raffel,Daniel Renshaw,Matthew Rocklin,Adriana Romero,Markus Roth,Peter Sadowski,John Salvatier,François Savard,Jan Schlüter,John Schulman,Gabriel Schwartz,Iulian Vlad Serban,Dmitriy Serdyuk,Samira Shabanian,Étienne Simon,Sigurd Spieckermann,S. Ramana Subramanyam,Jakub Sygnowski,Jérémie Tanguay,Gijs van Tulder,Joseph Turian,Sebastian Urban,Pascal Vincent,Francesco Visin,Harm de Vries,David Warde-Farley,Dustin J. Webb,Matthew Willson,Kelvin Xu,Lijun Xue,Li Yao,Saizheng Zhang,Ying Zhang +111 more
TL;DR: The performance of Theano is compared against Torch7 and TensorFlow on several machine learning models and recently-introduced functionalities and improvements are discussed.
2.6K
Convolutional Networks with Dense Connectivity
TL;DR: DenseNet as discussed by the authors proposes to connect each layer to every other layer in a feed-forward fashion, which alleviates the vanishing-gradient problem, strengthen feature propagation, encourage feature reuse, and substantially improve parameter efficiency.
References
Deep Residual Learning for Image Recognition
Kaiming He,Xiangyu Zhang,Shaoqing Ren,Jian Sun +3 more
- 27 Jun 2016
TL;DR: In this article, the authors proposed a residual learning framework to ease the training of networks that are substantially deeper than those used previously, which won the 1st place on the ILSVRC 2015 classification task.
•Posted Content
Deep Residual Learning for Image Recognition
TL;DR: This work presents a residual learning framework to ease the training of networks that are substantially deeper than those used previously, and provides comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth.
117.9K
Long short-term memory
TL;DR: A novel, efficient, gradient based method called long short-term memory (LSTM) is introduced, which can learn to bridge minimal time lags in excess of 1000 discrete-time steps by enforcing constant error flow through constant error carousels within special units.
99K
ImageNet classification with deep convolutional neural networks
TL;DR: A large, deep convolutional neural network was trained to classify the 1.2 million high-resolution images in the ImageNet LSVRC-2010 contest into the 1000 different classes and employed a recently developed regularization method called "dropout" that proved to be very effective.
•Proceedings Article
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky,Ilya Sutskever,Geoffrey E. Hinton +2 more
- 03 Dec 2012
TL;DR: The state-of-the-art performance of CNNs was achieved by Deep Convolutional Neural Networks (DCNNs) as discussed by the authors, which consists of five convolutional layers, some of which are followed by max-pooling layers, and three fully-connected layers with a final 1000-way softmax.
Related Papers (5)
Kaiming He,Xiangyu Zhang,Shaoqing Ren,Jian Sun +3 more
- 27 Jun 2016
Diederik P. Kingma,Jimmy Ba +1 more
- 01 Jan 2015