Open AccessPosted Content
Minimal Achievable Sufficient Statistic Learning
TL;DR: In a series of experiments, it is shown that deep networks trained with MASS Learning achieve competitive performance on supervised learning and uncertainty quantification benchmarks.
read more
Abstract: We introduce Minimal Achievable Sufficient Statistic (MASS) Learning, a training method for machine learning models that attempts to produce minimal sufficient statistics with respect to a class of functions (e.g. deep networks) being optimized over. In deriving MASS Learning, we also introduce Conserved Differential Information (CDI), an information-theoretic quantity that - unlike standard mutual information - can be usefully applied to deterministically-dependent continuous random variables like the input and output of a deep network. In a series of experiments, we show that deep networks trained with MASS Learning achieve competitive performance on supervised learning and uncertainty quantification benchmarks.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
The Information Bottleneck Problem and its Applications in Machine Learning
Ziv Goldfeld,Yury Polyanskiy +1 more
- 30 Apr 2020
TL;DR: The information bottleneck (IB) theory recently emerged as a bold information-theoretic paradigm for analyzing DL systems, and its recent impact on DL is surveyed.
170
•Posted Content
Bayesian Image Reconstruction using Deep Generative Models
TL;DR: This work uses a single pre-trained generator model to solve different image restoration tasks, i.e., super-resolution and in-painting, by combining it with different forward corruption models, which enable application of Bayes' theorem for many downstream reconstruction tasks.
28
Bottleneck Problems: Information and Estimation-Theoretic View
Shahab Asoodeh,Flavio P. Calmon +1 more
TL;DR: In this paper, a general family of optimization problems, termed as \textit{bottleneck problems}, were introduced by replacing mutual information in IB and PF with other notions of mutual information, namely $f$-information and Arimoto's mutual information.
Bottleneck Problems: An Information and Estimation-Theoretic View
Shahab Asoodeh,Flavio P. Calmon +1 more
TL;DR: This work investigates the functional properties of information bottleneck and privacy funnel through a unified theoretical framework, and introduces a general family of optimization problems, termed “bottleneck problems”, by replacing mutual information in IB and PF with other notions of mutual information, namely f-information and Arimoto’s mutual information.
10
•Journal Article
The Deterministic Information Bottleneck
DJ Strouse,David J. Schwab +1 more
TL;DR: This work introduces an alternative formulation of the information bottleneck method that replaces mutual information with entropy, which it is argued better captures this notion of compression and empirically finds that the DIB offers a considerable gain in computational efficiency over the IB, over a range of convergence parameters.
References
Deep Residual Learning for Image Recognition
Kaiming He,Xiangyu Zhang,Shaoqing Ren,Jian Sun +3 more
- 27 Jun 2016
TL;DR: In this article, the authors proposed a residual learning framework to ease the training of networks that are substantially deeper than those used previously, which won the 1st place on the ILSVRC 2015 classification task.
•Proceedings Article
Adam: A Method for Stochastic Optimization
Diederik P. Kingma,Jimmy Ba +1 more
- 01 Jan 2015
TL;DR: This work introduces Adam, an algorithm for first-order gradient-based optimization of stochastic objective functions, based on adaptive estimates of lower-order moments, and provides a regret bound on the convergence rate that is comparable to the best known results under the online convex optimization framework.
138.5K
•Posted Content
Deep Residual Learning for Image Recognition
TL;DR: This work presents a residual learning framework to ease the training of networks that are substantially deeper than those used previously, and provides comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth.
117.9K
•Dissertation
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky
- 01 Jan 2009
TL;DR: In this paper, the authors describe how to train a multi-layer generative model of natural images, using a dataset of millions of tiny colour images, described in the next section.
Automatic differentiation in PyTorch
Adam Paszke,Sam Gross,Soumith Chintala,Gregory Chanan,Edward Z. Yang,Zachary DeVito,Zeming Lin,Alban Desmaison,Luca Antiga,Adam Lerer +9 more
- 28 Oct 2017
TL;DR: An automatic differentiation module of PyTorch is described — a library designed to enable rapid research on machine learning models that focuses on differentiation of purely imperative programs, with a focus on extensibility and low overhead.
Related Papers (5)
Shachar Shayovitz,Meir Feder +1 more
- 01 Jun 2021
Maria-Florina Balcan,Eric Blais,Avrim Blum,Liu Yang +3 more
- 20 Oct 2012