Open AccessJournal Article
Information-geometric optimization algorithms: a unifying picture via invariance principles
TL;DR: A canonical way to turn any smooth parametric family of probability distributions on an arbitrary search space X into a continuous-time black-box optimization method on X, the information-geometric optimization (IGO) method, which achieves maximal invariance properties.
read more
Abstract: We present a canonical way to turn any smooth parametric family of probability distributions on an arbitrary search space X into a continuous-time black-box optimization method on X, the information-geometric optimization (IGO) method. Invariance as a major design principle keeps the number of arbitrary choices to a minimum. The resulting IGO flow is the flow of an ordinary differential equation conducting the natural gradient ascent of an adaptive, time-dependent transformation of the objective function. It makes no particular assumptions on the objective function to be optimized.
The IGO method produces explicit IGO algorithms through time discretization. It naturally recovers versions of known algorithms and offers a systematic way to derive new ones. In continuous search spaces, IGO algorithms take a form related to natural evolution strategies (NES). The cross-entropy method is recovered in a particular case with a large time step, and can be extended into a smoothed, parametrization-independent maximum likelihood update (IGO-ML). When applied to the family of Gaussian distributions on Rd, the IGO framework recovers a version of the well-known CMA-ES algorithm and of xNES. For the family of Bernoulli distributions on {0, 1}d, we recover the seminal PBIL algorithm and cGA. For the distributions of restricted Boltzmann machines, we naturally obtain a novel algorithm for discrete optimization on {0, 1}d. All these algorithms are natural instances of, and unified under, the single information-geometric optimization framework.
The IGO method achieves, thanks to its intrinsic formulation, maximal invariance properties: invariance under reparametrization of the search space X, under a change of parameters of the probability distribution, and under increasing transformation of the function to be optimized. The latter is achieved through an adaptive, quantile-based formulation of the objective.
Theoretical considerations strongly suggest that IGO algorithms are essentially characterized by a minimal change of the distribution over time. Therefore they have minimal loss in diversity through the course of optimization, provided the initial diversity is high. First experiments using restricted Boltzmann machines confirm this insight. As a simple consequence, IGO seems to provide, from information theory, an elegant way to simultaneously explore several valleys of a fitness landscape in a single run.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
•Proceedings Article
Optimizing Neural Networks with Kronecker-factored Approximate Curvature
James Martens,Roger Grosse +1 more
- 06 Jul 2015
TL;DR: K-FAC is an efficient method for approximating natural gradient descent in neural networks which is based on an efficiently invertible approximation of a neural network's Fisher information matrix which is neither diagonal nor low-rank, and in some cases is completely non-sparse.
Rethinking drug design in the artificial intelligence era
Petra Schneider,W. Patrick Walters,Alleyn T. Plowright,Norman Sieroka,Jennifer Listgarten,Robert Alan Goodnow,Jasmin Fisher,Jasmin Fisher,Johanna M. Jansen,José S. Duca,Thomas S. Rush,Matthias Zentgraf,John Edward Hill,Elizabeth Krutoholow,Matthias Kohler,Jeff Blaney,Kimito Funatsu,Kimito Funatsu,Chris Luebkemann,Chris Luebkemann,Gisbert Schneider +20 more
TL;DR: The views of a diverse group of international experts on the ‘grand challenges’ in small-molecule drug discovery with AI are presented, including obtaining appropriate data sets, generating new hypotheses, optimizing in a multi-objective manner, reducing cycle times and changing the research culture.
Training restricted Boltzmann machines
Asja Fischer,Christian Igel +1 more
TL;DR: This tutorial introduces RBMs from the viewpoint of Markov random fields, starting with the required concepts of undirected graphical models and reviewing the state-of-the-art in training restricted Boltzmann machines from the perspective of graphical models.
546
•Posted Content
New insights and perspectives on the natural gradient method
TL;DR: This paper critically analyze this method and its properties, and shows how it can be viewed as a type of approximate 2nd-order optimization method, where the Fisher information matrix can be view as an approximation of the Hessian.
Measures, Integrals and Martingales
TL;DR: The corresponding chapters of this volume range from the foundations of parametric, semi-parametric, and nonparametric inference through current work on stochastic modeling in financial econometrics, epidemics, and eco-hydrology to the prognosis for the treatment of breast cancer, with excursions on biostatistics, social statistics, and statistical computing.
177
References
•Book
Elements of information theory
Thomas M. Cover,Joy A. Thomas +1 more
- 01 Jan 1991
TL;DR: The author examines the role of entropy, inequality, and randomness in the design of codes and the construction of codes in the rapidly changing environment.
A simplex method for function minimization
John A. Nelder,R. Mead +1 more
TL;DR: A method is described for the minimization of a function of n variables, which depends on the comparison of function values at the (n 41) vertices of a general simplex, followed by the replacement of the vertex with the highest value by another point.
30.6K
A fast learning algorithm for deep belief nets
TL;DR: A fast, greedy algorithm is derived that can learn deep, directed belief networks one layer at a time, provided the top two layers form an undirected associative memory.
Representation Learning: A Review and New Perspectives
TL;DR: Recent work in the area of unsupervised feature learning and deep learning is reviewed, covering advances in probabilistic models, autoencoders, manifold learning, and deep networks.