Time consistent discounting
Tor Lattimore,Marcus Hutter +1 more
- 05 Oct 2011
- Vol. 6925, pp 383-397
TL;DR: In this paper, the authors generalize the usual discounted utility model to one where the discount function changes with the age of the agent and show the existence of a rational policy for an agent that knows its discount function is time-inconsistent.
read more
Abstract: A possibly immortal agent tries to maximise its summed discounted rewards over time, where discounting is used to avoid infinite utilities and encourage the agent to value current rewards more than future ones. Some commonly used discount functions lead to time-inconsistent behavior where the agent changes its plan over time. These inconsistencies can lead to very poor behavior. We generalise the usual discounted utility model to one where the discount function changes with the age of the agent. We then give a simple characterisation of time- (in)consistent discount functions and show the existence of a rational policy for an agent that knows its discount function is time-inconsistent.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Optimistic agents are asymptotically optimal
Peter Sunehag,Marcus Hutter +1 more
- 04 Dec 2012
TL;DR: This work uses optimism to introduce generic asymptotically optimal reinforcement learning agents that achieve, with an arbitrary finite or compact class of environments, asymptonically optimal behavior.
The Multi-slot Framework: A Formal Model for Multiple, Copiable AIs
Laurent Orseau,Laurent Orseau +1 more
- 01 Aug 2014
TL;DR: A novel “multi-slot” framework for dealing with multiple intelligent agents, each of which can be duplicated or deleted at each step, in arbitrarily complex environments is introduced.
8
UGAE: A Novel Approach to Non-exponential Discounting
TL;DR: In this article , the authors propose the Universal Generalized Advantage Estimation (UGAE), which allows for the computation of GAE advantage values with arbitrary discounting, and introduce a continuous interpolation between exponential and hyperbolic discounting to increase flexibility in choosing a discounting method.
Algorithmic Learning Theory
Eiji Takimoto,Kohei Hatano +1 more
TL;DR: Some recent results on universal and efficient implementations of low-regret algorithmic frameworks such as Follow the Regularized Leader FTRL and Follow the Perturbed Leader FPL are surveyed.
•Journal Article
Rationality, optimism and guarantees in general reinforcement learning
Peter Sunehag,Marcus Hutter +1 more
TL;DR: This article introduces a framework for general reinforcement learning agents based on rationality axioms for a decision function and an hypothesis-generating function designed so as to achieve guarantees on the number errors, and introduces a notion of a class of environments being generated by a set of laws.
References
Time Discounting and Time Preference: A Critical Review
TL;DR: In this paper, the authors discuss the discounted utility (DU) model, its historical development, underlying assumptions, and "anomalies" -the empirical regularities that are inconsistent with its theoretical predictions.
Neuro-Dynamic Programming.
Dimitri P. Bertsekas
- 01 Jan 2009
TL;DR: In this article, the authors present the first textbook that fully explains the neuro-dynamic programming/reinforcement learning methodology, which is a recent breakthrough in the practical application of neural networks and dynamic programming to complex problems of planning, optimal decision making, and intelligent control.
4.7K
Myopia and Inconsistency in Dynamic Utility Maximization
TL;DR: In this article, the authors present a problem which has not heretofore been analysed and provide a theory to explain, under different circumstances, three related phenomena: (1) spendthriftiness; (2) the deliberate regimenting of one's future economic behaviour, even at a cost; and (3) thrift.
3.6K
Some empirical evidence on dynamic inconsistency
TL;DR: In this paper, individual discount rates for losses and gains were estimated from survey evidence and they were found to vary inversely with the size of the reward and the length of time to be waited.
2.5K
Related Papers (5)
Ming Li,Paul M. B. Vitányi +1 more
- 01 Sep 1993
Rakesh K. Sarin
- 01 Jan 1992