Time consistent discounting
Tor Lattimore,Marcus Hutter +1 more
- 05 Oct 2011
- Vol. 6925, pp 383-397
TL;DR: In this paper, the authors generalize the usual discounted utility model to one where the discount function changes with the age of the agent and show the existence of a rational policy for an agent that knows its discount function is time-inconsistent.
read more
Abstract: A possibly immortal agent tries to maximise its summed discounted rewards over time, where discounting is used to avoid infinite utilities and encourage the agent to value current rewards more than future ones. Some commonly used discount functions lead to time-inconsistent behavior where the agent changes its plan over time. These inconsistencies can lead to very poor behavior. We generalise the usual discounted utility model to one where the discount function changes with the age of the agent. We then give a simple characterisation of time- (in)consistent discount functions and show the existence of a rational policy for an agent that knows its discount function is time-inconsistent.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
•Posted Content
Hyperbolic Discounting and Learning over Multiple Horizons
TL;DR: It is demonstrated that a simple approach approximates hyperbolic discount functions while still using familiar temporal-difference learning techniques in RL and a surprising discovery is made that simultaneously learning value functions over multiple time-horizons is an effective auxiliary task which often improves over a strong value-based RL agent, Rainbow.
A Survey on Reinforcement Learning Methods in Character Animation
Ariel Kwiatkowski,Eduardo Alvarado,Vicky Kalogeiton,C. Karen Liu,Julien Pettr'e,Michiel van de Panne,Marie-Paule Cani +6 more
TL;DR: This paper surveys the modern Deep Reinforcement Learning methods and discusses their possible applications in Character Animation, from skeletal control of a single, physically‐based character to navigation controllers for individual agents and virtual crowds, and describes the practical side of training DRL systems.
Universal knowledge-seeking agents
TL;DR: A new model of intelligence is proposed, the knowledge-seeking agent (KSA), halfway between Solomonoff induction and AIXI, that defines a completely autonomous agent that does not require a teacher, and a proof of strong asymptotic optimality for a class of horizon functions shows that this agent behaves according to expectation.
25
Asymptotically optimal agents
Tor Lattimore,Marcus Hutter +1 more
- 05 Oct 2011
TL;DR: In this article, the authors define two versions of asymptotic optimality and prove that no agent can satisfy the strong version while in some cases, depending on discounting, there does exist a non-computable weak optimal agent.
•Posted Content
Beyond Exponentially Discounted Sum: Automatic Learning of Return Function.
Yufei Wang,Qiwei Ye,Tie-Yan Liu +2 more
TL;DR: This paper proposes to use a general mathematical form for return function, and employs meta-learning to learn the optimal return function in an end-to-end manner, and results clearly indicate the advantages of automatically learning optimal return functions in reinforcement learning.
12
References
•Posted Content
General Discounting versus Average Reward
TL;DR: It is shown that asymptotically U for m →∞ and V for k→∞ are equal, provided both limits exist, and if the effective horizon grows linearly with k or faster, then the existence of the limit of U implies that thelimit of V exists.
19
General discounting versus average reward
Marcus Hutter
- 07 Oct 2006
TL;DR: In this article, the authors compare the average reward U from cycle 1 to m (average value) with the future discounted reward V from cycle k to ∞ (discounted value).
•Book
Neuro-dynamic programming
Dimitri P. Bertsekas,John N. Tsitsiklis +1 more
- 01 Jan 1996
TL;DR: This is the first textbook that fully explains the neuro-dynamic programming/reinforcement learning methodology, which is a recent breakthrough in the practical application of neural networks and dynamic programming to complex problems of planning, optimal decision making, and intelligent control.
Subgame-perfect equilibria of finite- and infinite-horizon games☆☆☆
Drew Fudenberg,David K. Levine +1 more
TL;DR: In this paper, it was shown that subgame-perfect equilibria of infinite-horizon games arise as limits, as the horizon grows long and epsilon small.
Related Papers (5)
Ming Li,Paul M. B. Vitányi +1 more
- 01 Sep 1993
Rakesh K. Sarin
- 01 Jan 1992