One-shot learning and behavioral eligibility traces in sequential decision making

Author Accepted Manuscript

PDF only version. The full online version will follow soon.

Download
Cite
Share
CommentOpen annotations (there are currently 0 annotations on this page).

Version of Record published: December 6, 2019 (Go to version)
Accepted Manuscript published: November 11, 2019 (This version)
Accepted: November 1, 2019
Received: April 5, 2019

1. Of interest
Similar excitability through different sodium channels and implications for the analgesic efficacy of selective drugs

Yu-Feng Xie, Jane Yang ... Steven A Prescott

Research Article Apr 30, 2024
Further reading

Abstract
Data availability
Article and author information
Metrics

Abstract

In many daily tasks we make multiple decisions before reaching a goal. In order to learn such sequences of decisions, a mechanism to link earlier actions to later reward is necessary. Reinforcement learning theory suggests two classes of algorithms solving this credit assignment problem: In classic temporal-difference learning, earlier actions receive reward information only after multiple repetitions of the task, whereas models with eligibility traces reinforce entire sequences of actions from a single experience (one-shot). Here we show one-shot learning of sequences. We developed a novel paradigm to directly observe which actions and states along a multi-step sequence are reinforced after a single reward. By focusing our analysis on those states for which RL with and without eligibility trace make qualitatively distinct predictions, we find direct behavioral (choice probability) and physiological (pupil dilation) signatures of reinforcement learning with eligibility trace across multiple sensory modalities.

Data availability

The datasets generated during the current study are available on Dryad, at the following address https://dx.doi.org/10.5061/dryad.j7h6f69

The following data sets were generated

1. Lehmann M
2. Xu HA
3. Liakoni V
4. Herzog MH
5. Gerstner W
6. Preuschoff K
(2019) Data from: One-shot learning and behavioral eligibility traces in sequential decision making
Dryad Digital Repository, doi:10.5061/dryad.j7h6f69.

https://dx.doi.org/10.5061/dryad.j7h6f69

Article and author information

Author details

Marco P Lehmann

School of Life Sciences, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland

For correspondence
marco.lehmann@alumni.epfl.ch

Competing interests
The authors declare that no competing interests exist.

"This ORCID iD identifies the author of this article:" 0000-0001-5274-144X
He A Xu

School of Life Sciences, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland

Competing interests
The authors declare that no competing interests exist.
Vasiliki Liakoni

School of Life Sciences, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland

Competing interests
The authors declare that no competing interests exist.
Michael H Herzog

School of Life Sciences, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland

Competing interests
The authors declare that no competing interests exist.
Wulfram Gerstner

School of Life Sciences, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland

Competing interests
The authors declare that no competing interests exist.
Kerstin Preuschoff

Swiss Center for Affective Sciences, University of Geneva, Genève, Switzerland

Competing interests
The authors declare that no competing interests exist.

Funding

Schweizerischer Nationalfonds zur Förderung der Wissenschaftlichen Forschung (CRSII2 147636 (Sinergia))

Marco P Lehmann
He A Xu
Vasiliki Liakoni
Michael H Herzog
Wulfram Gerstner
Kerstin Preuschoff

Schweizerischer Nationalfonds zur Förderung der Wissenschaftlichen Forschung (CRSII2 200020 165538)

Marco P Lehmann
Vasiliki Liakoni
Wulfram Gerstner

Horizon 2020 Framework Programme (Human Brain Project (SGA2) 785907)

Michael H Herzog
Wulfram Gerstner

H2020 European Research Council (268 689 MultiRules)

Wulfram Gerstner

Horizon 2020 Framework Programme (Human Brain Project (SGA1) 720270)

Michael H Herzog
Wulfram Gerstner

The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.

Reviewing Editor

Thorsten Kahnt, Northwestern University, United States

Ethics

Human subjects: Experiments were conducted in accordance with the Helsinki declaration and approved by the ethics commission of the Canton de Vaud (164/14 Titre: Aspects fondamentaux de la reconnaissance des objets : protocole général). All participants were informed about the general purpose of the experiment and provided written, informed consent. They were told that they could quit the experiment at any time they wish.

Version history

Received: April 5, 2019
Accepted: November 1, 2019
Accepted Manuscript published: November 11, 2019 (version 1)
Version of Record published: December 6, 2019 (version 2)

Copyright

This article is distributed under the terms of the Creative Commons Attribution License permitting unrestricted use and redistribution provided that the original author and source are credited.