One-shot learning and behavioral eligibility traces in sequential decision making

  1. Marco P Lehmann  Is a corresponding author
  2. He A Xu
  3. Vasiliki Liakoni
  4. Michael H Herzog
  5. Wulfram Gerstner
  6. Kerstin Preuschoff
  1. École Polytechnique Fédérale de Lausanne, Switzerland
  2. University of Geneva, Switzerland

Abstract

In many daily tasks we make multiple decisions before reaching a goal. In order to learn such sequences of decisions, a mechanism to link earlier actions to later reward is necessary. Reinforcement learning theory suggests two classes of algorithms solving this credit assignment problem: In classic temporal-difference learning, earlier actions receive reward information only after multiple repetitions of the task, whereas models with eligibility traces reinforce entire sequences of actions from a single experience (one-shot). Here we show one-shot learning of sequences. We developed a novel paradigm to directly observe which actions and states along a multi-step sequence are reinforced after a single reward. By focusing our analysis on those states for which RL with and without eligibility trace make qualitatively distinct predictions, we find direct behavioral (choice probability) and physiological (pupil dilation) signatures of reinforcement learning with eligibility trace across multiple sensory modalities.

Data availability

The datasets generated during the current study are available on Dryad, at the following address https://dx.doi.org/10.5061/dryad.j7h6f69

The following data sets were generated

Article and author information

Author details

  1. Marco P Lehmann

    School of Life Sciences, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland
    For correspondence
    marco.lehmann@alumni.epfl.ch
    Competing interests
    The authors declare that no competing interests exist.
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0001-5274-144X
  2. He A Xu

    School of Life Sciences, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland
    Competing interests
    The authors declare that no competing interests exist.
  3. Vasiliki Liakoni

    School of Life Sciences, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland
    Competing interests
    The authors declare that no competing interests exist.
  4. Michael H Herzog

    School of Life Sciences, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland
    Competing interests
    The authors declare that no competing interests exist.
  5. Wulfram Gerstner

    School of Life Sciences, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland
    Competing interests
    The authors declare that no competing interests exist.
  6. Kerstin Preuschoff

    Swiss Center for Affective Sciences, University of Geneva, Genève, Switzerland
    Competing interests
    The authors declare that no competing interests exist.

Funding

Schweizerischer Nationalfonds zur Förderung der Wissenschaftlichen Forschung (CRSII2 147636 (Sinergia))

  • Marco P Lehmann
  • He A Xu
  • Vasiliki Liakoni
  • Michael H Herzog
  • Wulfram Gerstner
  • Kerstin Preuschoff

Schweizerischer Nationalfonds zur Förderung der Wissenschaftlichen Forschung (CRSII2 200020 165538)

  • Marco P Lehmann
  • Vasiliki Liakoni
  • Wulfram Gerstner

Horizon 2020 Framework Programme (Human Brain Project (SGA2) 785907)

  • Michael H Herzog
  • Wulfram Gerstner

H2020 European Research Council (268 689 MultiRules)

  • Wulfram Gerstner

Horizon 2020 Framework Programme (Human Brain Project (SGA1) 720270)

  • Michael H Herzog
  • Wulfram Gerstner

The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.

Ethics

Human subjects: Experiments were conducted in accordance with the Helsinki declaration and approved by the ethics commission of the Canton de Vaud (164/14 Titre: Aspects fondamentaux de la reconnaissance des objets : protocole général). All participants were informed about the general purpose of the experiment and provided written, informed consent. They were told that they could quit the experiment at any time they wish.

Copyright

© 2019, Lehmann et al.

This article is distributed under the terms of the Creative Commons Attribution License permitting unrestricted use and redistribution provided that the original author and source are credited.

Metrics

  • 2,954
    views
  • 420
    downloads
  • 15
    citations

Views, downloads and citations are aggregated across all versions of this paper published by eLife.

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Marco P Lehmann
  2. He A Xu
  3. Vasiliki Liakoni
  4. Michael H Herzog
  5. Wulfram Gerstner
  6. Kerstin Preuschoff
(2019)
One-shot learning and behavioral eligibility traces in sequential decision making
eLife 8:e47463.
https://doi.org/10.7554/eLife.47463

Share this article

https://doi.org/10.7554/eLife.47463

Further reading

    1. Neuroscience
    Gergely F Turi, Sasa Teng ... Yueqing Peng
    Research Article

    Synchronous neuronal activity is organized into neuronal oscillations with various frequency and time domains across different brain areas and brain states. For example, hippocampal theta, gamma, and sharp wave oscillations are critical for memory formation and communication between hippocampal subareas and the cortex. In this study, we investigated the neuronal activity of the dentate gyrus (DG) with optical imaging tools during sleep-wake cycles in mice. We found that the activity of major glutamatergic cell populations in the DG is organized into infraslow oscillations (0.01–0.03 Hz) during NREM sleep. Although the DG is considered a sparsely active network during wakefulness, we found that 50% of granule cells and about 25% of mossy cells exhibit increased activity during NREM sleep, compared to that during wakefulness. Further experiments revealed that the infraslow oscillation in the DG was correlated with rhythmic serotonin release during sleep, which oscillates at the same frequency but in an opposite phase. Genetic manipulation of 5-HT receptors revealed that this neuromodulatory regulation is mediated by Htr1a receptors and the knockdown of these receptors leads to memory impairment. Together, our results provide novel mechanistic insights into how the 5-HT system can influence hippocampal activity patterns during sleep.

    1. Neuroscience
    Ulrike Pech, Jasper Janssens ... Patrik Verstreken
    Research Article

    The classical diagnosis of Parkinsonism is based on motor symptoms that are the consequence of nigrostriatal pathway dysfunction and reduced dopaminergic output. However, a decade prior to the emergence of motor issues, patients frequently experience non-motor symptoms, such as a reduced sense of smell (hyposmia). The cellular and molecular bases for these early defects remain enigmatic. To explore this, we developed a new collection of five fruit fly models of familial Parkinsonism and conducted single-cell RNA sequencing on young brains of these models. Interestingly, cholinergic projection neurons are the most vulnerable cells, and genes associated with presynaptic function are the most deregulated. Additional single nucleus sequencing of three specific brain regions of Parkinson’s disease patients confirms these findings. Indeed, the disturbances lead to early synaptic dysfunction, notably affecting cholinergic olfactory projection neurons crucial for olfactory function in flies. Correcting these defects specifically in olfactory cholinergic interneurons in flies or inducing cholinergic signaling in Parkinson mutant human induced dopaminergic neurons in vitro using nicotine, both rescue age-dependent dopaminergic neuron decline. Hence, our research uncovers that one of the earliest indicators of disease in five different models of familial Parkinsonism is synaptic dysfunction in higher-order cholinergic projection neurons and this contributes to the development of hyposmia. Furthermore, the shared pathways of synaptic failure in these cholinergic neurons ultimately contribute to dopaminergic dysfunction later in life.