Correctness is its own reward: bootstrapping error signals in self-guided reinforcement learning

  1. Department of Neurobiology, Duke University, Durham, United States
  2. Department of Cell Biology, Duke University, Durham, United States
  3. Department of Electrical and Computer Engineering, Duke University, Durham, United States

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Daniel Takahashi
    Universidade Federal do Rio Grande do Norte, Natal, Brazil
  • Senior Editor
    Michael Frank
    Brown University, Providence, United States of America

Reviewer #1 (Public review):

Summary:

This manuscript addresses how internally generated evaluative signals can arise during self-guided learning in the absence of external reward. Using zebra finch song learning as a model system, the authors propose that tutor-song memorization and vocal performance evaluation are not separate processes, but instead emerge from a shared local circuit that learns to predictively cancel tutor-song-related auditory input. The comparison across several candidate circuit architectures, the quantitative comparison to experimental calcium imaging data, and the decomposition of the learned recurrent connectivity into modes shaping the error landscape are all strong aspects of the work. The final demonstration that the learned error signal can guide a downstream reinforcement learning agent also provides a useful proof of principle.

Strengths:

The idea that tutor-song memorization and performance evaluation can emerge from a shared predictive-cancellation circuit is interesting, and the combination of circuit modeling, comparison to experimental data, and error-landscape analysis is compelling.

Weaknesses:

(1) A central conclusion of the manuscript is that the E→I→E model best matches experimental data. This establishes model fit, but it does not yet explain why E→I and I→E plasticity are important for tutor-song cancellation and error-signal formation. Does E→I plasticity primarily teach the inhibitory population to represent tutor-song-related excitatory activity? Does I→E plasticity then implement the negative image required to cancel expected excitatory responses? Does the closed E/I loop primarily control gain, shift the minimum of the error landscape, or both? A useful analysis would be to compare models in which only E→I synapses are plastic, only I→E synapses are plastic, both are plastic, or neither is plastic.

(2) The analysis in Fig. 5 does not yet explain how the identified modes arise from the specific E→I/I→E plasticity mechanism. For example, are the landscape modes mainly produced by E/I gain-control dynamics? Are the memory modes related to an inhibitory negative image of the tutor song? Are these modes localized to particular blocks of the recurrent connectivity, such as E→I or I→E weights, or are they distributed across the full network?

(3) The manuscript emphasizes the emergence of sparse population error codes. However, in Fig. 6, the downstream actor-critic model uses the population mean excitatory activity as a scalar negative reward. This compresses the high-dimensional sparse population response into a single scalar. If the downstream system only uses the mean response, why is a sparse high-dimensional error code functionally important, beyond matching the observed response distribution? Conversely, if the sparse population pattern contains richer information about the direction or structure of vocal errors, how might downstream reinforcement pathways read out this information?

The manuscript should clarify whether sparsity is proposed to have a functional role in motor learning, or whether it is primarily a biological feature of the evaluative circuit. This point is particularly important because the broader framing of the paper concerns internal evaluative signals, whereas the final reinforcement learning demonstration uses a scalar reward.

(4) The actor-critic model in Fig. 6 is useful because it demonstrates that the learned error signal contains enough information to guide motor learning. However, the reinforcement learning module is attached downstream of the auditory circuit and is highly simplified. Therefore, it remains somewhat ambiguous whether Fig. 6 should be interpreted as a circuit model of song learning or as a demonstration of sufficiency. The latter interpretation seems appropriate and valuable, but the manuscript should state this more explicitly. The central contribution appears to be the bootstrapping of an internal evaluative signal, rather than a complete model of sensorimotor song learning. Clarifying this distinction would prevent overinterpretation of the actor-critic results.

Reviewer #2 (Public review):

Summary:

The paper proposes a network model that explains how birdsong learning can be guided by reinforcement signals.

Strengths:

It is well known that self-generated motor actions typically suppress their associated sensory input (for example, in the mammalian auditory cortex; see Eliades & Wang, 2003). This study presents a mechanism that effectively reverses this process. The theory posits that, initially, the motor signal generated in HVC, although not yet sufficient to produce an accurate song, nevertheless sends an efference copy to auditory areas, where it acts to cancel external auditory input from the tutor. This establishes a "scaffold," such that only an accurate replica of the tutor song can successfully suppress the corresponding auditory activity.

During learning, poorly generated plastic songs produce residual auditory activity that cannot be fully suppressed. This remaining activity then serves as an error signal that guides the refinement of motor output. The idea is elegant and is supported by experimental evidence.

Weaknesses:

The authors compare several possible sites of synaptic plasticity within the auditory network and conclude that the E-to-I-to-E model provides the best fit to the existing data. In this model, the auditory network consists of recurrent excitatory (E) and inhibitory (I) neurons, and Hebbian plasticity at E-to-I and I-to-E synapses is required to establish the cancellation pattern necessary to reproduce the tutor song.

However, the manuscript's presentation of the underlying plasticity mechanisms is somewhat puzzling. The authors repeatedly emphasize anti-Hebbian learning, even though their most successful model fundamentally relies on Hebbian plasticity. Although the resulting functional relationship may be described as anti-Hebbian, the biological learning mechanism implemented in the model is Hebbian. The repeated emphasis on anti-Hebbian learning therefore distracts from the central message and may confuse readers about the actual mechanism responsible for learning.

This emphasis may reflect an effort to distinguish the present work from previous anti-Hebbian models, but I suggest restructuring the manuscript. The authors should first present the optimal E-to-I-to-E model in detail, clearly explaining its mechanism and biological interpretation. Subsequent sections could then compare this model with the less successful alternative architectures. Such a reorganization would substantially improve the clarity and overall structure of the manuscript.

Finally, the abstract presents self-guided reinforcement learning as a novel concept, although this general idea has been described in previous work (e.g., Fiete et al., 2007). The abstract should therefore be revised to more precisely identify the specific novelty and contribution of the present study, rather than attributing novelty to the broader concept of self-guided reinforcement learning.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation