Dual pathway architecture in songbirds enables robust sensorimotor learning

  1. Univ. Bordeaux, Inria, IMN, UMR 5293, Bordeaux, France
  2. LMU Munich, Planegg, Germany
  3. Univ. Bordeaux, CNRS, IMN, UMR 5293, Bordeaux, France

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, public reviews, and a provisional response from the authors.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Long Ding
    University of Pennsylvania, Philadelphia, United States of America
  • Senior Editor
    Michael Frank
    Brown University, Providence, United States of America

Reviewer #1 (Public review):

The authors sought to devise a model of the song system that captures the essential features of that neural circuitry and couple it to a behavioral model that captures the challenges of motor learning while remaining tractable. They seek to use this model to explain known features of song learning and relate them to the general problems associated with learning via gradient ascent. Their syrinx model uses two control parameters-air sac pressure and syringeal labial tension-to generate birdsong-like spectrograms. Normalized spectrograms generated by the syrinx model are compared to a target spectrogram by computing Pearson's correlation coefficient. This correlation coefficient quantifies the performance of the model, and the goal of learning is to maximize it. The syringeal model, while simplified, is complex enough to generate multiple local maxima in the correlation coefficient within the 2D control space with wide variation in the magnitude of the maxima, making it challenging to find a good optimum via gradient ascent. They also use a more abstract motor model where local maxima are generated by randomly placing Gaussians within the 2D control space. These two approaches to modeling motor space are a major strength of this work.

The neural model is extremely generic and consists of units (each representing a population of excitatory and inhibitory neurons) with continuous-valued outputs ("firing rate") ranging from -1 to +1 with sigmoidal activation functions. Premotor HVC simply generates a fixed temporal sequence that drives activity in RA (analogous to the primary motor cortex) that constitutes commands to the motor controller that translates RA output into 2D control signals for the syrinx. A second, indirect, pathway from HVC to RA is represented by a single node ("BG"); in songbirds this pathway consists of 3 structures (one of them quite complex and heterogeneous) with recurrent connections (i.e., from LMAN back to Area X). Learning is primarily driven by performance-modulated Hebbian plasticity in HVC-BG connection weights. There is also Hebbian plasticity in HVC-RA weights that are in effect trained by the RA activity patterns driven by BG-RA connections.

I worry that this model is too simplified to capture essential features of the song system (and cortico-basal ganglia circuits more generally). That problem is most acute in the way the authors model (or fail to model) the anterior forebrain pathway (AFP) through X, DLM, and LMAN; their model in effect reduces the AFP to just LMAN. I do not think that invalidates this study, but there is a significant danger that this model will end up missing the mark in some important way relative to a more realistic model. However, there is another flaw that comes close to doing that-the connection weights between the units are allowed to vary from -1 to +1. That means, for example, that the connections between HVC and RA units can be inhibitory and can flip between excitation and inhibition during learning. I can imagine some hand-wavy justifications for this (e.g., the connection becomes "inhibitory" because excitation to inhibitory neurons becomes stronger than that to excitatory neurons within the unit), but I can't easily imagine one that I would find persuasive.

A key feature of this model is "synaptic volatility" in the connections between HVC and BG. In songbirds, performance often improves steadily throughout the day, then deteriorates after sleep, a feature which the authors suggest is an important method for avoiding getting stuck on relatively low-performing local optima. I find this suggestion to be intriguing and reasonably persuasive. To implement this in their model, after a simulated "day" of practice, the HVC-BG connection weights are partially randomized (synaptic volatility). There is nothing intrinsically wrong with this idea, but the authors imply that there is experimental support for this phenomenon, which they relate to "continuous remodeling of the cortico-striatal synapses with volatility over hours to days." None of the papers cited really support this kind of synaptic randomization. However, given the simplicity of the authors' neural model, the HVC-BG weights are probably the only place this randomization can be implemented. The authors' implementation does not just add noise to HVC-BG weights "overnight"; it is also scaled inversely by the magnitude of accumulated weight change through the day. The authors do not appear to provide a justification for this aspect of their synaptic volatility.

If we accept the authors' model as detailed and accurate enough for their purposes, it does explain several features of song learning and relates them to solving general problems of learning through gradient ascent. This is particularly true of the daily deterioration of performance and how that relates to escaping local maxima. They show that their model reproduces some of the known effects of lesioning the motor (HVC-RA) and anterior forebrain (HVC-BG-RA) pathways and how those effects depend on the current stage of song learning. They also explore the implications of delayed maturation of the HVC-RA pathway, exemplified by a gradual increase in HVC-RA weights. However, it is not clear that this actually happens; the one paper they cite in support of this contention (Mooney and Rao, 1994) shows no such thing (it does show that HVC axons enter RA later in development than LMAN axons do). Moreover, it's not clear how essential this could be given that some songbird species continue to show profound vocal plasticity throughout their lives and presumably long after their HVC-RA pathway has fully matured. The authors compare their dual pathway model to a single pathway model (an AFP-only model, in effect); a very illuminating comparison that demonstrates the advantage of the dual pathway. They close by showing that dual pathway performance is robust under variation of key parameters.

Reviewer #2 (Public review):

Summary:

The authors describe a computational model for the acquisition of sensorimotor skills and explore these dynamics using vocal learning in the zebra finch, a system rich in experimental data, to describe the developmental trajectory of vocal imitation by trial and error. They set up their model as a dual-pathway system, with a cortical pathway that drives the vocal effector and a basal ganglia (BG) pathway that uses dopamine-mediated reinforcement learning (RL) by gradient descent to optimize the vocal imitation process.

Strengths:

A key strength of the model, due in part to the fact that his model was generated by a computational laboratory that has also contributed significantly to the collection of experiment-driven empirical data, is that the model is biologically constrained and incorporates a considerable amount of experimental data, including some of the latest findings in the field. In addition to providing a compelling model for the acquisition of vocal learning, this biologically based RL model outperforms many current models. A key feature of this model, which makes it unique, is the implementation of a synaptic volatility variable within the BG pathway that aims to mimic published work showing that juvenile birds exhibit post-sleep deterioration. The model uses a motor output to drive a biophysical model of the avian vocal organ (syrinx) and explores not only the ability to copy song acoustic units (syllables) but also the underlying neural dynamics in both the cortical and BG pathways, showing that each converges onto the types of neural activity patterns that are observed experimentally. Because of the richness of experimental data in this system, the authors can perform "computational experiments" where they can block sleep-driven synaptic volatility or lesion various pathways to replicate experimental observations.

In addition to providing important computational insights to our understanding of vocal learning in the songbird, this study provides key insights into the general architectures that are optimal for RL by gradient descent. These include the conclusion that effective RL requires adaptive regulation of exploration and exploitation, that cortical consolidation must occur at a slower timescale than BG-driven exploration, and intriguingly that the introduction of synaptic volatility prevents RL models of incomplete learning by getting "stuck" in local minima.

Weaknesses:

In the methods section, the authors state "... HVC and RA layers are fully connected, as are the HVC and BG layers. Synaptic weights in these pathways are plastic, reflecting activity-dependent plasticity at RA and BG synapses." Unless I missed it, it is unclear how much the authors consider the synaptic differences between HVC and BG inputs to RA. This seems like an important feature to highlight, especially given that HVC-RA connections are primarily AMPA-mediated whereas those from LMAN are predominantly NMDA. The authors should be clearer about how they model these synapses and better highlight (and describe) the importance of these synaptic differences in their modeling efforts. Ideally, they should evaluate whether the differences in synapse type influence the outcome of their model. It would be interesting, for example, to test the effect on learning of synaptic conductance substitution (i.e., replacing NMADA with AMPA) on the LMAN-RA synapse.

The model focuses exclusively on the interaction of two converging pathways, and learning is based purely on acoustic feature properties of what seem like four independent syllables of similar or identical duration. For this model, this is fine. But it would be helpful for the authors to state more clearly that they are not modeling respiratory influences on syllable production, which include amplitude modulation of the syllables and expiratory pulse duration. It should be noted that the authors do not (unless I missed it) mention the existence or role of recurrent loops in song initiation (and possibly syllable sequencing). They should at least mention this in the discussion, perhaps as a limitation and item for future versions of the model.

Reviewer #3 (Public review):

Summary:

This study modeled vocal learning in zebra finches with a network of three components: a pathway with delayed/slow Hebbian learning that mimics the HVC-RA projection, a pathway with reinforcement learning that mimics the HVC-BG-RA projection, and a motor unit that mimics the syrinx and produces song output. The model convincingly reproduces key features of song learning and is a valuable step towards understanding how vocal learning is substantiated in the song system. Further examination of model assumptions and presentation of testable predictions would increase the impact of the study.

Strengths:

(1) The model reproduces several key features of song learning, including learning in a non-convex performance landscape, decreasing motor variability during learning, the relative importance of the HVC-RA and HVC-BG-RA pathways at different learning stages, and sleep-related deterioration.

(2) The model incorporates several key physiological properties of the song system, including performance-dependent dopamine signals to the BG, the neural variability in the system, and multiple global/local optima of the motor production landscape.

(3) The study convincingly demonstrates the advantages of a dual-pathway network over a single-pathway one.

(4) The model demonstrates the counterintuitive benefit of sleep-related performance deterioration for facilitating the escape from local optima to reach the global optimum.

(5) The model seems robust to some variations in model parameters and task structure.

Weaknesses:

(1) The study could be more impactful if the model can generate new, testable predictions. The predictions provided in the Discussion are not well justified. For example, with the HVC spine turnover, it seems unlikely that BG lesions would abolish overnight performance deterioration. Because the volatility term is inversely related to learning during the day, the changes in LMAN and RA during sleep are not necessarily larger than those during the day.

(2) The section "Neural activity patterns in the model parallels song system neurophysiology" seems fully anticipated because the model construction is based on the known neural activity patterns. Are there new testable predictions from the model?

(3) Certain model assumptions lack explanation or justification.
a) The authors treat "the delayed maturation of the cortical pathway" as an important component. However, it is unclear if/how this component was implemented in the model. If it was not included in the model, the authors should remove statements related to the delayed maturation idea.
b) How critical is the inverse relationship between learning and sleep-deterioration? If it is known that sleep-related deterioration is inversely related to learning during the previous day, a citation should be added. Similarly, it should be clarified if the spine turnover observed in HVC depends on previous plastic changes, like the assumed inverse relationship in the model.
c) Spine turnover has been demonstrated in HVC and not yet in Area X, but the model implements volatility only in the HVC-BG pathway. It would be important to compare the effects of sleep-related volatility in the HVC-RA and HVC-BG-RA projections.
d) In Table 1/Figure 8, the learning rate for HVC-BG is 10^4 times bigger than the learning rate for HVC-RA. What is the biological justification for this difference?

(4) Related to #3, it is unclear how changing those assumptions would affect model performance.

(5) It is not explained/shown why cross-day exploration (with sleep, sporadic) is better than continuous, non-sleep-related ones. Figure 7B presents results with different noise levels, which may approximate non-sleep-related, continuous volatility, but only for the single-pathway model. Comparable simulations by adding continuous volatility in the dual-pathway model would be helpful. For example, would just a bigger intrinsic noise in either pathway confer the same benefit in escaping local optima, e.g., epsilon-greedy exploration? If the authors can establish the advantages of sleep-specific volatility in its particular form (based on day learning) and relate it to other sensorimotor learning behaviors, it could increase the general impact of the study.

(6) The model description can be improved.
a) What are mBG and mRA in Eqs 6/7?
b) Where is the learning rate specified in Equations 1-8?
c) It is unexplained why Equations 1-3 are not in the same form: Equations 1 and 2 normalize by activity, but Equation 3 normalizes by weight (Equation 8 also normalizes by weight). The authors should confirm that these equations are correct.
d) How J_HVC is generated should be defined.
e) How is R calculated in Equation 10? Does this model maintain a PPE for each syllable or for the overall performance? Would it make any difference in learning?
f) Are the weight updates in Equations 9 and 13 added at different time points (e.g., immediately after each spike versus after a syllable). This should be clarified (perhaps a diagram would help).
g) How are the optimal threshold and slope determined for Equation 14?
h) Parameters in Equation 20 are not defined.

Author response:

We thank all three reviewers for their detailed and constructive reviews, which will help us improve the manuscript. Concerning the simplicity of the model, we will elaborate on the limitations that arise from the modelling choices and better justify these choices. With that in mind, we believe this level of abstraction is well suited to the study’s objective. As a rate-coded model, it is designed to capture interactions between multiple circuit pathways, and lays the foundation to delve into neuronal dynamics in the future for further insight, once the network-level questions we address here - regarding the exploration-exploitation tradeoff and evasion of sub-optimal performance convergence - are well examined. For this purpose, we believe the model’s simplified network-level perspective is not a bug, but a feature. We will also explore in depth why and how the dual pathway model performs better in non-convex sensorimotor landscapes, building on prior analytical work that speaks to the question (Sankar, Leblois & Rougier, ICDL 2022). We will further substantiate and discuss our choice concerning the potential mechanisms underlying the proposed overnight synaptic volatility. We will include comparable approaches (including relevant machine learning analogues) in the introduction. Finally, we will clarify our interpretation of the sign of synaptic weights. We hope to incorporate all valuable feedback by the reviewers in the revised manuscript.

References

Remya Sankar, Arthur Leblois, Nicolas P. Rougier (2022). Dual pathway architecture underlying vocal learning in songbirds. In IEEE International Conference on Development and Learning, ICDL 2022, London, United Kingdom, September 12-15, 2022. pages 265-271, IEEE, 2022.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation