The dual pathway model for vocal learning.

A. Timeline of song learning. The model learns to produce 4 syllables (50 ms each) per motif, and practices the motif 1000 times per day for 60 days (dph: days post hatch), similar to the sensorimotor phase of zebra finches. B. The neural substrates involved in vocal learning in zebra finches, the song system. The cortical pathway driving adult song production, in blue, with dotted lines representing its delayed maturation and gradual strengthening over the sensorimotor phase. The subcortical pathway, in yellow, necessary for song learning and maintenance, receiving information about performance quality through dopaminergic input. C. The architecture of the dual pathway model is derived from the song system. The cortical pathway, in blue, consists of two nuclei, HVC (used as a proper name) and the robust nucleus of the arcopallium (RA). The pathway strengthens during learning (dotted). The subcortical pathway, in yellow, consists of one basal ganglia (BG) neuronal population, representing the BG-thalamo-cortical loop (comprising Area X, medial part of the dorsolateral nucleus (DLM) and lateral magnocellular nucleus of the anterior neostriatum (LMAN)). The motor control (MC) layer transforms the RA neural commands to the appropriate vocalisation parameters (see Methods. D. Representative neural activity in each neuronal population of the model at the end of the sensorimotor phase. For each population, the top panel shows the firing rate generated for each 10 ms time bin during song, and the bottom panel shows illustrative spiking activity (see Methods), where each vertical black line denotes a spike. HVC neurons (black) fire sequentially, providing a timing scaffold for the song. By the end of the sensorimotor period, RA neuronal populations (blue) display sparse, stereotyped bursts across renditions. BG neuronal populations (yellow) exhibit high rendition-to-rendition variability in firing rate. E. Examples of performance landscapes generated using a biophysical model of the syrinx (see Methods) with four target syllables (inset). The performance metric at each point (determined by pressure, P, on y-axis, tension, T, on x-axis) on the landscape is indicated by the grey heatmap. The landscapes have a few global optima (performance metric ≈ 1) and several local optima (performance metric < .7). F. Examples of Gaussian-based landscapes from each density category. The landscape is generated by superimposing several Gaussian functions (see Methods). The superimposition results in several local (performance metric < .7) and one globally optimal solution (performance metric = 1). The landscapes are divided into four classes according to the number of local optima present: low (2-5), moderate (3-9), compact (8-18), high (11-25).

Song acquisition in the dual pathway model.

A. Target song composed of 4 syllables. B. Model simulation of the learning process for each syllable over the sensorimotor period. The model starts at a random initial position (square) and is able to find the optimal solution (cross) over the sensorimotor period of 60 days with 1000 iterations each day. The output at each iteration is shown as a black dot. Each panel shows the motor outputs produced for the corresponding syllable. C. Example trajectory of the model output. The two-dimensional MC output is fed as pressure and tension in the syrinx model, driving the production of each syllable. The model starts at a random initial position (square), produces several sub-optimal vocalisations, over the coarse of learning (smoothened trajectory in black) and eventually finds the globally optimal solution (cross). Two example output vocalisations are shown on the right, corresponding to a locally (top) and globally (bottom) optimal performance, with differing fundamental frequencies. D. Smoothened trajectories for the control parameters representing the syrinx pressure and tension for the example syllable simulation in C are shown in black. The target output for each control parameter is shown in grey horizontal dotted lines. The model explores different outputs (grey dots) in the two dimensions in the beginning of sensorimotor learning and eventually converges at the target output. E. Smoothened trajectory for the performance metric obtained by the model at every rendition of the example syllable over the sensorimotor period (in black). Grey dots represent performance metric at each rendition. F. Performance of the model over 100 simulations on the syrinx-based landscapes. Left: Terminal performance for all simulations (black dots), calculated as the average performance over the last 100 trials on the last day (see Methods). Performance metrics above the threshold (shaded portion above the black dotted line), lie close to the global optimum. The violin plot denotes the distribution of the terminal performance (black dots) across the 100 simulations. Right: Success rate for each syllable, i.e. the proportion of simulations that learned the optimal performance (achieved terminal performance > .7). G. Performance of the model over 100 simulations on the gaussian-based landscapes, with increasing density of local optima. Same conventions as in F.

Effects of inactivating the BG and cortical inputs to RA, respectively, at different stages of the sensorimotor learning period.

A. Performance before and after inactivation of BG input to RA at different stages of sensorimotor learning in representative simulations. The three columns show the effect of BG inputs being inactivated before (left), during (middle) and after (right) sensorimotor learning. Top panel: Performance (black) two days before and after lesion. Vertical dotted line denotes the lesion time. Middle panel: Performance (black) throughout the sensorimotor period. Bottom panel: Motor variability over the sensorimotor period. In all three panels, the grey dots denote the performance metric value at the end of the day before inactivation. The light blue dots denote the metric at the start of day after inactivation. The blue/yellow dots denote the metric at the end of the sensorimotor period. B. Distribution of success rate (top) and terminal performance (bottom) across 100 simulations before and after inactivation of BG input to RA performed at different stages of sensorimotor learning. Bottom: Terminal performance for all simulations (black dots, see Methods). Performance metrics above the threshold (shaded portion above the black dotted line), lie close to the global optimum. The violin plot denotes the distribution of the terminal performance (black dots) across the 100 simulations. Top: Success rate for each condition (proportion of simulation achieving terminal performance > 0.7). C. Motor variability across 100 simulations before and after inactivation of BG input to RA performed at different stages of sensorimotor learning. Black dots show the motor variability for each simulation (see Methods) while violin plot denotes the distribution of motor variability across the 100 simulations. D-F. Corresponds to A-C for the suppression of HVC input into RA (in yellow).

Neural activity patterns across the sensorimotor learning period (40 dph vs 100 dph) in the model.

A. Neurons in the HVC layer encode time during song and provide context information to the model. The activity pattern of HVC neurons (black) across one song rendition. Top panel: Firing rate of multiple neurons across time steps during one song rendition. Bottom panel: Poisson spike trains generated based on the above firing rate profiles (see Methods). Each vertical black line corresponds to one action potential. B. Neurons in the RA layer have variable firing rates across song renditions in the beginning of sensorimotor learning and stereotyped patterns towards the end of sensorimotor learning. Top panel: Firing rate of one RA neuron (shades of blue) across ten song rendition. Bottom panel: Poisson spike trains generated based on the above firing rate profiles. C. The neurons in the BG layer have variable firing rates across song renditions throughout sensorimotor learning, akin to LMAN neural activity. Top panel: Firing rate of one BG neuron (shades of blue) across 10 song renditions. Bottom panel: Poisson spike trains generated based on the above firing rate profiles.

Synaptic volatility induces overnight displacement in performance and aids sensorimotor learning.

The two columns show one representative example simulation of the learning of a single vocal gesture (single time step), with (right) and without (left) overnight synaptic volatility. A. Representative learning trajectory in the absence of synaptic volatility. Each black dot corresponds to one rendition of the motor output. The model converges (cross) at the sub-optimal solution closest to the initial position (square). B. Smoothened trajectory (black line) of the performance metric (black dots) over a representative simulation. Vertical grey lines denote day change. Inset: Performance metric over two days during the mid and late stages of sensorimotor learning displays a monotonic performance quality. C. Smoothened trajectory (yellow line) of the synaptic strength (yellow dots) of one representative HVC-BG weight. Vertical grey lines denote day change. Inset: Synaptic strength over two days shows gradual changes in the HVC-BG weight during the day and absence of overnight fluctuation. D. Representative learning trajectory with synaptic volatility, with same convention as in A. The model escapes sub-optimal solutions and converges (cross) at the globally optimal performance. E. Smoothened trajectory (black line) of the performance metric (black dots) over a representative simulation. Vertical grey lines denote day change. Inset: Performance metric over two days during the mid and late stages of sensorimotor learning displays gradual changes in performance during the day and sudden overnight fluctuations. The magnitude of overnight performance fluctuation reduces over the sensorimotor period. F. Smoothened trajectory (yellow line) of the synaptic strength (yellow dots) of one representative HVC-BG weight. Vertical grey lines denote day change. Inset: Synaptic strength over two days shows gradual changes in the HVC-BG weight during the day and strong fluctuations overnight. G. Change in HVC-BG synaptic strengths across the day vs across the night over one simulation with and and one without overnight synaptic volatility in the BG pathway. H. Terminal performance (left, black dots) and success rate (right) of the model over 100 simulations with and without overnight synaptic volatility in the BG pathway.

Delayed maturation of cortical pathway modulates the exploration-exploitation trade off.

A. Schema of delayed maturation of the HVC projections to RA and its growing strength over development in male zebra finches. B. Representative trajectory of the model output for one time-step (single vocal gesture) in the dual pathway model. Smoothened trajectories of the overall motor output (top, black line), contribution of the cortical pathway to motor output (middle, blue line) and contribution of the BG pathway to motor output (bottom, yellow line). The cortical contribution slowly consolidates a trace of the BG-led exploration of the motor space. Squares denote initial positions and crosses denote final positions. C. Learning trajectory along each dimension. The overall motor output (black) overlaps primarily with the BG contribution (yellow) in the beginning of the sensorimotor phase. Over time, it aligns increasingly with the cortical pathway contribution (blue), as the model approaches convergence. As the motor output reaches the target output (grey dotted line), the BG contribution slowly fades to zero. D. Change in motor output across different stages of sensorimotor tasks (Top: overall output, middle: cortical pathway contribution, bottom: BG pathway contribution). The histogram show the distribution of motor output (dots) on the landscape in the two dimensions. BG contribution (bottom panel) initially induces a strong displacement in the overall motor output (top panel) as the cortical contribution (middle panel) incorporates a slow trace. However, its effect on motor output fades over time, eventually contributing to moderate exploration around the mainly cortical-driven output. E. Growth of synaptic strength (shades of blue) in the cortical pathway over the sensorimotor period (each line represents a single synapse). The HVC-RA synapses, initalized at zero weight, drive a strong net excitatory or net inhibitory effect on the postsynaptic RA populations at the end of learning. F. Top panel: Motor variability reduces over the sensorimotor period. Middle panel: Variability of RA activity reduces over the sensorimotor period. Bottom panel: Variability of BG activity remains stable over the sensorimotor period.

Comparison of different strategies for sensorimotor learning.

In A-D, Terminal performance over 100 simulations (black dots) on artifical landscapes is depicted on the left for various levels of exploratory noise (violin plot: distribution of the terminal performance). Success rate, i.e. the proportion of simulations that learned the optimal performance (achieved terminal performance > .7), is depicted on the right for the same levels of exploratory noise. A. Performance of a single pathway architecture employing gradient descent based reinforcement learning with no overnight synaptic volatility (StdRL). B. Performance of a single pathway architecture employing gradient descent based reinforcement learning with exploratory variability reducing exponentially over development (DevRL). C. Performance of the dual pathway model, with overnight synaptic volatility and no explicit reduction in BG variability. D. Performance of a simulated annealing algorithm on sensorimotor learning on gaussian-based landscapes.

The dual pathway model is robust to fluctuations in key parameters that affect learning.

Here, the performance of the model is evaluated over 100 simulations for different values of each parameter considered. A. Synpatic noise into BG neurons. B. synaptic noise into RA neurons. C. Learning rate for plasticity at HVC-BG synapses. D. Learning rate of plasticity at HVC-RA synapses. E. Number of renditions over which prediction of performance () is computed (temporal integration window). F. Width of the global optimum gaussian in the gaussian-based performance landscapes. G. Scaling factor for overnight synaptic volatility at HVC-BG synapses. H. Probability of occurrence of overnight synaptic volatility at HVC-BG synapses.

Model parameters and their values used in all simulations unless otherwise stated.

Construction of sensorimotor performance landscapes.

A. Interface of model with syrinx-based landscapes. The neuronal activity pattern in RA is reduced in dimension at the MC layer. The output from the MC layer modulates 50 ms tension and pressure waves, which is fed as input to a biophysical model of the syrinx to produces zebra finch-like vocalizations. For each output of the model, the produced vocalization is compared with a target syllable to generate a performance metric (grey heatmap). Thus, the performance landscape represents the imitation quality for every model output. B. Four sample performance landscapes generated using four target zebra finch syllables (inset). The landscapes have a few global optima (performance metric ≈ 1) and several local optima (performance metric < .7). C. Interface of model with gaussian-based landscapes. The output from the MC layer is directly used as coordinates on these sensorimotor performance landscapes. D. Examples of gaussian-based landscapes from each density category. The landscape is generated by superimposing several gaussian hills. The superimposition results in several local (performance metric < .7) and one globally optimal solution (performance metric = 1).