Learning to cancel tutor song results in an error landscape sufficient for subsequent motor learning.

(A) Schematic of the general hypothesis. Neuronal responses to sensory input define a multidimensional, time-dependent “error landscape.” During the sensory learning phase, local learning rules increase the landscape’s curvature (sensitivity to error) and move its minimum toward the target tutor song (red dot). In the subsequent sensorimotor learning phase, this error serves as a negative reward, and reinforcement learning algorithms refine the bird’s own song by performing error minimization on this landscape. (B) General model structure. A sparse coding model maps tutor song to auditory input, which is sent on to meet song-locked premotor timing input in secondary auditory nuclei. Local plasticity rules in these regions gradually learn to perform predictive cancellation of this pattern, giving rise to sparse error codes.

Four classes of models for local learning.

(A) Feedforward model. Song-related premotor input directly drives excitatory neurons in secondary auditory areas, which receive both excitatory and inhibitory input from primary auditory areas. Plasticity takes place at the premotor→secondary auditory synapse (1). (B) Balanced excitation-inhibition (EI) networks with three potential sites of plasticity (24). As in (A), excitatory neurons in secondary auditory areas receive input from both primary auditory areas and premotor areas, as well as a local population of inhibitory interneurons. Numbers indicate the site(s) of plasticity in each of the three models. (C) Schematic of rate-based Hebbian and anti-Hebbian plasticity rules. Hebbian rules strengthen synaptic connections when postsynaptic excitation exceeds a threshold θ. Anti-Hebbian rules do the reverse. (D) Distributions of initial recurrent weights for the four types of connections. The bins at zero represent the 50% of connections in each model type that are set to 0 for sparsity. In general, inhibitory weights need to be stronger than excitatory weights to achieve EI balance. A detailed description of the parameters and construction of weights is in Methods and Table 1. (E) Illustration of EI balance. Left vertical axis: density of the distributions of E→E (blue), I→E (red), and their net inputs (gray); right vertical axis: output firing rate of the f-I curve (black). EI balance is achieved when the distribution of net inputs lies near the activation threshold of the f-I curve.

Summary of comparisons between models and experiments.

Neurons produce sparse error codes after learning to cancel tutor song auditory patterns.

(A) Two example neurons in the E→I→E model decrease firing over learning. Gray dashed vertical lines mark song onset. Other models are qualitatively similar. (B) Population mean firing rates decrease over training, relative to the initial population means. Feedforward and non-recurrent models exhibit the largest percent decreases. (C) Quality of cancellation (see SI Methods) in the four models as functions of premotor projection density, plasticity threshold θactive, and excitatory neuron activation threshold θE. Positive (or negative) quality of cancellation indicates lower (or higher) similarity of population activity with tutor song patterns after training. Widths of the shaded areas denote 1 s.d. across n = 5 initializations of each model. Models with recurrent plasticity perform best in the physiological regime of sparse inputs and low firing thresholds. (D) Learned singing responses under practice and perturbation. Raster plots show the responses of all excitatory neurons during singing under conditions of correct song (matching the tutor song), perturbation (added white noise), and deafening (no auditory input). Neurons are sorted by the difference between auditory input (bird’s own song) patterns and the sample-averaged tutor song pattern. (E) Population mean rates in the correct (black), deafened (grey), and perturbation (solid red) cases during singing. The purple bar in each subplot indicates the 50-ms white noise perturbation to the auditory feedback of the bird’s own song. Only the E→E model fails to exhibit a population response to error. (F) Percentages of active excitatory neurons (, see SI Methods) in different models, averaged across all samples, from the earliest perturbation onset to 100 ms after the latest perturbation offset. Except for the E→E model, a significantly denser set of neurons are active during perturbed singing than during correct singing (double asterisks indicate p < 10−3, one-sided permutation test). Note that the learning rate is set to zero in (D-F).

Comparison of experimental data to models with different sites of local plasticity.

(A) Left: schematic of the pan-neuronal calcium imaging experiments. Right: z-scored responses of four example units in the case of normal singing (dashed lines) and singing with auditory perturbation or deafening (solid lines). (B) Distributions of trial-averaged z-scored calcium fluorescence in correct versus perturbed singing in area CM of adult male zebra finches. Top row: scatterplot of white noise perturbation responses versus responses on correct singing trials (n = 397 ROIs). A sparse set of neurons are more activated by white noise perturbation (red arrow). Each dot represents the trial average of the normalized activity of one neuron over the song period (see SI Methods). The horizontal and vertical error bars at each dot represent the standard errors for the correct and perturbed or deafened trials. Bottom row: same as top row, but comparing responses post-deafening with correct singing responses pre-deafening (n = 161 ROIs). (C) Same as (B) for the four classes of plasticity models (with model firing rates instead of fluorescence). Only 300 randomly selected neurons were plotted for visualization purpose. (D) Wasserstein distances between the distributions from experimental data (B) and from models (C) indicate that the E→I→E model is the closest to experimental data.

The minimum and shape of the error landscape are encoded by different connectivity modes in the E→I→E model.

(A) Schematic of acoustic variation used to probe learning. Auditory feedback inputs were varied both in the song error direction (dark red arrow), which interpolates between the correct song and a random pattern, and song variation direction (dark blue arrow), which scales the correct song pattern between deafened and correct singing. The song error direction represents auditory input off the manifold of correct song (light blue region), while the song variation direction partially overlaps with the manifold. (B) Left: Mean relative singing responses to the auditory feedback in (B) the song error direction during training. 0: correct song; 1: random auditory feedback pattern. Relative responses are defined as the percent change in population firing rate compared to the correct singing responses. Learning progress increases from purple to yellow. Right: The midpoint slope is the slope of the line tangent to the curves at 0.5 on the left plot. Higher values indicate greater sensitivity to auditory feedback. Widths of the shaded areas denote 1 s.d. across n = 10 simulations. (C) Left: Mean relative singing responses to the auditory feedback in the song variation direction during training. 0: no auditory feedback (deafened); 1: correct song. The dot on each curve in the left panel marks its minimum. Right: Minimum location of the responses in the left panel. Higher values indicate minimal error responses to away from silence and closer to the tutor song. (D) Schematic of the singular value decomposition of the concatenated connectivity matrix. (E) The sorted singular value spectra before (blue) and after (orange) learning. The spectrum exhibits a clear kink that we use to differentiate between dynamics, landscape, and memory modes and “other modes.” (F) Left: Dissimilarity between the non-memory modes and their initial values over training time. The green curve is the dynamic mode, which barely changes, and the red curves are landscape modes, which change rapidly and quickly stabilize. Right: Strength of memory encoding for all modes during training. The blue curves are those with final memory encoding strengths larger than the maximum pre-learning memory encoding strength (i.e., noise correlation). (G, H) Altering landscape and memory modes alters population error responses. Comparisons among the original trained models (black); models with the top 15 landscape modes removed (red), with memory modes removed (blue); and with the top 15 non-memory, non-landscape modes removed (purple). Widths of the shaded areas denote 1 s.d. across n = 35 simulations of each model. Asterisks in the the right panels indicate significant differences between distributions (p < 0.01; two-sided Wilcoxon signed-rank test; colored asterisks: altered model versus original model; black asterisks: models with perturbed memory versus landscape modes). Perturbations of the landscape modes primarily affect the error landscape slope, while perturbations of the memory modes alter its minimum. (I) Schematic of changes to the error landscape as a result of learning.

Error codes produced by local learning rules can be used to train a motor policy.

(A) Schematic of the actor-critic reinforcement learning model. A syllable-based state variable s(t) indexes both the motor policy (actor a(s)) and the prediction of the expected reward (critic V (s)). Syllables y(s) are generated by linearly combining spectral basis elements B using the outputs of the policy a(s). The negative of mean error signal serves as the reward signal to be maximized. (B) Excitatory population firing rate (error signal) and temporal difference (TD) error plotted as a function of trial. Results shown are for the E→I→E model, but results for the feedforward and premotor→E models are similar (see Fig. S11). (C) Tutor syllable templates (top row), and model-produced song before, during, and after learning (last three rows). The model is able to reproduce the tutor song using only the learned population error code as feedback.