Author response:
The following is the authors’ response to the original reviews.
Public Reviews:
Reviewer #1 (Public Review):
Summary:
This study examines the context-dependent modulation of auditory cortical neurons in response to expected sensory input, either self-generated sounds or expected perturbations of self-generated sounds. Specifically, using songbirds, the authors ask whether social context (the presence of a female conspecific) affects 1) the response of auditory cortical neurons to the bird's own song when he is singing; and 2) the response of neurons to perturbations of auditory feedback that the bird has been trained to expect.
Strengths:
First, the authors report that across the population, the responses of the neurons does not differ when a male bird sings alone or if he sings to a female. A fraction of auditory cortical neurons, however, do show significant differences in the firing rate, precision, and/or degree of burst firing when males sing alone vs. when they sing to females. This finding is broadly consistent with the literature showing that sensory neurons (visual, auditory, somatosensory, etc.) can be rapidly reconfigured into different "information processing modes" depending on behavioral state (e.g., quiescence vs. vigilance).
For the perturbation experiments, the authors trained birds to expect distorted auditory feedback during a particular syllable. They found that some neurons showed greater responses during perturbation when a female was present (compared to when males were alone) while other neurons had smaller responses during perturbation when a female was present. In addition, the response of a small number of auditory cortical neurons were not affected by behavioral state. These results contrast with their prior report that the responses of midbrain dopaminergic neurons that project to the basal ganglia are "uniformly reduced" in the presence of a female, raising a question of how an evaluation signal is transformed in the circuit from the primary sensory region to the midbrain.
Weaknesses:
While the experiments and analysis are solid, the finding that social context can alter responses of auditory cortical neurons in a multitude of ways (increase, decrease or no change) raises several questions that can be examined with additional analysis. For example, do context-dependent differences in auditory responses derive from context-dependent differences in the songs? Are context-dependent differences present in all classes of neurons and throughout the auditory system?
The observed heterogeneity in the firing properties of auditory cortical neurons, both in response to self-generated sounds and during perturbations of auditory feedback, raises the question of which neurons are sensitive to social context (which likely can be addressed by the authors in a revision). The authors should provide additional details about the recordings:
(a) What are the locations of the recording sites? Prior work has shown that there is an organized map of spectrotemporal features of sounds in the auditory cortex of songbirds; spectral tuning widths change along the medial-lateral axis and temporal tuning widths differ between the input and output layers of Field L. Were the recordings primarily in Field L2 (thalamo-recipient region), L1 or L3? Were some recordings lateral to Field L in secondary auditory regions? Were the neurons that showed context-dependent changes in firing properties localized or distributed throughout Field L (i.e., were the context-dependent differences in neural responses truly brain-wide)? At a minimum, the authors should include a schematic showing the different regions of Field L and a summary of the location of the recording sites. Images of the processed tissue with electrolytic lesions would also be helpful.
We agree that the anatomical targeting and limits of localization should be made explicit. In the original manuscript, we referred broadly to recordings from "Field L" and described targeting coordinates in the Methods. In the revised manuscript, we have softened the anatomical claim from "Field L" to "auditory pallium" where appropriate, while explicitly stating that electrodes were aimed at Field L. We also added anatomical caveats and a new supplemental figure.
The revised title and abstract now reflect this more conservative anatomical framing. For example, the abstract now states: "Here we recorded neural activity from the auditory pallium in zebra finches practicing singing alone and directing courtship songs to females." In the Introduction, we now explicitly state both the intended target and the limitation: "We targeted our recording electrodes to Field L, a primary auditory pallial area that projects into multiple higher auditory areas that, in turn, project to VTA."
We then added the caveat: "Field L is composed of multiple subdivisions and surrounds the interfacial nucleus and because the implanted wire bundles spread in a small radius of up to ~0.5 mm, our recordings likely included large territories of the auditory pallium (Figure S1)."
We also added mechanistic/anatomical context for why these recordings may reflect activity shaped by broader auditory forebrain circuitry: "Although Field L is classically described as a primary auditory thalamorecipient region, its activity may also be shaped by contextual signals related to the courtship context, potentially via recurrent interactions with higher-order auditory forebrain regions such as the caudal mesopallium (CM) and caudomedial nidopallium (NCM) (Bauer et al., 2008; Figure S1C)."
Changes made in revision: We added Figure S1, which includes anatomical subdivisions, an example histological slice showing the area where the cannula was implanted, and auditory pathway connectivity. We also revised the wording throughout the manuscript from "Field L neurons" to more conservative phrasing such as "auditory pallium neurons" or "pallial auditory neurons" when appropriate. We did not claim layer-specific localization, because the revised manuscript explicitly states that we cannot make such claims.
(b) Was the context-dependent modulation limited to a particular class of neurons (distinguished by spike waveform shape, spontaneous firing rate, or other feature)?
We agree that identifying whether context-dependent modulation is associated with specific neuronal classes is important. In the revision, we added analyses examining relationships between spike width and various firing characteristics. We also looked for potential relationships between mean rate and DAF response, IMCC and DAF response, and found no clear trend. Overall, we did not observe any clear relationship between DAF response, DAF response modulation, and metrics like spike width or mean firing rate.
The revised manuscript states: "Action potential width of individual neurons has previously been used to classify putative interneurons or putative principal cells in the zebra finch auditory pallium (Calabrese and Woolley, 2015)."
We then describe the new analysis: "We tested if DAF-response scores, mean firing rates during singing, burst fraction, IMCC, and the change in all of these between undirected and directed singing was correlated with spike half width (spike half-width measured as peak-to-trough time; Figure S5)."
The revised result is: "DAF response in either condition, the change in DAF response across conditions, and the change in firing rate, burst fraction, and IMCC were not significantly correlated with spike width (Figure S5A-B)."
We also report that some general firing properties did correlate with spike width: "Consistent with previous literature, firing rate was significantly correlated with spike width (Pearson's correlation, p=0.003; Figure S5C). Interestingly, burst fraction (p=8.3x10-4) and IMCC (p=5.4x10-5) were also significantly correlated with spike width (Figure S5C)."
Changes made in revision: We added Figure S5 and associated text analyzing whether context-dependent changes in DAF response, firing rate, burst fraction, and IMCC were correlated with spike half-width. These analyses did not support the conclusion that context-dependent DAF modulation was restricted to a waveform-defined neuronal class.
(a) Prior work has shown that songs of zebra finches differ slightly when males sing alone compared to when they sing to females: songs are faster; pitch is less variable; and the number of introductory elements is greater when males sing to females. Do some of the observed social context-dependent differences in the responses of auditory neurons reflect differences in the songs in the two conditions? Did the authors of this study also find premotor activity in Field L, and if so, did it differ between the two social contexts? Might differences in Field L responses reflect motor/song differences?
The revised manuscript now addresses the issues of context-dependent changes in song in several ways. First, we explain why motif-aligned comparisons are meaningful: "The acoustic structure of undirected and directed motifs is highly similar in adult finches, enabling singing-related neural activity to be precisely aligned and compared across contexts."
Second, we added analysis and discussion of premotor-related activity. Changes made in revision: New results paragraph and new Fig S4. "A previous study recording from Field L in zebra finches reported neural activations prior to the onset of singing, consistent with premotor signaling (Keller and Hahnloser, 2009). We tested for context-dependent changes in premotor activity by examining peaks in neural activity aligned to motif onsets. Across the population neurons did not exhibit significant changes in the timing of motif onset-aligned activity (Figure S4)."
(b) For the perturbation experiments, this raises a question of whether perturbation amplitude is different when a male is alone and when a female is present. It would be useful to know if (and how much) perturbation amplitude varied depending on the location inside the cage as well as whether the sound pressure level of the underlying song was higher (e.g., Lombard effect).
We previously calibrated the perturbation amplitude in Roeser et al., 2023, in an identical recording setup. Two speakers deliver the feedback on either side of the bird's home cage. We acknowledge the possibility that the position and orientation of the bird can affect the way the sound hits either of the bird's eardrums and thus potentially affect a neural response. However, neural activations following the absence of distortion playbacks were a major feature of the dataset and were context-dependent in some cases.
The Methods state: "DAF was implemented with a custom LabVIEW acquisition program that analyzed song syllables in real-time and delivered syllable-targeted feedback." and "DAF (50 ms broadband noise bandpass filtered at 1.5-8 kHz to match frequency range of zebra finch song) was played over speakers in the recording chamber on top of a specific target syllable randomly on 50% of motif renditions."
The revised manuscript also makes clear that experiments occurred in the bird's home cage: "Experiments were carried out in the male's home cage, which was inside a sound isolation chamber."
Importantly, the revised Results show that not all DAF-related responses were simple activations to additional sound. Some neurons were activated by the absence of distortion: "Unexpectedly, some neurons were not activated by the song distortion but rather by the lack of target syllable distortion." and "These activations following undistorted renditions could also depend on the courtship context."
Changes made in revision: We clarified the DAF stimulus and recording setup in Methods and added Figure 3 showing neurons activated following undistorted renditions. These data argue that context-dependent responses are not simply explained by DAF sound amplitude, although we do not claim that position-dependent acoustic variation was fully eliminated.
Finally, it would be helpful if the authors could include a model and/or more discussion of how the uniform attenuation in midbrain dopaminergic neurons may arise given the heterogeneous responses in Field L.
The revised manuscript provides evidence for context-dependent retuning upstream of VTA, but does not offer a direct mechanistic explanation for the uniform attenuation seen in dopaminergic neurons. The revised Discussion states: "Because the main goal of this study was to test if courtship-associated reduction in DAF signaling, recently observed in VTA DA neurons (Roeser et al., 2023), resulted from a local process in VTA or reflected a retuning of auditory responsiveness, we explicitly tested for changes in DAF responsiveness between alone and female-directed singing."
It then explicitly contrasts auditory pallium and VTA: "Surprisingly, we discovered that Field L neurons could retune at the transition from lone to courtship singing in diverse ways, consistent with a more widespread process in the brain that does not fully explain the uniform DAF-signal attenuation observed in VTA."
Changes made in revision: We expanded the Discussion to explicitly state that auditory pallium retuning is heterogeneous and therefore does not fully explain the uniform attenuation observed in VTA. We do not present a formal circuit model, but we now more clearly frame the result as evidence for broader sensory retuning that is likely transformed downstream.
Reviewer #2 (Public Review):
Summary:
In the manuscript, Jones and Goldberg study auditory cortex in male zebra finches. They explore song-related responses in two different contexts, when the male is either alone or in the presence of a female. They find a heterogeneity of responses, in line with auditory cortical neurons computing the social modulation of responses found in VTA.
Weaknesses:
Stability of responses has not been studied: some neurons seem to have responses that slowly drift in time, which could lead to observed differences between alone and with-female conditions. Also, possible motor confounds and sound-of-audience confounds should be addressed. The language is often imprecise.
Stability and Reversal: It is a bit unfortunate that stability of effects seemingly has not been studied by reversing experimental conditions. The work would be much stronger if authors could show that audience-dependent tuning is robust in individual cells. Did they record from some neurons during reversal back to the alone condition?
We agree that recording stability is essential. A reversal experiment was not feasible for this dataset, as it is difficult to confirm whether song motifs produced immediately following female presence represent undirected singing or are directed to an unseen but recently present female. Instead, the revised manuscript adds a strict unit-stability criterion based on waveform similarity across conditions.
The revised Results state: "Importantly, because these neural recordings were performed over long time courses (~2-8 hours), a strict threshold for stability was imposed." The exact criterion is: "A Pearson's correlation coefficient of at least 0.99 between the average neural waveform during undirected and directed singing was required for a unit to be considered stable (Dickey et al., 2009; Figure S2)."
Changes made in revision: The strict waveform-stability inclusion criterion and new Figure S2 directly showcase unit stability across the time course of the experiments.
Motor responses: Does DAF playback change song? If so, especially if it applies only in one of the two conditions (audience/no audience), then the observed response differences could be motor-related rather than auditory responses.
We agree that motor confounds must be minimized. We previously found that DAF did not affect the acoustics of the subsequent syllable (Gadagkar et al., 2016). The revised manuscript clarifies that DAF and undistorted trials were randomly interleaved and analyzed by comparing matched renditions within conditions. Importantly, we only analyzed motif-aligned activity, ensuring that all syllables within the song motif are the same.
Changes made in revision: We clarified the DAF analysis framework and added a more conservative permutation-based analysis comparing distorted and undistorted trials within each context, then comparing those DAF-response vectors across contexts. We do not claim that all possible motor consequences of DAF are eliminated, but the analysis directly tests neural responses to randomly interleaved distorted versus undistorted renditions.
Similarly, motif-aligned spiking activity was time warped to the median duration of undirected or directed motifs. Could the shorter motifs during directed song lead to alignment differences that would account for the different error responses in alone/with-female conditions?
We agree this is an important technical point. The time-warping we conducted, standard in the field, compensates for the tempo differences between directed and undirected song. Importantly, our main analysis of change in error response no longer uses a 100 ms response window, but rather includes all windows in the motif.
Changes made in revision: We clarified that the revised DAF response analysis uses motif-aligned, time-warped spike trains. Importantly, the revised analysis moves away from relying on a single scalar response window and uses bin-wise permutation tests with family-wise error correction.
Audience versus sound of audience: Is it truly the audience that causes the difference in error responses or is it the sounds the audience makes?
We agree that the sensory cues defining "audience" cannot be fully separated in this experiment. The reviewer raises an important point that female zebra finches occasionally call at the male. We have excluded all song motifs from analyses that include an overlapping female call.
The revised Methods now explicitly state that motifs overlapping with female calls were excluded: "Any motifs that had overlapping time with a female call in directed motifs was excluded from analysis."
We also revised the Discussion to treat the mechanism by which auditory pallium receives information about the female as an open question: "An open question is how auditory pallium receives information about whether a female is present, and how this information influences neural activity."
Changes made in revision: We excluded motifs overlapping with female calls and added discussion explicitly acknowledging that how female presence is represented in auditory pallium remains unresolved. We do not claim to distinguish visual, auditory, social, or motivational components of the female-present condition.
Reviewer #3 (Public Review):
Summary:
In this study, Jones et al. examine how neural activity in a primary auditory area (field L) of singing male songbirds is modulated by the presence or absence of an audience (a female conspecific). Prior work has demonstrated that the presence of an audience attenuates the responses of dopaminergic neurons to distortions of auditory feedback (DAF). Here the authors report that even in a region that is primarily considered sensory, responses to DAF are also modulated by the audience, although in a heterogeneous manner. However, to be fully persuasive, additional analyses will be required to address how much of the apparent modulation by audience may be explained by other factors such as changes in recorded neurons or their properties over time.
(1) A central concern relates to whether the main reported effects associated with differences in singing directed versus undirected song reflect only those changes in conditions, versus contributions from changes in unit isolation or response properties over time.
We completely agree that unit stability is critically important in this study. To address this concern, we now quantify stability and apply strict inclusion criteria adopted from a study that assessed unit stability over days (Dickey et al., 2009). Additionally, we now include average waveform overlays for all example units across conditions as supplemental Figure S2.
Changes made in revision: We added: "Importantly, because these neural recordings were performed over long time courses (~2-8 hours), a strict threshold for stability was imposed." and "A Pearson's correlation coefficient of at least 0.99 between the average neural waveform during undirected and directed singing was required for a unit to be considered stable (Dickey et al., 2009; Figure S2)."
(2) A second concern has to do with the categorical definition of 'error neurons'. The authors define a subset of neurons as error responsive only if their responses to DAF exceed a specific threshold (2.5 standard deviations). The problem is that for some neurons categorically defined as being responsive to DAF in only one condition, there is almost certainly not a significant difference in the actual responses to DAF between conditions.
We overhauled our analyses characterizing DAF responses. Rather than relying only on a 2.5 z-score threshold, we now use a more conservative permutation-based approach that directly tests DAF responsiveness and context-dependent changes in DAF responsiveness.
The revised Results state: "Statistical tests defining auditory neurons as DAF-responsive or not in a binary fashion may not be suitable if the underlying population of DAF-related responses exist on a continuum from responsive to non-responsive."
The updated result is: "This more conservative approach identified 48/147 neurons as DAF-responsive in at least one condition, with 13 of those neurons exhibiting a significant modulation in their DAF response between undirected and female-directed singing."
(3a) Some discussion of what is already known about the auditory tuning of Field L, and the extent to which responses associated with distortion of feedback may reflect the frequency tuning of Field L neurons versus something that might be construed as more specifically as detecting an error in perceived feedback.
We agree that DAF-related changes in firing do not necessarily imply that neurons are explicitly detecting an "error" between predicted and actual feedback. Field L neurons can have spectrotemporal receptive fields and frequency tuning such that a broadband DAF stimulus could drive excitation or inhibition simply because the stimulus overlaps with excitatory or inhibitory regions of a neuron's receptive field. We therefore revised the manuscript to use more cautious language and to describe these responses as DAF-related or feedback-related signals rather than categorically as "error responses".
Changes made in revision: The title was changed from "Auditory cortical error signals retune during songbird courtship" to "Auditory cortical feedback signals are modulated during songbird courtship". We also added a sentence to the Discussion: "However, it is important to note that DAF-related changes in firing in auditory neurons do not necessarily imply that neurons compute sensory prediction errors. DAF-related responses could arise from ordinary auditory tuning to the broadband distortion stimulus."
(3b) It would also be useful to discuss further previous work on differences in auditory tuning or responses between conditions when subjects are vocalizing, versus when vocalizations are played back, and to what extent efference copy signals might contribute to the processing of feedback distortions.
We agree these are important points. Our experimental design did not include sufficient passive bird-own-song (BOS) playback trials to permit quantitative comparisons with vocalizing conditions, and we therefore cannot draw firm conclusions about the contribution of efference copy signals to the DAF responses described here. We did observe robust motif onset-associated neural activations, including some activity preceding motif onset, which were present across both social contexts (see new Figure S4). These observations are consistent with prior reports of premotor-related signals in Field L (Keller and Hahnloser, 2009), but whether such signals contribute differentially to DAF processing across contexts remains an open question that we now acknowledge in the Discussion.
(3c) To what extent did the current study control for any vocalizations or other sounds produced by females during the directed singing, and could this have contributed to differences in Field L activity between conditions?
Please see response R2.4 above, in which we describe the exclusion of all song motifs that overlapped in time with a female call. This exclusion criterion was applied throughout all analyses of directed singing.
Figure 1D: In the directed condition there are no spikes at all following the first handful of motif renditions. Were the directed and undirected recordings interleaved here?
Undirected and directed trials were not interleaved. The raster plots are presented in chronological order; however, for each behavioral condition, rows are sorted with the earliest renditions at the bottom and the most recent at the top. We have clarified this in the figure legend.
A minor issue: the raw example trace with male alone does not seem to have a corresponding set of points in the raster plot. For panel E, I also cannot find rasters that correspond to the example recordings shown at top.
In the original version, we randomly downsampled the condition with more trials to equalize trial counts across conditions in the example rasters, while performing all analyses on the full set of recorded trials. As a result, the example spike shown in the raw trace was drawn from one of the downsampled trials not displayed in the raster.
Changes made in revision: For greater transparency, we now include all trials from both conditions for each example neuron in Figure 1.
Figure 2A also shows a neuron that looks like it has non-stationarity; for the alone condition without altered feedback, the main peak has no spikes for the bottom half of the rasters.
In the original version, example neurons were selected to illustrate the DAF-response scoring method, which in some cases highlighted neurons with less stable response profiles. In the revised manuscript, we have replaced this example with neurons that exhibit more robust and stable DAF-related responses, and we now provide a broader set of example neurons illustrating both increases and decreases in DAF responsiveness across conditions.
Other figures show firing rate distributions that appear to be very non-Gaussian, with some motifs during which there is a lot of activity, and others in which there is little activity. Please consider applying non-parametric tests as appropriate.
We agree. In general, some neurons exhibited non-uniform firing rate distributions across trials. All of our main analyses are now conducted using non-parametric permutation tests, which do not assume a Gaussian distribution of trial-by-trial firing rates.
Approaches to addressing the non-stationarity issue could include more specifically indicating examples in which recordings from the alone condition and directed condition are interleaved and exhibit reversible changes in the pattern of responses.
Unfortunately, nearly all of our undirected and directed recording periods were not interleaved, as the experimental design required a block of undirected singing followed by directed singing with female presence. We find it informative, however, that DAF-response modulation was observed in both directions, with some neurons losing DAF responsiveness during directed song and others gaining it, a pattern that is difficult to attribute to a simple unidirectional drift in recording quality. We now provide additional examples illustrating both directions of modulation in Figures 2 and 3.
The methods and/or raster plots should include some further explanation of the time periods over which recordings were made in the alone versus directed conditions, and the extent to which they are interleaved or not.
We have clarified this in the revised Methods. In brief, recording began when the home cage lights came on each day, with the male left to sing alone until at least 40 undirected song motifs were collected. A female was then introduced in approximately 10-minute intervals until at least 40 directed song motifs were collected. The total recording duration on a given day ranged from 0.56 to 10.27 hours, reflecting variability across birds in the time required to elicit sufficient singing in each context. We have added this information to both the Methods and relevant figure legends.
It would be most helpful to assess the stability of waveforms and unit isolation across time.
We now apply strict inclusion criteria based on waveform stability, as described in R3.1 above. SNR was quantified as Vpp/(2*sigma_noise), where Vpp was the peak-to-peak amplitude of each filtered spike waveform and sigma_noise was estimated from the median absolute deviation of the filtered voltage trace. This combines the peak-to-peak normalization used by Nordhausen et al. (1996) with the robust noise estimator described by Rey et al. (2015). Waveform overlays for all included example units are provided in Figure S2.
It would be reassuring to see that significant differences between conditions are equally or more prevalent under the conditions of greatest unit isolation and recording stability.
The average SNR of neurons ultimately included in the analysis was 9.47 +/- 3.57, with a minimum of 4.69. Neurons that exhibited significant DAF-response modulation did not have a significantly different SNR than neurons that did not exhibit significant modulation (Wilcoxon rank-sum test, p=0.38). The mean SNR for significantly modulated neurons was 8.70, compared to 9.5 for non-modulated neurons, indicating that the detection of context-dependent modulation was not systematically biased toward neurons with lower recording quality.
One other way that the authors might be able to address the main concern would be to look at the stability of firing patterns within conditions.
We agree that stability of firing patterns within conditions is an important consideration, and this concern directly motivated the adoption of the permutation-based analysis described above. In this framework, the observed DAF-response difference between conditions is compared to a null distribution generated by shuffling condition labels across trials. This approach inherently accounts for within-condition trial-by-trial variability and does not assume stationarity of firing rates.
It would be helpful to have additional explanations of the criteria used for counting spikes, and assessing stability of recordings.
Spike waveforms were visually inspected for consistency using our custom MATLAB GUI on a 12-second file basis. Interspike interval violations below 1 ms were explicitly checked as an indicator of multi-unit contamination. Detection thresholds were manually set, and each recording file included in the analysis was independently inspected. We have added a more explicit description of these procedures to the Methods section.
For the specific examples shown in figures, it would be useful to indicate by small tick marks or otherwise which spikes were counted as single units.
We appreciate this suggestion. In the revised figures, we have improved the clarity of the example raw voltage traces by annotating the detection threshold and, where multiple units were present on a channel, indicating the waveform amplitude range corresponding to the isolated single unit. We believe this provides sufficient transparency regarding spike identity without requiring tick marks on every individual spike, which would substantially reduce legibility of the example traces.
What were the criteria for determining multi-unit versus single-unit activity?
In the context of this manuscript, "multi-unit activity" refers to channels on which no single neuron could be reliably distinguished from others based on waveform shape and amplitude. Units ultimately included in the study were those for which a single, consistent waveform cluster could be identified and isolated in the custom GUI. In cases where a second distinguishable unit was present on the same channel, it was manually excluded from the sorted single-unit record. We have clarified this distinction in the Methods.
Categorical scores: This definition results in cases where responses of 2.45 vs 2.55 are described as 'retuned', even if these responses are not significantly different. Retuning would be more persuasively demonstrated if the authors could provide a test of whether or not the responses for individual neurons differ significantly between conditions.
We completely agree, and thank the reviewer for motivating us to develop a more rigorous statistical approach. Our revised analysis uses a non-parametric permutation test that explicitly tests for significantly different DAF responses between undirected and directed singing conditions, with correction for multiple comparisons. This replaces the previous threshold-based categorical classification and directly addresses the concern that neurons near the threshold boundary were being treated as categorically different.
Recommendations for the authors:
Reviewer #1 (Recommendations For The Authors):
Minor comments:
(1) Please include a schematic of the brain, including the different subregions of Field L and the connections between auditory regions and the midbrain.
Done. Figure S1 has been added, including a schematic of Field L subdivisions and auditory pathway connectivity.
(2) The authors should include some additional information about the recordings, such as the proportion of Field L neurons that exhibited singing-related changes in firing rate. It would be helpful to include some examples of spontaneous activity when the bird is quiescent in Figs. 1-2, especially for cells that do not show firing locked to song.
We appreciate this suggestion. Given the scope of the current revision and the primary focus on DAF-response modulation, we have elected not to add spontaneous activity examples to Figures 1-2 at this time. We agree this would be a valuable addition in future work and have noted it as a limitation in the Discussion.
(3) Methods, p. 10: Surgery and awake-behaving electrophysiology: "The of the cannula" - this is the only mention of a cannula. Do the authors mean the ends of the probes?
Cannula placement and wire bundle extension from the end of the cannula has been clarified in the Methods.
(4) Bottom of p. 10: Fix reference for biorxiv paper: "ref andreas paper"
Fixed.
(5) Methods, p. 12: Redundant sentences regarding significant error response criteria.
Fixed. The redundant sentences have been removed.
Reviewer #2 (Recommendations For The Authors):
(1) The abstract is too vaguely formulated. Authors should try to quantify the statements already in the abstract.
We have reworded the abstract to align with the revision's more conservative claims regarding social context modulation of auditory feedback, and have added specific quantitative statements where possible.
(2) Authors repeatedly refer to 'perceived song errors' without performing experiments or reporting on behavioral readouts of how birds perceive the jamming sounds. The wording should be changed to something more neutral, e.g. 'DAF responses'.
We revised the manuscript throughout to use more neutral language centred on "DAF-related" or "feedback-related" responses rather than "perceived errors" or "mistakes". The title was changed from "Auditory cortical error signals retune during songbird courtship" to "Auditory cortical feedback signals are modulated during songbird courtship". We similarly revised the abstract and all relevant passages in the Results and Discussion.
(3) Authors write that 33 neurons were DAF responsive in both conditions. How should we interpret this overlap relative to independence and identity assumptions?
We agree that the original presentation made the interpretation of overlap across conditions unclear. The observed overlap is greater than expected under a strict independence assumption but smaller than expected if responsiveness were identical across conditions, consistent with partial but incomplete sharing of DAF responsiveness across social contexts. In the revised manuscript, however, we have moved away from this binary classification framework because DAF responsiveness appears to vary continuously across neurons. The permutation-based analysis now directly tests for changes in DAF responsiveness across contexts without requiring categorical assignment.
(4) Only 10 neurons were not affected by courtship state or only 10 error responsive neurons were not affected? I suggest authors do a multivariate analysis or use a mixed effect model and summarize the result as a table.
We agree that the categorical accounting of neurons across conditions was difficult to follow in the original manuscript. In the revised manuscript, we clarified the distinction between neurons responsive to DAF within a condition and neurons exhibiting significant modulation of DAF responsiveness across conditions. We now explicitly report: "This analysis identified 71/147 neurons as DAF responsive in at least one behavioral condition, whereas 76/147 were not responsive in either condition." and "This more conservative approach identified 48/147 neurons as DAF responsive in at least one condition, with 13 of those neurons exhibiting a significant modulation in their DAF response between undirected and female-directed singing."
(5) It would help if authors could define 'z-scored difference'. Better known is d prime, is this the same?
For each neuron, the z-scored DAF response was computed as the z-scored firing rate difference between distorted and undistorted trials. Importantly, our revised main analysis avoids any normalization such as z-scoring, and instead uses a permutation-based approach applied directly to spike counts.
(6) Is the 'retuning' assessment a bit conservative? Neurons could also retune by showing error scores greater than 2.5 in both conditions but a shifted response time.
We agree that neurons could retune by shifting the latency of DAF responses. Although potential latency shifts are beyond the scope of the current study, we did observe suggestive evidence of possible latency changes in some example neurons across conditions. We have noted this as an interesting direction for future analysis.
(7) Could the stability of DAF response across trials be described? E.g. as the ratio between intra versus inter condition variability?
We agree that stability of DAF responses across trials is an important concern. In addition to imposing strict waveform stability requirements, our permutation-based statistical test explicitly accounts for trial-by-trial variability by constructing null distributions from within-condition trial shuffles. We have also replaced the previously shown unstable example neuron with neurons that exhibit more consistent DAF-related responses across trials, and provide additional examples in Figures 2 and 3.
Minor:
(8) 'significant increase in burst fraction': specify effect size of t test in results section.
We now specify in the main text: "A small but significant increase in burst fraction was observed (paired t-test, p=9.3x10-6, n=138 neurons, mean +/- SEM: 0.11 +/- 0.006 vs 0.15 +/- 0.007, Figure 1J)."
(9) The IMCC parameter should be specified in the main text.
The Gaussian smoothing parameter (20 ms) has now been specified in the main text.
(10) Fig. 2: indicate the windows within which error scores are computed.
This is no longer applicable, as the revised permutation-based analysis does not rely on scoring error responses within a fixed window.
(11) In Fig. 2A, the neuron has an error score of -2.54 (significant), but the red and blue curves look almost the same.
We agree that the previous error score quantification did not always capture firing rate differences in an intuitive way. This example neuron has been replaced in the revised manuscript, and the new analysis avoids scalar error scores in favor of the permutation-based approach.
Reviewer #3 (Recommendations For The Authors):
Minor points:
(1) "(ref andreas paper)." Add reference here?
Fixed.
(2) Hessler and Doupe 1999 is a good reference for premotor signal re-tuning during courtship.
We agree. The reference has been included in the revised manuscript.
(3) Page 5: "discharge depended on courtship state, using" - should this be "depending"?
The original wording was intentional: "we tested how discharge depended on courtship state." We have verified this reads correctly in context and made no change.
(4) Page 9: "consistent with a brainwide process" - what is meant here?
We have revised this wording. The revised manuscript replaces "brainwide process" with clearer language describing a distributed modulation of auditory responsiveness that is not confined to a single nucleus.