Peer review process
Revised: This Reviewed Preprint has been revised by the authors in response to the previous round of peer review; the eLife assessment and the public reviews have been updated where necessary by the editors and peer reviewers.
Read more about eLife’s peer review process.Editors
- Reviewing EditorJoshua JohansenRIKEN Center for Brain Science, Saitama, Japan
- Senior EditorLaura ColginUniversity of Texas at Austin, Austin, United States of America
Reviewer #1 (Public review):
Summary:
The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure.
To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.
Comments on revised version.
The authors have addressed my previous concerns well, and the revised manuscript is substantially improved. In particular, the additional analyses strengthen the conclusion that prelimbic cortical activity reflects learned threat value rather than simply freezing behavior, while the revised framing and additional controls clarify the interpretation of the longitudinal neural dynamics. This paper represents an important contribution to our understanding of the neural mechanisms supporting aversive learning, memory, and generalization.
Reviewer #2 (Public review):
The authors have substantially revised the paper in response to the original review, which is greatly appreciated. It is clear that it will eventually make a nice contribution to the literature. This being said, the following points are somewhere between major and minor in term of their implications for interpretation of the study results. If they were to be addressed, the paper would be again improved.
There are a few remnants of the past language that are not helpful re interpretation of the study results: 1) "Specifically, the observed population gradients could emerge either from the pooled activity of frequency-selective neurons that respond to individual tones or from neuronal subpopulations that integrate information across tones to encode their learned threat-value."; and 2) "Together, these findings suggest that the PL integrates sensory similarity with learned threat value to generate stable representations that support adaptive generalization and discrimination." Neither of these statement follows what has been shown in the study, even with inclusion of the results from the GLM analysis (see point 4 below).
(1) This paragraph in the Discussion is difficult to follow: "Generalization has traditionally been explained by perceptual similarity (Shepard, 1987), whereby stimuli resembling a conditioned cue recruit overlapping sensory representations and evoke similar behavioral responses (Corches et al., 2019; Grosso et al., 2018). Although perceptual similarity clearly influences the extent of generalization, accumulating evidence indicates that it cannot fully account for generalized responding (Verra et al., 2026). More recent frameworks propose that associative learning assigns learned value to novel stimuli by integrating their sensory similarity with previous experience, allowing behavior to scale according to predicted biological significance (Verra et al., 2026; Zaman et al., 2023). Our findings provide a neural framework consistent with these ideas. Sensory similarity promoted consistent neuronal population responses across tones, whereas associative learning organized these responses into graded representations that tracked learned threat value across the stimulus continuum. Thus, sensory similarity appears to define the neuronal substrate upon which associative learning constructs value-based representations that support graded behavioral generalization."
While the revisions have removed the many unnecessary references to inference and integration, this paragraph seems like it is adhering to the original idea of how the authors wished to present their work. If the authors wished to talk about something more than perceptual similarity in the context of generalization, they should have used a task that lends itself to a more-than-perceptual-similarity explanation. Again, the inclusion of the GLM analysis is suggestive for some of what the authors wish to say, but doesn't justify the statements that: "Sensory similarity promoted consistent neuronal population responses across tones, whereas associative learning organized these responses into graded representations that tracked learned threat value across the stimulus continuum." In short, the analysis does not substitute for the design that could have and should have been used to assess learned threat value independently of sensory similarity.
(2) The next paragraph in the Discussion is also confusing. "Such reorganization has been proposed to provide flexibility by allowing new information to be incorporated into existing cortical representations while preserving stable behavioral performance (Mau et al., 2020; Zaki & Cai, 2024). Several mechanisms could contribute to this turnover, including systems consolidation, retrieval-induced reconsolidation or memory updating, and repeated nonreinforced stimulus exposure (Lacagnina et al., 2019; Mau et al., 2020; Sangha, 2015; Zaki & Cai, 2024). Although our experiments cannot distinguish between the first two possibilities, the behavioral data argue against extinction as the primary explanation. Extinction is generally associated with the formation of new CS+-safety associations (Bouton et al., 2021), whereas discrimination ratios increased across retrieval sessions, indicating that animals progressively improved their discrimination between threat-associated and safe stimuli rather than acquiring generalized safety responses. This pattern is consistent with previous work showing that discrimination learning sharpens stimulus representations and narrows behavioral generalization gradients (Dunsmoor & LaBar, 2013; Herzog et al., 2021; Jenkins & Harrison, 1960; Lommen et al., 2017). Importantly, turnover was not uniform across the population. Graded neurons retained remarkably consistent response profiles across retrieval sessions, and their activity remained more strongly associated with learned threat value than with freezing behavior. These observations indicate that stable components of the population code can coexist with extensive reorganization of surrounding neuronal ensembles."
The issue with repeated testing is *not* caused by extinction per se. The issue is that non-reinforcement across the repeated testing should differentially affect the CS+ and CS-. Specifically, it should extinguish responding to the CS- stimulus at a rate that matches its distance from the CS+, thereby sharpening the CS+ versus CS- discrimination in precisely the ways that have been observed. Ergo, the repeated testing *is* a problem for inferences that might be drawn about the way that generalization gradients change with time; and *is* a problem for statements regarding "dynamic reorganization of cortical activity patterns over time." There is nothing in the study that allows one to comment on the reorganization of cortical activity patterns over time. The reorganization can and should be attributed to the repeated testing, which is confounded with time. Nonetheless, the reorganization must be due to the repeated testing and NOT time as the present findings are inconsistent with the well-documented broadening of generalization gradients with time.
(3) In the next paragraph, the authors state: "At the same time, narrower generalization gradients and improved discrimination across retrieval sessions suggests ongoing memory updating. These observations are consistent with contemporary theories proposing that systems consolidation and retrieval-dependent updating are complementary processes through which memories continue to evolve after learning (Mau et al., 2020; Tome et al., 2024; Zaki & Cai, 2024)."
In general, I'm not sure why one would invoke systems consolidation or retrieval-induced reconsolidation as an explanation for any of the present findings: they are not explanations of much at all. In this specific text, the authors seem to be implying an updating process that occurs independently of what is learned across the repeated sessions of testing. Why? The changes that occur in the behaviour and neuronal representations are perfectly explicable in terms of additional learning that occurs - of the sort that I hope to have made clear in my previous comment. Why invoke more than what is needed to explain the observed pattern of results?
(4) Re the GLM analysis - The authors write that: "the fact that the GLM analysis indicates that these neurons reflect learned threat value more than freezing behavior, suggests that they encode an abstract property of the learned stimulus rather than simply mirroring behavioral output."
This is fine if freezing fully indexes the state of conditioned fear and there are no other behaviours in which animals express their fear. If, however, fear is expressed in a range of other behaviours that are likely coordinated by the PL (e.g., startle, vigilance, scanning, orienting to source of danger), this interpretation of the GLM analysis is unwarranted. This is an important point and would be worth noting somewhere in the paragraph where the statement appears.
Reviewer #3 (Public review):
Summary:
Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.
Strengths:
(1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.
(2) Neural coding of generalization is examined, which is under examined in the field.
Comments on revised version.
The authors have convincingly and thoroughly addressed my concerns. I have no further issues regarding this study.
Author response:
The following is the authors’ response to the original reviews.
Public review:
Reviewer #1 (Public review):
Summary:
The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure. To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.
Detailed Comments
(1) A general concern is that the repeated test procedure itself may contribute to extinction. Because the animals are exposed to multiple CS frequencies across multiple test days, and each tone is presented three times per session, some of the reported changes in behavior and neural activity across days could reflect extinction or repeated nonreinforced retrieval rather than the passage of time per se. This is especially relevant given that the manuscript makes claims about recent versus remote representations and representational drift over 30 days. At a minimum, the authors should discuss this limitation explicitly and temper claims about time-dependent changes. Ideally, they would include a control group in which animals are tested only once or twice (e.g., at an early and later time point with fewer CS frequencies), or a reduced-frequency testing design that minimizes extinction while still allowing evaluation of recent versus remote memory.
We agree with the reviewer that repeated testing is an inherent limitation of longitudinal memory studies and may itself contribute to neural changes across sessions. Repeated retrieval can induce memory updating (reconsolidation) or extinction, the latter involving the formation of a new association between the CS+ and safety. Although memory updating may have contributed to the ensemble reorganization observed here, several aspects of our findings argue against extinction as the primary explanation for the observed neural changes.
First, we observed substantial neuronal ensemble turnover beginning with the first retrieval session. This early turnover is consistent with previous observations in the prefrontal cortex [1, 2] and with growing evidence that cortical memory representations remain dynamic throughout systems consolidation [3, 4]. Longitudinal studies have shown that neurons are continuously recruited into and removed from cortical memory ensembles while memory expression remains stable [1-4].
Second, we calculated discrimination ratios to quantify discrimination of each tone relative to the CS+ across retrieval sessions (Figure S1). These analyses showed that discrimination increased, rather than decreased, over successive retrieval sessions, a pattern inconsistent with the behavioral profile expected if repeated testing had induced extinction.
Finally, one of the most novel findings of our study is that ensemble turnover does not affect all neuronal populations equally. The graded neurons identified by our clustering analysis maintained their identity and functional organization across retrieval sessions, and their activity was better explained by tone threat value than by freezing behavior (Figure 8). This selective stability indicates that ensemble reorganization is not a uniform process but instead preferentially affects specific neuronal subpopulations while preserving a stable threat-value generalization gradient. Thus, although repeated retrieval may contribute to ongoing ensemble reorganization, our results demonstrate that this process is selective and largely spares the neuronal subpopulations that encode graded threat-value representations.
Accordingly, we have revised the Discussion to explicitly acknowledge these points as follows:
“The ensemble turnover observed here is consistent with previous studies demonstrating dynamic reorganization of cortical activity patterns over time [1-3, 5]. Such reorganization has been proposed to provide flexibility by allowing new information to be incorporated into existing cortical representations while preserving stable behavioral performance [4, 6]. Several mechanisms could contribute to this turnover, including systems consolidation, retrieval-induced reconsolidation or memory updating, and repeated nonreinforced stimulus exposure [4, 6-8]. Although our experiments cannot distinguish between the first two possibilities, the behavioral data argue against extinction as the primary explanation. Extinction is generally associated with the formation of new CS+-safety associations [9], whereas discrimination ratios increased across retrieval sessions, indicating that animals progressively improved their discrimination between threat-associated and safe stimuli rather than acquiring generalized safety responses. This pattern is consistent with previous work showing that discrimination learning sharpens stimulus representations and narrows behavioral generalization gradients [10-13]. Importantly, turnover was not uniform across the population. Graded neurons retained remarkably consistent response profiles across retrieval sessions, and their activity remained more strongly associated with learned threat value than with freezing behavior. These observations indicate that stable components of the population code can coexist with extensive reorganization of surrounding neuronal ensembles.” Pg. 19
(2) More generally, some of the reported learning-related neural differences may be driven by behavioral differences, particularly freezing, rather than by learning or generalization per se. For example, animals that freeze more to certain frequencies may show corresponding neural response differences simply because freezing alters PL activity. The authors should examine this possibility more directly. Analyses testing whether recorded cells encode freezing behavior, or whether tone frequency-related neural differences remain robust when comparing high- and low-freezing epochs, would help determine whether the reported effects reflect learned stimulus value rather than behavioral state differences.
This is an important point, which was also highlighted by the other reviewers. To directly address this concern, we implemented the generalized linear model (GLM) analysis suggested by Reviewer 3. We modeled the neuronal activity time series using both tone identity and freezing behavior as simultaneous predictors. Because tone identity was fixed across trials whereas freezing varied from trial to trial, the GLM allowed us to dissociate their independent contributions to neuronal activity.
As described in the original submission, freezing was estimated from the miniscope's onboard inertial measurement unit (IMU), which measures body acceleration along three axes. Rather than classifying freezing using a fixed threshold, we estimated the continuous probability of freezing from the accelerometer signal using a Gaussian mixture model. This probabilistic estimate was incorporated directly into the GLM together with tone identity, providing a conservative test of whether neuronal activity was better explained by freezing behavior or by the auditory stimulus.
We applied the GLM both to all sound-responsive neurons contributing to the population response curves (Figure 4) and to the graded and frequency-selective neuronal subpopulations identified by our clustering analysis (Figure 8). Across both experimental groups and all analyses, the median regression coefficients (β) associated with tone identity were consistently larger than those associated with freezing, indicating that tone identity contributed more strongly to neuronal activity. Moreover, tone coefficients exhibited graded monotonic profiles that closely tracked the learned threat value of each tone, with graded neurons showing the strongest gradients (Figures 4a, 8a, and 8e). Consistent with previous reports [14, 15] freezing accounted for a modest but significant component of PL activity. However, only 6–8% of graded neurons were classified as freezing-dominant, indicating that for the vast majority of these neurons, tone identity was the stronger predictor. Together, these findings demonstrate that the graded representation of learned threat value persists after accounting for freezing behavior, supporting our conclusion that PL activity reflects learned threat value rather than merely the behavioral expression of fear.
(3) A central feature of the manuscript is the analysis of neural response properties over an extended period of time, up to 30 days after learning. However, aside from a brief mention in the Methods that spatial registration was used, the manuscript provides very little quantitative information about this critical aspect of the study. The paper would be strengthened by including explicit metrics describing longitudinal cell tracking, such as the number and proportion of ROIs retained across all sessions, distributions of spatial-footprint correlations or centroid distances across days, and representative examples of matched imaging fields over time. Without this information, it is difficult to assess how strongly the longitudinal claims are supported.
We thank the reviewer for this suggestion. We now include measures of registration quality in the resubmission. Specifically, we calculated shifts in centroid distances, proportion of ROIs retained across all sessions, and representative examples of matched imaging fields over time (Fig, S3).
(4) The text states that "Figs. 1c and 1d show GCaMP6f expression in PL, representative calcium footprints, and activity traces". However, the figure as presented does not clearly show all of these elements, at least not in a way that matches the description in the Results. The correspondence between text and figure should be corrected.
We corrected correspondence between text and Figure.
(5) The labeling of Figure 2a is insufficient for interpretation. The legend states that the panel shows raster plots of sound responsiveness, but the axes and scaling are not clearly defined. It is not clear from the figure what the x-axis represents, whether the y-axis corresponds to individual neurons, where the CS period occurs, or what the activity scale at the right denotes. Also, the term 'rasters' implies that spikes were analyzed. It seems that the spike inference approach (CASCADE) was only used for later analyses. Perhaps 'heat-plot' would be more accurate here? Generally, this figure should be annotated more clearly so that the reader can understand it without referring back to the Methods.
We clarified the labelling of the Figure 2a and call the graphs “activity-plots”.
(6) In relation to Figure 3, the analysis of population-averaged responses across tone frequencies is useful, but the manuscript would be stronger with additional statistical analyses across time and across groups. For example, if the authors want to argue that learning induces graded changes in neural responses and that these evolve across time, they should directly compare within-group responses across days and also compare matched frequencies between the conditioned groups and the no-shock controls. These analyses would help establish whether the observed differences are genuinely learning dependent and whether they change significantly over time.
For Figure 3, we maintained the previous one-way ANOVAs assessing changes in AUC per day to be able to note significance on the Figure panels. However, we added a three-way mixed-effects analysis, using group (CS15, CS3, no shocks), frequency (3, 7, 11, 15), and day of testing (2, 15, 30) as variables, with frequency and day of testing as repeated measures. The results were described as follows (statistical details Table S1):
“To determine how AUC varied across groups over time, we performed a three-way mixed-effects ANOVA with group (CS+15, CS+3, and no shock), frequency (3, 7, 11, and 15 kHz), and time (test days 1, 15, and 30) as factors, with repeated measures on frequency and time. For positive responder neurons, the analysis revealed significant main effects of group (p < 0.001) and time (p < 0.05), as well as a significant group × frequency interaction (p < 0.001), whereas the time × frequency and group × time × frequency interactions were not significant (p > 0.05; Table S2a). Tukey-corrected post hoc comparisons showed that, in the CS+15 group, AUC differed between all frequency pairs except 11 and 15 kHz (p < 0.05). In the CS+3 group, the AUC at 3 kHz differed from those at 7, 11, and 15 kHz (p < 0.05), whereas no significant frequency differences were observed in the no-shock controls (p > 0.05). For negative responder neurons, the only significant effect was a time × frequency interaction (p < 0.01). Tukey-corrected simple-effects analyses revealed that, on day 30, the AUC at 15 kHz differed from those at 3, 7, and 11 kHz (p < 0.05; Table S2b). Because this pattern was observed across all experimental groups, including the no-shock controls, it is unlikely to reflect associative learning. These results indicate that although the AUC exhibited modest changes over time, these changes were not group-specific and therefore do not support learning-dependent alterations in neuronal responses. Together, these results show that despite substantial neuronal turnover, PL population responses encode generalization gradients, closely matching behavioral expression.” Pg. 9
(7) The inclusion of two different CS+ frequencies and a no-shock control is a strength of the study and substantially improves the interpretation that graded neural responses are related to learning and generalization rather than to simple sensory processing or passage of time. That said, I am not entirely comfortable with the use of the term "inference" throughout the manuscript. What is being measured here appears closer to sensory generalization than inference in a stronger cognitive sense. The current task does not clearly require that animals infer hidden structure or stimulus value through abstract reasoning; rather, the generalized stimulus may simply be treated as similar to the conditioned cue. The terminology should therefore be reconsidered or softened.
We thank the reviewer for appreciating the strengths of the experimental design and for this thoughtful suggestion regarding terminology. We agree that the term inference may overstate the cognitive processes engaged by the current task. Accordingly, we revised the terminology throughout the manuscript to describe these effects as graded generalization of threat value across stimuli. The new GLM analyses further support this interpretation by demonstrating that, in the conditioned groups, neuronal activity at both the population and single-neuron levels is explained substantially better by tone identity than by freezing behavior (Figures 4 and 8). We therefore retained the term threat value, as our results indicate that PL activity primarily reflects learned threat value rather than simply the expression of freezing behavior, but removed inference.
(8) I also found the use of the term "valence" somewhat problematic. The manuscript appears to use valence to refer to graded responding across tones with different aversive significance, but valence typically refers more broadly to distinctions between appetitive and aversive value. Here, terms such as "threat value," "aversive value," may be more precise. The authors should consider revising this language throughout.
We corrected the language and replaced valence for “threat value”
Reviewer #2 (Public review):
Summary:
The following points are those that occurred to me across readings of the paper. They are listed in what I take to be the order of their significance. Many of the points relate to the loose use of language and invocation of concepts that are not warranted, given the study design and results obtained.
Major Comments:
(1) The concept of ensemble turnover is interesting - the way it is introduced and discussed implies some type of spontaneous change in the neural underpinnings of fear discrimination and generalization in the PL. But, of course, every trial involves an opportunity to learn about the threat CS or the generalization test stimuli, and I am troubled by the thought that stability in the neural underpinnings of fear discrimination and generalization will actually reflect the level of defensive behaviours evoked on different trial types and/or the discrepancy between those behaviours and the outcome of a given trial in the generalization test. That is, stability in the neural underpinnings may be related to an animal's certainty or uncertainty in the contingency between a stimulus and danger; or, put another way, an animal's confidence that danger will or won't occur given the presence of some stimulus. This is not uninteresting. It is, however, not considered anywhere in the paper, which is overloaded with references to inferred threat values and integration of information across different types of stimuli. The protocol is not one that requires inference about anything or integration across anything.
We thank the reviewer for this thoughtful comment. We agree that our original wording may have implied that turnover was a spontaneous process. Repeated retrieval provides opportunities for updating the learned contingencies associated with both the conditioned and generalization stimuli, and therefore changes in ensemble composition across sessions need not arise independently of experience. We also agree that the stability of graded neuronal representations may be related to the animal's certainty about the learned contingencies. However, in our data the graded neuronal population remained remarkably stable across retrieval sessions, whereas changes occurred primarily within the dynamic, frequency-selective neuronal populations. This suggests that stable ensembles preserve representations of learned threat value while updating is concentrated in a distinct neuronal subpopulation. We have now incorporated these ideas into the Discussion.
(2) I appreciate the link to Gu and Johansen in paragraph 3 of the Introduction, but the type of generalization under investigation here is not the same as the type of 'generalization' studied by Gu and Johansen [who used a sensory preconditioning protocol]. Nonetheless, the authors have forced the language used by Gu and Johansen into their paper, and this has created tension [at least for this reader] as the concepts introduced by Gu and Johansen [inference, integration] are simply not relevant given the generalization protocol used here. Here are a few examples of points where the tension might interfere with a reader's understanding:
We thank the reviewer for these specific criticisms. We revised the manuscript throughout to remove or redefine terms like "inferred valence" and "integration," replacing them with clearer, more accurate descriptions of gradient generalization of threat value. Below we address each point raised by the reviewer regarding terminology clarifications.
(a) 'We hypothesized that generalization to novel stimuli depends on stable subnetwork organization that enables comparisons between learned and inferred valence, as well as population-level features that reduce variability across related representations.'
I understand the words in the hypothesis, but can't form a representation of what is being said because of the reference to terms that stand in need of clarification [inferred valence, variability across related representations], but, ultimately, won't be clarified. This needs to be re-expressed so that the reader can appreciate what is being said.
(a) We hypothesized that the PL generates representations of learned threat value that support threat generalization and discrimination, and that these representations emerge from the coordinated activity of stable and dynamic neuronal subnetworks, preserving consistent relationships among stimuli despite ongoing cellular turnover.
(b) 'Our results show that stable cortical subnetworks integrate the emotional "gist" of memory and inferred valence for novel cues over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity across stimulus presentations determines threat generalization.'
Again, what does this mean? How is the gist of a memory integrated with inferred valence for novel cues over time? The statement simply doesn't make sense. This needs to be rewritten for clarity.
(b) The summary statement was rewritten: " Together, these findings provide a neural framework for understanding how the PL supports adaptive threat generalization and discrimination.” pg. 4
(c) 'In CS+ 15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting the contingency learned valence as well as the inferred valence of novel tones across testing days...'.
Can this be rewritten as 'In CS+15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization.'? The overloading of the text with references to 'contingency learned valence' and 'inferred valence' is unnecessary and makes it much harder to understand what has been shown in the results.
We adopted the reviewer's suggested rewording: " In CS+ 15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting learned contingency value across testing days" pg. 9
We will systematically review the entire manuscript to ensure consistency with this revised framing.
(3) Re the same passage of text as in 2c:
Is it the case that these neurons are simply tracking the expression of freezing to the various tones? The same question applies to the results obtained for the CS+3 mice. If this is the case, then why should the results be taken to support the banner statement that 'Sound-modulated PL population responses encode learned and inferred valence' - these analyses do not support that statement. And, as indicated, I don't believe that the language of learned and inferred valence is appropriate to such statements, given the nature of the protocol used and results obtained. It is a study looking at how populations of neurons in the PL respond during presentations of auditory stimuli that were subject to discriminative conditioning, and during tests of generalized freezing to other [intermediate] auditory stimuli.
The reviewer is correct that the graded population responses observed in PL could reflect freezing behavior across tone frequencies rather than encoding an abstract threat-value representation. This important concern was also raised by other reviewers. To address it directly, we followed Reviewer 3’s suggestion and implement a Generalized Linear Model (GLM) using the time series activity derived from the Ca2+ signals, with both tone identity and freezing behavior included as predictors. This analysis allowed us to dissociate the respective contributions of tone frequency and freezing to the graded neural responses. Based on the outcome of this analysis, we concluded that tone identity was a stronger predictor of neuronal activity than freezing. These results are summarized in Figures 4 for all cells contributing to population responses and Figure 8 for the main neuron types identified in the clustering analysis (frequency-selective and graded neurons). All details of this extensive new analysis are shown in red in the revised resubmission.
In addition, we revised the text to remove the terminology of “learned and inferred valence” throughout the manuscript.
(4) It is stated that:
'In no-shock controls, although both positive and negative responses were present, population activity was not modulated by tone frequency or valence'.
What does this mean? I can understand that population activity was not modulated by tone frequency. But what does it mean to say that it was not modulated by valence? Why should it have been when none of the tones were conditioned in this group and, hence, mice were responding to all the tones equally? And given that this is true, I don't understand the use of 'valence' here, or the subsequent statements in this paragraph that 'graded responses require associative learning' and that 'PL population responses encode graded sound-valence associations that reflect both learning and inference, closely matching behavioral generalization.' The latter statement is particularly unwarranted and, again, highlights a major issue with the paper. It could and should be rewritten as 'PL population responses reflect behavioral generalization.' There is nothing in the additional language that adds to the reader's understanding of what has been shown. The reference to 'graded sound-valence associations that reflect both learning and inference' is completely unwarranted, given the nature of this study. It is anathema to the vast literature on stimulus generalization. If the authors wished to make statements of this sort, they should have taken a different approach, perhaps using protocols like those featured in Gu and Johansen.
We thank the reviewer for this helpful comment. We agree that our use of the term valence in describing the no-shock controls was imprecise. Because none of the tones was associated with reinforcement in this group, there was no learned valence that could modulate neuronal activity. Our intention was simply to convey that, although both positive and negative sound-responsive neurons were present, the population responses did not vary systematically across tone frequencies. We have revised this section accordingly.
We also agree that our original wording overstated the interpretation of the graded population responses. Our data do not demonstrate that associative learning is required for sound responsiveness itself; rather, they show that associative learning is required for the emergence of graded population responses that distinguish tones according to their learned threat value. We have revised the text to make this distinction explicit.
Finally, we agree that our previous references to "learning and inference" were not justified by the behavioral paradigm. We have removed this language throughout the manuscript and now describe the findings more directly as graded representations of learned threat value that closely parallel the observed behavioral generalization gradients.
(5) The section titled, 'Consistently active neurons preserve valence representations as newly recruited neurons sharpen remote memory traces' ends with the following summary:
'Together, these results indicate that consistently active neurons maintain stable representations of learned and inferred sound associations across time, whereas neurons recruited after conditioning progressively acquire graded tuning at later retrieval stages. This dynamic refinement suggests that cortical memory representations become increasingly selective during systems consolidation, while a stable neuronal subpopulation preserves the core emotional content of the memory.'
Once again, the summary is not in keeping with the results obtained. The 'dynamic refinement' of representations is far more likely to reflect the repeated testing across days 1, 15, and 30 rather than anything to do with systems consolidation - at the very least, it is the simplest interpretation of the results. The impact of repeated testing is evident in the sharpening of generalization gradients over time, which is contrary to what is otherwise observed in the literature - the incredibly well -documented broadening of generalization gradients with time. Given this impact of repeated testing, surely the changes in the neuronal population that underlie performance are more likely to reflect the learning that occurs on days 1, 15, and 30, which is reflected in reduced freezing to the non-conditioned tones. If this is a reasonable take on the results, then I don't see the basis for invoking systems consolidation at all, and I don't see the basis for inferring a stable neuronal subpopulation that preserves the emotional content of the memory. Rather, non-reinforced presentations of 'never-reinforced' tones result in recruitment of additional neurons that result in suppression of freezing responses to those stimuli.
We thank the reviewer for this thoughtful comment. We agree that repeated retrieval is an inherent limitation of longitudinal memory studies and that repeated non-reinforced presentations of the tones provide opportunities for memory updating. Accordingly, we have revised the Discussion to explicitly acknowledge that repeated retrieval may contribute to the ensemble reorganization observed across sessions through memory updating or reconsolidation processes (Discussion, pg. 19).
We also agree that the progressive sharpening of the behavioral generalization gradients across retrieval sessions is consistent with memory updating. Both the behavioral data (increased discrimination ratios) and the neuronal data (progressively sharper population generalization gradients among neurons active after conditioning) indicate that the memory representation became more precise over time. We now discuss this possibility explicitly in the revised Discussion. We also agree that fear generalization often broadens with time; however, this is not universal. Under discriminative conditioning paradigms, repeated retrieval can instead produce progressively narrower generalization gradients [11]. We have revised the Discussion to clarify this distinction and added the appropriate references (pg. 19).
While the reviewer's interpretation is therefore plausible, we do not believe it fully accounts for our observations. If repeated non-reinforced presentations were the sole driver of the observed neuronal changes, one might expect a more uniform reorganization across the neuronal populations engaged by the task. Instead, the reorganization was highly selective. Neurons encoding graded threat value remained remarkably stable across retrieval sessions, whereas neuronal turnover occurred primarily within the frequency-selective subpopulations. Thus, although repeated retrieval may update the memory representation, the neuronal substrate supporting graded threat-value coding is largely preserved while refinement occurs within a distinct neuronal subpopulation.
Moreover, we observed substantial neuronal turnover beginning with the first retrieval session, consistent with previous longitudinal studies showing that cortical memory ensembles remain dynamic despite stable memory [1-4]. This early emergence of turnover suggests that repeated testing alone is unlikely to account for the continuous population dynamics observed throughout the experiment.
Rather than viewing these findings as evidence exclusively for either memory updating or systems consolidation, we believe they are more consistent with current models proposing that these processes occur in parallel. Several influential frameworks argue that memories are continuously modified through retrieval while simultaneously undergoing systems-level reorganization [4, 6, 16, 17].We have therefore revised the Discussion to interpret the longitudinal changes more conservatively as reflecting the combined influence of retrieval-dependent memory updating and systems-level reorganization.
In summary, we have revised the manuscript to better acknowledge the contribution of repeated retrieval while emphasizing what we believe is the principal finding of our study: despite substantial turnover within the overall ensemble, the neuronal population encoding graded threat value remained remarkably stable, whereas refinement occurred primarily within dynamic frequency-selective neuronal populations.
(6) In the section titled, 'Population vector similarity at stimulus onset determines degree of generalization', it is stated that:
'Because population similarity peaked shortly after stimulus onset, we quantified similarity during the first 5 s after tone onset relative to the CS+. In CS+15 mice, population similarity was highest for 15/15 and 15/11 tone pairs with no differences between them.'
Isn't this consistent with the view that the population response in the PL simply reflects the level of freezing? Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained. That is, these results appear to clearly indicate that neuronal responses in the PL reflect the degree of stimulus generalization, as evidenced in freezing behavior. Given all that we know about the involvement of the PL in expressing fear responses, it is not appropriate to claim that 'population vector similarity at stimulus onset *determines* the degree of generalization. The PL responses simply reflect the varying levels of performance displayed to the different types of tones. What have I missed that could be taken to support additional statements?
We agree that, because population similarity is highest for the 3/3, 15/15, and 15/11 tone pairs and freezing is also greatest for these same stimuli, the neural data could, in principle, reflect a correlate of behavioral expression rather than an independent representation of learned threat value.
To directly address this possibility, we implemented a generalized linear model (GLM) to dissociate the contributions of tone identity and freezing behavior to neuronal activity. Across all analyses, tone identity consistently explained substantially more variance in neuronal activity than freezing behavior. Importantly, this finding held not only for the full population of sound-responsive neurons used to generate the population similarity analyses (Figure 4), but also for both the stable graded neurons and the dynamic tone-selective neuronal populations identified by our clustering analysis (Figure 8). Thus, although freezing behavior contributes modestly to PL activity, it cannot account for the enhanced similarity of population vectors across stimulus presentations or the graded population responses that form the basis of our conclusions.
In addition, the temporal dynamics of the population vector similarity analysis are not entirely consistent with the interpretation that PL activity simply reflects the expression of freezing behavior. Population vector similarity peaked during the first 5 seconds following tone onset, whereas freezing occurred intermittently throughout the tone presentations. Although this temporal relationship does not establish causality, it is consistent with the interpretation that PL activity reflects the learned threat value associated with each tone rather than merely tracking the magnitude of freezing.
Finally, we have revised the manuscript to more clearly acknowledge the correlational nature of these analyses. Specifically, we now state that population vector similarity is associated with, rather than determines, the degree of threat generalization.
Later in the same section, it is stated that 'population-level similarity at stimulus onset scales with behavioral threat generalization and is maximal for tones associated with robust threat responses.' For simplicity and, therefore, clarity, this should be rewritten as 'population-level similarity at stimulus onset reflects behavioral threat generalization.'
We made this correction. (“These findings indicate that population-level similarity at stimulus onset scales with behavioral threat generalization”. pg. 13)
(7) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:
'Our previous analyses show that learned and inferred associations are represented at the population level. However, these results do not resolve whether graded responses arise from pooled activity of frequency-selective neurons or from subnetworks encoding integrated learned valence across tones.'
What does it mean to say 'integrated learned valence across tones'? As it presently stands, the meaning of the phrase is unclear. It only makes sense if one supposes that generalized freezing responses to the 11 and 7 kHZ tones reflect separate associations between those tones and the aversive foot shock US. This supposition is inconsistent with the rich literature on generalization of Pavlovian conditioned fear responses. Specifically, it is inconsistent with the many theories of fear generalization, which attribute the reduction in fear as one moves away from the specific conditioned stimulus to a decrement in the ability of the test stimulus to activate the trained CS-US association. My strong impression is that the authors would do well to ground their findings in theories of stimulus/fear generalization, of which there are many. This would better serve the results obtained [and the reader's appreciation of them] - at present, the unnecessary invocation of concepts does very little to enhance the reader's appreciation or understanding of what has been found in the study.
We agree that the phrase "integrated learned valence" is unnecessarily opaque and we replaced it with more precise language “Our previous analyses demonstrated that threat-value generalization gradients are represented at the population level. However, these findings do not reveal how these representations arise. Specifically, the observed population gradients could emerge either from the pooled activity of frequency-selective neurons that respond to individual tones or from neuronal subnetworks that integrate information across tones to encode their learned threat-value.” (Pg. 13)
(8) Another example of what has been a common theme in this review:
'...we hypothesized that the PL active ensemble segregates into functionally distinct subnetworks: one encoding tone-specific sensory features with dynamic characteristics, and another responding to all frequencies encoding stable core memory content and inferred emotional valence.'
What does it mean to say 'all frequencies encoding stable core memory content and inferred emotional valence'? Do the authors mean to say '...and another that tracks freezing/defensive responses regardless of whether they were elicited by the trained CS or one of the generalization test stimuli'?
We thank the reviewer for pointing out that this section was unclear. We agree that our original wording was imprecise and could be interpreted as implying cognitive processes that were not directly tested in the present study. Accordingly, we have revised the terminology throughout the manuscript. We no longer refer to "inferred emotional valence" or "core memory content" and instead describe these neurons more specifically as exhibiting graded representations of learned threat value.
This is not the interpretation we intended. To determine whether these neurons primarily reflected defensive behavior rather than learned stimulus value, we implemented a Generalized Linear Model (GLM) that dissociates the contributions of tone identity and freezing behavior to neuronal activity. Across the entire neuronal population, as well as within the stable graded and dynamic tone-selective neuronal subpopulations, tone identity consistently explained substantially more variance than freezing behavior (Figures 4 and 8). Furthermore, after accounting for freezing, the regression coefficients of the graded neurons continued to follow the learned threat value of the tones, exhibiting opposite monotonic gradients in the CS+15 and CS+3 groups. If these neurons simply tracked defensive behavior irrespective of the stimulus presented, this relationship would not be expected to persist after accounting for freezing. We therefore conclude that the activity of this stable neuronal subpopulation is better explained by graded representations of learned threat value than by defensive behavior alone, and we have revised the manuscript accordingly.
(9) It is stated that - 'Graded clusters encode emotional valence but constitute only a fraction of the active population; yet valence coding at the population level remains accurate and precise. This indicates that neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.'
What does this mean? Are the authors trying to say that - 'Some clusters of PL neurons track freezing responses. In spite of the fact that these are only a fraction of the total active neuronal population, the population-level response of PL neurons also tracks the levels of fear to the trained tone and its variants used in the test for generalization.' If this is what one wants to say, then the final statement in the reproduced section does not follow. That is, there is no indication that 'neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.' As noted, the characteristics of other ensembles that become active across the repeated tests on days 1, 15, and 30 are more likely to reflect learning from non-reinforcement that occurs within and across those sessions. Perhaps this is what is meant by the phrase, 'shaped by associative processes'? If so, it should be stated explicitly instead of left to the reader to work out.
We thank the reviewer for highlighting that this section was unclear. We agree that the original phrasing was insufficiently precise. Our intention was to convey that only a subset of PL neurons displays graded tuning that tracks behavioral generalization across tones. Nevertheless, despite constituting only a fraction of the total active population, this graded coding is also reflected at the population level. This observation led us to hypothesize that neurons recruited into the active population after conditioning— likely dynamic, frequency-selective neurons—also contribute to these graded population responses through modulation of their firing rates.
The reviewer correctly notes that the phrase "shaped by associative processes" was too vague. By this we meant that the firing properties of these neurons are modified by the animal's associative history, including both the original conditioning experience and any retrieval-dependent updating that may occur during subsequent test sessions. We have revised the manuscript to make this interpretation explicit rather than leaving it to the reader to infer.
To test this hypothesis, the GLM analysis we implemented dissociated the contributions of tone identity and freezing behavior to neuronal activity. After accounting for freezing, tone identity (i.e., learned threat value) remained a significant predictor of neuronal responses. Importantly, this was also true for the dynamic, frequency-selective neurons (Fig. 8e–f), indicating that these neurons contribute to population-level representations of learned threat value through firing-rate modulation rather than simply reflecting defensive behavior.
To clarify our interpretation, we have rewritten the relevant section as follows:
"Graded clusters encode generalization gradients but constitute only a subset of the active neuronal population. Nevertheless, population-level representations, which incorporate all active neurons, remain robust and accurately preserve these gradients. This observation led us to hypothesize that neurons recruited over time (e.g., dynamic, frequency-selective cells) also contribute to threat-value representations. Consistent with findings in the hippocampus showing that neurons can encode task contingencies through firing-rate modulation despite responding selectively to a single location (Gagliardi et al., 2024; Huxter et al., 2003; Sanders et al., 2019), we tested whether dynamic, frequency-selective clusters exhibited firing-rate differences proportional to learned threat value." (page 15)
Regarding the reviewer's suggestion that the characteristics of the newly recruited neurons may reflect learning during repeated non-reinforced test sessions, we agree that retrieval-dependent memory updating likely contributes to the reorganization of the dynamic neuronal population, and we now explicitly acknowledge this possibility in the Discussion. However, we do not believe that our findings are fully explained by repeated non-reinforced retrieval alone. First, no-shock control animals underwent the same repeated testing but failed to develop graded neuronal representations, indicating that repeated exposure in the absence of associative learning is insufficient to account for the observed changes. Second, both behavioral discrimination and the corresponding population-level neural gradients became progressively sharper over time, consistent with refinement of learned threat representations rather than an effect of repeated testing alone, which must lead to extinction.
In summary, we thank the reviewer for highlighting both the ambiguity of our original wording and an important alternative interpretation. In response, we have clarified the text to explicitly define what we mean by associative processes, added a GLM analysis demonstrating that the newly recruited neurons encode learned threat value beyond freezing behavior, and revised the Discussion to acknowledge that retrieval-dependent memory updating likely contributes to the reorganization of the dynamic neuronal population.
(10) The following points all relate to the Discussion and reiterate many of the points above.
(a) 'A subset of neurons remains consistently active across sessions, preserving core components of the memory trace and supporting inference of emotional valence for novel sounds, while neurons recruited after conditioning progressively acquire valence selectivity at remote time points.'
'Inference of emotional valence' is unclear and unwarranted for all of the reasons provided above regarding the use of language.
We modified the language as stated in the prior points.
(b) '...Our data reconcile these views by demonstrating that cortical representations of emotional valence emerge rapidly after learning and persist within stable subnetworks, even as the broader population undergoes substantial turnover. This architecture preserves core mnemonic content while allowing flexibility in the surrounding ensemble.'
These statements assume that the PL neuronal responses reflect something more than the levels of freezing behavior to the different stimuli; what are the grounds for this assumption?
We incorporated the new GLM analysis to address this point and conclusions.
(c) 'Importantly, these subnetworks encode both learned contingencies and the inferred valence of novel stimuli along a graded representational axis, suggesting that strong recurrent connectivity provides a stable scaffold for emotional memory representations.'
What is a graded representational axis, and what part of the first statement suggests that 'strong recurrent connectivity provides a stable scaffold for emotional memory representations'? If the authors' goal was to make statements about emotional memory representations vis-à-vis emotional memory content, they should have used protocols that allowed them to probe such content. The auditory fear conditioning protocol used here [followed by tests for generalization to other auditory stimuli that differ in frequency from the conditioned tone] is not one that lends itself to analysis of emotional memory representations or content.
We agree that the term "graded representational axis" was insufficiently defined and could be interpreted in multiple ways. Because this terminology was not essential to our conclusions, we have removed it and instead describe the observed phenomenon as a graded population representation of learned threat value across tone frequencies. We also removed the statement suggesting that recurrent connectivity provides a stable scaffold for these representations, as this mechanistic interpretation is not directly supported by our data.
We also agree that some sections of the manuscript overstated the scope of our conclusions and have revised the wording accordingly. Our study uses neuronal activity recorded during memory retrieval after learning, an approach widely used in studies of systems consolidation to infer how learned information is represented within neural populations. Accordingly, we have revised the manuscript to explicitly state that our findings pertain to neural representations of learned threat value during memory retrieval rather than the broader content of emotional memories.
Finally, we agree that our data are correlational and do not establish the causal role of the neuronal representations we identify. Throughout the manuscript, we now refer more precisely to population- and single-neuron correlates of learned threat value during memory retrieval following auditory fear conditioning.
(d) 'Dynamic tone-selective responsive neurons emerge independently of learning, as they are present in both control and experimental mice, reflecting pre-existing PL sensory-driven properties (Hockley & Malmierca, 2024; Zikopoulos & Barbas, 2006).'
Maybe. They are also likely to have developed as a consequence of the repeated testing on days 1, 15, and 30, which involved intermixed exposures to the tones of different frequencies. That is, rather than 'pre-existing PL sensory-driven properties', the responses of these neurons might reflect the emergence of discrimination between the various tones across testing, and greater suppression of freezing to the non-trained tones compared to the trained tone across the various test intervals.
We thank the reviewer for this thoughtful comment. Our interpretation that these neurons reflect preexisting sensory-driven properties of PL cortex is based on two observations. First, tone-selective neuronal clusters were present in both conditioned and no-shock control animals, consistent with previous reports of sensory responsiveness in PL cortex [18, 19]. Second, these responses were already present during the first retrieval session, when the intermediate frequencies were presented for the first time. Thus, they cannot be explained by repeated exposure to those tones across subsequent test sessions.
We therefore interpret the frequency-selective response properties as pre-existing features of PL circuitry that are present independently of conditioning. In contrast, associative learning modifies the firing activity of these neurons, allowing them to contribute to graded representations of learned threat value. This interpretation is supported by our GLM analysis, which showed that, after accounting for freezing, tone identity significantly predicted the activity of frequency-selective neurons in conditioned animals but not in no-shock controls. Thus, while the frequency-selective response properties are present independently of learning, associative learning modifies how these neurons encode learned threat value. We have revised the manuscript to clarify this distinction.
Reviewer #3 (Public review):
Summary:
Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP-positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.
Strengths:
(1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.
(2) Neural coding of generalization is examined, which is under-examined in the field.
We thank the reviewer for appreciating our design to track ensembles over time and the relevance of studying the neural substrates of generalization.
Weaknesses:
(1) Difficult to determine if responses treated as encoding stimulus valence are driven instead by the behavior that the stimulus elicits, freezing.
We thank the reviewer for this thoughtful and constructive comment. We agree that an alternative interpretation is that the graded neuronal responses may partially reflect freezing-related activity rather than representations of learned threat value. In the revised manuscript, we acknowledge that previous studies have identified PL neurons whose activity tracks freezing independently of stimulus identity or associative content. To directly address this possibility, we implemented the reviewer's suggestion by fitting a generalized linear model (GLM) to the neuronal activity time series derived from the Ca2+ signals, using tone identity and freezing behavior as predictors. Because tone identity is fixed across trials, whereas freezing varies both during tone presentation and across trials (see below our answer to the Recommendations to Authors), this approach allowed us to dissociate their respective contributions to neuronal activity. We are grateful for this excellent suggestion, which has substantially strengthened both the manuscript and the conclusions that can be drawn from our data. The new analyses are summarized in Figures 4 and 8.
In the points below we summarize the new findings.
(2) The study implies that the identified ensembles are causally related to valence memory, but no experimental interventions are performed to justify this.
We appreciate the reviewer's point. We agree that our data are correlational in nature and that establishing a causal relationship between identified ensembles and valence memory would require experimental interventions such as combinations of optogenetic and two-photon manipulations, which are beyond the scope of the present study but represent an important direction for future work.
We examined inter-individual variability in freezing relative to the proportion of graded cells but the number of mice used in this study (CS+3= 5 and CS+15=7) did not give us enough power to reach significance.
Therefore, we modified the manuscript terminology accordingly, replacing causal language with phrasing that accurately reflects the correlational nature of our conclusions.
Recommendations for the authors:
Reviewer #2 (Recommendations for the authors):
Many sections of the paper should be rewritten along the lines that I have suggested in my public review and below.
Minor Comments:
(1) INTRO - 'This broad accessibility reduces spatial specificity and increases learning variability...'.
Broad accessibility of what, exactly? And how does the 'broad accessibility' reduce spatial specificity and increase learning variability? That is, I do not understand what the terms 'reduced spatial specificity' and 'increased learning variability' refer to at this point in the first paragraph...
We rewrote the introduction and discussion to address the points raised by the reviewer.
(1) INTRO - 'The prelimbic cortex (PL) contributes to the expression (Burgos-Robles et al., 2009; SierraMercado et al., 2011; Sotres-Bayon & Quirk, 2010) and the proper discrimination and generalization of threat memories (Rosas-Vidal et al., 2025; Stujenske et al., 2022).'
What is achieved by calling it 'proper' discrimination and generalization? Can't one simply say that the PL contributes to the expression, discrimination, and generalization of threat memories?
This was corrected.
(3) INTRO - '... and that population-level firing rate similarity across stimulus presentations determines threat generalization'.
Or, alternatively, that generalization of conditioned freezing responses from the tone CS to variants along the dimension of Hz values is reflected in systematic changes in the firing rate of PL neuronal ensembles; when the test stimulus is similar to the conditioned stimulus, the two elicit similar behavioural responses and evoke a similar population-level firing rate in the respective PL neuronal ensembles.
We have revised the Introduction as stated above. However, as discussed in our detailed responses, freezing behavior cannot fully account for the observed patterns of PL activity.
(4) METHODS - 'Memory retrieval was tested on days 1, 15, and 30 after conditioning to probe early, long-term, and remote memory (Bontempi et al., 1996). During retrieval, mice were tested in a novel context with the CS+, CS1-, and two intermediate frequencies (7 and 11 kHz), presented in semirandom order, with each tone repeated three times (Fig. 1a).'
Why was testing conducted in a different context than that of conditioning? This is likely to result in an underestimation of generalization to the different tones...
In tone fear conditioning, it is always customary to test in a different context to dissociate conditioning to the context vs conditioning to the tones, which usually take place simultaneously in the same context [20]. Therefore, testing generalization in a novel context gives the correct estimate of generalization to the tones in the absence of contextual conditioning confounds. Please note that while overall freezing levels may be lower in a novel context due to the absence of contextual conditioning, the relative generalization gradient across tones — which is what your study measures — is unlikely to be systematically distorted by context change.
(5) RESULTS - 'No-shock control mice showed no significant differences in freezing across frequencies on any testing day (p > 0.05; Fig. 1b, right), confirming that freezing reflected associative learning.'
The inference doesn't follow from the result described. Was there more freezing among animals in the shocked groups compared to those in the no-shock group? I presume so - my point is that this comparison is the one that most directly speaks to the presence or absence of associative learning.
Experimental animals exhibited not only higher overall freezing but also graded freezing responses across tone frequencies. It is important to note that no-shock controls did not display this pattern, ruling out the possibility that the different frequencies themselves elicited graded behavioral responses. To clarify this point, we revised the sentence as follows: "No-shock control mice showed no significant differences in freezing across frequencies on any testing day (p > 0.05; Fig. 1b, right), confirming that the graded freezing patterns resulted from associative learning rather than the acoustic properties of the tones." (Pg. 6)
(6) 'Across animals and sessions, we identified distinct neuronal populations showing positive modulation, negative modulation, mixed responses, or no consistent response to sound (Fig. 2b)...'
To be clear, do you mean to say that there were distinct neuronal populations that consistently [i.e., across all three sessions] increased their responses to the tones [positive modulation], decreased their responses to the tones [negative modulation], showed variable responses to the tones [mixed responses], and did not respond to tones [not modulated]?
The sentence refers to neuronal populations identified within each recording session based on their responses to the tones, not to neurons that maintained the same response profile across all three sessions. We have revised the text to make this distinction explicit. The only stable patterns across sessions were observed in graded neurons that were stable across retrieval.
“Across animals, we identified distinct neuronal subpopulations showing positive modulation, negative modulation, mixed responses, or no consistent response to sound in each session (Fig. 2b)” Pg. 7
(7) What does 'active' mean in relation to Figure 2? Does this refer to neurons that displayed either positive responses, negative responses, and/or mixed responses? In the text, it is stated that 'Sound responder neurons were classified using a test that detected modulation based on magnitude relative to baseline variability, allowing reliable identification of both transient and sustained responses while remaining robust to noise...'
I can't work out if this is the same classification criteria used for the determination of positive modulation, negative modulation, and mixed responding.
We thank the reviewer for pointing out this ambiguity. In Figure 2, the term "active" referred to neurons that exhibited significant sound-evoked modulation and were subsequently classified as showing positive, negative, or mixed responses. Thus, active and sound-responsive refer to the same population of neurons. We removed the word active to avoid confusion.
The reference to transient and sustained responses describes the temporal profile of the calcium signals rather than separate response categories. Some neurons exhibited brief calcium transients that rose and decayed rapidly, whereas others displayed sustained activity throughout the tone presentation. The sound-response detection algorithm was designed to reliably identify both temporal response profiles. We have revised the manuscript to make these definitions explicit. We modified the sentence as follows: “Sound-responsive neurons were identified using a statistical test that detected activity modulation relative to baseline variability, allowing reliable identification of responses while remaining robust to noise. This approach was effective for neurons exhibiting either brief calcium transients that rose and decayed rapidly or sustained activity throughout the tone presentation.” Pg. 7-8
(8) 'A moderate proportion of neurons was present across all retrieval sessions, with no differences between groups (p > 0.05).'
Do you mean to say that 'A moderate proportion of neurons was ACTIVE across all retrieval sessions, with no differences between groups (p > 0.05)'?
We replaced the word present and replaced it with “active”. Pg. 8
(9) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:
'If neurons encoding graded responses carry core mnemonic information, they should exhibit enhanced stability over time. To test this hypothesis, we quantified the proportion of registered neurons that retained their cluster identity across at least two retrieval sessions and compared these values to a shuffled null distribution (10,000 iterations), with multiple comparisons controlled using the BenjaminiHochberg procedure.'
What does the comparison to the shuffled null distribution tell us exactly? I accept that some neurons were stable positive responders across at least two sessions. The comparison to the shuffled null distribution creates a false impression about the robustness of this stability or the 'enhanced stability over time'.
Our intention in comparing the observed stability to a shuffled null distribution was to evaluate whether the proportion of neurons retaining cluster identity exceeded chance levels expected from random assignment. The shuffled distribution therefore provides a statistical baseline against which the observed degree of stability can be evaluated. We agree, however, that the wording “enhanced stability over time” may be confusing regarding this finding. We rephrased this paragraph to clarify that a subset of neurons retained cluster identity across all retrieval sessions at levels greater than expected by chance as follows:
“These data demonstrate that graded clusters remain consistently active at levels exceeding chance, preserving their cellular identity and providing a stable representation of learned contingencies and generalization gradients.” Pg. 15.
(10) ABSTRACT. The abstract states that, 'Stimulus-evoked population similarity scaled precisely with behavioral generalization, and consistent population states emerged only for tones associated with shock or those eliciting strong generalized freezing, indicating that population-level similarity predicts inferred threat.'
I believe that the sentence could be rewritten as, 'Stimulus-evoked population similarity reflected the degree of generalization, and consistent population states emerged only for tones associated with shock or those eliciting strong generalized freezing.'
We revised the text according the reviewer’s suggestion; however, we had to shorten the sentence due to word limits. “Population similarity tracked behavioral generalization, whereas consistent population states emerged only for shock-associated or highly generalized tones.” Pg. 2
Reviewer #3 (Recommendations for the authors):
Major points:
(1) The ensembles with graded activation in proportion to stimulus valence are described at various points in the manuscript as "maintaining the emotional 'gist'", "preserving core components of the memory trace", and "preserving core components of the memory trace". This conclusion is premature because there is an alternative interpretation. The graded response ensembles would also be consistent with coding for the freezing behavior itself, irrespective of the specific memory or stimulus association that drives it. An ensemble that encodes a behavior in this way would not be considered mnemonic, just as motor neurons in the spinal cord are not, even if they may fire during a conditioned response. Indeed, previous work has identified neurons in the prelimbic cortex that encode freezing independently from the stimuli that signal an aversive outcome (e.g., Kyriazi, Headley, and Pare 2020; Casanova, Pouget, ..., Vetere 2024).
There are two ways the authors can address this point.
(a) Fit a generalized linear model to the time series of inferred spiking activity from the Ca2+ signal and include stimuli and freezing as predictors. Since freezing behavior is inconsistent across trials, while stimulus presence is fixed, they can be disassociated. If, after accounting for freezing, responsiveness neurons still show a graded coding of stimuli that agrees with inferred aversiveness, this would strengthen their claim that they have identified an ensemble that corresponds with mnemonic or salience aspects of the stimuli.
(b) Conduct no further analysis but cover the issue in the discussion as a limitation to their study and to dampen some of the language throughout the manuscript that implies that a memory trace has been identified.
We thank the reviewer for this thoughtful and constructive comment. We agree that an important alternative interpretation is that graded-response ensembles could reflect freezing-related activity rather than representations of learned threat value. To directly address this possibility, we implemented a Generalized Linear Model (GLM) analysis, as suggested by the reviewer. The GLM was fitted to the activity of every sound-responsive neuron included in the population analyses and simultaneously incorporated tone identity and continuous freezing probability (derived probabilistically from miniscope acceleration) as predictors, allowing us to quantify their independent contributions to neuronal activity.
We want to note that freezing was quantified from the miniscope's inertial measurement unit (IMU) using a two-component Gaussian mixture model applied to the log-transformed body-acceleration signal. Rather than classifying freezing with a binary threshold, we used the posterior probability of the low-movement state as a continuous freezing regressor. This approach captures graded variations in immobility and provides a more conservative test of tone encoding, because it accounts for more behaviour-related variance than a binary classifier, making it more difficult to detect an independent contribution of tone identity.
We applied the GLM both to all sound-responsive neurons contributing to the population response curves and separately to the identified frequency-selective and graded neuronal subpopulations. Across all analyses, tone identity consistently explained neuronal activity better than freezing. Furthermore, the freezing-corrected tone β coefficients scaled with learned threat value, with graded neurons exhibiting the strongest monotonic gradients, indicating that they provide the most robust representation of learned threat value. These findings demonstrate that the graded coding of learned threat value persists after accounting for freezing behavior and therefore cannot be explained simply by the behavioral expression of fear. The new analyses are presented in Figures 4 and 8. Notably, although freezing-dominant neurons were present in both the tone-selective and graded populations, they represented only a small fraction of each group and were least prevalent among graded neurons (6–8%), further supporting the conclusion that graded neurons primarily encode learned threat value.
In addition, we revised the manuscript to avoid language implying that these neuronal populations constitute a mnemonic trace. Instead, we consistently describe them as encoding learned threat value, a more accurate interpretation that is directly supported by the new GLM analyses.
(2) The title makes a seemingly causal claim by using the term 'arise', "Learned and inferred valence arise from interactions between stable and dynamic subnetworks". While it is true that the authors show that both stable and dynamic ensembles encode valence, they do not demonstrate that the behavioral expression of valence depends on these codes, nor their interaction. Experimentally testing this is beyond the scope of this study (holographic two-photon stimulation of transient and stable ensembles?), but they may be able to get closer to it by examining inter-individual variability. The authors could measure the proportion of neurons in each subject that participate in the stable (graded responding) and dynamic (stimulus-specific) ensembles, and see if they predict individual differences in the expression of freezing behavior or its generalization. Indeed, this correlation may change across testing days.
We agree that the term “arise” in the title may imply a stronger causal relationship than is directly supported by the present data. We modified the title in the resubmission as follows: “Complementary stable and dynamic prelimbic ensembles encode learned threat value underlying generalization and discrimination”
The new LGM analysis confirms that a large proportion of neural activity can be predicted by tone threat value; therefore, we think this title fully captures our findings.
We also appreciate the reviewer's suggestion to examine inter-individual variability. In the revised manuscript, we tested whether the proportion of graded neurons correlated with freezing behavior. However, the limited number of experimental animals in each experimental group provided insufficient statistical power to reliably assess this relationship. Accordingly, we revised the manuscript to clarify that our conclusions are based on correlational observations rather than causal inferences.
In summary, we revised the title and related language throughout the manuscript to avoid implying causal mechanisms beyond the scope of the current experiments.
Minor points:
(1) I was surprised by the absence of an ensemble in the No-shock group that responded uniformly to all stimuli. Can the authors confirm this?
Yes, we confirm this finding. It was unexpected to us as well. We would like to clarify, however, that some control neurons may have responded to more than one frequency, but these responses were too infrequent or too weak to be classified as a distinct graded neuronal population by our clustering algorithm. Thus, while broadly responsive neurons may have been present in the control group, they did not form a robust, identifiable ensemble comparable to that observed after fear conditioning.
(2) Several different approaches were used to analyze the same Ca2+ responses to stimuli across testing days. These were the "Sound responder classification", "Average stimulus-aligned trace procedure", "Population similarity over time across tone pairs", and the construction of "Stimulus response vectors". These feature differing alignment/binning/interpolation, normalization, and response quantification procedures, and it is unclear why they cannot all be in agreement, at least when it comes to alignment and normalization.
We thank the reviewer for this careful reading of our Methods. All analyses were performed on the same underlying calcium imaging dataset, but they were designed to address different aspects of the data and therefore required different preprocessing steps. The analyses share a common initial pipeline leading to the calcium traces (all z-scored across the session). Differences in subsequent processing (e.g., use of ΔF/F versus CASCADE-deconvolved activity, normalization, baseline correction, temporal binning, and interpolation) were introduced only when required by the specific analysis method.
To make this clearer, we have substantially revised the Methods. We added a new overview of preprocessing section that summarizes the common preprocessing pipeline and explicitly distinguishes the shared steps from those that are analysis-specific. We also included a summary table describing the input signal (ΔF/F or CASCADE-deconvolved activity), normalization procedure, and temporal processing used for each analysis. Finally, the individual Methods sections were revised to eliminate redundancies and more clearly describe the steps to avoid confusion. We hope these revisions make the rationale for the different preprocessing procedures and the overall analytical workflow more transparent (Pg. 22-23)
(3) In the methods section "Window-wise response quantification" the Ca2+ signal was baselinesubtracted and divided by the standard deviation in the baseline across trials, but that data was already presumably z-normalized to the baseline of each trial ("Data alignment and normalization"). This second step of normalization seems excessive. Why is it not sufficient to just take the average peri-stimulus response across the z-normalized trials from the "Data alignment and normalization" section? This is simpler and would capture the effect size of the response relative to baseline.
We thank the reviewer for this careful observation.
The two operations are also not the same normalization applied twice; they standardize different sources of variability. During the alignment step, each trial is z-scored relative to its own baseline by dividing by the standard deviation of that trial's baseline across time. This places all trials on a common within-trial scale before averaging. In the window-wise step, the trial-averaged, baseline-subtracted response is expressed relative to the standard deviation of the per-trial baseline levels across trials—a distinct quantity that reflects trial-to-trial baseline stability rather than within-trial fluctuations. The purpose of this second term was to down-weight windows in cells with unstable baselines across trials, and it entered the analysis only as a significance criterion; the magnitude threshold defining a sound responder was applied to the trial-averaged baseline-relative response itself. We have revised the Methods to clarify the distinct roles of these two normalization steps.
To further address this concern, we re-ran the sound-responder classification after removing the second (between-trial) normalization step, so that responder detection depended only on the per-trial baselinerelative response magnitude and its temporal persistence. Across all cells, tones, and sessions (n = 89,504 cell–tone–session classifications), the two procedures agreed on 95.3% of labels. The small fraction of cells whose labels changed were almost exclusively those lying immediately at the detection threshold: 86.6% of changes involved cells moving into or out of the "modulated" category, whereas direct reversals between excitatory and inhibitory classification occurred in only 8 of 89,504 cases (0.009%). Consistent with the between-trial standard deviation being a less stable quantity when few trials are available, label changes were approximately twice as frequent in the three-trial retrieval sessions (5.2%) as in the ten-trial conditioning sessions (2.1%). Overall responder proportions changed only minimally (positive responders +2.2%, negative responders +4.3%), and all population-level findings—including the graded threat-value gradient across tones, its absence in no-shock controls, and its persistence after controlling for freezing in the , as analysis—were unaffected. These analyses demonstrate that our conclusions are robust to this methodological choice.
(4) It would increase confidence in the tracking of neurons across days if the authors showed some example images of neurons tracked across days.
We added an example in the Supplement. Additionally, we now provide measures of registration quality (Fig. S3)
(5) Table S2 is a bit confusing. I take it that Graded A/B were only for CS15, and Graded C/D/E were only for CS3. Also, Common B1-4 were the cells with positive responses to individual stimuli, and Common C1-4 were the cells with negative responses to stimuli. If this is the case, it should be explained in the figure legend (or even better, clusters should be named and numbered consistently in all figures.
Thank you for pointing this out, we corrected the Table to indicate which test corresponds to which figure and cluster, specifying which ones were positive or negative modulated. Please note old Table 2 is now Table 5
(6) The term network and subnetworks implies some connectivity between neurons, but in this study, it is used to refer to the ensembles of cells activated in a similar manner. Since connectivity is never assessed, it would be better if the authors stuck to the terms ensembles or populations.
We changed the wording and now use ensembles or populations
(7) The legend for Figure S3 has the text 'eded', which seems to be a typo.
We corrected this typo.
References
(1) Kitamura, T., et al., Engrams and circuits crucial for systems consolidation of a memory. Science, 2017. 356(6333): p. 73–78.
(2) DeNardo, L.A., et al., Temporal evolution of cortical ensembles promoting remote memory retrieval. Nat Neurosci, 2019. 22(3): p. 460–469.
(3) Tome, D.F., et al., Dynamic and selective engrams emerge with memory consolidation. Nat Neurosci, 2024. 27(3): p. 561–572.
(4) Mau, W., M.E. Hasselmo, and D.J. Cai, The brain in motion: How ensemble fluidity drives memory-updating and flexibility. Elife, 2020. 9.
(5) Gallego, J.A., et al., Long-term stability of cortical population dynamics underlying consistent behavior. Nat Neurosci, 2020. 23(2): p. 260–270.
(6) Zaki, Y. and D.J. Cai, Memory engram stability and flexibility. Neuropsychopharmacology, 2024. 50(1): p. 285–293.
(7) Lacagnina, A.F., et al., Distinct hippocampal engrams control extinction and relapse of fear memory. Nat Neurosci, 2019. 22(5): p. 753–761.
(8) Sangha, S., Plasticity of Fear and Safety Neurons of the Amygdala in Response to Fear Extinction. Front Behav Neurosci, 2015. 9: p. 354.
(9) Bouton, M.E., S. Maren, and G.P. McNally, Behavioral and Neurobiological Mechanisms of Pavlovian and Instrumental Extinction Learning. Physiol Rev, 2021. 101(2): p. 611–681.
(10) Jenkins, H.M. and R.H. Harrison, Effect of discrimination training on auditory generalization. J Exp Psychol, 1960. 59: p. 246–53.
(11) Dunsmoor, J.E. and K.S. LaBar, Effects of discrimination training on fear generalization gradients and perceptual classification in humans. Behav Neurosci, 2013. 127(3): p. 350–6.
(12) Herzog, K., et al., Reducing Generalization of Conditioned Fear: Beneficial Impact of Fear Relevance and Feedback in Discrimination Training. Front Psychol, 2021. 12: p. 665711.
(13) Lommen, M.J.J., et al., Training discrimination diminishes maladaptive avoidance of innocuous stimuli in a fear conditioning paradigm. PLoS One, 2017. 12(10): p. e0184485.
(14) Casanova, J.P., et al., Threat-dependent scaling of prelimbic dynamics to enhance fear representation. Neuron, 2024. 112(14): p. 2304–2314 e6.
(15) Kyriazi, P., D.B. Headley, and D. Pare, Different Multidimensional Representations across the Amygdalo-Prefrontal Network during an Approach-Avoidance Task. Neuron, 2020. 107(4): p. 717–730 e5.
(16) McKenzie, S. and H. Eichenbaum, Consolidation and reconsolidation: two lives of memories? Neuron, 2011. 71(2): p. 224–33.
(17) Winocur, G. and M. Moscovitch, Memory transformation and systems consolidation. J Int Neuropsychol Soc, 2011. 17(5): p. 766–80.
(18) Hockley, A. and M.S. Malmierca, Auditory processing control by the medial prefrontal cortex: A review of the rodent functional organisation. Hear Res, 2024. 443: p. 108954.
(19) Zikopoulos, B. and H. Barbas, Prefrontal projections to the thalamic reticular nucleus form a unique circuit for attentional mechanisms. J Neurosci, 2006. 26(28): p. 7348–61.
(20) Phillips, R.G. and J.E. LeDoux, Differential contribution of amygdala and hippocampus to cued and contextual fear conditioning. Behav Neurosci, 1992. 106(2): p. 274–85.