A Functional Influence Based Circuit Motif That Constrains the Set of Plausible Algorithms of Cortical Function

  1. Friedrich Miescher Institute for Biomedical Research, Basel, Switzerland
  2. Faculty of Science, University of Basel, Basel, Switzerland

Peer review process

Revised: This Reviewed Preprint has been revised by the authors in response to the previous round of peer review; the eLife assessment and the public reviews have been updated where necessary by the editors and peer reviewers.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Rui Ponte Costa
    University of Oxford, Oxford, United Kingdom
  • Senior Editor
    Timothy Behrens
    University of Oxford, Oxford, United Kingdom

Reviewer #1 (Public review):

Vasilevskaya and Keller test different models of cortical function through the lens of predictive processing, a powerful framework for the brain to learn and predict the statistics of the world via generative internal models. The authors use a clever combination of behavioral perturbations in closed-loop and open-loop visuomotor virtual reality assays, a paradigm the Keller lab pioneered and used effectively in the past decade, in conjunction with two photon imaging of neuronal calcium responses and targeted optogenetic perturbations of activity. They specifically put to test proposed hierarchical vs. non-hierarchical circuit implementations of predictive processing by analyzing the logic of inter-lamina interactions (superficial vs. deep; L2/3 vs. L5/6).

The authors conclude that both versions of predictive processing architectures they analyze are likely invalid and instead formulate an alternative novel model of cortical function based on a recently developed machine learning algorithm for self-supervised learning (joint embeddings of predictive architectures, JEPA) and its further refinements. JEPA borrows elements from predictive processing engaging two encoder networks and training the output of one network to predict the output of the other. In their new model of cortical computations, prediction errors neurons in L2/3 compare the deep layers (L5/6) activity, which is taken as a teaching signal, to a local, L2/3 prediction of this latent representation.

Specifically, the authors build on their previous work and reports from other groups that different sets of L2/3 neurons compute positive prediction errors (fire when sensory stimuli appear unexpectedly with respect to the movements of the animal; e.g., grating onsets in the absence of locomotion) and respectively negative prediction errors (fire when sensory stimuli are absent, while the brain expected them to be present; e.g. mice locomote but visual flow is suddenly halted - visuomotor mismatches). These L2/3 positive and negative prediction error neurons exchange messages with neurons in the deeper cortical layers that, the authors propose, build an internal representation (R) of the sensory stimuli given the animals' movements.

In the hierarchical model, internal representation neurons (R) are supposed to act as a teaching signal for both types of prediction error neurons; the output of the positive prediction error neurons is assumed to suppress activity of R such that the error between the teaching signal and the prediction is minimized; similarly, in the non-hierarchical version, R serves as a prediction for the prediction error neurons, and in turn it receives excitatory drive from the positive prediction error neurons and negative input from the negative prediction error neurons.

The authors find that the functional impact of L5 neurons to L2/3 neurons is not compatible with the non-hierarchical architecture they and other groups proposed, but rather in accordance with the hierarchical model. At the same time, the functional impact of L2/3 neurons (positive vs. negative prediction error neurons) on L5 neurons (internal representation) appears not compatible with the hierarchical model, but rather in accordance with the non-hierarchical implementation.

They further hypothesize that L2/3 prediction error neurons don't use sensory input, but rather the L5 activity as a teaching signal, and test it using perturbations (halts) of optogenetic stimulation of L5 neurons coupled with locomotion (Fig.7).

All in all, the question is topical, and the new model addresses a decades-long quest to develop a unifying model of cortical function. The findings reported here transform our understanding of cortical computations, opening new exciting avenues for future investigation. The experimental design and execution are rigorous; the arguments are clearly laid out (in spite of ample potential for confusion given the numerous loops and sign flips). These include a discussion of why the non-hierarchical model proposed by the same group does not hold, as well as potential caveats in interpreting the results and novel testable proposed experiments emerging from the JEPA-like model.

Comments on revised version

I commend the authors for nicely provided answers and addressing my concerns and clarifying their points on the relationship of their findings to JEPA and current state of understanding.

In particular for Q5 -I meant if there is also a positive correlation between optomotor mismatch response and visuomotor mismatch response when looking only at neurons that the authors identify as PE-?
The authors provided the answer on point.

Overall, I think this is a fundamental study and the strength of evidence for the claims is exceptional as an exemplary use of existing approaches.

Reviewer #2 (Public review):

This manuscript reveals functional connectivity of two different classed of cortical neurons that respond in opposite ways to mismatches between sensory and top-down inputs. These data are very valuable because different theories of information processing in the cortex make different predictions on the patterns of connectivity of these neurons. Therefore, these data strongly constrain possible theories of cortical processing.

Comments on revised version.

I thank the Authors for answering my questions and updating the manuscript.

Congratulations on this important work!

Reviewer #3 (Public review):

Vasilevskaya and Keller set out to experimentally distinguish between two variants of predictive processing: a hierarchical and a non-hierarchical variant. The hierarchical variant assumes a hierarchical organization in which internal representation neurons (believed to be a subset of layer 5 excitatory neurons) serve as a source of a teaching signal for local prediction error neurons as well as for the next higher level of the hierarchy, while simultaneously providing prediction signals to the preceding lower level. In contrast, the non-hierarchical variant posits that these layer 5 internal representation neurons provide local predictions to layer 2/3 prediction error neurons.

The interaction between internal representation neurons and prediction error neurons differs fundamentally between the two variants. In the hierarchical variant, internal representation neurons excite positive prediction error neurons and inhibit negative prediction error neurons, while at the same time being inhibited by positive prediction error neurons and excited by negative prediction error neurons. In the non-hierarchical variant, this pattern of connectivity is reversed.

This work is very exciting, timely, and carefully executed. The authors functionally, and later molecularly, identify layer 2/3 prediction error neurons in V1 and probe their interactions with genetically defined neuron types in cortical layers 5 and 6 using optogenetics. They demonstrate that the functional influence of putative prediction error neurons in layer 2/3 onto layer 5 is incompatible with the hierarchical variant, whereas the influence of layer 5 onto putative prediction error neurons in layer 2/3 is incompatible with the non-hierarchical variant. They then test an alternative hypothesis, in which layer 2/3 responses resemble prediction errors with respect to perturbations of artificial layer 5 activity patterns. To investigate this, they designed an experiment in which optogenetic activation of L5 IT neurons was closed-loop coupled to the mouse's locomotion speed in the absence of visual feedback, allowing them to probe the causal influence of L5 activity on layer 2/3 responses.

Finally, the authors hypothesize that their data are more consistent with a joint embedding predictive architecture (JEPA) and outline experimentally testable predictions arising from this framework.

While the work is overall convincing and provides important insights into the circuit-level implementation of predictive processing, I think the connection to JEPA networks would benefit from a more in-depth discussion of its relationship to recently proposed and implemented models. Below, I address the specific points raised by the authors (flanked by ' ... ' to make the author's statements stand out), in particular in relation to the model proposed by Nejad et al. (2025):

- 'The two proposals indeed share similarities in assuming that bottom-up input for both L2/3 and L5 arrives from thalamus, and that representations formed in L2/3 are used for predicting the activity of L5. However, there are a few important differences between the JEPA implementation proposal formulated here and the Nejad et al. model.

(1) There is no proposed mapping of computations in the Nejad et al. model onto different JEPA networks. We assume that the suggested mapping would be L4 and L5 as encoder networks, and L2/3 as a predictor network? In that case, it is different to our proposal, in which L2/3 is part of the encoder network.'

I think there is some confusion here. In Nejad et al., both L2/3 and L5 function as encoder networks. Each receives sensory input (with L2/3 receiving this input delayed via L4) and computes a latent representation of that input, denoted z_{L2/3} and z_{L5}, respectively. The prediction is obtained by comparing the output of L2/3 with L5 latent representations (via W_{L2/3->L5} * z_{L2/3}). In other words, L5 is the target in the learning objective.

Although, they did not explicitly state the mapping with JEPA, their model has the same fundamental property - predictive learning happens in the latent space. Also, this appears very similar to the roles assigned to L2/3 and L5 in your Figure 9B. The fact that you made this more explicit and the new data included, is in my view, a very interesting contribution. However, from an architectural perspective, it appears that your proposal and the model of Nejad et al. are conceptually very similar, and I do not see a substantial difference between the two. This should be made more clear in the Discussion.

- '2. Our proposal contains explicit prediction error neuron cell types within L2/3, while prediction errors in Nejad et al. are encoded in the gradients, and the layer origin of these signals is hypothesized to be L5 ('the learning-driving error signal originates in L5'). Hence, also the role of L5-L2/3 connection is distinct between the two proposals. In Nejad et al. this connection serves the role of error propagation and update for predictions in L2/3, while in our proposal this connection contains teaching signal (target representations) that are compared to predictions within L2/3. Similarly, the functional role of L2/3-L5 connection is also different, since in Nejad et al, it is supposed to carry predictions of L5 activity, whereas in our proposal we expect it to drive plasticity in L5 encoder.'

Indeed, in Nejad et al., the layer-dependent mismatch responses are modeled as gradients with respect to neuronal activity, and the model does not explicitly include prediction error neurons. However, this appears to be a modeling choice rather than a fundamental aspect of the proposal, and it does not preclude an implementation with explicit prediction error neurons. In fact, the authors explicitly acknowledge this possibility in the Discussion:

"The second approach would be to recast our model within a predictive coding framework... Predictive coding jointly optimizes both model parameters and neuronal activities, which could naturally lead to prediction errors observable in the activity of both L2/3 and L5 neurons. Note that these two views are not mutually exclusive."

While Nejad et al. hypothesize that the learning-driving error signals (that are distinct from their mismatch responses) originate in L5, the abstract loss function itself does not uniquely specify where the underlying comparison between the predicted representation (W_{L2/3 -> L5} z_{L2/3}) and the target representation (z_{L5}) must be implemented. The proposed biological implementation places this computation in L5, but from my understanding, the computational objective itself does not require this specific localization.

That said, I agree that your proposed model introduces a genuine difference. The functional roles assigned to the vertical projections are effectively reversed: in Nejad et al., the L2/3->L5 projection carries the prediction, whereas the L5->L2/3 projection conveys the error/gradient. In your architecture, by contrast, the L5->L2/3 projection carries the teaching signal (target). This is, in my view, a real and testable interpretational divergence that is worth stating clearly.

Therefore, I think the novelty lies less in the computational architecture itself and more in committing to a particular biological implementation-one that adds cell-type-specific detail to an implementation that Nejad et al. hypothesised as being compatible with their framework.

- '3. The difference outlined above also makes it evident that the two proposals should differ in how deep and superficial layers are expected to influence the activity of one another. Indeed, the proposal in Nejad et al. is based on the cortical column idea, and according to eq. 2 and 3 in the Methods, activity in L5 is a function of activity in L2/3, while activity in L2/3 is not a function of activity in L5. Our proposal is based on idea of layers forming parallel networks, where horizontal communication is the dominant mode of cortico-cortical interactions, and activity in deep layers serve as a teaching signal for L2/3. In our case, we expect the opposite - that activity in L2/3 depends on activity of L5, while activity of L5 is not immediately dependent on activity of L2/3 (only via plasticity route). This led us to propose one of direct tests for our framework - silencing L2/3 in a familiar setting should result in no immediate changes to L5 activity and behavior of the animal.'

My reading of Nejad et al. is consistent with your interpretation of the equations. Specifically, Eq. 3 makes L5 activity depend on L2/3 activity (albeit weakly, with a = 0.3), whereas Eq. 2 contains no L5 term, so L2/3 activity does not depend directly on L5 activity. In that model, the L5->L2/3 pathway carries the learning gradient rather than contributing to the activity dynamics. By contrast, in your proposed model, L2/3 activity depends on L5 activity, whereas L5 activity is not immediately dependent on L2/3 activity (except indirectly through learning/plasticity). So, you state that "silencing L2/3 in a familiar setting should result in no immediate changes in L5 activity.
[...].

However, I am unsure how to reconcile this prediction with the results shown in Fig. 6. If I understand the figure correctly, optogenetic activation of Rrad-positive (positive prediction error) L2/3 neurons produces a small increase in L5 activity, whereas activation of Adamts2-positive (negative prediction error) L2/3 neurons produces a decrease in L5 activity. Although these experiments involve activation rather than silencing, they nevertheless suggest that perturbing L2/3 activity can have an immediate effect on L5 activity. Could you clarify how this is consistent with the proposed model? In other words, what aspect of the proposed circuitry makes activation effective while silencing is predicted to have no immediate consequence?

For comparison, Nejad et al. performed a related perturbation analysis in Fig. S15 by scaling the output of L2/3 neurons exhibiting positive mismatch signals (defined through the activity gradients), which increased L5 activity, whereas scaling neurons with negative mismatch signals produced the opposite effect. I am not entirely sure how directly these simulations map onto the experiments shown in your Fig. 6, since the Nejad simulations were performed during mismatch conditions, if I have understood them correctly.

- '4. The proposal in Nejad et al. relies on input reconstruction or variance maximization within the L5 autoencoder network to avoid collapse. Instead, our proposal has no reconstruction objective.'

Nejad et al. only require two encoders (like in JEPA), how these two are learnt can be done in several ways. While Nejad et al. focus on using a reconstruction loss to learn the L5 target, they also show that it works equally well with non-reconstruction objectives. Therefore, I do not think the presence or absence of a reconstruction objective constitutes a fundamental distinction between the two proposals.

As your current work presents a conceptual architecture rather than a fully implemented learning algorithm (in a model), the mechanism that would prevent representational collapse has not yet been defined. From my understanding, every joint-embedding approach must address this issue, whether through reconstruction, variance/covariance regularization, stop-gradient or EMA mechanisms, or other approaches. Thus, the absence of a reconstruction objective (or another anti-collapse mechanism) is not, in itself, a distinguishing feature of the proposed architecture, but rather an as-yet unspecified design choice within the learning objective.

- '5. Lastly, there is time-delay between inputs to L5 and L2/3 that is proposed in Nejad et al., while this is not something inherent to our proposal.'

I agree that the temporal delay introduced by L4 is a key component of the Nejad et al. model and is currently absent from your proposal. However, I would expect temporal delays to emerge naturally in your framework as well, given the multisynaptic and highly parallel organization of cortical circuits. More generally, implementing predictive learning over time (as in JEPA) requires comparing representations at times t and t+1, which seems to require some form of temporal delay. How else would you suggest this is done?

In general, I think the manuscript would benefit from a clearer discussion of its relationship to the model proposed by Nejad et al. (perhaps following the discussion above), including both the shared conceptual claims and the aspects that genuinely differ between the two frameworks. At present, some of the claims are presented as novel, although at least some of these core ideas have already been proposed in Nejad et al.

For example, the authors state: "Thus, we propose that layer 2/3 functions to predict layer 5 activity, not sensory input per se, hence making predictions in the internal representation space, not input space." This appears to be exactly what Nejad et al. proposed as discussed above - in their model L2/3 predicts L5 activity (purely in the latent space), as they state in the abstract.

That said, there are some interesting differences, and I think the community would greatly benefit from making these clear, including the roles assigned to interlaminar connections and the interpretation of the signals carried by these pathways. These differences are interesting and potentially testable, and I think the manuscript would be strengthened by explicitly distinguishing which aspects are in line with the ideas already present in Nejad et al. and which aspects represent new contributions.

Author response:

The following is the authors’ response to the original reviews.

We thank you for the time you took to review our work and for your feedback! We have performed the following additional analyses:

(1) We analyzed the optomotor mismatch response as a function of time spent in the optogenetic closed-loop session.

(2) We show the correlation between optomotor mismatch response and visuomotor mismatch response for functionally identified PE neurons.

The main changes to the manuscript are:

(3) Clarification of our terminology (moved and refined definition of the teaching signals).

(4) A new figure panel summarizing the functional influence patterns we identified.

(5) Clarification of our argumentation for the JEPA-inspired proposal.

All comments are addressed individually in the following.

Public Reviews:

Reviewer #1 (Public review):

Vasilevskaya and Keller test different models of cortical function through the lens of predictive processing, a powerful framework for the brain to learn and predict the statistics of the world via generative internal models. The authors use a clever combination of behavioral perturbations in closedloop and open-loop visuomotor virtual reality assays, a paradigm the Keller lab pioneered and used effectively in the past decade, in conjunction with two-photon imaging of neuronal calcium responses and targeted optogenetic perturbations of activity. They specifically put to test proposed hierarchical vs. non-hierarchical circuit implementations of predictive processing by analyzing the logic of inter-lamina interactions (superficial vs. deep; L2/3 vs. L5/6).

The authors conclude that both versions of predictive processing architectures they analyze are likely invalid, and instead formulate an alternative novel model of cortical function based on a recently developed machine learning algorithm for self-supervised learning (joint embeddings of predictive architectures, JEPA) and its further refinements. JEPA borrows elements from predictive processing, engaging two encoder networks and training the output of one network to predict the output of the other. In their new model of cortical computations, prediction error neurons in L2/3 compare the deep layers (L5/6) activity, which is taken as a teaching signal, to a local, L2/3 prediction of this latent representation.

Specifically, the authors build on their previous work and reports from other groups that different sets of L2/3 neurons compute positive prediction errors (fire when sensory stimuli appear unexpectedly with respect to the movements of the animal; e.g., grating onsets in the absence of locomotion) and respectively negative prediction errors (fire when sensory stimuli are absent, while the brain expected them to be present; e.g. mice locomote but visual flow is suddenly halted - visuomotor mismatches). These L2/3 positive and negative prediction error neurons exchange messages with neurons in the deeper cortical layers that, the authors propose, build an internal representation (R) of the sensory stimuli given the animals' movements.

In the hierarchical model, internal representation neurons (R) are supposed to act as a teaching signal for both types of prediction error neurons; the output of the positive prediction error neurons is assumed to suppress activity of R such that the error between the teaching signal and the prediction is minimized; similarly, in the non-hierarchical version, R serves as a prediction for the prediction error neurons, and in turn it receives excitatory drive from the positive prediction error neurons and negative input from the negative prediction error neurons.

The authors find that the functional impact of L5 neurons on L2/3 neurons is not compatible with the non-hierarchical architecture they and other groups proposed, but rather in accordance with the hierarchical model. At the same time, the functional impact of L2/3 neurons (positive vs. negative prediction error neurons) on L5 neurons (internal representation) appears not compatible with the hierarchical model, but rather in accordance with the non-hierarchical implementation.

They further hypothesize that L2/3 prediction error neurons don't use sensory input, but rather the L5 activity as a teaching signal, and test it using perturbations (halts) of optogenetic stimulation of L5 neurons coupled with locomotion (Figure 7).

All in all, the question is topical, and the new model addresses a decades-long quest to develop a unifying model of cortical function. The findings reported here transform our understanding of cortical computations, opening new, exciting avenues for future investigation. The experimental design and execution are rigorous; the arguments are clearly laid out (in spite of ample potential for confusion given the numerous loops and sign flips). These include a discussion of why the non-hierarchical model proposed by the same group does not hold, as well as potential caveats in interpreting the results and novel testable proposed experiments emerging from the JEPA-like model.

I have several questions about the interpretations of some of the claims and suggestions for potential additional experiments and analyses.

We thank the reviewer for their comments. We address them below.

(1) Some of the pieces of the puzzle remain to be identified and demonstrated: the existence of internal representation neurons in L2/3 and ascertaining that the L5/6 neurons analyzed function indeed as internal representation neurons. The authors find that stimulation of L2/3 positive prediction error neurons enhances activity of L5 neurons...If L5 neurons hold a latent representation that serves as a teaching signal for L2/3 neurons (as the authors posit), wouldn't one expect that the input they receive from the positive prediction neurons be suppressive, such that the error is further minimized?

Not necessarily - this depends on the model one has for how cortex works. In the hierarchical predictive processing model, PE+ neurons are expected to suppress the local internal representation neurons in L5. Our data, however, are not consistent with this model. This is one of the key arguments we build the idea on that JEPA is a better model for cortex than hierarchical predictive processing. In JEPA we would not expect the L2/3 prediction error to update L5 directly, but instead update the prediction of L5 activity (the source of this signal remains to be identified; see our speculations on the origin of prediction signals in the comments to question 3), and drive plasticity in the local L5 encoder network. That something acts like a teaching signal for the L2/3 comparator does not, by itself, determine how the outcome of the comparison influences the source of the teaching signal.

(2) Do the authors envision any specific differences between the representations of the two encoder networks posited to exist in L2/3 and L5 in the JEPA-like implementation? Are they synchronous/offset in their temporal representations, or any other features?

Given theoretical work (Mohammadi et al., 2025), one would expect to find differences in learning rates between the predictor and the encoder networks. Assuming the predictor network is also implemented in L2/3, we would expect to see faster learning rates in L2/3 compared to L5. Implementations inspired by related theoretical works also predict differences in learning rate between the two encoder networks (Grill et al., 2020), again with higher learning rates expected in L2/3 compared to L5. Beyond that, however, we are not aware of any experimentally observable differences one might expect to find. We are hoping the computational community will remedy this soon.

(3) Where is the prediction coming from onto L2/3 neurons? Is it emerging locally in L2/3 from the putative internal representation neurons, or is it long-range - as work from the authors previously proposed? Or a mix of both?

We expect the predictions to come from long-range inputs. In classical (hierarchical) JEPA one would expect these to be the lateral communication within L2/3. In a non-hierarchical implementation that is capable of operating on arbitrary graphs (as would be necessary for it to work in cortex), we suspect that the L2/3 network will use both long-range L2/3 and long-range L5 input. In JEPA terminology: the encoder A network uses information from non-local sources of both networks to predict local activity of encoder B network (assuming that information is statistically useful in predicting that activity).

(4) What is the role of the indiscriminate L4 input that appears to enhance activity of both positive and negative prediction error neurons in L2/3?

The short answer is, we don’t know. We would have indeed expected to find some asymmetry of influence. There are a few options: A) We might be failing to activate a specific interneuron that mediates feedforward inhibition on this pathway by the artificial stimulation of Scnn1a neurons. B) We know that Scnn1a neurons are only a subset of L4 neurons – there are other populations of L4 neurons that might exhibit the opposing influence. C) The Scnn1a population is more than 1 layer of a JEPA away from the comparator and not part of the predictor that forms the actual representation compared against L5 (e.g. L4 could provide an input to L2/3 internal representation neurons, and thus only indirectly influence L2/3 PE neurons). D) The JEPA analogy is wrong.

(5) Does Figure 7D change in a meaningful manner if the authors plot the correlation between optomotor mismatch response and visuomotor mismatch response specifically for the negative prediction error neurons in L2/3 (Adamts-2) rather than for all L2/3 cells sampled?

We might be misunderstanding. If the reviewer means genetically identified negative prediction error neurons (Adamts2), we do not have the data to address this question, as we did not perform any recordings of molecularly defined Adamts2 population in L2/3. If the reviewer means functionally identified negative prediction neurons, this would be the two rightmost data points in Figure 7D (the x-axis in this panel is the visuomotor mismatch response strength we use to functionally identify PE- neurons). These two data points on the right-hand side of the plot correspond exactly to what we classify as negative prediction error neurons throughout the rest of the manuscript (15% of the most responsive neurons to visuomotor mismatch).

Does the reviewer mean, is there also a positive correlation between optomotor mismatch response and visuomotor mismatch response when looking only at neurons that we identify as PE-? If so, the answer is yes (Author response image 1). Interestingly, while there is a positive correlation, there is also nonuniformity in response patterns. We think that this is expected, given that our visuomotor coupling paradigm captures only a very small subspace of stimuli that prediction error neurons are tuned to. Bulk stimulation of L5, in contrast, might work better to separate all PE neurons. In other words, we speculate that L5 stimulation is a better predictor of the functional role of an L2/3 neuron.

Author response image 1.

Optomotor mismatch response as a function of visuomotor mismatch response for neurons that are functionally classified as PE. Red line shows a linear fit estimated with a bootstrap approach.

(6) Do the optomotor mismatch responses in L2/3 neurons depend on how long the closed-loop coupling of optogenetic stimulation of Tlx3 L5 neurons and locomotion speed has been in place for?

No, not that we can measure. We performed an analysis in which we split the optogenetic closed-loop session into two equal parts, early and late. We then quantified the average optomotor mismatch response independently for early and late parts of the session (Author response image 2A). Based on this quantification, we find no evidence of a change as a function of time in the closed-loop session. We also performed a sliding window analysis on a shorter timescale. While it did look like the responses may be smaller in the first few minutes, we do not have sufficient data to address this. None of the differences in response size were significant (Author response image 2B).

Author response image 2.

Optomotor mismatch response as a function of experience with artificial closed-loop coupling. (A) Mean L2/3 population response to optomotor mismatch in the first half (Early MM) and in the second half (Late MM) of the optogenetic closed-loop session. (B) Mean L2/3 population response to optomotor mismatch as a function of time in the optogenetic closed-loop session. Error bars indicate SEM. Differences between the mean values were estimated by hierarchical bootstrap and are not significant.

Reviewer #2 (Public review):

This manuscript reveals the functional connectivity of two different classes of cortical neurons that respond in opposite ways to mismatches between sensory and top-down inputs. These data are very valuable because different theories of information processing in the cortex make different predictions on the patterns of connectivity of these neurons. Therefore, these data strongly constrain possible theories of cortical processing.

We thank the reviewer for their comments. We address them below.

General comments:

(1) The methods of statistical testing are insufficiently described. I did not understand the description in lines 1105-1119. The authors should provide sufficient details so the reader can reproduce their analyses. For example, it may be helpful to provide specific details of the testing procedure for one of the comparisons (e.g. the first comparison in Table S1).

We assume the reviewer is not familiar with hierarchical bootstrapping in general. If so, the explanation below would summarize the procedure. This is the procedure with particular emphasis on its application to neuroscience data is described in the paper we reference in that part of the methods (Saravanan et al., 2020). Given that the analysis has become relatively standard (and is described in the reference provided), we think it might be an overkill to add the full explanation below to the manuscript. In addition to the general procedure, there are only 2 pieces of information relevant to fully reconstructing the analysis:

(1) What are the “levels” (mice, recording sites, neurons, trials)?

(2) What is the number of bootstrap samples used.

Thus, we think all information is already provided in the methods. We now also explicitly added the levels when describing the nested structure of the data in the manuscript to increase clarity.

Hierarchical bootstrap analysis:

Our data are naturally nested: multiple neurons are recorded within a single mouse, and multiple mice are tested within an experimental group. Standard bootstrapping (sampling with replacement from the entire pool of neurons) fails because it assumes all observations are independent. In reality, neurons from the same mouse are more similar to each other than to neurons from a different mouse. The hierarchical bootstrap (or multi-level bootstrap) preserves this nested structure, ensuring your confidence intervals are not artificially narrow due to pseudoreplication.

To illustrate the problem, assume you have 100 neurons from Mouse A and 10 neurons from Mouse B, a simple bootstrap will be heavily biased toward Mouse A. Furthermore, the simple bootstrap ignores the fact that the true variance in your population comes from two sources:

(1) Between-mouse variance (differences in surgery, genetics, or behavior).

(2) Within-mouse variance (differences in tuning or activity between individual cells).

Hierarchical bootstrap addresses this problem, and is implemented as follows: To estimate the mean response while accounting for different sample sizes per mouse, a two-level resampling scheme is used

(1) Resample the higher level (Mice)

First, you account for the variability between animals.

Suppose you have N mice in total.

Randomly draw N mice with replacement from your original pool.

Note: Because this is with replacement, a single mouse’s data might be included multiple times in one bootstrap iteration, while another mouse might be left out entirely.

(2) Resample the lower level (Neurons)

For each mouse selected in Step 1, you must now account for the variability within that specific animal.

Look at the number of neurons actually recorded from that mouse (let’s call it ki).

Randomly draw ki neurons with replacement from that mouse’s specific pool of recorded cells.

This step is crucial: you always resample the same number of neurons that were originally recorded for that specific mouse. This maintains the "weight" or "influence" that animal had in the original dataset.

(3) Calculate the resampled statistic

Calculate the mean of all neurons collected in this “bootstrap sample”.

(4) Iterate

Repeat Steps 1–3 many times (typically B = 1,000 or 10,000 iterations).

The distribution of these B bootstrap means represents your sampling distribution and is used to calculate confidence intervals and p-values. 

(2) The authors should clarify how the problem of multiple comparisons was addressed for comparisons performed in multiple moments of time, where significance is indicated by a black bar (e.g. in Figure 2F).

There is no family-wise error correction in cases of comparing response time courses implemented in our analysis. If the reviewer has a good suggestion for how to implement family wise error correction, we would be happy to implement it. We are not aware of anything that is better than what we currently do (the time bin-wise comparison). To briefly explain the problem: In most neuroscience papers, response curves are compared by choosing a time window (e.g. 0.5 to 1 s following a trigger) and calculating mean values of the curves in these windows. This hides the problem of multiple comparisons that arises from the fact that the experimenter is free to choose a response window used for analysis. Note, this also creates a strong incentive to “optimize” choice of an analysis window – a part of the analysis that a reader is typically completely blind to. We could of course also choose an analysis window that sounds reasonable and yields significant differences for our analyses. However, to provide a more unbiased image of the data, we have come to do bin-wise comparisons with a fixed p-value (typically 0.05). This is used in all of our papers at the moment. Given that samples from neighboring timepoints are correlated via a combination of actual responses and a subset of noise sources, the samples are not independent. We now also implemented a correction for spurious positive values by requiring at least 2 neighboring bins to have a p-value below 0.05 to be shown. If the samples were independent, this would mean a false positive rate of 0.0025. Given that they are not, this is a lower bound only. Additionally, any family-wise error correction would be a function of the number of time bins we show in the plot. This would mean that our choice of the time window shown in a plot (-1s to +4s, or +5s, etc.) would change the p-value we consider significant. Thus, there is no explicit family-wise error correction, and we have come to the conclusion that the bin-wise comparison with a fixed p-value is the most unbiased representation of the data we can provide. The alternative would be to additionally plot z-scores or the p-values as a function of time, but in our experience these types of plots are even harder to read for readers not used to it.

(3) It would be helpful to add a figure in the Discussion summarising the functional connectivity suggested by all experiments.

We now added a panel that summarizes the functional connectivity observed in our experiments to Figure S10. 

(4) Throughout the manuscript, the authors use the term "teaching signals", but I am unclear what they mean by it: after reading the definition in lines 45-46, I thought that they corresponded to values (as they are compared to sensory signals). Later (428-430), the text suggests that they correspond to error neurons. But then lines 605-607 say it is not an error signal. The authors should define teaching signals very precisely or remove this term.

The formal definition of the teaching signal was in footnote 1 of the manuscript. We assume the reviewer may have missed this. We now moved this definition into the main text to increase clarity.

We use the term teaching signal to mean exactly this definition throughout the manuscript. We suspect, a second source of confusion may come from ambiguity in regards to anatomical vs. functional definitions. We have attempted to emphasize that the definition of teaching signal is a functional one, not an anatomical one (as is the case for ‘prediction’ – predictions are functionally defined, not anatomically – hence it makes sense to ask questions of the form “what are the potential sources of predictions” etc.). A teaching signal is a signal that is compared against a prediction (the ‘ground truth’ the prediction is compared and trained against). In different circuit implementations of predictive processing, different inputs function as predictions and teaching signals. In the hierarchical implementation, activity of the internal representation neuron at the lower level serves as a teaching signal for a prediction signal that is formed by the internal representation neuron from the higher level. In the non-hierarchical implementation, external inputs from the thalamus or other cortical areas serve as teaching signals for the respective prediction signals that are formed by internal representation neurons. This terminology is most intuitive when thinking from the perspective of a prediction error neuron, since a prediction error neuron computes the difference between two inputs signals – one of which functions as a teaching input, and the other one as a respective prediction. Hence, lines 428-430 specify that layer 5 input onto prediction error neurons of layer 2/3 serves as a teaching input.

Reviewer #2 (Public review):

Vasilevskaya and Keller set out to experimentally distinguish between two variants of predictive processing: a hierarchical and a non-hierarchical variant. The hierarchical variant assumes a hierarchical organization in which internal representation neurons (believed to be a subset of layer 5 excitatory neurons) serve as a source of a teaching signal for local prediction error neurons as well as for the next higher level of the hierarchy, while simultaneously providing prediction signals to the preceding lower level. In contrast, the non-hierarchical variant posits that these layer 5 internal representation neurons provide local predictions to layer 2/3 prediction error neurons.

The interaction between internal representation neurons and prediction error neurons differs fundamentally between the two variants. In the hierarchical variant, internal representation neurons excite positive prediction error neurons and inhibit negative prediction error neurons, while at the same time being inhibited by positive prediction error neurons and excited by negative prediction error neurons. In the non-hierarchical variant, this pattern of connectivity is reversed.

This work is very exciting, timely, and carefully executed. The authors functionally, and later molecularly, identify layer 2/3 prediction error neurons in V1 and probe their interactions with genetically defined neuron types in cortical layers 5 and 6 using optogenetics. They demonstrate that the functional influence of putative prediction error neurons in layer 2/3 onto layer 5 is incompatible with the hierarchical variant, whereas the influence of layer 5 onto putative prediction error neurons in layer 2/3 is incompatible with the non-hierarchical variant. They then test an alternative hypothesis, in which layer 2/3 responses resemble prediction errors with respect to perturbations of artificial layer 5 activity patterns. To investigate this, they designed an experiment in which optogenetic activation of L5 IT neurons was closed-loop coupled to the mouse's locomotion speed in the absence of visual feedback, allowing them to probe the causal influence of L5 activity on layer 2/3 responses.

Finally, the authors hypothesize that their data are more consistent with a joint embedding predictive architecture (JEPA) and outline experimentally testable predictions arising from this framework.

We thank the reviewer for their comments. We address them below.

While the work is overall convincing and significantly advances our understanding of the circuit-level implementation of predictive processing, there are a few weaknesses that should be addressed or discussed:

(1) The authors define putative positive prediction error neurons as the 15% of neurons most responsive to grating onset and putative negative prediction error neurons as the 15% most responsive to visuomotor mismatch. While this selection would be expected to overlap with negative and positive prediction error neurons, the criterion is not sufficiently stringent (independent of the exact percentage chosen). In particular, classification of a neuron as a prediction error neuron should ideally be accompanied by evidence that it does not exhibit a significant increase in activity when the prediction matches the sensory input or teaching signal.

We understand the reviewer’s intuition. We can indeed use other stimuli to identify prediction error neurons, like the relative suppression of responses in closed-loop running onset vs open-loop running onsets. This was the reason behind including Figure S1, to show that our selection results in expected pattern of running onset responses. We don’t typically use the running onset responses as they have an additional confound we have not fully understood. This is that running onset always tends to result in an increase of calcium activity in all neurons. This could have a variety of reasons: A) Contamination of hemodynamic occlusion signals (blood vessels tend to constrict at running onset, making it appear like an increase in calcium activity) – see Yogesh et al., 2025. B) The virtual coupling in our VR is not good enough to provide a true “closed loop” experience. Humans typically notice lags of larger than 30ms – in our VR it is approximately 100 ms. C). Running onset in head-fixed animals is not accompanied by a vestibular input. D) Predictive processing is wrong. We tend to think it is a combination of the three.

We can also use combinations of the two criteria to select neurons - if the reviewer has a specific selection criteria in mind (top XXX% MM responsive AND top XXX% closed-loop suppressed, etc.) we are happy to repeat the analysis for that specific set of criteria, but the fundamental problem that we are using a functional response to select these neurons does not go away. We know that our functional selection criteria mean we select a population of neurons that is enriched for prediction error neurons. If the enrichment is too weak, we would expect to find no effects in terms of functional influence. It is hard to explain, however, how a weak enrichment could result in a strong effect on functional influence. Our arguments in more lengthy form, for why the visuomotor mismatch is a good stimulus to identify negative prediction error neurons can be found here: Attinger et al., 2017; Jordan and Keller, 2020; Leinweber et al., 2017; O’Toole et al., 2023; Vasilevskaya et al., 2022; Zmarz and Keller, 2016.

(2) The authors "speculate that the prediction error responses in layer 2/3 may not be computed with respect to sensory input, but with respect to layer 5 activity as a teaching signal." However, it is unclear how this perspective differs from earlier statements in the manuscript. In the Introduction, the authors note that "these signals, typically referred to as sensory signals, we will refer to as teaching signals," and later describe the hierarchical variant as one "in which internal representation neurons act as a source of the teaching signal." Given this framing, it is difficult to identify what is conceptually novel in the updated view. Is the key distinction that layer 2/3 neurons are now proposed to generate predictions in an internal representation space rather than in sensory input space, as briefly suggested in the Discussion? Or are the authors introducing a distinction between an external (sensory) and an internal (cortical) teaching signal? If so, this distinction should be made explicit. Clarifying this point would considerably strengthen the manuscript.

There might be a misunderstanding regarding our usage of the term teaching signal. In hierarchical predictive processing the teaching signal is typically referred to as a sensory signal, as e.g. in: “prediction error neurons compare predictions to sensory input”. In non-hierarchical predictive processing, or far away from the sensory input (think prefrontal cortex), or for cross-modal interactions “sensory” input is misleading. Also, in non-predictive-processing type models (like JEPA), sensory input has a different functional role. Thus, we operationally define teaching signal as the signal that is compared against the prediction by prediction error neurons.

The two primary options we are comparing are:

(1) Is the teaching signal to the L2/3 comparator a bottom-up input to V1 (as one would expect in predictive processing). 

(2) Is the teaching signal to the L2/3 comparator L5 input (as one would expect in JEPA).

Our data argue in favor of option 2. We have rephrased parts of the manuscript to try to make this clearer. 

(3) The authors propose that "L2/3 neurons predict L5 activity, hence making predictions in the internal representation space rather than the input space," and further suggest that, since both deep and superficial cortical layers receive thalamic input, the cortex may function like a JEPA. This idea appears closely related to the model introduced by Nejad et al. (2025), which effectively implements a JEPA-like architecture: L5 activity serves as a target against which L2/3 predictions are compared in a selfsupervised manner, with both L5 and L2/3 (via L4) receiving thalamic input. It would be helpful for the authors to clarify how their framework differs from that model, and to specify the key conceptual or mechanistic distinctions between the present proposal and the approach described by Nejad et al.

The two proposals indeed share similarities in assuming that bottom-up input for both L2/3 and L5 arrives from thalamus, and that representations formed in L2/3 are used for predicting the activity of L5. However, there are a few important differences between the JEPA implementation proposal formulated here and the Nejad et al. model.

(1) There is no proposed mapping of computations in the Nejad et al. model onto different JEPA networks. We assume that the suggested mapping would be L4 and L5 as encoder networks, and L2/3 as a predictor network? In that case, it is different to our proposal, in which L2/3 is part of the encoder network.

(2) Our proposal contains explicit prediction error neuron cell types within L2/3, while prediction errors in Nejad et al. are encoded in the gradients, and the layer origin of these signals is hypothesized to be L5 (‘the learning-driving error signal originates in L5’). Hence, also the role of L5-L2/3 connection is distinct between the two proposals. In Nejad et al. this connection serves the role of error propagation and update for predictions in L2/3, while in our proposal this connection contains teaching signal (target representations) that are compared to predictions within L2/3. Similarly, the functional role of L2/3-L5 connection is also different, since in Nejad et al, it is supposed to carry predictions of L5 activity, whereas in our proposal we expect it to drive plasticity in L5 encoder.

(3) The difference outlined above also makes it evident that the two proposals should differ in how deep and superficial layers are expected to influence the activity of one another. Indeed, the proposal in Nejad et al. is based on the cortical column idea, and according to eq. 2 and 3 in the Methods, activity in L5 is a function of activity in L2/3, while activity in L2/3 is not a function of activity in L5. Our proposal is based on idea of layers forming parallel networks, where horizontal communication is the dominant mode of cortico-cortical interactions, and activity in deep layers serve as a teaching signal for L2/3. In our case, we expect the opposite - that activity in L2/3 depends on activity of L5, while activity of L5 is not immediately dependent on activity of L2/3 (only via plasticity route). This led us to propose one of direct tests for our framework – silencing L2/3 in a familiar setting should result in no immediate changes to L5 activity and behavior of the animal.

(4) The proposal in Nejad et al. relies on input reconstruction or variance maximization within the L5 autoencoder network to avoid collapse. Instead, our proposal has no reconstruction objective.

(5) Lastly, there is a time delay between inputs to L5 and L2/3 that is proposed in Nejad et al., while this is not something inherent to our proposal.

We expect that the most useful future models should move beyond JEPA, with the emphasis on nonhierarchical models capable of operating on arbitrary graphs. We think cortex functions according to principles of a JEPA (predictions in latent space), and that the role of cell types and the exact computational organization remain to be constrained.

REFERENCES

Attinger, A., Wang, B., Keller, G.B., 2017. Visuomotor Coupling Shapes the Functional Development of Mouse Visual Cortex. Cell 169, 1291-1302.e14. https://doi.org/10.1016/j.cell.2017.05.023

Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P.H., Buchatskaya, E., Doersch, C., Pires, B.A., Guo, Z.D., Azar, M.G., Piot, B., Kavukcuoglu, K., Munos, R., Valko, M., 2020. Bootstrap your own latent: A new approach to self-supervised Learning. https://doi.org/10.48550/arXiv.2006.07733

Jordan, R., Keller, G.B., 2020. Opposing Influence of Top-down and Bottom-up Input on Excitatory Layer 2/3 Neurons in Mouse Primary Visual Cortex. Neuron 108, 1194-1206.e5. https://doi.org/10.1016/j.neuron.2020.09.024

Leinweber, M., Ward, D.R., Sobczak, J.M., Attinger, A., Keller, G.B., 2017. A Sensorimotor Circuit in Mouse Cortex for Visual Flow Predictions. Neuron 95, 1420-1432.e5. https://doi.org/10.1016/j.neuron.2017.08.036

Mohammadi, A.G., Halvagal, M.S., Zenke, F., 2025. Understanding cortical computation through the lens of joint-embedding predictive architectures. https://doi.org/10.1101/2025.11.25.690220

O’Toole, S.M., Oyibo, H.K., Keller, G.B., 2023. Molecularly targetable cell types in mouse visual cortex have distinguishable prediction error responses. Neuron 111, 2918-2928.e8. https://doi.org/10.1016/j.neuron.2023.08.015

Saravanan, V., Berman, G.J., Sober, S.J., 2020. Application of the hierarchical bootstrap to multi-level data in neuroscience. Neurons Behav. Data Anal. Theory 3, https://nbdt.scholasticahq.com/article/13927-application-of-the-hierarchical-bootstrap-tomulti-level-data-in-neuroscience.

Vasilevskaya, A., Widmer, F.C., Keller, G.B., Jordan, R., 2022. Locomotion-induced gain of visual responses cannot explain visuomotor mismatch responses in layer 2/3 of primary visual cortex. https://doi.org/10.1101/2022.02.11.479795

Yogesh, B., Heindorf, M., Jordan, R., Keller, G.B., 2025. Quantification of the effect of hemodynamic occlusion in two-photon imaging of mouse cortex. eLife 14, RP104914. https://doi.org/10.7554/eLife.104914

Zmarz, P., Keller, G.B., 2016. Mismatch Receptive Fields in Mouse Visual Cortex. Neuron 92, 766–772. https://doi.org/10.1016/j.neuron.2016.09.057

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation