Peer review process
Revised: This Reviewed Preprint has been revised by the authors in response to the previous round of peer review; the eLife assessment and the public reviews have been updated where necessary by the editors and peer reviewers.
Read more about eLife’s peer review process.Editors
- Reviewing EditorRui Ponte CostaUniversity of Oxford, Oxford, United Kingdom
- Senior EditorTimothy BehrensUniversity of Oxford, Oxford, United Kingdom
Reviewer #1 (Public review):
Vasilevskaya and Keller test different models of cortical function through the lens of predictive processing, a powerful framework for the brain to learn and predict the statistics of the world via generative internal models. The authors use a clever combination of behavioral perturbations in closed-loop and open-loop visuomotor virtual reality assays, a paradigm the Keller lab pioneered and used effectively in the past decade, in conjunction with two photon imaging of neuronal calcium responses and targeted optogenetic perturbations of activity. They specifically put to test proposed hierarchical vs. non-hierarchical circuit implementations of predictive processing by analyzing the logic of inter-lamina interactions (superficial vs. deep; L2/3 vs. L5/6).
The authors conclude that both versions of predictive processing architectures they analyze are likely invalid and instead formulate an alternative novel model of cortical function based on a recently developed machine learning algorithm for self-supervised learning (joint embeddings of predictive architectures, JEPA) and its further refinements. JEPA borrows elements from predictive processing engaging two encoder networks and training the output of one network to predict the output of the other. In their new model of cortical computations, prediction errors neurons in L2/3 compare the deep layers (L5/6) activity, which is taken as a teaching signal, to a local, L2/3 prediction of this latent representation.
Specifically, the authors build on their previous work and reports from other groups that different sets of L2/3 neurons compute positive prediction errors (fire when sensory stimuli appear unexpectedly with respect to the movements of the animal; e.g., grating onsets in the absence of locomotion) and respectively negative prediction errors (fire when sensory stimuli are absent, while the brain expected them to be present; e.g. mice locomote but visual flow is suddenly halted - visuomotor mismatches). These L2/3 positive and negative prediction error neurons exchange messages with neurons in the deeper cortical layers that, the authors propose, build an internal representation (R) of the sensory stimuli given the animals' movements.
In the hierarchical model, internal representation neurons (R) are supposed to act as a teaching signal for both types of prediction error neurons; the output of the positive prediction error neurons is assumed to suppress activity of R such that the error between the teaching signal and the prediction is minimized; similarly, in the non-hierarchical version, R serves as a prediction for the prediction error neurons, and in turn it receives excitatory drive from the positive prediction error neurons and negative input from the negative prediction error neurons.
The authors find that the functional impact of L5 neurons to L2/3 neurons is not compatible with the non-hierarchical architecture they and other groups proposed, but rather in accordance with the hierarchical model. At the same time, the functional impact of L2/3 neurons (positive vs. negative prediction error neurons) on L5 neurons (internal representation) appears not compatible with the hierarchical model, but rather in accordance with the non-hierarchical implementation.
They further hypothesize that L2/3 prediction error neurons don't use sensory input, but rather the L5 activity as a teaching signal, and test it using perturbations (halts) of optogenetic stimulation of L5 neurons coupled with locomotion (Fig.7).
All in all, the question is topical, and the new model addresses a decades-long quest to develop a unifying model of cortical function. The findings reported here transform our understanding of cortical computations, opening new exciting avenues for future investigation. The experimental design and execution are rigorous; the arguments are clearly laid out (in spite of ample potential for confusion given the numerous loops and sign flips). These include a discussion of why the non-hierarchical model proposed by the same group does not hold, as well as potential caveats in interpreting the results and novel testable proposed experiments emerging from the JEPA-like model.
Comments on revised version
I commend the authors for nicely provided answers and addressing my concerns and clarifying their points on the relationship of their findings to JEPA and current state of understanding.
In particular for Q5 -I meant if there is also a positive correlation between optomotor mismatch response and visuomotor mismatch response when looking only at neurons that the authors identify as PE-?
The authors provided the answer on point.
Overall, I think this is a fundamental study and the strength of evidence for the claims is exceptional as an exemplary use of existing approaches.
Reviewer #2 (Public review):
This manuscript reveals functional connectivity of two different classed of cortical neurons that respond in opposite ways to mismatches between sensory and top-down inputs. These data are very valuable because different theories of information processing in the cortex make different predictions on the patterns of connectivity of these neurons. Therefore, these data strongly constrain possible theories of cortical processing.
Comments on revised version.
I thank the Authors for answering my questions and updating the manuscript.
Congratulations on this important work!
Reviewer #3 (Public review):
Vasilevskaya and Keller set out to experimentally distinguish between two variants of predictive processing: a hierarchical and a non-hierarchical variant. The hierarchical variant assumes a hierarchical organization in which internal representation neurons (believed to be a subset of layer 5 excitatory neurons) serve as a source of a teaching signal for local prediction error neurons as well as for the next higher level of the hierarchy, while simultaneously providing prediction signals to the preceding lower level. In contrast, the non-hierarchical variant posits that these layer 5 internal representation neurons provide local predictions to layer 2/3 prediction error neurons.
The interaction between internal representation neurons and prediction error neurons differs fundamentally between the two variants. In the hierarchical variant, internal representation neurons excite positive prediction error neurons and inhibit negative prediction error neurons, while at the same time being inhibited by positive prediction error neurons and excited by negative prediction error neurons. In the non-hierarchical variant, this pattern of connectivity is reversed.
This work is very exciting, timely, and carefully executed. The authors functionally, and later molecularly, identify layer 2/3 prediction error neurons in V1 and probe their interactions with genetically defined neuron types in cortical layers 5 and 6 using optogenetics. They demonstrate that the functional influence of putative prediction error neurons in layer 2/3 onto layer 5 is incompatible with the hierarchical variant, whereas the influence of layer 5 onto putative prediction error neurons in layer 2/3 is incompatible with the non-hierarchical variant. They then test an alternative hypothesis, in which layer 2/3 responses resemble prediction errors with respect to perturbations of artificial layer 5 activity patterns. To investigate this, they designed an experiment in which optogenetic activation of L5 IT neurons was closed-loop coupled to the mouse's locomotion speed in the absence of visual feedback, allowing them to probe the causal influence of L5 activity on layer 2/3 responses.
Finally, the authors hypothesize that their data are more consistent with a joint embedding predictive architecture (JEPA) and outline experimentally testable predictions arising from this framework.
While the work is overall convincing and provides important insights into the circuit-level implementation of predictive processing, I think the connection to JEPA networks would benefit from a more in-depth discussion of its relationship to recently proposed and implemented models. Below, I address the specific points raised by the authors (flanked by ' ... ' to make the author's statements stand out), in particular in relation to the model proposed by Nejad et al. (2025):
- 'The two proposals indeed share similarities in assuming that bottom-up input for both L2/3 and L5 arrives from thalamus, and that representations formed in L2/3 are used for predicting the activity of L5. However, there are a few important differences between the JEPA implementation proposal formulated here and the Nejad et al. model.
(1) There is no proposed mapping of computations in the Nejad et al. model onto different JEPA networks. We assume that the suggested mapping would be L4 and L5 as encoder networks, and L2/3 as a predictor network? In that case, it is different to our proposal, in which L2/3 is part of the encoder network.'
I think there is some confusion here. In Nejad et al., both L2/3 and L5 function as encoder networks. Each receives sensory input (with L2/3 receiving this input delayed via L4) and computes a latent representation of that input, denoted z_{L2/3} and z_{L5}, respectively. The prediction is obtained by comparing the output of L2/3 with L5 latent representations (via W_{L2/3->L5} * z_{L2/3}). In other words, L5 is the target in the learning objective.
Although, they did not explicitly state the mapping with JEPA, their model has the same fundamental property - predictive learning happens in the latent space. Also, this appears very similar to the roles assigned to L2/3 and L5 in your Figure 9B. The fact that you made this more explicit and the new data included, is in my view, a very interesting contribution. However, from an architectural perspective, it appears that your proposal and the model of Nejad et al. are conceptually very similar, and I do not see a substantial difference between the two. This should be made more clear in the Discussion.
- '2. Our proposal contains explicit prediction error neuron cell types within L2/3, while prediction errors in Nejad et al. are encoded in the gradients, and the layer origin of these signals is hypothesized to be L5 ('the learning-driving error signal originates in L5'). Hence, also the role of L5-L2/3 connection is distinct between the two proposals. In Nejad et al. this connection serves the role of error propagation and update for predictions in L2/3, while in our proposal this connection contains teaching signal (target representations) that are compared to predictions within L2/3. Similarly, the functional role of L2/3-L5 connection is also different, since in Nejad et al, it is supposed to carry predictions of L5 activity, whereas in our proposal we expect it to drive plasticity in L5 encoder.'
Indeed, in Nejad et al., the layer-dependent mismatch responses are modeled as gradients with respect to neuronal activity, and the model does not explicitly include prediction error neurons. However, this appears to be a modeling choice rather than a fundamental aspect of the proposal, and it does not preclude an implementation with explicit prediction error neurons. In fact, the authors explicitly acknowledge this possibility in the Discussion:
"The second approach would be to recast our model within a predictive coding framework... Predictive coding jointly optimizes both model parameters and neuronal activities, which could naturally lead to prediction errors observable in the activity of both L2/3 and L5 neurons. Note that these two views are not mutually exclusive."
While Nejad et al. hypothesize that the learning-driving error signals (that are distinct from their mismatch responses) originate in L5, the abstract loss function itself does not uniquely specify where the underlying comparison between the predicted representation (W_{L2/3 -> L5} z_{L2/3}) and the target representation (z_{L5}) must be implemented. The proposed biological implementation places this computation in L5, but from my understanding, the computational objective itself does not require this specific localization.
That said, I agree that your proposed model introduces a genuine difference. The functional roles assigned to the vertical projections are effectively reversed: in Nejad et al., the L2/3->L5 projection carries the prediction, whereas the L5->L2/3 projection conveys the error/gradient. In your architecture, by contrast, the L5->L2/3 projection carries the teaching signal (target). This is, in my view, a real and testable interpretational divergence that is worth stating clearly.
Therefore, I think the novelty lies less in the computational architecture itself and more in committing to a particular biological implementation-one that adds cell-type-specific detail to an implementation that Nejad et al. hypothesised as being compatible with their framework.
- '3. The difference outlined above also makes it evident that the two proposals should differ in how deep and superficial layers are expected to influence the activity of one another. Indeed, the proposal in Nejad et al. is based on the cortical column idea, and according to eq. 2 and 3 in the Methods, activity in L5 is a function of activity in L2/3, while activity in L2/3 is not a function of activity in L5. Our proposal is based on idea of layers forming parallel networks, where horizontal communication is the dominant mode of cortico-cortical interactions, and activity in deep layers serve as a teaching signal for L2/3. In our case, we expect the opposite - that activity in L2/3 depends on activity of L5, while activity of L5 is not immediately dependent on activity of L2/3 (only via plasticity route). This led us to propose one of direct tests for our framework - silencing L2/3 in a familiar setting should result in no immediate changes to L5 activity and behavior of the animal.'
My reading of Nejad et al. is consistent with your interpretation of the equations. Specifically, Eq. 3 makes L5 activity depend on L2/3 activity (albeit weakly, with a = 0.3), whereas Eq. 2 contains no L5 term, so L2/3 activity does not depend directly on L5 activity. In that model, the L5->L2/3 pathway carries the learning gradient rather than contributing to the activity dynamics. By contrast, in your proposed model, L2/3 activity depends on L5 activity, whereas L5 activity is not immediately dependent on L2/3 activity (except indirectly through learning/plasticity). So, you state that "silencing L2/3 in a familiar setting should result in no immediate changes in L5 activity.
[...].
However, I am unsure how to reconcile this prediction with the results shown in Fig. 6. If I understand the figure correctly, optogenetic activation of Rrad-positive (positive prediction error) L2/3 neurons produces a small increase in L5 activity, whereas activation of Adamts2-positive (negative prediction error) L2/3 neurons produces a decrease in L5 activity. Although these experiments involve activation rather than silencing, they nevertheless suggest that perturbing L2/3 activity can have an immediate effect on L5 activity. Could you clarify how this is consistent with the proposed model? In other words, what aspect of the proposed circuitry makes activation effective while silencing is predicted to have no immediate consequence?
For comparison, Nejad et al. performed a related perturbation analysis in Fig. S15 by scaling the output of L2/3 neurons exhibiting positive mismatch signals (defined through the activity gradients), which increased L5 activity, whereas scaling neurons with negative mismatch signals produced the opposite effect. I am not entirely sure how directly these simulations map onto the experiments shown in your Fig. 6, since the Nejad simulations were performed during mismatch conditions, if I have understood them correctly.
- '4. The proposal in Nejad et al. relies on input reconstruction or variance maximization within the L5 autoencoder network to avoid collapse. Instead, our proposal has no reconstruction objective.'
Nejad et al. only require two encoders (like in JEPA), how these two are learnt can be done in several ways. While Nejad et al. focus on using a reconstruction loss to learn the L5 target, they also show that it works equally well with non-reconstruction objectives. Therefore, I do not think the presence or absence of a reconstruction objective constitutes a fundamental distinction between the two proposals.
As your current work presents a conceptual architecture rather than a fully implemented learning algorithm (in a model), the mechanism that would prevent representational collapse has not yet been defined. From my understanding, every joint-embedding approach must address this issue, whether through reconstruction, variance/covariance regularization, stop-gradient or EMA mechanisms, or other approaches. Thus, the absence of a reconstruction objective (or another anti-collapse mechanism) is not, in itself, a distinguishing feature of the proposed architecture, but rather an as-yet unspecified design choice within the learning objective.
- '5. Lastly, there is time-delay between inputs to L5 and L2/3 that is proposed in Nejad et al., while this is not something inherent to our proposal.'
I agree that the temporal delay introduced by L4 is a key component of the Nejad et al. model and is currently absent from your proposal. However, I would expect temporal delays to emerge naturally in your framework as well, given the multisynaptic and highly parallel organization of cortical circuits. More generally, implementing predictive learning over time (as in JEPA) requires comparing representations at times t and t+1, which seems to require some form of temporal delay. How else would you suggest this is done?
In general, I think the manuscript would benefit from a clearer discussion of its relationship to the model proposed by Nejad et al. (perhaps following the discussion above), including both the shared conceptual claims and the aspects that genuinely differ between the two frameworks. At present, some of the claims are presented as novel, although at least some of these core ideas have already been proposed in Nejad et al.
For example, the authors state: "Thus, we propose that layer 2/3 functions to predict layer 5 activity, not sensory input per se, hence making predictions in the internal representation space, not input space." This appears to be exactly what Nejad et al. proposed as discussed above - in their model L2/3 predicts L5 activity (purely in the latent space), as they state in the abstract.
That said, there are some interesting differences, and I think the community would greatly benefit from making these clear, including the roles assigned to interlaminar connections and the interpretation of the signals carried by these pathways. These differences are interesting and potentially testable, and I think the manuscript would be strengthened by explicitly distinguishing which aspects are in line with the ideas already present in Nejad et al. and which aspects represent new contributions.

