Abstract
Animals must learn and make predictions in a world that is not only uncertain, but shaped by unmodelled contingencies: ‘unknown unknowns’. Despite this, current theories of how this is achieved are based on optimal (Bayesian) strategies in simple settings, where it is possible to fully represent a model of the world and infer its state based on sensory data. This could be impractical in situations where an animal has limited prior experience and cognitive capacity, particularly in short-lived animals with small brains. Using the neural circuitry of Drosophila as inspiration, we derive a learning framework, which we call ‘Hetlearn’, for coping in such situations. Hetlearn sacrifices asymptotic optimality on idealised inference tasks, making it less fragile to abrupt changes in environmental statistics and providing more accurate state information than standard Bayesian procedures when data are limited. Hetlearn explains key physiological and anatomical features of the Drosophila Mushroom body (MB), and predicts behavioural features of learning that are conserved across species, such as second-order conditioning and reversal learning.
Introduction
Animals learn from experience to predict future events. Existing theory frames this as state inference: updating a probabilistic map of relevant aspects of the environment, called states, and predicting how these states evolve as a function of time, context, actions and external events [35, 7, 15]. Simple, familiar situations such as repeated laboratory tasks can provide a mathematically tractable environment where an agent can represent all possible states, actions, and consequences, allowing for optimal probabilistic reasoning [14, 19, 31].
Despite their utility in establishing normative principles for reasoning and decision making, idealised tasks whose causal structure can be fully modelled and learned occupy what is known as a ‘Small World’ [43]. This contrasts with the real world, or ‘Large World’: a messy reality of unforeseen contingencies (“unknown unknowns”), and open-ended state spaces where optimal inference is intractable.
Animal brains operate in the Large World setting [29]: they inhabit the real world, and even simpler, restricted situations (such as laboratory tasks) may present animals with too little data to perform optimal inference. Under these conditions, an optimal Small World learning procedure may be less relevant for understanding biological learning, particularly if the optimal procedure is fragile to violations in modelling assumptions [22].
Using well-established features of Drosophila learning circuitry and behaviour, we devised a novel learning algorithm ‘Hetlearn’ (for heterologous learners) designed for rapid, approximate inferences that are robust to unmodelled causes. Hetlearn incorporates a key feature of the mushroom body: parallel, associative learning modules that operate independently on fast and slow timescales. This feature addresses a well-known issue with Bayesian inference by effectively mitigating ambiguities between stochasticity and volatility [6, 41]. Remarkably, in a low to moderate data regime of hundreds of samples, Hetlearn outperforms asymptotically optimal, Bayesian algorithms despite being suboptimal in the long run. This regime is likely relevant to animals making decisions with little data in volatile environments.
Hetlearn recreates Drosophila behaviour across several chosen tasks: second-order, extinction, and reversal conditioning. It explains both behaviour and the heterologous dynamics of dopaminergic signalling in the mushroom body. Furthermore, it provides new, testable hypotheses on the roles of poorly-understood inter-compartmental mushroom body interactions during these tasks. More abstractly, it provides a normative principle for learning in truly novel environments where there is uncertainty in the values of represented states (as in probabilistic inference), but also uncertainty in which states to represent in a cognitive map.
Results
State inference in a Large World
We first motivate Hetlearn abstractly, using the standard probabilistic inference framework as a comparison. In doing so, we will expose two well-known issues with probabilistic inference which can be detrimental to biological learning in naturalistic settings, and that Hetlearn addresses.
A learner undertaking probabilistic state inference at time t holds and updates a probability distribution P [xt, θ | y1:t] on dynamic states xt and static parameters θ that collectively represent the environment, given an observation history y1:t. The optimal update for this distribution uses Bayes’ rule, which requires a pre-specified generative model (or cognitive map [5]) of the environment comprising probability distributions for state evolution: Px(xt|xt−1, θ), and observation: Py(yt |xt, θ). This generative model is variously assumed to be learned from a mixture of prior experience and developmental processes that embed innate assumptions about the world.
Generative models provide probabilistic outcomes. This is assumed to account for the effects of countless unmodelled ‘large-world’ factors zt, which collectively make the future unpredictable from the perspective of the learner. Specifically, we could specify ground-truth models of state update and observation as xt = f (xt−1, zt), and yt = g(xt, zt). Probabilistic inference assumes that


where the distributions Px and Py are known ahead of time as part of the generative model, and the parameters θ are learnt during the task.
This theoretical framework raises two profound questions that form the theoretical motivation for Hetlearn. First, even if the generative model specified in Eqs. 1 accurately reflects ground-truth, its use in state estimation may be inherently fragile to approximation errors, particularly in situations where data to calibrate states is limited. This is inevitable for algorithms implemented in biological circuits, and could leave an animal with a predictive model that is both algorithmically complicated and poorly performing. There may be algorithmic tricks by which an inherently suboptimal, simpler inference can yield predictions that are more accurate in the short-term. Hetlearn implements these tricks.
Second, systematic ‘large-world’ changes in the environment can fundamentally invalidate the animal’s internal generative model. If the true data-generating process shifts, the animal’s cognitive map and prior assumptions are decoupled from reality. Consequently, the theoretical guarantees underpinning the reliability of Bayesian inference (such as the evidence lower bound [52]) cease to apply, as they remain tied to an outdated representation of the world. Again, there is no question that animals can fail to adapt their behaviour to new ecological niches. However, a simple algorithmic means of minimising the fragility of an animal’s predictive model to Large World effects would clearly offer a reproductive advantage. Hetlearn offers a means of achieving this.
We will next construct and analyse a simple example that illustrates how Hetlearn addresses the first issue by outperforming a Bayesian inference procedure in the low-data regime. Later sections will tackle abrupt changes to generative models.
Basic Hetlearn algorithm for learning via prediction errors
Many models of learning posit that animals update state estimates from prediction errors (PEs) [45]. States are taken to represent physical or abstract variables, from properties of visual stimuli to reward. Notable examples of learning algorithms that are proposed to model biological state estimation are Kalman filtering [11, 25, 20] and temporal difference reinforcement learning [44, 46, 39, 9]. A common dilemma faced by any learning algorithm is that there is a tradeoff in deciding to what extent a PE is attributable to stochastic fluctuations that should be ignored, versus systematic changes (often known as volatility) that should be adapted to.
We first propose a simple algorithm for updating to PEs and claim it has advantages over classical probabilistic inference in biologically relevant scenarios: low data, limited algorithmic resources. Our setting is a learner attempting to infer a single state from ambiguous observations over repeated trials. For concreteness, we consider the state as a reward 


This models a general situation where large-world causes of systematic change (volatility, Cp(zt)) and unsystematic, noisy fluctuations (stochasticity, Cm(zt)) both affect rt. The learner holds a state estimate 



An excessive learning rate (αt ≈1) erases valuable accumulated experience by overreacting to stochastic fluctuations. An insufficient learning rate (αt ≈ 0) cannot keep up with systematic volatility, or reduce a large estimation error. The truly optimal learning rate (zero estimation error on trial t + 1, see SI S1) is

This learning rate in inaccessible to any learning algorithm: Cm, Cp, and et will fluctuate on each timestep, and their values cannot be disambiguated from the scalar PE signal which sums their individual effects.
To deal with this, and in agreement with neural circuit properties that we discuss in detail later, we propose the learner maintains multiple parallel valence estimates (sublearners), each with intrinsic, fixed guesses of α* within extremal values of 0 and 1. To form an estimate, the learner weights the sublearners according to their recent predictive performance.
A Hetlearn strategy for learning from prediction errors
We formulate a simple Hetlearn strategy for PE updates: Take n parallel sublearners, each with their own estimates: 


where α(i) is fixed and heterogeneous across sublearners. The overall prediction is a weighted vote: 

To determine the weights, the agent uses an estimate of recent error for each sublearner. One simple estimate is:

a moving average of PE magnitude, with recency bias 0 ≤ τ ≤ 1. Weights are then 

Hetlearn differs importantly from probabilistic inference, where large-world effects must be modelled as probability distributions to infer learning rate (Eq. (1)). We now demonstrate this on a simple didactic example, where we compare Hetlearn to the Kalman Filter (KF): a popular model for how animals learn from prediction errors [11, 25, 20], with biologically plausible neural circuit implementations [12, 54]. The KF assumes fixed Gaussian distributions:



With these assumptions, baseline reward is a Gaussian distribution whose mean updates with a dynamic, Bayes-optimal learning rate. The learning rate calculation, however, requires knowledge of 
In Figure 1, we simulate a large-world where 

The Kalman Filter requires seeding with values for reward stochasticity/volatility.
It is seeded with the initial values, for which it is statistically optimal. When the values change due to unmodelled, unrepresentable largeworld effects, it performs poorly. A Bayes-optimal learner can be fragile to large-world effects that break its foundational assumptions. Hetlearn is suboptimal for any particular, fixed reward stochasticity/volatility, but robust to changes.
More generally, any KF requires direct access to the correct stochasticity/volatility values, which stretches the plausibility of a biological implementation. Consequently, studies that propose the KF as a model of biological learning often hardcode guesses into a model [11, 12, 20]. Figure 1 suggests that a nonoptimal, non-probabilistic strategy could be preferable where such guesses, which form part of the generative model structure, are unreliable or subject to change.
The standard KF deployed in Figure 1 could be improved by incorporating stochasticity and volatility as parameters to be concurrently estimated. More generally, generative models can be improved by adding flexibility through more parameters. Previous work has thus proposed algorithms that learn stochasticity and volatility online [6, 41]. This clearly results in greater algorithmic complexity. More importantly, we claim that it necessarily results in worse short-term performance than Hetlearn.
To demonstrate, we first implemented an (augmented) Kalman learner that additionally estimates 



The performance of the Kalman learner is shown in Figure 2. For over 100 trials, Hetlearn markedly outperforms in mean-square prediction error, with cumulative error lower for Hetlearn until past the 500th trial. Hetlearn superiority is also robust to choices of the (fixed) recency bias parameter, τ (see SI S5). In the long run, as expected, the Kalman learner converges on the true stochasticity/volatility values (Figure 2A,C), and asymptotically approaches Bayes-optimal performance.

A: Convergence of ALS algorithm’s stochasticity and volatility estimates over time, compared to their (dotted) true values. One run plotted for illustration. B: Faster convergence of the weightings of Hetlearn sublearners over time. One run plotted for illustration C: 100 samples of the final ALS estimates of stochasticity and volatility, after 1000 timesteps. Estimation errors are anticorrelated. D: Mean MSE of Kalman learner vs Hetlearn over 100 repeats. Bands represent standard error. E: Mean trajectory (100 repeats) of the difference in cumulative MSE for the Kalman learner and Hetlearn. Green (positive) values indicate that the Kalman learner had higher cumulative MSE than Hetlearn, while purple (negative) values indicate the opposite. F: Learning rate of the two algorithms as compared to the optimal formula of Eq. (3) over a single trial.
Hetlearn possesses a key property that provides robustness to changing conditions and limited data. The (expected) estimated error of a sublearner with fixed learning rate α converges to its asymptotic value 

during any period of constant stochasticity and volatility (see SI S6). This exponential convergence means that traces of previous environmental conditions are quickly washed out. Consequently sublearner weights (which are calculated from these estimated errors) converge exponentially fast to reasonable values, as seen in the first few timesteps of Figure 2B. Remaining weighting inaccuracy is caused by fluctuations around the expectations in (5), as seen in later timesteps of Figure 2B.
The exponential convergence of Hetlearn weightings stands in contrast to the slow convergence of stochasticity/volatility estimates (Figure 2A and C), which result in poor initial performance of the Kalman learner. This slow convergence is due to the fact that stochasticity and volatility explain each other away [42]: a contribution from either noise term has the same effect on PE magnitude on any given trial. Only long-term patterns of PE autocorrelations can disambiguate the two, which is precisely how the ALS algorithm estimates them. More generally, good estimates of second-order statistics are inherently data-hungry to accurately estimate. Overall imprecise approximates of the stochasticity/volatility distributions (such as the point-estimate provided by ALS), are far worse in the low-data regime. Biological implementations of such Bayesian procedures, which are inevitably approximate, would likely suffer more.
Competing sublearners should stay simple
We now strengthen the claims of Figure 1: that Hetlearn gains robustness to unmodelled environmental changes by assuming less of the environment. To do so, we compare Hetlearn against the algorithm proposed in [42], which we call the ‘Piray learner’. This belongs to a class of probabilistic algorithms known as particle filters, but was specifically formulated to address state inference under conditions of changing stochasticity and volatility. Like Hetlearn, they use parallel sublearners with different beliefs, weighted on predictive performance. Unlike Hetlearn, each sublearner estimates stochasticity and volatility: necessary for statistically optimal reward prediction. Specifically, each sublearner (/particle) is a Kalman Filter, with stochasticity/volatility estimates that stochastically diffuse to track changing environments.
We use a test environment similar to the previous section, in that the state follows Eqs. (2) and stochasticity/volatility are Gaussian ((4)). However, the variances of the (Gaussian) Cm(zt) and Cp(zt) change over time, following the schedules in Figure 3 (top row). These model a variety of large-world changes that can’t be predicted by the learner.

A: Volatility and stochasticity for four different environments not encompassed by a single generative model. B: Hetlearn (top) and Piray et al (bottom) tracking the the optimal learning rate across all four environments. C: Cumulative error over 10 trials for Hetlearn and Piray et al. D: Cumulative mean squared error (MSE) over 10 trials for Hetlearn and Piray et al. For B, C, and D, and the average of 10 runs were plotted for Hetlearn and Piray et al.
Hetlearn outperformed the Piray learner both in overall predictive performance (Figure 3E), and in tracking the correct learning rate over time (Figure 3D). Furthermore, Hetlearn used fewer sublearners (3, versus 10 for Piray).
The previously mentioned issues of stochasticity/volatility taking lots of data to estimate accurately compound in a situation where they can drift. Note that estimating the state as a probability distribution requires this estimation. Hetlearn, by avoiding it, cannot be reformulated as a probabilistic inference algorithm (see SI S4).
Inferring a hidden state subject to unmodelled changes that could be both first-order (i.e. from stochasticity and volatility) and second-order (i.e. from changes to stochasticity/volatility) is a generic problem. It is relevant for animals that incompletely model their environments due to environmental novelty or neural circuit limitations. Figure 3 demonstrates that chasing optimal probabilistic inference is not a sensible approach in these circumstances, and that Hetlearn is a simple, superior alternative.
Relating Hetlearn to Drosophila learning circuitry
We now discuss the connection between the circuitry and physiology of the Drosophila Mushroom Body (MB) and Hetlearn (Figure 4C). This will enable us to test whether Hetlearn recapitulates known features of biological learning in the final section of the Results. The MB contains several compartments, which individually make predictions, and adapt their internal prediction errors with intrinsic, compartment-specific characteristics. Different compartments have different intrinsic learning rates, and are weighted dynamically. More detailed aspects of circuitry are considered in a subsequent section.

A: Schematic of an individual MBON compartment. The effect of sensory representations, transmitted through KC axons, on MBON activity, is adapted through plasticity induced by co-occurring DAN activation. Adapted from Figure 1A of [34]. B: Topology of the 16 compartments comprising the MB. KC axons arising from the three KC cell types (α/β, α′/β, γ) are depicted as blue lines, segregated according to anatomical layer. Example MBON and DAN included. Adapted from Figure 1C of [3] C: Schematic of proposed qualitative mechanism. Separate MBON compartments receive highly overlapping sensory representations from KCs. However, their valences are different, and are updated by anatomically and functionally distinct DAN teaching signals. Overall valence is a weighted sum of the individual compartments’ opinions.
Sensory representations in the MB are formed in the Kenyon Cell (KC) population which receives inputs from sensory projection neurons (Figure 4A). Tuning curves of the (≈ 2000) KCs are input-selective and sparse [50], suggesting an architecture optimised for decorrelating representations of similar sensory inputs [8, 28]. Conversely, KC outputs converge onto only 34 MB output neurons (MBONs), arranged in 16 compartments [3, 27] (Figure 4B).
MBON responses to similar sensory inputs, unlike KCs, are highly correlated [23]. Conversely, inputs that carry opposing valence (aversive vs appetitive) are well separated in MBON responses [23]. This suggests that MBON activity reflects the behavioral value of a stimulus rather than its low level features. Indeed, photoactivation/suppression of specific MBONs leads to attraction or avoidance behavior [40, 3].
Behavioral changes can occur via associative learning that updates the valence of a stimulus or set of behavioral states. These changes are driven by plasticity in the KC-MBON synapses, which is in turn induced by a ‘teaching signal’ from Dopaminergic Modulatory neuron (DAN) activity (Figure 4A,C) that communicates perceived reward or value.
The architecture of the MB constrains the abstract learning mechanisms its neural circuitry implements. Individual KC axons traverse and connect to multiple distinct compartments. In contrast, each of the 20 DAN types innervates only one or two compartments (Figure 4B). The instructive role of the DAN input to these compartments was demonstrated by [2], by photoactivating specific DANs, in conjunction with an olfactory stimulus, to induce appetitive conditioning. Subsequent behavioural changes could thus be attributed to a specific combination of teaching signals, compartments and stimuli, extracting the physiological chain of events leading to a conditioned stimulus (CS).
Differences in behaviour produced by activating different compartments were then compared, by observing the flies’ subsequent preference for the CS over time. The consequent mechanistic effects mirror the Hetlearn algorithm:
Compartments can bi-directionally update the valences they attribute to conditioned stimuli, even through modulation of the same DAN (Unconditioned Stimulus / US).
Compartments exhibited intrinsic heterogeneity in the degree and manner they updated their valence estimates of stimuli in response to new experiences. In particular, different compartments had intrinsically different learning rates, and held memories for different amounts of time.
The contribution of the memories of different modules to overall behaviour is dynamic: they ‘vote’, but the importance of different voters changes over time. Co-activating DANs to the (appetitive) α1 and (aversive) γ1 pedc compartments with an odour resulted in an odour valence that was initially strongly aversive, and then appetitive after one day. This is consistent with the known feature of slower memory decay in the α1 compartment. Moreover, the initially strongly aversive response implied a nonlinear, dynamic summation of the compartmental valences.
Hetlearn for complex conditioning experiments
So far we have considered how learning-rate parallelism aids inference of a single state from observations directly pertaining to the state. We now consider how adding parallelism on an alternative dimension aids learning when animals are faced with a continuous stream of stimuli. Here, a surprising outcome could be due to an as-yet unidentified pattern in preceding stimuli, or a fundamentally unidentifiable large-world effect. This difficulty is highlighted in three standard conditioning experiments (see Figure 6A,B,C), involving stimuli S1 and S2, a reward R and a punishment P :
Extinction: An initial association (e.g. {S1 → R} repeated), is subsequently extinguished ({S1 → ∅}repeated).
Reversal (e.g. [57, 32]): Flipping between repeated bouts of {S1 → R} and {S1 → P }.
Second-order conditioning (e.g. [49]): Initial Pavlovian conditioning (e.g. {S1→R}repeated) is followed by unrewarded presentation of a similar stimulus: ({S2 → S1 → ∅} repeated).
The abrupt environmental changes in extinction and reversal conditioning are due to unsensed large-world effects that hold for a limited (reversal) or indefinite (extinction) period. Conversely in second-order conditioning, prediction errors arise from an initially deficient predictive model: the animal initially has no way of knowing that S2 signals a context where S1 is unrewarded. Once this knowledge is acquired, a learner can perfectly predict rewards from stimuli.
Many experiments have shown that Drosophila cache adaptations to different large prediction errors in separate engrams, which each possess unique dynamics. We propose that this serves a purpose: the organism can ‘hedge their bets’ on the appropriate adaptation to prediction errors through parallelism. We now provide a Hetlearn model of MB that enacts this goal, while explaining organismal behaviour and making testable mechanistic predictions.
Our model first projects presented stimuli into a high-dimensional representation modelling the Kenyon-cell layer, denoted xt at the tth presentation. Representations decay on a slower timescale than stimulus presentation. Hence, the representation x2 after experiencing S2 → S1 holds traces of the S2 representation. Each representation xt is normalised, e.g. by the recurrent inhibitory action of the APL neuron on the KC layer.
Our model learns to predict the current xt by adapting a matrix A representing linear, recurrent Kenyoncell connections. This adaptation follows the simple online-least-mean-squares rule, which can be implemented with Hebbian plasticity [53]:

We use the learning rate α = 0.4. Biologically, structural changes in the calyx, presynaptic to Kenyon cells, do accompany the consolidation of olfactory conditioned memories [10, 4].
Conditioned valences are held in parallel engrams, which form in the KC-MBON connections. The ith engram is defined by a ‘template’: e(i) = xt−1, corresponding to the Kenyon cell representation that maximally activates it, and a binary effect v(i), which is ±1 for an appetitive/aversive valence. The engram has fast and slow components held in different MBON compartments, and defined by intrinsic learning rates, which we set as αf = 0.6, αs = 0.2. Each component has a dynamic strength 


On each stimulus presentation, the system makes a prediction ŷt+1 of the next unconditioned stimuli. These are later compared against a vector of experienced unconditioned stimuli yt, represented in separate circuits such as the lateral horn. The prediction sums engram activations:

If the PE ∥yt+1 − ŷt+1∥2 exceeds a threshold (> T1 = 0.3) and all extant engram activations ⟨e(i), xt⟩ are low (< T2 = 0.3), then a new engram is formed with template xt. If the PE was an unsystematic fluke, this engram will not contribute constructively to future predictions and will die (described later). If not, it will crystallise without interfering with different existing engrams.
Engram strengths adapt to each PE. To minimise interference of each adaptation with different conditioned responses, engrams compete to claim and adapt to errors, a process we call belief competition. We take belief as

where 

Slow strengths only update once the fast strength has crossed a threshold (Tf = 0.4) replicating the repeated conditioniong required to form engrams in slow-learning MBON compartments such as α1. If a fast strength dips below 0.1 before a slow strength has activated, the engram is extinguished.
In reinforcement learning, the value of an environmental state is the long-term predicted reward, which incorporates the reward obtained at future states likely to be visited. Correspondingly, our model values a KC representation xt by evolving the representation through the recurrent matrix A, which predicts future representationas as 

with overall value 
Overall our model relies on simple computations: soft-max, correlation, and recurrent linear dynamics, which can all be enacted with biologically plausible learning rules.
Hetlearn explains and recreates multiple behaviours and their mechanistic under-pinnings
In Drosophila, extinction of a conditioned memory is supported by the acquisition of an entirely new memory towards the conditioned stimulus, which competes with the initial memory to govern extinction [17, 18]. Similarly, our model added a new, punishing engram at the onset of extinction (Figure 6C), whose increasing strength induced extinction. The original, rewarding engram is stably maintained meanwhile. As such, if extinction were a one-off fluke, or a temporary circumstance, it does not corrupt the original, rewarding memory.
Analogously, our model created four parallel memories during reversal conditioning protocols, as depicted in Figure 6B. The preference index of Figure 6 demonstrates how the speed of initial learning was governed by the slow/fast learning compartments within a context. Conversely, the speed of switching preferences during reversals was governed by competition between contexts, leading to changing degrees of conditioning across reversal blocks. The first reversal resulted in a weaker memory, while subsequent reversals built stronger memories, in agreement with experimental findings (e.g. Figure 1A of [56]).
Finally, our model recreated the behavioural dynamics observed in second-order conditioning.The valence of S2 initially rose due to an association formed with S1, whose representation formed the template of a rewarding context. A punishing context was formed to account for reward omission after S2 →S1, whose increasing strength accounted for slower-timescale extinction of this initially positive valence. These behavioural dynamics are highly conserved across animal phyla [21], and we provide a mechanistic and functional explanation.
We can test our model on mechanistic and behavioural data provided in [58], and contrast it against a different computational model of second-order conditioning provided in the same paper (Figure 5D) to capture parallel engram recruitment. There, a slow, rewarding compartment (MBON-α1) learns the first order association between S1 and reward. It then acts as a ‘teacher’ to ‘student’ compartments (γ5 and β′2a) that learn the second order association. The transience of this second-order memory therefore reflects the intrinsic fast learning/forgetting nature of these student compartments, rather than competition between engrams, as in our model, or in extinction condtioning. Teacher-to-student instruction occurs via the critical SMP108 interneuron, which forms part of a 2-hop connection recurrently connecting these two compartments. However, two key experimental observations are inconsistent with the Yamada student-teacher model.

A: Hetlearn model for behavioural conditioning in Drosophila. Different engrams strengths compete to update from a prediction error, reducing interference with previously learned information. Predictions themselves do not invoke this competition, allowing the system to make rich, context-dependent predictions that make use of the representional capacity of multiple engrams. B: Proposed model of mechanism underlying second order conditioning. Compare to C: Mechanisms underlying second-order conditioning as proposed in [58]
First, the putative student compartments were able to act as teachers. The authors observed that coactivation of the DANs PAM-γ5 and PAM β′2a with the S1 stimulus implanted a first order memory (Figure 2B of [58]). Subsequent behavioural conditioning then allowed for a second order association. The flexibility of the underlying circuit is thus more consistent with bidirectional signalling between fast and slow learners, and also the connectivity of SMP108.
Secondly, the student-teacher model predicts complete elimination of second-order conditioning if the link between student and teacher is lost. However, optogenetic inactivation of the SMP108 interneuron halves, but doesn’t eliminate, second-order conditioning from the MBON-α1 implanted memory (Figure 6D, pane 3/4).

Model predictions vs behavioural experiments.
A: Extinction protocol adapted from [18]: S1 is associated with punishment (reward of −1) 30 times, followed by 30 punishment omissions. Stimulus preference of the model to S1 over time is graphed, along with the time at which the first (rewarding) and second (omitting) contexts are automatically formed in the model. B: Reversal protocol taken from [56]. Each block depicts ten pairings of S1 or S2 with punishment (reward of −1) or omission. Bar plot depicts difference between valences for S1 and S2 at the end of each block. C: Experimental protocol used in [58], and our model, for second-order conditioning. We modelled DAN activation as a reward on the timestep after the stimulus was presented. In context 2, S2 and S1 were presented on successive timesteps. D: Drosophila preference to S1 and S2 over time. Experimental data (Figure 1H of [58]) on leftmost pane, depicting preference index. Our model graphs valence on other panes, depicting (left) standard conditioning, (centre) conditioning under SMP108 knockout, and (right) conditioning when S1 is presented before S2 in context 2.
Our model requires belief competition across MBON compartments. We propose that SMP108 enacts this competition. This is consistent with its status as an anatomical hub for bidirectional connections between MBON compartments [58]. As such, we modelled SMP108 inactivation as the cessation of belief competition, by removing the softmax in Eq. (6). This halved second-order conditioning (Figure 6, Pane 3/4), consistent with experimental data.
Overall, while both models account for basic second-order conditioning, our alternative account of the role of SMP108 explains additional artefacts of behaviour and mechanism that are inconsistent with the Yamada model. It also makes distinct, testable predictions:
SMP108 inactivation should have less of an effect on second-order conditioning during its early phases, when S2 →S1 → ∅ has been presented fewer times, and the strength of the punishing engram is thus weak.
SMP108 inactivation should reduce the preference to S1 after a bout of second-order conditioning. The lack of belief competition increases the adaptation of the rewarding engram to reward omission after S2 → S1, thus reducing its strength more.
Discussion
In this work we introduced a biologically motivated design principle, Hetlearn, for learning in scenarios where the causes of systematic environmental change can neither be represented, nor treated as noise. This provided a model of learning that recreates, interprets, and predicts Drosophila neurophysiology and behaviour higher-order conditioning experiments, without per-experiment tuning. Notably, Hetlearn trades asymptotic optimality for performance in low-data regimes and robustness to unmodelled, ‘Large world’ deviations in environmental statistics.
We first considered a fundamental component of learning across animal species: updating estimates of model states of the world through prediction errors. Existing work bases animal learning strategies on Kalman Filtering [25, 1], and Reinforcement learning [44, 37]. Here, the learning rate (i.e. sensitivity to prediction errors) is a key parameter determining estimation performance. Across tasks, both the optimal and actual values can vary dynamically. As such, neuroscientists have long considered how animals could approximate a Bayes-optimal learning rate [11, 25, 20].
In the simplest setting, two factors: stochasticity and volatility, must be estimated to set an optimal learning rate, and to represent the state as a probability distribution. Extensive literature proposes how this estimation might be performed across species [30, 33, 36, 41, 42, 47]. However, we showed that optimal learning strategies are fragile to two inevitable scenarios: drift in environmental stochasticity/volatility, and approximation error when estimating these factors from limited data.
We therefore propose that the costs of Hetlearn, both in terms of anatomical parallelism and asymptotic sub-optimality, are more than compensated by its robustness to unmodelled environmental change and its superior performance in the low-data regime.
State inference through parallelism is well known engineering principle, underpinning algorithms such as mixture-of-experts [24] and particle filtering [13]. Previous work proposed learning rate parallelism for the prediction error-based learning we considered [55]. A key difference in Hetlearn is that it avoids a specific generative model of environmental change (e.g. through abrupt, discontinuous changepoints, as in [55]) and instead incorporates recency bias in sub-learner weighting. This feature is shared with existing algorithms that deal with nonstationary large-world drift for classification problems [16] and means that weightings remain reactive to unmodelled environmen-tal changes indefinitely.
The basic Hetlearn algorithm justified the core idea of using parallelism to implicitly infer environmental quantities that are speedily constrained by data (i.e. learning rate to PEs). This is superficially consistent with detailed Drosophila MB circuitry, which displays intrinsically heterogeneous learning rates in compartments that learn behavioural valence from prediction errors. Future work could explore whether detailed circuit models strongly constrain circuitry to operate in a manner consistent with Hetlearn, though degeneracy and gaps in neurophysiological data would make this extremely challenging.
We proposed a simple, high-level model of MB learning that used parallelism to cache adaptations to surprising prediction errors before they could be confidently attributed to fluke or systematic, large-world environmental change. Our model recreated Drosophila behaviour (which is shared across phylla [21]) on key conditioning tasks where surprises arose either from fundamentally unpredictable large-world effects (extinction, reversal) or insufficient information on the relationship between observable stimuli (second-order conditioning), without requiring task-specific calibration. Given the emergence of these traits from a minimal model, we propose they reflect functional constraints on learning that enable animals to cope with truly novel situations.
Our model is qualitatively distinct from reinforcement learning algorithms, which are often used to describe conditioning [44, 48]. An RL-algorithm requires a preformed, fixed subdivision of the entire environment into behaviourally relevant states, whether these correspond directly to stimuli or emerge from pre-trained neural network [26, 51]. It then holds and updates values for each state. Our model instead spawns and removes states (i.e. engrams in the KC–MBON synapses) dynamically, after episodic experiences involving reward or punishment. This reflect parsimonious use of circuitry because there are fewer engrams to assign value to. It is also flexible, because it enables contextual modifiers of reward (such as S2 in second-order conditioning) on a biological timescale.
Many organisms, including mice and humans, share with Drosophila both behaviour on the conditioning tasks we considered, and the normative difficulty of flexibly building and acting on environmental representations in a truly novel environments. Our model ties these together through a simple, non-probabilistic algorithm exploiting parallelism for multiple purposes. While our algorithm had a plausible implementation in the mushroom body, the generality of the problems it mitigates and the behaviours it explains suggests its relevance to neural circuit computations across species.
Methods
Kalman Filter equations
The Kalman Filter updates valences from observed PEs δt according to the following equations:



This requires seeding with 
This also requires seeding with 

ALS Inference
ALS requires a separate learner that updates from PEs with a fixed learning rate α ∈ (0, 1). We set α = 0.5 in all simulations. So vt+1 = vt + 0.5δt+1.
A hyperparameter is the lag j, which we set at j = 10. A large lag requires a longer burn-in before inference can occur. A smaller lag reduces best-case inference accuracy.
Let Yt = rt − vt be the reward estimation error of this ‘fixed learner’. Note the difference from the PE: δt+1 = rt+1 − vt. Yt is observable.
The ALS method builds estimates 



Here, Aj,• represents the jth row of A, and bj represents the sample estimate of E [Yt+jYt].
For t < j, the ALS method cannot provide the Kalman learner with estimates. For these first 10 timepoints, we used α = 0.5. After, the Kalman learner updated according to the optimal equations of Eq. (7), using the ALS-derived estimates of

Piray et al implementation
We implemented the particle filtering algorithm described in [42]. Each particle (/sublearner) updates valences according to the Kalman Filter (Eqs. (7)), but has private estimates of 

Private estimates of stochasticity and volatility update stochastically and identically. For volatility, the update equation is:

where
ηp ∈ [0, 1] is a constant hyperparameter. Larger values indicate less change per trial.
ϵt ∼β(0.5ηp(1−ηp)−1, 0.5) ∈ [0, 1] (a beta distribution)
.
The ith particle calculates probabilities for the next reward by assuming

When an actual reward arrives, the relative weights of the particles are taken as

So particles whose probabilistic model predicted the actual reward with higher probability are weighted more. The weights are then normalised so that 

We follow [42] in using the Systematic Resampling procedure, to resample particles (deleting low-probability particles and duplicating high-probability particles) if their weight inequality exceeds 0.5.
Data availability
All code for simulations in this work can be obtained at the public repositories: https://github.com/VRaman-Lab/HetlearnComplexConditioningModel.git (For Figures 1, 2, and 3) https://github.com/VRaman-Lab/HetlearnRewardPrediction.git.
Additional files
Additional information
Funding
European Research Council
References
- [1]Uncertainty, neuromodulation, and attentionNeuron 46:681–692https://doi.org/10.1016/j.neuron.2005.04.026PubMedGoogle Scholar
- [2]Dopaminergic neurons write and update memories with cell-type-specific ruleseLife 5:e16135https://doi.org/10.7554/eLife.16135PubMedGoogle Scholar
- [3]Mushroom body output neurons encode valence and guide memory-based action se-lection in drosophilaeLife 3:e04580https://doi.org/10.7554/eLife.04580PubMedGoogle Scholar
- [4]Circuit reorganization in the drosophila mushroom body calyx accompanies memory consolidationCell reports 34https://doi.org/10.1016/j.celrep.2021.108871PubMedGoogle Scholar
- [5]What is a cognitive map? organizing knowledge for flexible behaviorNeuron 100:490–509https://doi.org/10.1016/j.neuron.2018.10.002PubMedGoogle Scholar
- [6]Learning the value of information in an uncertain worldNature neuroscience 10:1214–1221https://doi.org/10.1038/nn1954PubMedGoogle Scholar
- [7]Deep reinforcement learning and its neuroscientific implicationsNeuron 107:603–616https://doi.org/10.1016/j.neuron.2020.06.014PubMedGoogle Scholar
- [8]Sparse synaptic connectivity is required for decorrelation and pattern separation in feedforward networksNature communications 8:1–11https://doi.org/10.1038/s41467-017-01109-yPubMedGoogle Scholar
- [9]A distributional code for value in dopamine-based reinforcement learningNature 577:671–675https://doi.org/10.1038/s41586-019-1924-6PubMedGoogle Scholar
- [10]Learning and memory using drosophila melanogaster: a focus on advances made in the fifth decade of researchGenetics 224:iyad085https://doi.org/10.1093/genetics/iyad085PubMedGoogle Scholar
- [11]Explaining away in weight spaceAdvances in neural information processing systems 13Google Scholar
- [12]Optimal sensorimotor integration in recurrent cortical networks: a neural implementation of kalman filtersJournal of neuroscience 27:5744–5756https://doi.org/10.1523/jneurosci.3985-06.2007PubMedGoogle Scholar
- [13]A tutorial on particle filtering and smoothing: Fifteen years laterIn:
- Crisan D
- Rozovskii B
- [14]Bayesian brain: Probabilistic approaches to neural codingMIT press Google Scholar
- [15]What do reinforcement learning models measure? interpreting model parameters in cognition and neuroscienceCurrent opinion in behavioral sciences 41:128–137https://doi.org/10.1016/j.cobeha.2021.06.004PubMedGoogle Scholar
- [16]Incremental learning of concept drift in nonstationary environmentsIEEE transactions on neural networks 22:1517–1531https://doi.org/10.1109/tnn.2011.2160459PubMedGoogle Scholar
- [17]Reevaluation of learned information in drosophilaNature 544:240–244https://doi.org/10.1038/nature21716PubMedGoogle Scholar
- [18]Integration of parallel opposing memories underlies memory extinctionCell 175:709–722https://doi.org/10.1016/j.cell.2018.08.021PubMedGoogle Scholar
- [19]The free-energy principle: a unified brain theory?Nature reviews neuroscience 11:127–138https://doi.org/10.1038/nrn2787PubMedGoogle Scholar
- [20]A unifying probabilistic view of associative learningPLoS computational biology 11:e1004567https://doi.org/10.1371/journal.pcbi.1004567PubMedGoogle Scholar
- [21]Using pavlovian higher-order conditioning paradigms to investigate the neural substrates of emotional learning and memoryLearning & Memory 7:257–266https://doi.org/10.1101/lm.35200PubMedGoogle Scholar
- [22]Heuristic decision makingAnnual review of psychology 62:451–482https://doi.org/10.1146/annurev-psych-120709-145346PubMedGoogle Scholar
- [23]Plasticity-driven individualization of olfactory coding in mushroom body output neuronsNature 526:258–262https://doi.org/10.1038/nature15396PubMedGoogle Scholar
- [24]Adaptive mixtures of local expertsNeural computation 3:79–87https://doi.org/10.1162/neco.1991.3.1.79PubMedGoogle Scholar
- [25]The dynamics of memory as a consequence of optimal adaptation to a changing bodyNature neuroscience 10:779–786https://doi.org/10.1038/nn1901PubMedGoogle Scholar
- [26]What learning systems do intelligent agents need? complementary learning systems theory updatedTrends in cognitive sciences 20:512–534https://doi.org/10.1016/j.tics.2016.05.004PubMedGoogle Scholar
- [27]The connectome of the adult drosophila mushroom body provides in-sights into functioneLife 9:e62576https://doi.org/10.7554/eLife.62576PubMedGoogle Scholar
- [28]Optimal degrees of synaptic connectivityNeuron 93:1153–1164https://doi.org/10.1016/j.neuron.2017.01.030PubMedGoogle Scholar
- [29]A signal-detection analysis of fast- and-frugal treesPsychological review 118:316https://doi.org/10.1037/a0022684PubMedGoogle Scholar
- [30]A bayesian foundation for individual learning under uncertaintyFrontiers in human neuroscience 5:39https://doi.org/10.3389/fnhum.2011.00039PubMedGoogle Scholar
- [31]Deep learning, reinforcement learning, and world modelsNeural Networks 152:267–275https://doi.org/10.1016/j.neunet.2022.03.037PubMedGoogle Scholar
- [32]Dopaminergic mechanism underlying reward-encoding of punishment omission during reversal learning in drosophilaNature Communications 12:1115https://doi.org/10.1038/s41467-021-21388-wPubMedGoogle Scholar
- [33]Functionally dissociable influences on learning rate in a dynamic environmentNeuron 84:870–881https://doi.org/10.1016/j.neuron.2014.10.013PubMedGoogle Scholar
- [34]The drosophila mushroom body: from architecture to algorithm in a learning circuitAnnual review of neuroscience 43:465–484https://doi.org/10.1146/annurev-neuro-080317-0621333PubMedGoogle Scholar
- [35]Midbrain dopamine neurons encode decisions for future actionNature neuroscience 9:1057–1063https://doi.org/10.1038/nn1743PubMedGoogle Scholar
- [36]An approximately bayesian delta-rule model explains the dynamics of belief updating in a changing environmentJournal of Neuroscience 30:12366–12378https://doi.org/10.1523/jneurosci.0822-10.2010PubMedGoogle Scholar
- [37]Reinforcement learning in the brainJournal of Mathematical Psychology 53:139–154https://doi.org/10.1016/j.jmp.2008.12.005Google Scholar
- [38]A new autocovariance least-squares method for estimating noise covariancesAutomatica 42:303–308https://doi.org/10.1016/j.automatica.2005.09.006Google Scholar
- [39]Temporal difference models and reward-related learning in the human brainNeuron 38:329–337https://doi.org/10.1016/s0896-6273(03)00169-7PubMedGoogle Scholar
- [40]Activity of de-fined mushroom body output neurons underlies learned olfactory behavior in drosophilaNeuron 86:417–427https://doi.org/10.1016/j.neuron.2015.03.025PubMedGoogle Scholar
- [41]A simple model for learning in volatile environmentsPLoS computational biology 16:e1007963https://doi.org/10.1371/journal.pcbi.1007963PubMedGoogle Scholar
- [42]A model for learning based on the joint estimation of stochasticity and volatilityNature communications 12:1–16https://doi.org/10.1038/s41467-021-26731-9PubMedGoogle Scholar
- [43]The foundations of statisticsCourier Corporation Google Scholar
- [44]A neural substrate of prediction and re-wardScience 275:1593–1599https://doi.org/10.1126/science.275.5306.1593PubMedGoogle Scholar
- [45]Neuronal coding of prediction errorsAnnual review of neuroscience 23:473–500https://doi.org/10.1146/annurev.neuro.23.1.473PubMedGoogle Scholar
- [46]A causal link between prediction errors, dopamine neurons and learningNature neuroscience 16:966–973https://doi.org/10.1038/nn.3413PubMedGoogle Scholar
- [47]Discounting future reward in an uncertain worldDecision https://doi.org/10.1037/dec0000219PubMedGoogle Scholar
- [48]Reinforcement learning: An introductionMIT press Google Scholar
- [49]Second-order conditioning in drosophilaLearning & memory 18:250–253https://doi.org/10.1101/lm.2035411PubMedGoogle Scholar
- [50]Olfactory representations by drosophila mushroom body neuronsJournal of neurophysiology 99:734–746https://doi.org/10.1152/jn.01283.2007PubMedGoogle Scholar
- [51]Prefrontal cortex as a meta-reinforcement learning systemNature neuroscience 21:860–868https://doi.org/10.1038/s41593-018-0147-8PubMedGoogle Scholar
- [52]Maximum likelihood estimation of misspecified modelsEconometrica: Journal of the econometric society :1–25https://doi.org/10.2307/1912526Google Scholar
- [53]Adaptive switching circuitsIn: IRE WESCON convention record pp. 96–104Google Scholar
- [54]A neural implementation of the kalman filterAdvances in neural information processing systems 22Google Scholar
- [55]A mixture of delta-rules approximation to bayesian inference in change-point problemsPLoS computational biology 9:e1003150https://doi.org/10.1371/journal.pcbi.1003150PubMedGoogle Scholar
- [56]The gabaergic anterior paired lateral neurons facilitate olfactory reversal learning in drosophilaLearning & Memory 19:478–486https://doi.org/10.1101/lm.025726.112PubMedGoogle Scholar
- [57]A new paradigm for operant conditioning of drosophila melanogasterJournal of Comparative Physiology A 179:429–436https://doi.org/10.1007/bf00194996PubMedGoogle Scholar
- [58]Hierarchical architecture of dopaminergic circuits enables second-order conditioning in drosophilaeLife 12:e79042https://doi.org/10.7554/eLife.79042PubMedGoogle Scholar
Article and author information
Author information
Version history
- Preprint posted:
- Sent for peer review:
- Reviewed Preprint version 1:
Cite all versions
You can cite all versions using the DOI https://doi.org/10.7554/eLife.112618. This DOI represents all versions, and will always resolve to the latest one.
Copyright
© 2026, Raman et al.
This article is distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use and redistribution provided that the original author and source are credited.
Metrics
- views
- 0
- downloads
- 0
- citations
- 0
Views, downloads and citations are aggregated across all versions of this paper published by eLife.