Accumulation of neural state transitions in dorsomedial striatum predicts patch foraging decisions

  1. Department of Neuroscience, Johns Hopkins University School of Medicine, Baltimore, United States
  2. Kavli Neuroscience Discovery Institute, Johns Hopkins University, Baltimore, United States

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, public reviews, and a provisional response from the authors.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Alicia Izquierdo
    University of California, Los Angeles, Los Angeles, United States of America
  • Senior Editor
    Kate Wassum
    University of California, Los Angeles, Los Angeles, United States of America

Reviewer #1 (Public review):

Summary:

In this study, Shuler and colleagues record neurons from the DMS in mice performing a patch foraging task. In this task, mice had the choice between harvesting rewards from 2 ports - one the time-investment port where the rate of reward declined over time and the other a context port where the rate of reward was either high or low. Mice performed the task appropriately, switching between ports as the rate of reward declined in the time-investment port and switching more rapidly when the context port delivered high versus low rewards. The behavior of the mice was also strongly driven by time since the most recent reward receipt, in conflict with normative accounts of patch foraging. Individual DMS neurons showed bistable firing patterns, transitioning to high rates of activity at various times from reward. Overall, the population tiled the temporal space, and the accumulation of the number of neurons in the high firing state was predictive of patch exit. The rate of accumulation varied with things that also affected behavior.

Strengths:

Overall, the aims of the study were clear and important, the experiment directly addresses them, and the results are clear and provide compelling support for the authors' conclusions.

Weaknesses:

I have only a few comments and questions to consider, none of which are criticisms of what was done, really.

(1) Probably my chief question, alluded to in the discussion, is what the evidence is that DMS plays a causal role in generating these correlates and the resulting behavior, in light of the lack of causal evidence here. What are other options? Could such information depend on upstream areas such as OFC or mPFC, with DMS just a pass-through? And while I would not ask for causal data, is there a specific prediction? That is, if the area were inactivated, would mice stay longer or shorter? Not do the task? If I wanted to do a causal test of the authors' idea regarding the contribution of DMS to this behavior, what would be predicted, and what result would invalidate the hypothesis? Speculating on this a bit, beyond just saying DMS is involved, would be useful.

(2) Not much is said about the suboptimal strategy. Would DMS continue to play the same role if the mice showed no effect of recent reward and instead performed appropriately? Or is some other area doing that job? Or is this not important? I thought it was interesting that the mice basically did not treat the game quite like they were supposed to. Is it important to go back and look at what is happening in DMS under normative conditions to really know how this area contributes to proper foraging?

(3) Do these neurons also track time in the context port? Or do they only exhibit this behavior in the port where rewards are depleting? This seems like an interesting question. Do they show the same profile in different ports, if so?

Reviewer #2 (Public review):

Summary:

Here, Sutlief et al. use a novel patch-foraging task to investigate the role of dorsomedial striatal (DMS) neurons in determining when animals disengage from a resource. They show that mice, contrary to canonical optimal-foraging predictions, adopt a strategy in which reward receipt resets timing behavior, with decisions further shaped by both cumulative time spent in a patch and the overall quality of the environment. The authors further demonstrate that a subset of DMS neurons exhibits step-like activity patterns during task performance. Importantly, the accumulation of these state transitions across the neuronal population predicts the timing of patch-leaving decisions on a trial-by-trial basis, providing a potential neural mechanism underlying decisions about when to abandon a currently exploited resource.

Strengths:

This study addresses an important question using a well-designed, interesting behavioral task. The finding that mice employ a reward-triggered exit-timing policy is particularly interesting, as it is pertinent to the many patch foraging-style tasks that have been developed for use in mice, where rewards are delivered as discrete events. The identification of step-like activity in DMS neurons is mostly compelling, and the authors' trial-by-trial analysis linking this activity to behavior provides some support for its relevance to patch-leaving decisions.

Weaknesses:

A key interpretational issue is whether the DMS signal reflects timing specifically, rather than movement initiation or other task-related factors. The authors argue that once a sufficient number of neurons transition, the animal exits the time-investment port. However, it remains unclear whether this population threshold reflects a timing computation that determines when to leave in the more abstract sense, or a signal more directly related to movement onset (that may also be initiated after some proportion of the population has changed its activity). An important control would be to examine neural activity while animals are engaged at the context port. In this epoch, animals presumably do not need to time their departure in the same way, but they still eventually initiate movement. If the DMS signal reflects timing rather than movement, one would not expect the same accumulation-to-threshold pattern of step-like transitions at the context port.

It would also be helpful for the authors to clarify the behavioral definition of the leaving decision. Can mice return to the time-investment port after exiting it if they do not subsequently enter the context port? How exactly is "exit" defined: as withdrawal from the time-investment port, entry into the context port, or some other behavioral event? Is there variability in the latency between time-investment port exit and context-port entry, and if so, is this latency related to DMS step-like activity? These details are important for interpreting whether the neural activity is aligned with a timing decision, movement initiation, or the execution of a transition between task states.

The classification approach for identifying step-like activity seems generally reasonable, and the low false-positive rate against homogeneous Poisson controls is reassuring. However, one potential issue is that the identification of trial-by-trial state transitions is not independent of the session-level characterization of each neuron. The algorithm first fits a sigmoid to the pooled session data and then uses the resulting high- and low-firing-rate states to constrain interval-level fits. This may bias the analysis toward finding step-like transitions in neurons whose activity is only approximately step-like at the session level, effectively reducing the space of alternative solutions available to the interval-level fits. As implemented, the approach therefore functions more as a detector of consistency with a session-defined step model than as an unbiased test of whether individual intervals are better described by discrete state transitions versus alternative dynamics such as ramps or gradual drifts. This concern could be addressed by comparing the constrained sigmoid model against alternatives, such as constant-rate or ramping models, on held-out intervals, or by deriving state parameters from an independent subset of trials and testing classification on the remaining trials.

The inclusion threshold for the accumulation analysis is difficult to evaluate. Sessions were included if they contained at least seven simultaneously recorded step-like units, but this number is hard to interpret without knowing the total number of recorded units per session and the fraction classified as step-like. Seven units may be sufficient for fitting a population accumulation trajectory, but because the cutoff is based on an absolute number rather than a proportion of the recorded population, it is unclear whether included sessions reflect robust population-level step-like dynamics or a relatively small selected subset of DMS activity. Reporting the number and fraction of step-like units per session, as well as the sensitivity of the accumulation results across different inclusion thresholds, would help clarify this point.

Reviewer #3 (Public review):

Sutlief and colleagues report behavioral and neural results from mice performing a patch foraging task. Behaviorally, they argue that time since last reward is a major determinant of when mice decide to leave a patch. In the brain, they find neurons in the dorsomedial striatum that show step-like changes in their firing rate at a range of times following reward. Population analyses show that the cumulative fraction of neurons that have undergone such a step-like change in firing rate can be used to predict patch-leaving times with impressive accuracy.

Overall, this is an interesting set of results that has been analyzed in a principled way. The manuscript is well written, the results are explained clearly, and the evidence supporting the authors' conclusions is strong. The manuscript is therefore a potentially valuable contribution to the growing literature assessing how the brain solves stopping problems like the patch foraging scenario. I have suggestions for the authors to consider that might further increase the rigor of their results, and a few suggestions for improving the clarity of the work for readers.

(1) I don't quite understand how the behavioral task works. Are mice rewarded for making discrete nose poke responses in the investment and context ports? Or are they required to nose poke and hold? Is reward given with some probability per response (which decreases with time in the patch), or is the reward probability a function of elapsed time in the patch, time since last response, or dwell time in the port? Also, exactly what equation defines how reward probability changes over time for the high- and low-value contexts? I couldn't find these details anywhere in the manuscript, and they would be helpful for better understanding the behavior and the later neural results.

(2) How was the optimal strategy determined? Several features of the author's task violate the assumption of the marginal value theorem, so computing the optimal residence time is not a straightforward application of the classic model. There's a diagram in Figure 1h that depicts an MVT-like graphical solution, but the conventions of the plot are not familiar to me, and there's no description of how it works in the results or methods. More detail here would be much appreciated. In a similar vein, the authors report that mice generally exceeded optimal residence times in patches, but no statistical comparison is provided to back up that statement. There should be some formal test of this if it is to be included in the results.

(3) The authors argue that time since last reward is the predominant determinant of patch leaving time. However, as the authors note, time since last reward is correlated with other task variables (patch reward rate, time in patch, etc.). I don't trust that SVM coefficients can be interpreted as straightforward measures of a variable's importance for classification performance in the case of correlated predictors. A better approach would be to assess how well the model performs as subsets of variables are added or removed from the model.

(4) For the SVM analysis, I'm not quite understanding how or why the authors are using 5 s after mice left the patch as additional "Leave" examples. For instance, is time since entry computed for the investment patch, or the context patch that mice enter after they leave the investment patch? Similarly, is the time since the last reward relative to the investment patch, or the reward the mouse is likely to receive at the context patch? Moreover, I'm not sure it's safe to assume that because the mouse left at time t, time t+1 necessarily reflects conditions on which the mouse would definitely leave again. If we're thinking about the stay/leave decision as something that is being repeated sequentially on a fast time scale to determine how long mice stay in the patch, it doesn't follow that observing a mouse leave means that any patch conditions after that would necessarily result in the same decision. If that were the case, it would mean that seeing a mouse leave a patch after 2 s would preclude ever observing a residence time longer than 2 s, which is clearly not compatible with the authors' data. Ultimately, it's only possible to observe one decision to leave per trial; including data points beyond that as additional leave examples seems overly speculative to me.

(5) The authors validate their approach for quantifying step-like changes in firing rate using simulations of constant-rate Poisson spiking and observe a low false positive rate. This is encouraging, but it doesn't seem like the only way in which their method could go awry, or even the most concerning way. I would be much more interested in seeing the false positive rate for continuous, ramp-like changes in firing rate, which would be much more likely to trip up the authors' approach and are also the major relevant alternative hypothesis to step-like changes in firing rate. Random walks in firing rate might also be worth testing.

(6) The finding that cumulative "transitioned" neurons is predictive of patch leaving is interesting. However, I can't help but wonder how truly informative this variable is for predicting patch leaving. It seems as though neurons can only transition firing rates one time. That means that as time in the patch increases, the fraction of transitioned neurons naturally increases. Similarly, all visits must eventually end with the mouse leaving the patch, so the hazard rate of leaving increases with time in the patch. Given that, can the authors be certain that the cumulative transitioned neurons are really what's predicting patch leaving time, or would any generically increasing function perform roughly the same? An interesting test would be to mismatch the neural predictor and behavior at the level of trials. If this mechanism is really specific, rather than something that captures the general structure of an increasing hazard rate of leaving, then prediction of leaving time should work substantially better when the neural predictor is correctly matched to behavior on the trial for which it was recorded.

Author response:

Reviewer #1

(1) Causality and the role of DMS; a specific, falsifiable prediction. We agree that the paper should not leave the causal question implicit, and we will expand the Discussion to state a concrete prediction rather than a general claim of involvement. Briefly, if the accumulation signal we describe carries the animal's intended departure time, then suppressing DMS during patch occupancy should not simply shift exit times in one direction but should degrade their structure: exit-time variability should increase, and exit timing should lose its systematic dependence on reward-rate context and on the time of the most recent reward. A plausible alternative outcome is disengagement from the task altogether, which would be uninformative and would need to be controlled for. The result that would falsify our hypothesis is the one we will state explicitly: exit timing that remains as predictable, and as sensitive to context and reward history, under DMS suppression, as without it. We will also discuss the alternative the reviewer raises, that these signals are inherited from cortical inputs such as OFC or mPFC with DMS acting as a relay, and note that our data cannot presently distinguish this from a locally generated signal.

(2) Behavior under a normative strategy. This is an interesting question and we will address it in the Discussion. Our expectation, which we will frame as a prediction rather than a result, is that an animal timing from patch entry rather than resetting at each reward would show accumulation that begins at entry and proceeds to a context-dependent threshold at the reward-rate-optimal time, rather than the reward-triggered resets we observe. In this view, the reset structure of the neural signal is a reflection of the behavioral policy rather than a property of the region. We will make clear that this is a testable prediction that our current dataset does not address.

(3) Do these neurons also track time at the context port? We intend to answer with new analysis, and it converges with Reviewer #2's suggested analysis (below), so we treat the two together there.

Reviewer #2

(a) Timing versus movement initiation: activity at the context port. We take this to be a central interpretational concern. We will examine whether the step-like DMS activity extends to the context port, testing the interval between the final context-port reward and departure for the same step-like transitions and accumulation we observe in the time-investment port. We will apply the same comparison to the context-port inter-reward intervals, which addresses Reviewer #1's third point about whether these neurons also track time at the context port.

We want to flag one feature of the task that bears on how the outcome should be read. The context port is not a timing-free epoch. Its four rewards are delivered at predictable, regularly spaced intervals, and the interval between the final reward and the animal's departure is self-timed. Departure from the context port is therefore also a self-timed action, and observing accumulation there would not by itself indicate that the signal reflects movement initiation rather than timing. What the comparison can inform is whether the accumulation is specific to a decision about when to disengage from a depleting resource, or is a more general feature of self-timed departures. This is a meaningful distinction either way, and one we will report and interpret whichever direction the result falls.

(b) Operational definition of leaving; the exit-to-entry latency. We agree these details are necessary for interpretation and their absence is our omission. Exit is the final withdrawal from the time-investment port preceding the next context-port visit, and we will make that clear in the revised methods. Mice can and occasionally do re-enter the time-investment port without an intervening context-port visit (especially early in training). Such re-entries are not counted as exits. We will also examine whether the latency between time-investment-port exit and context-port entry relates to the accumulation slope on the corresponding interval, to test whether the neural signal relates to the decision or to the execution of the transition.

(c) Independence of interval-level fits from the session-level model. This is a fair characterization of the procedure, and we accept the distinction the reviewer draws between a detector of consistency with a session-defined step model and an unbiased test of discrete versus continuous dynamics. We will address it with a held-out validation: estimating each unit's state parameters and transition time from one half of its intervals and testing whether the transition times recovered from the withheld half agree. The discrete-versus-continuous comparison is addressed directly by the ramp simulations under Reviewer #3's point (5) below.

(d) The inclusion threshold for the accumulation analysis. We will add a supplementary figure reporting the total number of recorded units per session and the fraction classified as step-like, so that the seven-unit criterion can be evaluated against the recorded population rather than in the abstract. Yield varied substantially across sessions, from a handful of units to roughly one hundred, and we will show this distribution directly. We will also report the accumulation results across a range of inclusion thresholds spanning approximately five to eight simultaneously recorded step-like units, so that readers can assess sensitivity to the choice.

Reviewer #3

(1) Specification of the task. We agree the task description was insufficient, and we will correct this at the front of the Results and in the Methods. The time-investment port operates on a poke-and-hold basis: the mouse maintains its head in the port and rewards are delivered stochastically over time for as long as it remains, with no requirement to withdraw and re-poke. Reward delivery follows an exponentially decaying rate in time since port entry, with a time constant of eight seconds, integrating to an expected eight rewards of one microliter each (8 µL total) for indefinite occupancy; we will give the explicit function. The reward probability function in the time-investment port is identical across blocks. The high- and low-reward-rate contexts are properties of the context port alone (four rewards over five seconds versus four rewards over ten seconds), and we will make this contrast unambiguous, since it is the manipulation on which the design rests.

(2) Derivation of the optimum, Figure 1h, and a formal test of overstaying. We appreciate this comment. The optimal residence time in our task is not obtained by the classical Charnov tangent construction. It is computed by explicit maximization of the overall reward rate over the full cycle, following the framework in Sutlief et al. (2025) and shown graphically in Figure 1h. We will expand the legend of Figure 1h so its conventions are stated explicitly, give the reward-rate-maximizing derivation as an explicit equation in the Methods, and reframe the surrounding text around reward-rate maximization as the normative principle, with MVT identified as the special case it is. We will also add the formal statistical comparison of observed residence times against the computed optimum, which the reviewer correctly notes was asserted rather than tested.

(3) Interpretation of SVM coefficients with correlated predictors. We accept this criticism. We will not rest the ordering of predictors on coefficient magnitudes alone. We will add a variable inclusion-and-ablation analysis, reporting cross-validated classification performance as each predictor is added to and removed from the model, so that the contribution of time since last reward is assessed by its effect on performance rather than by its normalized weight. We will additionally add a complementary analysis of the leave hazard that estimates the contribution of each variable without requiring the classification framework.

(4) The five-second post-exit window. The reviewer is right that we did not explain this choice, and right that it rests on an assumption. Our reasoning was that a single exit moment per trial leaves the decision boundary badly under-constrained, and that treating the moments immediately following an exit as conditions under which the animal would also have left is licensed by the fact that within-patch reward rate declines monotonically with time, so conditions in the counterfactual continued visit would have been strictly less favorable than those already rejected. We accept that this is an assumption rather than an observation and will state it as such. We will also report the analysis across a range of window durations so that the independence of the result on this choice is visible. To the reviewer's specific questions: both time since entry and time since last reward are computed with respect to the time-investment port throughout, and we will state this explicitly.

(5) False positive rate against ramps and random walks. We agree this is the more informative validation, and that continuous ramping is the most relevant alternative to ours. We will generate simulated units with continuous ramp-like rate changes, matched to the firing rates and interval structure of our recorded units, and pass them through the identical classification pipeline to obtain false positive rates comparable to the Poisson analysis already reported. We will retain the flat-rate Poisson simulation and present the ramp results as additional panels of the same supplement. We will also explore random-walk dynamics. Together with the held-out validation of transition times described under Reviewer #2(c), this converts the step characterization from a single-null validation into a comparison against the relevant continuous alternatives.

(6) Specificity of the accumulation signal versus a generic increasing function. This is a valuable challenge and we will address it directly. We will implement the trial-mismatch control the reviewer proposes, randomly reassigning accumulation trajectories to reward-to-exit intervals within session and showing the extent to which predictive performance degrades relative to the correctly matched case.

We would also note two features of the existing results that speak to this concern, and which we will bring forward in the revision because we did not make them salient enough. First, a signal that merely tracked elapsed time would be expected to shift its starting level as well as its rate across trials with different exit times; instead, the accumulation slope is strongly related to exit time (mean r = -0.551) while the intercept is not (mean r = 0.013), indicating a variable rate from a stable origin. Second, and more to the point, the accumulation arrives at a common level at the moment of exit whether the animal leaves early or late. The rate of accumulation shifts with the animal’s policy on that trial such that the threshold is met at the intended time. 

What makes this predictor non-trivial is its trial-by-trial correspondence to behavior, not simply that it increases over time. We will make this argument explicitly alongside the shuffle control.

Summary

To summarize the planned additions: (i) analysis of step-like activity and its accumulation at the context port, with the interpretive caveat noted above; (ii) validation of step detection against ramping alternatives, together with held-out estimation of transition times; (iii) a trial-mismatch control for the specificity of the accumulation predictor; (iv) an inclusion-and-ablation analysis of the behavioral predictors and a complementary hazard model of the leave decision; (v) reporting of unit yield, step-like fraction, and sensitivity of the accumulation results to the inclusion threshold; (vi) analysis of the exit-to-context-entry latency in relation to the neural signal; (vii) a formal statistical test of overstaying relative to the computed optimum; and (viii) substantial clarification of the task specification, the operational definition of exit, and the derivation of the optimal residence time, including an expanded Figure 1h legend.

Our aim in the revision is to meet the specificity concern raised in the assessment as directly as the existing data allow, and we hope the revised manuscript will warrant reconsideration of the strength-of-evidence characterization.

We are grateful to the reviewers for the care evident in their reports, and to you both for handling the manuscript.

References

Charnov, E. L. (1976). Optimal foraging, the marginal value theorem. Theoretical Population Biology, 9(2), 129–136.

Sutlief, E., Walters, C., Marton, T., & Hussain Shuler, M. G. (2025). The value of initiating a pursuit in temporal decision-making. eLife. https://doi.org/10.7554/eLife.99957.2.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation