Task structure and behavioral strategy.

(a) Schematic of the behavioral box (left) and structure of the task (right). Mice alternate between collecting rewards from the “context port” and the “time-investment port”. In the context port, four rewards are delivered at regular intervals before the port becomes inactive, indicated by a light cue turning on. In the time-investment port rewards are delivered stochastically with exponentially decreasing probability. The context port is reset when the mouse enters the time-investment port, while the time-investment port is reset after the mouse receives all four rewards in the context port. Created in BioRender. Sutlief, E. (2025) https://BioRender.com/PLACEHOLDER. (b) Observed exit times from the time-investment port across mice (n = 10 mice, 134 sessions, 6,970 trials), measured from the time of entry into the port (two-sample t-test, p = 1.9 × 10−232). Arrows indicate the theoretically optimal exit times for each context (see panel h). (c) The probability (hazard rate) of exiting the time-investment port based on the time since port entry (x-axis) and the context port reward rate (hue), (d) Same as b, except measured from the time of the final reward received in the time-investment port (two-sample t-test, p = 8.6 × 10−111). (e) The probability (hazard rate) of exiting the time-investment port based on the time since the most recent reward (x-axis) and the context port reward rate (hue), (f) The probability (hazard rate) of exiting the time-investment port based on the time since the most recent reward (x-axis) and the time of the most recent reward relative to time of port entry (hue), (g) Support vector machine (SVM) coefficients for predicting time-investment port occupancy. All three predictors are significantly different from zero (one-sample t-tests, all p < 0.001) and all pairwise comparisons are significant (paired t-tests, all p < 0.001). (h) The theoretically optimal times to exit the time-investment port (vertical lines) for each context block based on the maximum achievable overall reward rates (dashed lines) in accordance with the Marginal Value Theorem (MVT).

Dorsomedial striatum neurons exhibit step-like activity patterns following rewards.

(a) Example activity of a recorded unit during a single trial in the time-investment port. Top: Example unit spike raster. Bottom: Firing rate trace calculated by binning spikes (10 ms bins) and smoothed with a Gaussian kernel (σ = 100 ms), (b) Anatomical localization of recorded units in DMS for all recorded neurons (1,849 units across 10 mice), (c) Four example units showing reward responsive patterns. Top: Spike rasters for each reward-reward and reward-exit interval, ordered by the duration of the interval. Bottom: Corresponding mean firing-rate traces aligned to reward delivery (time = 0). (d) Same as c, except showing intervals where each unit exhibits a state-transition, a discrete step in activity post reward, identified by the inflection point of a sigmoid function fit to the firing rate. Rasters and mean firing rate traces are aligned to the detected step times (yellow), (e) Heatmap showing the normalized firing rates of all DMS units that make step transitions following rewards (n = 898 step-like units). Units are ordered by the direction of their activity step and their mean step time. Example units are marked with arrows, (f) Distribution of the mean activation times of the units displayed in e.

Neural state-transition times correlate with projected exit times.

(a) Plots for each example DMS unit (same as Figure lc,d) showing the probability of the unit being in the “on” state in the time following reward, separated by quartiles of the projected behavioral exit time (see Figure 1-figure supplement 2). (b) Correlation between unit state-transition times and projected exit times for each example unit. Hue and vertical dashed lines indicate the quartile groupings used in a. (c) Population mean state-transition times (n = 898 step-like units) shown as a cumulative function for each of the four projected exit time ranges, including only units with measurements in every time range. All groups differ significantly (Kolmogorov-Smirnov, all p < 0.008). (d) Same as c, except separated by the time of the last reward. Only comparisons involving the “>7s” group differ significantly (Kolmogorov-Smirnov; “<2s” vs “>7s”: p = 1.4 × 10−5; “2-4s” vs “>7s”‘. p = 3.7 × 10−5; “4-7s” vs “>7s”: p = 0.021), while remaining comparisons are not significant (all p > 0.06). (e) Same as c, except separated by the reward rate in the context port. Groups differ significantly (Kolmogorov-Smirnov, p = 5.3 × 10−5).

Accumulation of neural state transitions in DMS predicts exit timing.

(a) Four example trials from a single session (ES025, 56 trials, 19 step-like units), showing the cumulative step transitions of simultaneously recorded units. Vertical lines indicate entry (green), exit (red), and reward events (blue), (b) The accumulation lines for all reward-to-exit intervals from the same session as the examples in a. Lines are colored based on the duration of time between the reward and exit events. Exits are marked by maroon dots capping each line. The gray dashed horizontal line indicates the session-specific leave threshold (median final cumulative count at exit), (c) The slope of each accumulation line in b fit with a linear model, plotted against the exit time (R2 = 0.64, p =1.9 × 10−13, OLS regression), (d) The intercept of each accumulation line in b fit with a linear model, plotted against the exit time (R2 = 0.04, p = 0.13, n.s.). (e) Histogram of per-session Pearson correlation coefficients between accumulation slopes and exit times for all sessions with at least 7 step-like units (n = 53 sessions from 6 mice; mean r = −0.551, one-sample t-test vs. 0: t = −20.5, p = 1.4 × 10−26; 48 of 53 sessions individually significant at p < 0.05). (f) Histogram of per-session Pearson correlation coefficients between accumulation y-intercepts and exit times (n = 53 sessions; mean r = 0.013, t = 0.35, p = 0.73; 11 of 53 individually significant), (g) The relationship between exit time predicted by the accumulate-to-threshold model and true exit times for all reward-to-exit intervals pooled across sessions (n = 2,418 intervals; R2 = 0.386, p = 3.1 × 10−258, OLS regression), (h) Histogram of per-session correlation coefficients between accumulation slopes and behaviorally projected exit times during inter-reward intervals (n = 48 sessions; mean r = −0.325, t = − 12.0, p = 5.8 × 10−16; 23 of 48 individually significant), (i) Histogram of per-session correlation coefficients between accumulation y-intercepts and behaviorally projected exit times during inter-reward intervals (n = 48 sessions; mean r = 0.039, t = 1.38, p = 0.17). (j) The relationship between exit time predicted by the accumulate-to-threshold model and behaviorally projected exit times for all inter-reward intervals (n = 2,622 intervals; R2 = 0.178, p = 2.4 × 10−113, OLS regression).

Session-specific behavioral correlates of patch exit timing.

a−f) Linear regression analysis of potential predictors of patch exit behavior from an example session, a) Leave time from port entry plotted against the time of the last reward relative to entry, b) Leave time from the final reward plotted against the time of the last reward relative to entry, c) Leave time from the final reward plotted against the duration of the final inter-reward interval, d) Leave time from the final reward plotted against the total number of rewards received in the time-investment port during that trial, e) Leave time from the final reward plotted against the reward rate experienced in the time between port entry and the final reward, f) Leave time from the final reward plotted against the reward rate experienced in the time between patch entry and patch exit, g) Distribution of correlation strengths (R2) across all recording sessions for each of the six behavioral predictors analyzed in panels a−f. Each data point represents the R2 value from the linear regression for a single session. Box plots show median, quartiles, and range of correlation strengths.

Session-specific linear regression models predict behavioral exit policy following rewards.

(a) Leave time from the final reward plotted against the time of the final reward relative to port entry for an example session. Data points are separated by context block reward rate with fitted linear regression lines, (b) Predicted leave times from port entry calculated using session-specific linear regression coefficients plotted against observed leave times across all sessions (R2 = 0.75). (c) Comparison of regression line slopes between high and low reward-rate context blocks across sessions (n = 134 sessions, p = 8.5 × 10−8, paired t-test). (d) Comparison of regression line intercepts between high and low reward-rate context blocks across sessions (n = 134 sessions, p = 2.6 × 10−27, paired t-test).

Lick patterns at the context and time-investment ports.

(a) Example single-trial lick pattern during a visit to the context port. Black tick marks indicate individual lick times. Dashed vertical lines mark port entry (green), reward deliveries (blue), and port exit (red), (b) Same as (a) for a visit to the time-investment port, (c) Mean lick rate as a function of time from port entry at the context port, plotted separately for low (green) and high (orange) reward-rate contexts. Shading indicates ±1 SEM across intervals, (d) Same as (c) for the time-investment port. At both ports, licking begins shortly after entry and is sustained throughout the visit. At the context port, lick rate is higher and more sustained in the low reward-rate context, consistent with the longer visits observed in that condition.

Step time distributions and variability scaling in DMS neurons.

(a-d) Histograms showing the distribution of step times for the four example DMS units from Figures 2 and 3. (e) Relationship between the mean step time and coefficient of variation (CV) of step times across all step-like units (n = 898 units). Linear regression line is shown in black (R2 = 0.281, p = 2.7 × 10-66, OLS regression). Later-stepping units tend to have lower temporal variability relative to their mean step time.

Classification and overlap of DMS unit response types during patch-foraging behavior.

Sunburst diagram showing the distribution of response typos among all recorded DMS units (n = 1,849). Stop-like units exhibit detectable stop transitions in firing rate during the post-reward period in the time-investment port. Movement-related units show modulated activity during locomotion into or out of either port, which could be for general travel or port-specific entry/exit responses. Lick-responsive units display firing modulation phase-locked to the lick cycle during continuous licking behavior throughout port occupancy. Context port-related units exhibit activity modulation during the fixed inter-reward intervals in the context port or responses to the depletion cue. Block-indicating units modulate their firing rate during the initial 1.25 seconds following context port entry, before reward delivery, based on the expected reward-rate block from the previous trial. Non-rosponsivo units show no detectable modulation to any measured task variables.

Poisson simulation validates step detection specificity.

(a) SVM classification scatter plot showing real step-like units (black, n = 898), real non-step units (gray, n = 939), and simulated flat-rate Poisson neurons (orange-red, n = 1,627) in CV-MCC feature space. The dashed line indicates the linear SVM decision boundary. Simulated neurons cluster below the boundary in the low-MCC, high-CV region. False positive rate: 1.35% (22/1,627; 95% CI [0.85%, 2.04%]). (b) Step-style histogram comparing MCC distributions across the three populations. Median MCC values: simulated = 0.29, real non-step = 0.34, real step-like = 0.45 (Mann-Whitney U, simulated vs. step-like: p = 2.97 × 10−273). For each of the 1,849 recorded units, a matched homogeneous Poisson process neuron was generated with the same mean firing rate and reward interval structure but zero modulation. Simulated spike trains were processed through the identical classification pipeline (sigmoid fitting, CV and MCC computation, SVM classification). Of 1,849 simulated units, 222 were excluded at the sigmoid fitting stage due to insufficient valid intervals (predominantly very low firing rate units, median 0.42 Hz).

Correlation strength between neural step times and task variables varies by mean step time of DMS units.

Box plots showing the distribution of R2 values quantifying the relationship between individual unit step times and three task variables: projected exit time (time the mouse would wait following each reward based on session-specific behavioral models), time of reward relative to port entry, and context block reward rate. Units are grouped by their mean step time. Individual R2 values for each unit are overlaid as points using a swarm plot distribution to avoid overlap.