Non-decision time: the elephant in the executive control room

  1. Cardiff University Brain Research Imaging Centre (CUBRIC), School of Psychology Cardiff University, Cardiff, United Kingdom

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, public reviews, and a provisional response from the authors.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Emilio Salinas
    Wake Forest University School of Medicine, Winston-Salem, United States of America
  • Senior Editor
    Tamar Makin
    University of Cambridge, Cambridge, United Kingdom

Reviewer #1 (Public review):

This interesting manuscript challenges the current interpretation of the well-established stop signal reaction time task (SSRT), commonly used in many research areas. SSRT has been traditionally thought of as primarily a measure of inhibitory control. In this work, the authors argue that this is influenced significantly by sensory and motor transmission times, and that these low-level processes may systematically confound SSRT estimates both in individual groups and also in many clinical populations.

Conceptually, this raises an important and significant question regarding the construct validity of one of the more widely used behavioral measures of response inhibition. The authors provide a clear theoretical framing and importantly address the overlooked assumptions in SSRT modelling, the underexplored source of variability.

Further, direct evidence separating peripheral sensory and motor contributions from central inhibitory processes is certainly needed, as well as a more balanced interpretation of prior methodological refinements in the field. Is a correlation between T0 and SSRT sufficient to conclude that SSSRT may be predominantly driven by peripheral delays? What proportion of SSRT variance remains unexplained after one accounts for T0? If T0 and inhibitory processes share neural processing speed, does that mean they covary? If one corrects T0, do the group differences become smaller or perhaps disappear?

Furthermore, how robust is the T0 estimate in noisy environments of all sorts and especially across modalities (visual and motor)? The authors argue that this relationship is clearer in "good quality data"; how do you define that, and what happens with noisy data? For instance, the authors refer to a range of clinical disorders such as ADHD and PD where noise is abundantly present, partly due to the disease itself or due to treatment.

In conclusion, this is certainly a thought-provoking and potentially influential contribution in the literature that raises important questions about the interpretation of the stop signal reaction time task.

Reviewer #2 (Public review):

Continuing their work distinguishing sensory latencies of "cognitive" processes, the authors turn their attention to "response inhibition". The "square quotes" are being used to highlight how this manuscript aims to challenge previous descriptions of performance data and inferred computational processes. The authors assert that previous descriptions of the measure known as "stop signal reaction time" (SSRT) are flawed because they did not account for sensory latencies empirically or theoretically.

Enthusiasm for the manuscript cannot be high in light of many weaknesses countering the possible strengths. Strengths include offering an opportunity to more carefully characterize the quantity SSRT and a specific empirical approach offered to the research community. However, these strengths are countered by the following structural, theoretical, and empirical weaknesses:

As announced by the elephant in the title, the writing could be described as excessively polemical. However, the characterization and interpretation of previous empirical and theoretical work is disputable.

The major theoretical claim regarding sensory delays inherent in SSRT is not novel. The authors assert, "...this corpus of work may have been misinterpreted because the SSRT is systematically influenced by low level sensory and motor transmission times, arguably more so than by inhibition or cognitive processes." This was certainly recognized by Logan and Cowan in their original work. They wrote, "An act of control, like any other act, must take time. The theory provides methods for measuring the latency of control even when the act of control is not directly observable." (page 298) Also, "... the estimate of stop-signal reaction time includes the latency of the internal response to the stop signal and the duration of the ballistic process." (page 316-317). Moreover, subsequent computational and empirical work, some noted by the authors, has distinguished the sensory encoding interval from the interval during which the STOP process interrupts the GO process.

The theoretical suggestion that an accounting for sensory delays undermines the functional interpretation of SSRT mischaracterizes the original literature. For example, in the Abstract the authors write "Sensory and motor contributions must be ruled out before linking SSRT results to inhibition or cognition". The original Logan and Cowan theory was about what happens at the end of SSRT, and that was described only as an "act of control", in perfectly positivist fashion. For example, Logan and Cowan wrote, "Estimates of stop-signal reaction time provide a measure of the latency of control." (page 315). Thus, the authors are misstating what was meant originally by SSRT. In addition, the authors offer no specific or formal definition to specify what they mean by "inhibition or cognitive processes".

Confidence in the new empirical conclusions of the manuscript must be low because the new performance data are of questionable quality. The first issue is that the stopping accuracy (or inhibition functions in original terminology) shown in Figure S3 is very problematic for the interpretation of the authors' empirical work in this manuscript. There are two problems. First, these plots should span from nearly 0% to nearly 100%. It is not possible to resolve the span of each individual in the figure, but it is clear that many, if not most, in both the Manual and Saccadic data span just 20-30%. Second, the plots should span the 50% success value. It is clear that the maximum or minimum values for many participants do not reach the 50% value. These two problems indicate that many (most?) participants were not really sensitive to the stop signal.

The second issue concerns the pattern of response times (RTs) on "ignore" trials. The authors portray performance as exemplifying a "pause-then-go" strategy. This is not uncommon, but it is not the only way participants perform. Many participants across multiple studies of selective stimulus stopping produce RTs on "Ignore" trials essentially indistinguishable from RTs on no-stop trials. The authors must acknowledge and account for such individual variability. In fact, the "T_s" value is measured by the difference in distributions of RT on no-signal and ignore trials. If these distributions are not different, then the measurement and interpretation of this quantity is questionable.

Related, the distributions of RT on stop trials, particularly for saccade responses, are portrayed with a second mode in the schematic illustrations and clearly peaking at SSRT in Figure S1. This second mode is not observed in other saccade stop signal studies. This indicates that the participants in this study were in a peculiar mode of performance.

Finally, given the pivotal role of measures of differences of RT distributions and the pronounced variation of stopping accuracy (Figure S3), the authors must show the distributions for all of their new participants. The authors' claim to higher resolution obliges them to reveal every step of analysis.

In its current form, this manuscript is unlikely to change the thinking of modelers or practitioners of the stop signal task.

Reviewer #3 (Public review):

Summary:

Statham and colleagues test an assumption underpinning a very large literature: that the stop-signal reaction time (SSRT) indexes the speed or efficacy of top-down inhibitory control. They argue instead, and support their claims with a total of eight datasets, that SSRT is substantially occupied by visuomotor deadtime (i.e., incompressible sensory and motor delays common to all visually guided responses), which varies across individuals, conditions and populations in ways that mimic effects usually attributed to inhibitory control. They propose two remedies: subtracting an independent estimate of visuomotor deadtime (T₀) from SSRT, and a new index, the selective stopping delay (ΔT), from the stimulus-selective stopping task.

Strengths:

The paper's principal strength is the combination of these components. That SSRT must contain peripheral delays is not itself new, as the authors point out (Boucher et al., 2007; Salinas and Stanford, 2013; Bompas et al., 2020). What is new is the quantification of the problem at scale, across seven archival datasets and a preregistered replication, together with the demonstration that T₀ can be recovered from existing stop-task data. That is important, as it provides a diagnostic that can be applied to data already collected. The authors' offer to assist others in doing so is exemplary. The supplementary analyses of trial numbers and participant pooling are very useful, and the paper provides important sanity checks, notably confirming that stop and ignore signals produce indistinguishable initial interference before pooling them.

Weaknesses:

The evidence for the central claim is strong but presented in a way that overstates it. Figure 2 reports 85% and 80% shared variance between SSRT and T₀, but these pool across datasets and, more critically, across response modality: manual and saccadic estimates from the same participants are plotted together with a single regression line through both. Because manual and saccadic deadtimes differ by roughly 130 ms, the resulting correlation largely reflects a between-condition difference rather than covariation among individuals. The numbers that speak to individual differences are more modest (40% for manual responses; 7% for saccades). The manual result is convincing and consequential; the saccadic result is not, and the explanation in terms of restricted range, while plausible, is offered after the fact and is directly testable by reporting the reliability of saccadic T₀ or correcting the correlation for attenuation. This limitation is arguably good news for the paper's practical message, since it implies saccadic measures are relatively protected, but the manuscript should make clear (including in the abstract) that the strong individual-differences case rests on the manual data, where motor execution delay is the main driver.

A related point concerns interpretation rather than analysis. Since SSRT is, on the authors' own account, approximately the sum of T₀ and a decision-related component, covariation between the two is expected on structural grounds; the preregistered correlation with reaction times from separate speeded blocks mitigates this, but the finding is less surprising than its current framing implies. What would determine whether past conclusions must be revised is not whether SSRT correlates with T₀ across individuals, but whether the decision-related component tracks the independent variable in any given study. The alcohol reanalysis could be a test case for this: the authors show that alcohol raises T₀ commensurately with SSRT and conclude the effects are "consistent with these effects being fully driven by visuomotor delays," yet (unless I missed something) they do not report the corrected measure for these data, while they do so for signal contrast and response modality. Running that analysis, and stating plainly what Campbell et al. (2017) would have concluded under the proposed treatment, would be an important demonstration.

The case for ΔT is the least developed part of the paper. ΔT is a difference between two independently estimated, individually noisy quantities, extracted by a non-trivial procedure (see also below), and no reliability estimates are reported for T₀, TS or ΔT. This would be possible based on the two-session design (and the group has prior work on the reliability of cognitive control measures). This matters because the argument that ΔT is superior rests, to some extent, on null findings: ΔT does not correlate with SSRT, with stopping accuracy, or with the differential response to stop and ignore trials. These null correlations are interpreted as freedom from confounds, but an unreliable measure would produce the same pattern, and the seven participants with implausible negative ΔT values indicate that noise is not negligible.

In addition, the subjective correction of dip onsets ("Departure points were visually inspected and adjusted if it was deemed that the algorithm had placed them in inappropriate places"), which is critical to the paper's central measurement, should be blinded to condition or show inter-rater agreement. Since T₀ and TS are compared across conditions and ΔT is their difference, this introduces researcher degrees of freedom.

This is particularly critical when a dip is not easy to extract. Figure 3 depicts an idealized ignore-trial distribution with a clean, deep dip. Real distributions are unlikely to look like this, and dip depth should depend on the behavioral relevance and salience of the ignored event; published work on rapid manual inhibition indicates that dips to behaviorally irrelevant events can be very shallow. Since ΔT is extractable only where the dip is resolvable, the generality of the method can be questioned. Ideally, the empirical distributions underlying every dataset analyzed should be shown to alleviate this concern.

One uncontrolled procedural difference also deserves comment. Corrective feedback about stopping too often or stopping too rarely was given after manual blocks only; saccadic blocks received none, and fixed rather than staircased delays were used throughout. Since the manual-saccadic contrast carries much of the argument, and saccadic blocks yielded both lower stopping accuracy (36% vs 52%) and far more exclusions (8 vs 1 of 37), this asymmetry offers an alternative to the interpretation that saccades are simply harder to inhibit.

A final point concerns the comparison between response modalities. Raw SSRT suggests that saccadic inhibition is faster than manual (174 vs 266 ms), while both proposed corrections reverse this, with SSRT−T₀ and ΔT each indicating that saccadic inhibition is slower (the latter consistently across nearly every participant). This is one of the clearest illustrations of the paper's thesis, but it is not taken up in the discussion, which returns to modality only to note that saccadic T₀ varies little (the one reference to variation across action modalities appears in the modeling section, without stating its direction). It would also benefit from a caveat. Both corrected measures subtract the same T₀, and manual and saccadic T₀ differ by roughly 130 ms, so the two do not corroborate one another independently (TS is itself longer for manual responses, and yields a shorter ΔT only once the larger manual T₀ is removed). The accuracy of the subtraction therefore matters here: if the manual regression slope of 0.75 reflects sub-additivity rather than attenuation, subtracting the full T₀ would overcorrect manual responses more than saccadic ones, and could produce the reversal on its own.

These concerns qualify rather than undermine the contribution. The core observation is robust, the diagnostic is practical and immediately applicable, and the case that a large body of work requires re-examination is well made. If the corrected analyses are carried through on the datasets already in hand, this will be an important paper for anyone who uses the stop-signal task.

Author response:

We are grateful to the editor and reviewers for providing their time and expertise in the assessment of this article. We are glad that the overall evaluation is supportive.

Reviewer #1 (Public review):

This interesting manuscript challenges the current interpretation of the well-established stop signal reaction time task (SSRT), commonly used in many research areas. SSRT has been traditionally thought of as primarily a measure of inhibitory control. In this work, the authors argue that this is influenced significantly by sensory and motor transmission times, and that these low-level processes may systematically confound SSRT estimates both in individual groups and also in many clinical populations.

Conceptually, this raises an important and significant question regarding the construct validity of one of the more widely used behavioral measures of response inhibition. The authors provide a clear theoretical framing and importantly address the overlooked assumptions in SSRT modelling, the underexplored source of variability.

Further, direct evidence separating peripheral sensory and motor contributions from central inhibitory processes is certainly needed, as well as a more balanced interpretation of prior methodological refinements in the field. Is a correlation between T0 and SSRT sufficient to conclude that SSSRT may be predominantly driven by peripheral delays? What proportion of SSRT variance remains unexplained after one accounts for T0? If T0 and inhibitory processes share neural processing speed, does that mean they covary? If one corrects T0, do the group differences become smaller or perhaps disappear?

We thank the reviewer for their positive assessment of our work. We agree all these questions are important. Our revision will provide repeatability estimates for all our indices, and use the repeatability of SSRT and T0 to estimate the proportion of SSRT variance that remains unexplained. We will mention that it is theoretically possible that shared neural processing speed between T0 and inhibitory processes contributes to the reported effect but, since the neural pathways are anatomically largely distinct, the available cognitive neuroscience literature suggests such contribution is unlikely to be strong enough to explain the observed relationship. The impact of correcting group differences in SSRT for T0 will entirely depend on where differences in SSRT between these groups come from. If it comes exclusively from peripheral delays, then the effect may indeed disappear. If it is central, then stronger differences may be revealed after correction. We appreciate these are key questions we would like answers to already, but identifying and reanalysing suitable datasets to answer it will be the purpose of future papers.

Furthermore, how robust is the T0 estimate in noisy environments of all sorts and especially across modalities (visual and motor)? The authors argue that this relationship is clearer in "good quality data"; how do you define that, and what happens with noisy data? For instance, the authors refer to a range of clinical disorders such as ADHD and PD where noise is abundantly present, partly due to the disease itself or due to treatment.

By good quality data, we mean enough trials at the optimal RT and SOA combinations, and an adequate preprocessing pipeline that excludes non-standard trials (poor fixation, pre-emptive responses, large undershoot …). Participants with increased intra-individual variance or low compliance will need more trials overall to get enough trials around divergence times for these to be accurately estimated. Supplementary figure 5 illustrates how low trial numbers lead to an overestimation of T0, which can be mitigated by pooling across participants. We expect noisy data to have a similar effect, although its impact may differ across clinical conditions based on where the additional variability comes from. We shall have more clarity on the reliability of T0 in clinical populations once relevant datasets have been reanalysed, and use this to define constraints for future data collection. Again, this will need to wait for future papers.

In conclusion, this is certainly a thought-provoking and potentially influential contribution in the literature that raises important questions about the interpretation of the stop signal reaction time task.

We thank the reviewer for their in-depth and thoughtful comments and suggestions

Reviewer #2 (Public review):

Continuing their work distinguishing sensory latencies of "cognitive" processes, the authors turn their attention to "response inhibition". The "square quotes" are being used to highlight how this manuscript aims to challenge previous descriptions of performance data and inferred computational processes. The authors assert that previous descriptions of the measure known as "stop signal reaction time" (SSRT) are flawed because they did not account for sensory latencies empirically or theoretically.

Enthusiasm for the manuscript cannot be high in light of many weaknesses countering the possible strengths. Strengths include offering an opportunity to more carefully characterize the quantity SSRT and a specific empirical approach offered to the research community. However, these strengths are countered by the following structural, theoretical, and empirical weaknesses:

As announced by the elephant in the title, the writing could be described as excessively polemical. However, the characterization and interpretation of previous empirical and theoretical work is disputable.

The major theoretical claim regarding sensory delays inherent in SSRT is not novel. The authors assert, "...this corpus of work may have been misinterpreted because the SSRT is systematically influenced by low level sensory and motor transmission times, arguably more so than by inhibition or cognitive processes." This was certainly recognized by Logan and Cowan in their original work. They wrote, "An act of control, like any other act, must take time. The theory provides methods for measuring the latency of control even when the act of control is not directly observable." (page 298) Also, "... the estimate of stop-signal reaction time includes the latency of the internal response to the stop signal and the duration of the ballistic process." (page 316-317). Moreover, subsequent computational and empirical work, some noted by the authors, has distinguished the sensory encoding interval from the interval during which the STOP process interrupts the GO process.

We agree that our manuscript should acknowledge that Logan and Cowan (1984) explicitly stated that internal and ballistic delays contribute to SSRT and thank the reviewer for highlighting the need to clarify the relationship between our work and the original Logan and Cowan framework. We will clarify that the novelty of our claim is not that peripheral delays contribute to SSRT, which is a logical necessity, but that differences in peripheral delays (across conditions or people) contribute to differences in SSRT, sometimes to a large extent. Such differences in SSRT are very widely assumed to reflect inhibitory control in the large corpus of work that followed this initial literature. This corpus has essentially ignored the message about stimulus processing and ballistic delays and their implications for individual differences or changes across conditions. Therefore, we maintain that SSRT differences may have been widely misinterpreted. We reference the articles where these implications were clearly spelled out: these empirical and modelling studies were based on a few monkeys or human participants, and therefore unable to provide the large-scale demonstration we provide here.

The theoretical suggestion that an accounting for sensory delays undermines the functional interpretation of SSRT mischaracterizes the original literature. For example, in the Abstract the authors write "Sensory and motor contributions must be ruled out before linking SSRT results to inhibition or cognition". The original Logan and Cowan theory was about what happens at the end of SSRT, and that was described only as an "act of control", in perfectly positivist fashion. For example, Logan and Cowan wrote, "Estimates of stop-signal reaction time provide a measure of the latency of control." (page 315). Thus, the authors are misstating what was meant originally by SSRT. In addition, the authors offer no specific or formal definition to specify what they mean by "inhibition or cognitive processes".

We thank the reviewer for pointing out that their original approach was mechanistically agnostic, which we will explicitly clarify in revision. We will remove the words “top-down” from our 4th sentence and reword the quoted sentence into "Changes in peripheral delays must be ruled out before linking changes in SSRT to inhibition or cognition". It remains the case that many hundreds of studies have since interpreted “ability to inhibit” and “latency of control” as specific to inhibition and control, and therefore have assumed that changes in SSRT directly reflect an inhibitory cognitive process. We agree that we do not currently offer a definition of what we mean by cognitive processes, except that they do not include incompressible sensory and motor delays. Based on the reviewer’s clarification, this common shortcut in the literature appears inconsistent with the spirit of the initial work, which we seek to rectify.

Confidence in the new empirical conclusions of the manuscript must be low because the new performance data are of questionable quality. The first issue is that the stopping accuracy (or inhibition functions in original terminology) shown in Figure S3 is very problematic for the interpretation of the authors' empirical work in this manuscript. There are two problems. First, these plots should span from nearly 0% to nearly 100%. It is not possible to resolve the span of each individual in the figure, but it is clear that many, if not most, in both the Manual and Saccadic data span just 20-30%. Second, the plots should span the 50% success value. It is clear that the maximum or minimum values for many participants do not reach the 50% value. These two problems indicate that many (most?) participants were not really sensitive to the stop signal.

We acknowledge our inhibition functions are narrow, but we do not believe this undermines the main conclusions, for several reasons. On the question of sensitivity to the stop signal, average spans for inhibition functions after participant exclusions were 37% for manual and 29% for saccades. This limited span is mainly attributable to our fixed SOA design, which was a necessary feature of a direct comparison between manual and saccadic behaviours. Figure S3 covers only 80 ms spread of SOA. The slopes are commensurate with most other studies, which cover much wider differences in SOA. If one selects the central 100 ms (where the slopes are steepest) from the figures in most previous papers, one will find stopping accuracy changes of around 30%. Therefore, sensitivity is similar.

On the question of some functions not crossing 50%, the correlations between SSRT and T0 remain the same if we only keep those participants who crossed the 50% point (R(26)=0.64 for manual, R(11)=0.4 for saccades, same statistical significance levels). We will additionally rerun our SSRT and T0 correlation with stopping accuracy as a covariate. As stopping accuracy affects SSRT but not T0, it is unlikely to drive our results.

The second issue concerns the pattern of response times (RTs) on "ignore" trials. The authors portray performance as exemplifying a "pause-then-go" strategy. This is not uncommon, but it is not the only way participants perform. Many participants across multiple studies of selective stimulus stopping produce RTs on "Ignore" trials essentially indistinguishable from RTs on no-stop trials. The authors must acknowledge and account for such individual variability. In fact, the "T_s" value is measured by the difference in distributions of RT on no-signal and ignore trials. If these distributions are not different, then the measurement and interpretation of this quantity is questionable.

We agree that the issue raised by the reviewer would be important if no difference were present between the distributions. In our dataset, however, all participants showed a measurable distributional difference. We believe the difference in perspective that ‘many participants…produce RTs on ignore trials essentially indistinguishable from RTs on no-stop trials’ may be attributed to differences in the way we analyse data (RT distributions versus mean RT).

Nearly all participants in our final sample had mean RTignore – RTgo > 10 ms (significant at the individual level), except for 2 in the manual condition (and none for saccades). Following Bisset & Logan (2014), a lack of clear mean RTignore - RTgo difference in these 2 participants might have been interpreted as reflecting a different strategy. However, all our participants showed clear dips between go and ignore RT distributions when locked on signal onset, and these two manual participants were no exception. Therefore, accounting for response probability and RT at each SOA in our distributional analysis revealed clear ignore versus go differences, masked when relying on mean RT. The lack of mean RT difference for these two participants in the manual modality had no impact on our main hypothesis testing because T0 was extracted by comparing signal-absent and signal-present trials (pooling ignore and stop), while TS was extracted by comparing ignore and stop (not ignore and go).

In terms of interpretation and whether participants employ a pause-then-go strategy, we understand performance in the selective stopping task as reflecting a combination of automatic activation and interference, and endogenous activation and inhibition. Our interpretation is that the pause component primarily reflects automatic interference triggered by stimulus onset, although strategic factors may also contribute in some circumstances. Individual differences in mean RTignore - RTgo could reflect both automatic interference and endogenous pausing (both of which would increase the difference between ignore and go trials), as well as subsequent failing to go on ignore trials (omissions, which would decrease differences by removing longer latency responses from ignore distributions just as in stop signal distributions). Although some of this can be described as strategic, some won’t be, and we therefore refrain from inferring strategy based on mean RTignore – RTgo.

Related, the distributions of RT on stop trials, particularly for saccade responses, are portrayed with a second mode in the schematic illustrations and clearly peaking at SSRT in Figure S1. This second mode is not observed in other saccade stop signal studies. This indicates that the participants in this study were in a peculiar mode of performance.

The second mode indicates that participants occasionally ignore the stop signal, which is why it peaks at the same latency as the rebound for the ignore distribution. These are not unusual, in particular in selective stopping designs, but are not as easily seen on cumulative functions, which are the standard way of plotting the results in this field.

Finally, given the pivotal role of measures of differences of RT distributions and the pronounced variation of stopping accuracy (Figure S3), the authors must show the distributions for all of their new participants. The authors' claim to higher resolution obliges them to reveal every step of analysis.

We will save figures showing the individual distributions in the OSF folder. Note that these figures can be produced by running the code we shared, so each step is already fully transparent, but we will create tidy versions that also highlight manually corrected indices.

In its current form, this manuscript is unlikely to change the thinking of modelers or practitioners of the stop signal task.

We thank the reviewer for their in-depth and thoughtful comments and suggestions, so that the paper can be revised to address the concerns.

Reviewer #3 (Public review):

Summary:

Statham and colleagues test an assumption underpinning a very large literature: that the stop-signal reaction time (SSRT) indexes the speed or efficacy of top-down inhibitory control. They argue instead, and support their claims with a total of eight datasets, that SSRT is substantially occupied by visuomotor deadtime (i.e., incompressible sensory and motor delays common to all visually guided responses), which varies across individuals, conditions and populations in ways that mimic effects usually attributed to inhibitory control. They propose two remedies: subtracting an independent estimate of visuomotor deadtime (T₀) from SSRT, and a new index, the selective stopping delay (ΔT), from the stimulus-selective stopping task.

Strengths:

The paper's principal strength is the combination of these components. That SSRT must contain peripheral delays is not itself new, as the authors point out (Boucher et al., 2007; Salinas and Stanford, 2013; Bompas et al., 2020). What is new is the quantification of the problem at scale, across seven archival datasets and a preregistered replication, together with the demonstration that T₀ can be recovered from existing stop-task data. That is important, as it provides a diagnostic that can be applied to data already collected. The authors' offer to assist others in doing so is exemplary. The supplementary analyses of trial numbers and participant pooling are very useful, and the paper provides important sanity checks, notably confirming that stop and ignore signals produce indistinguishable initial interference before pooling them.

Weaknesses:

The evidence for the central claim is strong but presented in a way that overstates it. Figure 2 reports 85% and 80% shared variance between SSRT and T₀, but these pool across datasets and, more critically, across response modality: manual and saccadic estimates from the same participants are plotted together with a single regression line through both. Because manual and saccadic deadtimes differ by roughly 130 ms, the resulting correlation largely reflects a between-condition difference rather than covariation among individuals. The numbers that speak to individual differences are more modest (40% for manual responses; 7% for saccades). The manual result is convincing and consequential; the saccadic result is not, and the explanation in terms of restricted range, while plausible, is offered after the fact and is directly testable by reporting the reliability of saccadic T₀ or correcting the correlation for attenuation. This limitation is arguably good news for the paper's practical message, since it implies saccadic measures are relatively protected, but the manuscript should make clear (including in the abstract) that the strong individual-differences case rests on the manual data, where motor execution delay is the main driver.

We will add separate regression lines and R-values for manual and saccadic modalities on Fig.2A and an inset showing the variance only driven by individual differences across all archival data (i.e. z-scored per condition and datasets, R(215)=0.43, p<0.001). We will also state more explicitly that the overall correlation may not be the relevant one for researchers specifically interested in individual differences.

While a large portion of SSRT literature is about individual differences, there are also many studies about differences between conditions, including comparing different response modalities. Therefore, it is a general question whether differences of any kind in SSRT reflect differences in inhibitory control or differences in sensory-motor delays. At a conceptual level, most users of the SSRT are intending to measure control, and have a conceptual model in which control is separate from the modality of response or the exact characteristics of stimulus delivery. Thus, it is important to point out that their measure of control is very much dependent on these things, and in fact to a much larger degree than the more subtle individual differences, group differences or conditions of interest.

We agree that repeatability is critical for interpreting null results and will provide split-half repeatability for all our indices, including T0 and SSRT, and use these to correct their correlations. We thank the reviewer for this suggestion.

A related point concerns interpretation rather than analysis. Since SSRT is, on the authors' own account, approximately the sum of T0 and a decision-related component, covariation between the two is expected on structural grounds; the preregistered correlation with reaction times from separate speeded blocks mitigates this, but the finding is less surprising than its current framing implies. What would determine whether past conclusions must be revised is not whether SSRT correlates with T0 across individuals, but whether the decision-related component tracks the independent variable in any given study. The alcohol reanalysis could be a test case for this: the authors show that alcohol raises T0 commensurately with SSRT and conclude the effects are "consistent with these effects being fully driven by visuomotor delays," yet (unless I missed something) they do not report the corrected measure for these data, while they do so for signal contrast and response modality. Running that analysis, and stating plainly what Campbell et al. (2017) would have concluded under the proposed treatment, would be an important demonstration.

We fully agree that the presence of a correlation is indeed entirely expected and obvious in our own account, as conveyed early on in the manuscript (“From Fig. 1D, it seems clear that SSRT and T0 are inevitably connected”). We agree the main question is what is left for SSRT to explain. We will update our analysis of the Campbell et al. (2017) alcohol study as suggested. Future work can then focus on other “independent variables”.

The case for ΔT is the least developed part of the paper. ΔT is a difference between two independently estimated, individually noisy quantities, extracted by a non-trivial procedure (see also below), and no reliability estimates are reported for T0, TS or ΔT. This would be possible based on the two-session design (and the group has prior work on the reliability of cognitive control measures). This matters because the argument that ΔT is superior rests, to some extent, on null findings: ΔT does not correlate with SSRT, with stopping accuracy, or with the differential response to stop and ignore trials. These null correlations are interpreted as freedom from confounds, but an unreliable measure would produce the same pattern, and the seven participants with implausible negative ΔT values indicate that noise is not negligible.

We fully agree with all this and will update the wording surrounding the lack of correlation between ΔT and the other measures in light of its repeatability

In addition, the subjective correction of dip onsets ("Departure points were visually inspected and adjusted if it was deemed that the algorithm had placed them in inappropriate places"), which is critical to the paper's central measurement, should be blinded to condition or show inter-rater agreement. Since T0 and TS are compared across conditions and ΔT is their difference, this introduces researcher degrees of freedom.

All divergence times (T0, T0,stop, T0,ignore and TS) were confirmed by two of the authors. Each index for each modality is plotted on a separate figure (showing all individuals). It is technically easy to compare, say, T0 and TS, for one individual, but we refrained from doing this (and indeed ended up with many T0 > TS). We agree that blinding and inter-rater reliability are important safeguards and that our current analysis fell short of this. We will explore whether this can be done retrospectively, and report on this exercise alongside guidance on criteria used for manual corrections.

This is particularly critical when a dip is not easy to extract. Figure 3 depicts an idealized ignore-trial distribution with a clean, deep dip. Real distributions are unlikely to look like this, and dip depth should depend on the behavioral relevance and salience of the ignored event; published work on rapid manual inhibition indicates that dips to behaviorally irrelevant events can be very shallow. Since ΔT is extractable only where the dip is resolvable, the generality of the method can be questioned. Ideally, the empirical distributions underlying every dataset analyzed should be shown to alleviate this concern.

We will make figures available in the OSF folder with individual distributions that supported the extraction of each index, flagging those that got manually corrected. This will make apparent that the vast majority of dips were very clear, for both T0 and TS. Unclear dips led to missing indices, and were therefore excluded from our hypothesis testing.

One uncontrolled procedural difference also deserves comment. Corrective feedback about stopping too often or stopping too rarely was given after manual blocks only; saccadic blocks received none, and fixed rather than staircased delays were used throughout. Since the manual-saccadic contrast carries much of the argument, and saccadic blocks yielded both lower stopping accuracy (36% vs 52%) and far more exclusions (8 vs 1 of 37), this asymmetry offers an alternative to the interpretation that saccades are simply harder to inhibit.

Indeed, blockwise feedback would have been hard to implement reliably for saccades. As manual and saccadic blocks were interleaved, our hope was that participants could use the feedback received for manual to adjust their strategy for both modalities. We will check how often the feedback was triggered for manual blocks and, if more than negligible, we will note the reviewer’s suggestion as a possible driver for modality differences.

A final point concerns the comparison between response modalities. Raw SSRT suggests that saccadic inhibition is faster than manual (174 vs 266 ms), while both proposed corrections reverse this, with SSRT−T0 and ΔT each indicating that saccadic inhibition is slower (the latter consistently across nearly every participant). This is one of the clearest illustrations of the paper's thesis, but it is not taken up in the discussion, which returns to modality only to note that saccadic T0 varies little (the one reference to variation across action modalities appears in the modeling section, without stating its direction). It would also benefit from a caveat. Both corrected measures subtract the same T0, and manual and saccadic T0 differ by roughly 130 ms, so the two do not corroborate one another independently (TS is itself longer for manual responses, and yields a shorter ΔT only once the larger manual T0 is removed). The accuracy of the subtraction therefore matters here: if the manual regression slope of 0.75 reflects sub-additivity rather than attenuation, subtracting the full T0 would overcorrect manual responses more than saccadic ones, and could produce the reversal on its own.

We agree with all this. We will use the repeatability of SSRT and T0 to disattenuate the slopes and consider alternatives to subtraction to correct for visuo-motor deadtime. Before we can elaborate on the modality effect on the speed of inhibition, we need to simulate the effect of motor variability on T0, TS and SSRT. If motor noise affects TS or SSRT more or less than it affects T0, this will affect our conclusions. We will explore this issue in the revision and report any analyses that bear on the robustness of the modality effect.

These concerns qualify rather than undermine the contribution. The core observation is robust, the diagnostic is practical and immediately applicable, and the case that a large body of work requires re-examination is well made. If the corrected analyses are carried through on the datasets already in hand, this will be an important paper for anyone who uses the stop-signal task.

We thank the reviewer for their in-depth and thoughtful comments and suggestions

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation