Peer review process
Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.
Read more about eLife’s peer review process.Editors
- Reviewing EditorNathan FaivreCentre National de la Recherche Scientifique, Grenoble, France
- Senior EditorMichael FrankBrown University, Providence, United States of America
Reviewer #1 (Public review):
Summary:
A central question in decision-making is whether confidence arises from the same evidence-accumulation process that led to the choice or from a separate second process. The manuscript addresses this question using a random-dot motion (RDM) task, with an initial choice followed by a confidence report, and a time-pressure manipulation on the confidence report. The authors fit a family of four models differing on two dimensions: the source of confidence (a continuation of the choice accumulator vs. a distinct accumulation process) and the stopping rule (time-based vs. boundary-based). The models are fitted to behavior, and then their simulated dynamics are compared with the CPP signal associated with evidence accumulation (not included in the fit). The main methodological contribution is the use of a neural signal to decide between the two boundary-based models, which are nearly indistinguishable behaviorally. The authors conclude that boundary-based stopping rules outperform time-based rules, and that the Boundary-Single model reproduces the certainty-related CPP dynamics better than the Boundary-Distinct model, supporting a single accumulation process for both choice and confidence.
Strengths:
The main strength is methodological: using a neural signal (CPP) as an out-of-sample arbiter between the two boundary-based models (which are nearly equivalent behaviorally). This addresses the model-identifiability problem: when behavior does not distinguish between competing models, a neural signal not included in the fit can provide external evidence.
The work is also thorough and empirically rigorous. The finding that CPP amplitude predicts subsequent certainty ratings several hundred milliseconds before the initial choice response is statistically supported and interesting in its own right, independent of its interpretation. The small-N/many-trials design (2,160 trials per participant) is well suited to capturing change-of-mind trials and is accompanied by thorough model and parameter recovery, with the generating model recovered in most simulations. The study also includes preregistration of the design and planned behavioral and neural analyses, and it replicates a previously reported behavioral pattern.
Weaknesses:
The central conclusion may well be correct, but in my view the current evidence does not fully support it. The behavioral comparison between the competing models did not resolve the issue and in fact showed a slight preference for the distinct-process model, so the weight of the decision falls mainly on the neural comparison.
(1) The neural evidence supports access to pre-choice variation, but does not necessarily establish a single continuous process. The critical difference between the single and distinct models is the presence of trial-to-trial variation in the evidence at choice commitment, which is inherited by the confidence process: the single model preserves it (the confidence process begins from the trial-specific endpoint of the choice DV), while the distinct model does not inherit it (z2 fixed). Thus, the neural test examines whether confidence has access to the state of evidence accumulation before the choice response but does not establish that the same accumulation process must continue seamlessly to determine confidence. The authors also test a model in which a distinct post-choice process is initialized using information from the endpoint of the choice process. However, its failure rules out one specific implementation of information transfer, rather than the broader class of two-stage models in which a distinct confidence process receives a readout of the decision state and may additionally integrate other metacognitive cues.
(2) A broader class of two-stage metacognitive models that combine a decision-state readout with additional cues is not tested. The distinct-process model implemented here captures only a limited subset of possible metacognitive architectures. Its starting point is independent of the choice-DV endpoint, and it remains driven by the available sensory evidence, potentially within a different reference frame. It does not capture second-order accounts in which confidence combines a readout of the decision state with additional cues not explicitly represented in the choice accumulator, such as response time, motor conflict, subjective stimulus clarity or attention. Moreover, the paradigm used in the manuscript provides few independently manipulated information sources that would allow such a process to be identified separately from the choice accumulator.
(3) The neural distinction between models is not quantitatively evaluated. The claim that the Boundary-Single model better reproduces the CPP rests primarily on a visual/qualitative comparison, without a numerical measure of the discrepancy between each model and the neural data. Although the plotted β coefficients (Figure 4D) provide estimates of the neural effects, no scalar summary of model-to-CPP fit is reported. Because the neural comparison carries much of the inferential weight, a quantitative comparison would strengthen the conclusion substantially.
(4) The fitted parameters raise a question about the post-choice process. Confidence responses were very fast (average median confidence RT = 243 ms), with no minimum RT threshold. The Boundary-Single model estimated a post-choice drift rate more than twice the pre-choice rate (1.69 vs. 0.75). In the model, confidence accumulation begins at commitment, before the initial response is executed. The measured confidence RT therefore does not capture the full accumulation window, which also includes the motor delay. Still, the sharp rise in drift rate at commitment requires an explanation. It may reflect stronger weighting of the still-available evidence, as the authors suggest. But it is also consistent with a fast readout of a decision state that was largely set before the initial response. The authors could compare the current model against a readout model, or against an intermediate variant that permits only a brief, bounded period of post-choice accumulation.
(5) The participant-level distribution of model preferences would clarify the comparison. Models were fit separately per participant and condition, but the comparison is summarized as mean BIC (a small average preference for Boundary-Distinct). A mean cannot distinguish two different situations: a consistent, weak preference for one model across all participants, versus a mixture in which some participants clearly favor one architecture and others the opposite. Reporting the distribution of per-participant ΔBIC and the number of participants favoring each model would clarify the result.
Reviewer #2 (Public review):
Summary:
Overall, the authors aimed to provide evidence that clarifies two debates within metacognition research concerning subjective confidence reports:
(1) Does the post-decision confidence report arise from the same process that drives the initial decision, or does a separate, independent process support confidence computation?
(2) How do we stop accumulating evidence for the post-decision confidence report? Is it based on a self-imposed time limit, or on accumulated evidence crossing a boundary?
For the investigation, the authors constructed four models (2 × 2 factorial) to compare each combination of processes to account for random-dot motion tasks data with speed/accuracy manipulations. The models are generally embedded in the drift diffusion model framework, retaining basic parameters such as drift rate, boundary separation, starting point, and non-decision time, while adding linearly collapsing boundaries to model the initial choice. For the single vs. distinct process dimension, the difference lies in whether post-decision evidence accumulation is referenced to the endpoint of the initial decision process or restarts from a new, freely estimated starting point. For the time- vs. boundary-based stopping rule dimension, the key difference is that post-decision evidence accumulation stops either at a deadline sampled from a normal distribution or when the accumulated evidence hits a collapsing boundary.
Based on model comparison, the boundary-based stopping rule clearly outperformed the time-based stopping rule. However, models with the boundary-based stopping rule performed similarly regardless of whether a single or distinct process was used. Here, the authors drew additional insights from EEG recordings during the task, focusing on the centro-parietal positivity (CPP), which has been proposed as a neural correlate of the evidence accumulation process. By simulating evidence accumulation trajectories (with additional assumptions) and comparing the patterns of those trajectories with observed ERP waveforms, the authors argued that the single-process model provided a better match to the CPP findings and was therefore preferred. This was specifically demonstrated by the model's superior ability to match the pre-response CPP amplitude differences conditioned on the post-decision confidence-related variables.
Strengths:
(1) The authors translated existing theories into computational models of decision-making and systematically compared different cognitive processes by assessing model fits to the data. This provides strong evidence supporting the idea that post-decision confidence reports could be better explained by boundary crossing rather than a self-imposed deadline to respond.
(2) Beyond model evidence, an important result is that CPP amplitude predicted confidence before the initial choice was reported, which is a unique prediction of the single-process model. The use of EEG as an independent validation measure provided additional evidence in favour of this model.
(3) Combining points 1 and 2, this study successfully addressed the two key debates with solid evidence to favour one theory over another.
(4) Another strength of this study is the data quality. The high number of trials provided a strong foundation for model inference as well as ERP analysis. The experiment also contained a speed-accuracy manipulation to evaluate model performance across diverse situations.
Weaknesses:
I have two main concerns around the modelling work and neural analyses, which in my opinion could have limited the interpretation of the findings. My responses here will be lengthier, but this reflects the nature of the modelling work rather than implying stronger criticisms than those suggested by the strengths discussed above.
(1) There are a few assumptions in the models that lack psychologically meaningful interpretations, and this study placed more effort into model comparison while lacking discussion of the cognitive processes inferred from parameter estimates.
To start, I think some of the parameterisations were not properly justified. For the boundary models, it is not very clear why the upper and lower boundaries were different and collapsed at different rates for confidence decisions, given that a single boundary parameter and collapse rate were used for the initial decision. This allows more flexible shifts in the model's predictions of confidence ratings without strong justification. Specifically, it is unclear why the boundary-single model has such an implementation while the boundary-distinct model was only equipped with one boundary parameter (a2, compared to a2up and a2down).
Similarly, the inclusion of metacognitive noise creates another layer of flexibility in the predictions of confidence ratings. In most existing evidence accumulation models with a diffusion process, noise comes from two sources: within-trial noisy evidence accumulation and across-trial variability (e.g., drift rate variability). Beyond these, such models almost always assume that the decision is made deterministically once the evidence reaches a specific boundary. The inclusion of metacognitive noise here sounds more like a noisy decision-to-action mapping.
I also have similar doubts about allowing the non-decision time parameter for confidence accumulation in the distinct model to be negative. The authors argued that confidence accumulation may begin during initial evidence accumulation. However, this is a flawed implementation, as the non-decision time was simply added to the evidence accumulation time rather than being incorporated within it. Allowing negative non-decision times may achieve similar predictions, but it is ad hoc.
The inclusion of a collapsing boundary mechanism in the post-decision confidence accumulator helped the model reach more diverse levels of accumulated evidence and ultimately improved predictions of confidence ratings. However, no strong argument is presented for this implementation beyond the observation that the model performs worse without it. The collapsing boundary mechanism has traditionally been interpreted as reflecting a sense of urgency. For the boundary models, I noticed that the collapse rate of the upper boundary differed significantly between speed and accuracy conditions, which is consistent with the urgency interpretation. Overall, I would like to see more discussion of the specific model mechanisms included by the authors, interpreted in light of parameter estimates.
(2) While the ERP findings provided external evidence and validation of the modelling results, I find the simulation practices not particularly useful and potentially misleading for naive readers. Specifically, the authors attempted to draw a parallel between patterns of simulated evidence accumulation traces and observed CPP waveforms. While the CPP has received support as a correlate of the evidence accumulation process, the DDM is by no means a neural model capable of generating predictions of neural observations. To my understanding, the superior fit of the boundary-single model was primarily due to the fact that pre-response CPP amplitude predicts post-decision confidence ratings. Therefore, as the boundary-distinct model did not connect the two phases of evidence accumulation, it would fail to account for this observation. I think this point could be clearly demonstrated without the need to introduce additional assumptions into the model simulations in order to directly compare averaged trajectories with averaged ERP waveforms. While the authors did not explicitly claim otherwise, this approach creates an illusion that the model can mechanistically account for ERP data. I would like the authors to provide explicit clarification on this point.
Appraisal:
Overall, the authors have provided solid evidence in support of their research aims. The findings contribute to longstanding debates with insights from model mechanisms and neural findings that should not be overlooked by future studies on this topic. This study also offers a good starting point for future model development and refinement in broader contexts of confidence reporting, such as paradigms involving simultaneous initial decisions and confidence judgements. The high quality of the behavioural and EEG data will make a valuable contribution to future research.