Introduction

Metacognition refers to the ability to reflect on one’s own cognition and behaviour (Fleming, 2024) and encompasses constructs such as metacognitive sensitivity (recognising whether individual decisions and actions are correct) and metacognitive bias (overall confidence in performance). These facets of metacognition are crucial for correcting mistakes, improving performance, and deciding whether to engage in activities (Bénon et al., 2024; Fleming & Dolan, 2012). Distortions in metacognition are thought to play a role in several mental health conditions (Hoven et al., 2019), with under-confidence emerging as a central characteristic of depression. Under-confidence has consistently been reported in case-control comparisons of depressive disorders (Hoven et al., 2019) and those scoring higher on an anxious-depression dimension in the general population (Benwell et al., 2022; Hoven et al., 2023; Katyal et al., 2025; Rouault et al., 2018; Seow et al., 2025; Seow & Gillan, 2020).

Recent longitudinal studies have extended these cross-sectional findings showing that depression symptoms reduce over time as confidence increases, both following therapy (Fox et al., 2023) or naturalistically after a two-year follow-up (Agnoli et al., 2023). These findings suggest that lower confidence may be less of a trait vulnerability marker and is instead more of a state-dependent feature of depression, but do not expose how this relationship unfolds over time. Understanding these linkages may have important implications for mechanistic accounts of depression and for developing new avenues for treatment. For example, if low confidence temporally precedes periods of depression, then interventions that address metacognition might be useful for disrupting the cycle of depression. Beyond the temporal dynamics of this relationship, there are also gaps in our understanding of the mechanisms through which confidence in day-to-day decisions contribute to, or are influenced by, broader concepts like self-esteem.

A promising theoretical framework posits that metacognition is structured as a hierarchy. At the lowest level sits our ‘local’ confidence in specific decisions (e.g. individual trials of a cognitive task), which in turn inform ‘global’ self-performance estimates (SPEs; e.g. confidence in our overall ability on a task). At even more global levels sit more general beliefs about one’s ability or worth as a person (Seow et al., 2021). Recent cross-sectional studies have found that depression is linked more strongly to global measures of confidence and self-beliefs (Hoven et al., 2023; Seow et al., 2025), and that reduced global confidence might act as a prior that negatively biases local confidence (Van Marcke et al., 2024). These findings are consistent with theoretical accounts of reciprocal interactions between levels of a metacognitive hierarchy (Rouault et al., 2019; Seow et al., 2021). Another possibility, however, is that depression may distort the way that local confidence is integrated into global self-beliefs. For example, negative expectations may downweigh positive information when updating higher-order beliefs (i.e., cognitive immunisation; Kube et al., 2020), resulting in persistent negative beliefs, a core feature of depression (Beck, 1979; Kube & Rauch, 2025). Recent experimental work corroborates this view, showing weaker integration of high local to global confidence during a single task session in those with higher levels of anxious-depression (Katyal et al., 2025).

The present study aimed to extend this work and test these ideas using a microlongitudinal design with ecological momentary assessments (EMA) (Myin-Germeys et al., 2018) of metacognition and mental health spaced over 8 weeks. Specifically, we gathered momentary depression ratings twice daily, and metacognitive performance on a gamified perceptual (dot difference) decision making task (Fox et al., 2024) once every second day, via smartphone from N=162 participants. Our primary goal was to isolate within and between-person relationships between depression, local and global confidence, and test for temporal precedence or ‘Granger causality’ (Epskamp et al., 2018). Thus far, only one study has utilised EMA to investigate confidence and mood (da Fonseca et al., 2023). Using an implicit confidence measure (decision to opt out of a trial), they did not find evidence of cofluctuations between confidence and mood, either contemporaneously or temporally. However, the nonsignificant findings may be due to the small study size conducted over just 10 days, the lack of a subjective confidence measure, and or control for objective accuracy. Interestingly, they reported that metacognition and mood fluctuated over slightly different timescales, on the order of 2.5 and 1 days, respectively. Therefore, a study with a longer overall duration and more widely spaced assessments may be necessary to capture cofluctuations. A secondary aim of the study was to test for any alterations in metacognitive sensitivity in depression. In contrast to confidence biases, studies examining this have typically yielded mixed results, with some studies finding no difference between patients with depression and healthy participants (Agnoli et al., 2023; Culot et al., 2022), while others show better (Rouault et al., 2018) or reduced (Katyal & Fleming, 2026) metacognitive sensitivity along a depression continuum in a healthy sample. We now know that estimating metacognitive sensitivity is more difficult than metacognitive bias and requires hundreds of trials to achieve reliable estimates (Guggenmos, 2021). We therefore sought to leverage our dense repeated samples of metacognition to address this as a secondary study aim; that is, combining data from repeat plays of the task over 8 weeks to achieve the required power to perform robust computational modelling of metacognitive sensitivity and test for associations with depression.

In the present study, we thus investigated i) the temporal dynamics of metacognitive confidence (local and global) and depressive mood; ii) the influence depression has on the integration of local to global confidence; and iii) whether these associations differ at the state and trait level or depending on specific depression symptoms. Extending previous findings, we hypothesised that contemporaneous within-person (i.e. state) and between-person (i.e. trait) depression would be negatively associated with confidence. We further hypothesised that temporal relationships (suggestive of causation) would also exist but did not have a strong a priori hypothesis about directionality. Assessing both local and global confidence allowed us to investigate whether depression influences the integration of local to global confidence at the state and trait level across time. Based on Katyal et al. (2025) findings, we hypothesised that depression is linked to poorer integration of local to global confidence. Finally, we leveraged our large volume of per-participant datapoints to fit a computational model to the entire 8-week time series to estimate metacognitive efficiency and biases in post-decisional processing biases in relation to depression (Desender et al., 2022; Katyal & Fleming, 2026).

Results

Depression is associated with lower local and global confidence between-person

We analysed data from 162 citizen scientists (75% female, Mage = 54.4 ± 11.3 [19 – 82]) who tracked symptoms of depression twice per day (Fig. 1a) and played a perceptual decision-making task every 2 days over an 8-week period (Fig. 1b). Total depression score (i.e. average of all EMA symptom items) was associated with lower local (Partial correlations (Pcorr) = -0.239, Ps < .05) and global (Pcorr = -0.262, Ps < .05) confidence.

Microlongitudinal design and undirected associations between depression and metacognitive confidence.

(a) Design of the response scale for EMA item and summary of observations used for analysis. 14 depressive symptoms were assessed twice daily (morning and evening) on an 11-point visual analogue scale, while metacognitive confidence was assessed bidaily in the evenings for 8 weeks. (b) Meta-mind procedure. For each assessment, players are instructed to navigate the ship to the side that has more dots. At the start of each meta-mind round (2 rounds with 20 trials each) players are shown a screen with no stimuli and then two icons with varying dot densities then appear at the top of the screen. After the dots disappear, participants move the spaceship and select the icon containing more dots. At the end of each trial, players are asked to rate their local confidence in the accuracy of their choice. After 20 trials, players are asked to evaluate their global confidence in the overall accuracy for the round. (c) Between-person and contemporaneous associations. Plots display results from multilevel autoregressive models examining the relationship between overall depression and its symptoms, with metacognitive local (diamond) and global (circle) confidence. Between-person partial correlation coefficients (in navy, ‘Trait’) indicate the association between the mean of each EMA item and the mean of confidence, controlling for temporal (autoregressive and cross-lagged) effects as well as age, gender, and education. Within-person contemporaneous partial correlation coefficients (in blue, ‘State’) indicate how well each EMA item relates to metacognitive confidence in the same time window. Vertical lines reflect the averaged 95% confidence interval. Post hoc analyses of individual symptom level associations were corrected for multiple comparisons using the Benjamini-Hochberg adaptive linear step-up procedure. Filled shapes indicate significant results using an alpha level of P < .05.

To examine which individual symptoms might drive this effect, we repeated this analysis for all 14 items, correcting for multiple comparisons using the adaptive Benjamini-Hochberg correction (PBH). There was no clear driver of the mean effect, all individual symptoms (except for tiredness) were significantly associated with lower local (-0.263 < Pcorr < -0.170 PsBH < .05) and global (-0.294 < Pcorr < -0.2, PsBH < .05) confidence (Fig. 1c). See Supplementary Table 1 and 2 for details. Descriptively, trouble concentrating, sadness and worry had the strongest relationships with both local and global confidence, while feeling guilty, restless or tired consistently showed the weakest association.

Older age was linked to lower trait total depression (β = -0.177, SE = 0.07, t = -2.515, p = .013), and lower symptom severity for 11 out of 14 individual depressive symptoms (-0.205 < P < -0. 130, PsBH < .05). Age was also associated with greater average local (β = 0.166, SE = 0.072, t = 2.308, p = .022) and global (β = 0.204, SE = 0.064, t = 3.179, p = .002) confidence. We found no effects of gender or education (PsBH > .05).

Within-person fluctuations in depressive mood are associated with changes in global, not local, confidence

We found a significant contemporaneous association between depression and global confidence (Pcorr = -0.081, ps<.023), such that when a person’s mood drops, so does their global confidence in the task. This was not the case for local confidence (Pcorr = -0.035, ps>.115). A follow up analysis of individual depression symptoms did not reveal any significant effects, after correcting for multiple comparisons (PsBH > .05). See Fig. 1c and Supplementary Table 3 and 4 for more details. We found significant autoregressive effects for local (β = 0.278, SE = 0.029, t = 9.352, p < .001) and global (β = 0.098, SE = 0.024, t = 4.001, p < .001) confidence, and to a lesser extent, total depression (β = 0.049, SE = 0.018, t = 2.347, p = .018). Individual depression symptoms also showed significant autoregressive effects (0.11 < p < 0.236, PsBH < .001). Turning to the question of Granger causality, crosslagged analyses did not reveal any significant relationships between depression and local or global confidence. Follow-up analyses of individual symptoms showed that increases in stress significantly predicted future decreases in global confidence (β = -0.063, SE = 0.022, t = -2.960, PBH = .041). We found no other temporal associations between depression and confidence in either direction (PsBH > .05; Supplementary Fig 1, Fig 2, Supplementary Table 5 - 8).

Trait depression impairs the integration of local to global confidence

Same-level interactions

Having established that global confidence fluctuates with depression within-person, next we investigated more complex candidate dynamics between mood and different hierarchical levels of confidence that might shed further light on this mechanism. One possibility is that depressed mood diminishes the extent to which local confidence updates or informs global confidence. We tested this by examining the interaction between momentary depression and local confidence in models predicting global confidence at the same time point. We carried out Bayesian analysis as frequentist models failed to converge when using random slopes.

We did not find a moderating effect of momentary depression on the extent to which local confidence related to global confidence contemporaneously (β = 0.268, 90% CrI = [-0.071, 0.607], BF10 = 1/9). In an exploratory analysis of individual symptoms, we did however see a contemporaneous interaction effect with worthlessness (β = 0.416, 90% CrI = [0.138, 0.698], BF10 = 1/161), feeling guilty (β = 0.301, 90% CrI = [0.036, 0.565], BF10 = 1/32), and trouble concentrating (β = 0.463, 90% CrI = [0.231, 0.699], BF10 = 1/1141). Contrary to our hypothesis, these factors increased the strength of the association between local and global confidence. Similarly, there was anecdotal to no evidence supporting the moderating effect of trait depression and its symptoms on the association between local and global confidence cross-sectionally (β = 0.024, 90% CrI = [-0.106, 0.155], BF10 = 1; -0.058 < p < 0.132, BF10 < 1/3). See Supplementary Table 9 and 10 for details.

We next explored how these dynamics unfold over time, by testing how momentary depression impacts the forward propagation of local confidence to global confidence at the next time point. We did not find any evidence for an effect of depression (β = 0.185, 90% CrI = [-0.226, 0.599], BF10 = 1/3). However momentary increases in feelings of indecisiveness (β = 0.463, 90% CrI = [0.136, 0.795], BF10 = 1/89), guilt (β = 0.305, 90% CrI = [0.003, 0.609], BF10 = 1/20), and worthlessness (β = 0.443, 90% CrI = [0.089, 0.787], BF10 = 1/49) specifically promoted the integration of local to future global confidence. See Supplementary Table 11 for details.

Cross-level interactions

The previous results concerned momentary fluctuations in depressed mood from one’s own set point. Next, we examined how trait depression relates to how local confidence integrates into global confidence. We did not find evidence for contemporaneous associations (Supplementary Table 12 for details), but for lagged analyses, we found significant and extreme evidence that trait total depression (β = -0.279, 90% CrI = [-0.467, -0.093], BF10 = 156) was associated with reduced integration of local confidence into global confidence at the next time point. This finding was consistent across all depressive symptoms (-0.366 < β < -0.199, BF10 > 20). See Fig. 2a and Supplementary Table 13 for details.

Trait depression weakens the carry-over from local to global confidence, and this operates via a moderated-mediation pathway.

(a) Cross-level interaction between trait depression/symptoms (between-person) and lagged local confidence (within-person deviation) predicting global confidence. Point estimates of standardised beta interaction coefficients from Bayesian mixed-effects models for total depression and each symptom, with horizontal lines indicating the 90% credible interval. Filled shapes indicate 90% credible intervals are outside of 0. (b) Moderated mediation path diagram for total depression. Trait total depression interacts with within-person lagged local confidence to predict within-person concurrent local confidence (Path a) and in turn global confidence (Path b). The same interaction term has a total effect on global confidence (path c) and a direct effect controlling for concurrent local confidence (Path c’). Values under each path reflects the standardized path coefficients with 90% credible intervals. ΔACME above the mediator denotes the index of moderated mediation, i.e., the difference in the indirect effect (a x b) at ±1 SD of trait total depression. (c) Exploratory moderated-mediation test. Plots show the difference in the posterior distributions of index of moderated mediation (ΔIMM) at ±1 SD of trait depression. Negative ΔIMM values indicate weaker indirect effects at higher depression. Fill indicates the Bayesian evidence ratio (Savage–Dickey Bayes factor), with lighter greens denoting stronger evidence. For total depression and symptoms (except for tiredness), ΔIMM < 0 with strong–extreme Bayesian support, consistent with depression blunting local-confidence autocorrelation and, in turn, its integration into global confidence.

One possible mechanism for this may be that trait depression makes experiences of local confidence less sustained, i.e. local confidence is more quickly pulled down towards a set point of trait under-confidence, thus weakening subsequent integration of local into global confidence. Indeed, a two-way interaction found depression was associated with a significantly weaker autoregression of local confidence (β = -0.072, 90% CrI = [-0.123, - 0.022], BF10 = 97). With the exception of tiredness and slowness, this effect was strong and consistent across depressive symptoms (-0.092 < β < -0.057, BF10 > 23). We additionally modelled residual within-person variance in local confidence as a function of depression while including a participant-level random intercept. This tested whether depression was associated with greater within-person variability in local confidence, above and beyond the variability at the participant level. However, this model resulted in similar LOO-CV performance to a random-intercept-only model, suggesting depression did not meaningfully explain variability in within-person local confidence. See Supplementary Table 14 and Note 2 for details.

Moderated mediation analysis

To further dissect the mechanism underlying this effect, we conducted a moderated mediation analysis (Hayes, 2015). See Fig. 2b for details. We tested if the significant cross-level interaction between trait depression and lagged local confidence (Path c), could explained by an indirect pathway where depression moderated the autocorrelation of local confidence (Path a) together with a cascading effect on the integration of contemporaneous local to global confidence (Path b), while controlling for lagged depression (Path c’).

To test this, we calculated the difference in the index of moderated mediation (ΔIMM) for depression at ±1 SD (i.e., difference in indirect effects at different levels of depression), and hypothesised a negative ΔIMM (reflecting poorer integration that scales with depression) across total depression and its symptoms. The results indicated moderate to extreme evidence in support of the moderated mediation process for total depression (ΔIMM = -0.310, 90% CrI = [-0.533, -0.090], BF10 = 95) and individual symptoms (-0.393 < ΔIMM < -0.272, BF10 > 20), except for slowness and tiredness. See Fig. 2c and Supplementary Table 15 for details. This appeared as a complete mediation for total depression and most symptoms except for restless (β = -0.201, 90% CrI = [-0.346, -0.059], BF10 = 100) and stressed (β = -0.15, 90% CrI = [-0.294, -0.011], BF10 = 25) which showed a strong evidence of partial mediation process through Path c’. See Supplementary Table 16 for details.

Overall, the findings suggest that trait depression blunts the autocorrelation of local confidence, which subsequently impairs how local confidence is integrated into global confidence. See Fig 2b for details on the process. This is consistent with a mechanistic explanation for how negative beliefs in depression are formed and maintained over time.

To interrogate this further, we performed a three-way interaction between trait depression, lagged local confidence, and valence (i.e. whether confidence was above or below a person’s mean). We found moderate but insufficient evidence supporting a weaker carry-over effect of local confidence when it fluctuated above the mean in total depression (β = -0.016, 90% CrI = [-0.046, 0.015], BF10 = 4), and this effect was similar across most symptoms (-0.009 < β < 0.006, BF10 < 11). This suggests that for total depression and most symptoms, the impairment is not dependent on the valence of local confidence fluctuations and is best described as a regression to the mean effect. However, there was very strong evidence for an interaction in models of slowness (β = -0.045, 90% CrI = [-0.078, 0.011], BF10 = 70) and trouble concentrating (β = -0.044, 90% CrI = [-0.076, 0.012], BF10 = 77), indicating a weaker sustained fluctuation of above-mean local confidence specifically at higher levels of these symptoms. Finally, we tested whether the autoregressive effect of local confidence was valence dependent using a two-way interaction, and only found moderate evidence that autocorrelation was stronger following positive compared to negative within-person fluctuations in local confidence (β = 0.019, 90% CrI = [-0.009, 0.048], BF10 = 6), indicating at most, weak asymmetry.

Under-confidence in depression is associated with computational markers signalling negative expectations

Finally, we investigated how aberrations in post-decisional processing in depression might generate distorted local confidence. We first examined metacognitive efficiency using v-ratio, i.e., the ratio of the amount of evidence used to estimate confidence, to that used to make the perceptual decision. A v-ratio of 1 would imply that the evidence used to make perceptual decisions and estimate confidence is equal, i.e. perfect metacognitive efficiency. We also investigated distortions in the accumulation of post-decisional evidence (v-bias), additive distortions in the shift of confidence criterion (a-bias), and multiplicative shifts in the distortions of confidence criterion (m-bias).

We did not find any association between v-ratio and depression (β = -0.041, SE = 0.08, t = - 0.518, p = 0.605) or its symptoms (PsBH > .05). Instead, we found that depression was associated with greater a-bias (β = 0.381, SE = 0.120, t = 3.167, p = .002), which corresponds to a more conservative threshold for reporting high confidence (Fig 3). With the exception of feeling guilty or tired, the same relationship with a-bias was observed for all individual symptoms of depression (0.276 < β <0.378, PsBH < .05) in the following order of effect size: trouble concentrating, feeling worried, lack of interest, feeling stressed, lack of enjoyment, feeling slow, feeling sad, feeling irritated, feeling indecisive, feeling anxious, feeling restless, feeling worthless. We did not find associations between overall depression and other candidate distortions (v-bias, a negative bias in post-decision drift rate) or (m-bias, a slower rate of change from evidence to confidence; PsBH > .05). See Fig. 3d for details and Supplementary Table 17 for symptom, parameter, and demographic associations.

Computational accounts of distorted post-decision evidence accumulation for the case of under-confidence in relation to depression.

(a) v-bias: a distorted accumulation of evidence over time. A negative bias in post-decisional drift rate (red) results in lower evidence favouring the decision made, resulting in lower confidence. (b) a-bias: a distorted shift in confidence criterion. An additive shift in the confidence criterion (red) - represented by a lateral shift in the sigmoid function translating evidence into bounded confidence, results in the same amount of evidence translating into lower confidence. (c) m-bias: a multiplicative shift in confidence criterion. A multiplicative shift in the confidence criterion (red) – represented by a flatter sigmoid function, results in a slower rate of change from accumulated evidence into bounded confidence. (d) Bootstrapped mixed effects models predicting depression from a-bias. Ridgelines show bootstrapped standardised beta coefficient distribution for parameter a-bias predicting depression, with 5th/50th/95th quantile lines. Shading of the ridges indicates frequentist significance. Analyses included all computational parameters and controlled for age, gender, and education as covariates. Symptom-level associations (but not total depression) were corrected for multiple comparisons.

Overall, these findings suggest that participants with higher depression require more evidence (higher expectations or more conservative evidence threshold) to achieve the same level of confidence as healthier participants with the same objective performance.

Discussion

Distortions in metacognition are a characteristic feature of several mental health disorders (Hoven et al., 2019). In depression, these distortions manifest as poorer metacognitive beliefs, global under-confidence, and local under-confidence (Benwell et al., 2022; Hoven et al., 2019, 2023; Rouault et al., 2018; Seow et al., 2025; Seow & Gillan, 2020), signaling a pervasive negative bias which manifests across levels of a metacognitive hierarchy. Prior work suggests that depression may not only vary over time with under-confidence (Agnoli et al., 2023; Fox et al., 2023) but also impair the integration of local confidence into global confidence (Katyal et al., 2025). However, the mechanistic processes underlying these relationships remain unclear as they have thus far only been assessed within the context of a single task. Here, we used a microlongitudinal design to ask i) whether within-person fluctuations in depressive mood precede under-confidence or vice versa, and ii) whether depression symptoms impair the integration of local confidence to global confidence across time. We found that within-person increases in depressive mood were accompanied by underconfidence at the same time point, but neither change preceded the other. Additionally, we found that depression was associated with an increase in confidence threshold, characteristic of defensive pessimism (Norem & Cantor, 1986). Crucially, individuals with higher levels of depression exhibited weaker autocorrelation of local confidence, which cascaded into poorer concurrent integration of local to global confidence. In other words, pervasive global underconfidence in depression may be driven by difficulties sustaining high levels of local confidence over time. Overall, our findings highlight how local and global under-confidence are characterised by distinct mechanisms operating at different time scales.

There is causal evidence from experimental inductions of sad (happy) mood that suggest alterations to mood can cause changes in metacognitive under-confidence (overconfidence) (Culot et al., 2021; Koellinger & Treffers, 2015). Here, we sought to extend this work using real world depression fluctuations (rather than mood induction) and tested for an important pre-requisite for causation, temporal precedence. Contrary to our expectations, we found little to no evidence that depression and confidence predicted each other at a two-day lag. Instead, we observed a small within-person covariation between total depression and global underconfidence at the same time point (Pcorr = -0.081). The absence of Granger causality may be explained by a mismatch between the temporal resolution. Put simply, temporal relationships at the scale of days may be too large to reliably establish Granger causality between mood and under-confidence. Supporting this, we found that depressive mood and metacognitive confidence exhibited different autocorrelations, in line with prior work (da Fonseca et al., 2023). Depression showed the weakest carry-over (β = 0.049), followed by global (β = 0.098) and then local confidence (β = 0.278). Therefore, it is likely that any top-down mood-driven changes in confidence decay rapidly as new mood states emerge. Alternatively, day-to-day mood fluctuations in an unselected general population sample may be too small and brief (Thompson et al., 2021) to exert a lasting effect on metacognitive confidence. For this reason, future studies should consider a clinical sample with larger affective variability. Future work should also explore the resolution with which depression and confidence co-fluctuate. More broadly, the significant contemporaneous association between global confidence and total depression, together with similar carry-over magnitude between these constructs, supports the view that depression is more tightly linked to global metacognitive constructs than to local confidence (Hoven et al., 2023; Seow et al., 2025).

Between-person analyses indicate that higher levels of overall depression and most depressive symptoms (except tiredness) were associated with lower average local and global confidence. This pattern is consistent with prior cross-sectional findings in clinical and non-clinical samples (Hoven et al., 2019).

Leveraging the large number of trials gathered per person over the 8 weeks of the study, we applied computational modelling (Katyal & Fleming, 2026) to gain insight into biases in post-decisional processing that may contribute to under-confidence. Higher levels of overall depression and its symptoms (except for tired and guilty) were associated with a more conservative confidence criterion, that is, requiring more evidence to report high confidence. This pattern is consistent with accounts of defensive pessimism, whereby more depressed individuals set higher standards or lower expectations to avoid disappointment (Norem & Cantor, 1986; Villano et al., 2023; Villano & Heller, 2024), and with predictive coding accounts of depression emphasising persistent negative expectation (more rigid negative priors) in depression (Kube et al., 2019, 2020). Consequently, it is plausible that underconfidence in depressed individuals may be addressed by targeting these rigid negative expectations. Nonetheless, because the computational model was fit to each participant’s full time series collapsed across days, we could only extract between-person snapshots of post-decisional dynamics. Future studies should investigate how these processes unfold at the within-person level over time.

Past work has established that local confidence integrates into global beliefs (Rouault et al., 2019), and one study suggested that depression may impair this integration during a single testing session (Katyal et al., 2025). Expanding on this finding, we tested whether the integration of local to global confidence would be similarly blunted over time. Using a crosslagged interaction and moderated mediation analysis, we found moderate to extreme evidence that higher trait depression and most symptoms (except worthlessness) were associated with a reduced carry-over of local confidence from the last assessment, subsequently impairing the contemporaneous integration of local into global confidence. This provides a plausible mechanism for how under-confidence and negative self-beliefs are maintained over time in depression. Specifically, a low autocorrelation means that momentary surges in local confidence are unlikely to persist and accumulate into positive global beliefs in those with higher depression. This is consistent with hypotheses of overprecise negative priors in predictive coding accounts of depression (Kube et al., 2020) - whereby such priors may constrain positive within-person fluctuations of local confidence back toward trait underconfidence, manifesting as a “regression to the mean” effect. Alternatively, individuals with higher depression may have more volatile internal representations of confidence over time, consistent with accounts of heightened sensitivity to integrate postdecision evidence (Moses-Payne et al., 2019). To test this, we additionally specified within-person residual variance of local confidence as a function of trait total depression in a separate model but did not find a meaningful improvement in the model’s predictive performance compared to a random-intercept only model. This suggests a pattern that is consistent with a stronger “regression to the mean” effect in high depression.

We then aimed to investigate whether impaired integration disproportionately targets positively valenced information (Kube et al., 2020), as shown in prior work (Hobbs et al., 2022; Katyal et al., 2025; Korn et al., 2014; Villano & Heller, 2024). However, our test for a three-way interaction did not replicate a valence-specific impairment that scales with depression, except in relation to symptoms of slowness and trouble concentrating. The findings suggest that depression may be responsible for a broader impairment to the temporal persistence of local confidence, rather than a valence specific deficit. There was also limited evidence to support stronger carry-over effects for positive fluctuations of within-person local confidence. Although modest, this pattern is broadly consistent with characterizations of optimism-bias in the general population (Lefebvre et al., 2017; Sharot et al., 2011; Sharot & Garrett, 2016; Villano et al., 2023).

We conducted exploratory symptom-level analyses and observed a surprisingly high degree of homogeneity across symptoms. This adds to the current conversation on the validity of the construct of depression and the extent to which individual symptom dynamics ought to be modelled (Fried et al., 2016; Fried & Nesse, 2015; Zimmerman et al., 2015), at least in the context of confidence. In terms of demographics, we also found older age to be associated with lower total depression, and greater local and global confidence. This is in line with previous literature noting improved mental health (Thomas et al., 2016) in older adulthood, but contrasts with other findings of lower confidence with age (Katyal & Fleming, 2026; McWilliams et al., 2023). We did not observe gender differences (women compared to men) in under-confidence either, in contrast to prior work (Hoogervorst et al., 2024; Katyal & Fleming, 2026). These discrepancies may reflect our relatively small, older (Mage = 54.4 ± 11.3) and predominantly female (75%) demographic which contrasts the larger, younger, and more balanced gender demographics typical of metacognition research (Xue et al., 2024).

Several limitations should be considered when interpreting these findings. Firstly, limited variability in mood fluctuations in the healthy general population sample may have attenuated within-person associations between under-confidence and depression. This is especially relevant given that the inclusion criteria required participants to be compliant and engaged in an 8-week study. Patient samples will exhibit a greater range and more variation in symptoms and this could reveal associations with metacognition which are distinct from healthier participants. Second, our data can only speak to fluctuations occurring across a two-day window; future research should explore whether the temporal dynamics of depression and confidence can be better characterized across alternative time scales. Third, because metacognitive constructs dynamically interact with each other across hierarchical levels (Rouault et al., 2019; Seow et al., 2021), the absence of high-level metacognitive beliefs (i.e., self-beliefs) in our analyses may, at worst, confound our associations, and at best, provide an incomplete account of the dynamics between metacognition and depression. Finally, our within-person “depression” variable reflects short-term fluctuations in symptom ratings around each individual’s mean, whereas “depressive disorders” are characterised by their persistence over time (Fifth Edition, Text Revision, DSM-V-TR). In this sense, our within-person analyses may not accurately reflect or generalize to longer-term trajectories associated with the onset, maintenance, and remission of depression and its symptoms.

In summary, our findings suggest that higher trait-level depression is associated with underconfidence partly because local confidence is less stable over time, and therefore less able to accumulate and integrate into later global confidence. This presents a temporal perspective into the emergence and persistence of global under-confidence in depression. More broadly, within-person fluctuations in mood may be better understood as a contemporaneous correlate of under-confidence rather than a causal component. At the trait level, higher depression was robustly associated with local and global under-confidence, and the former may be explained by a more conservative confidence criterion characteristic of defensive pessimism. Together, these findings suggest that preventative measures aimed at depression may benefit from stabilizing positive local confidence fluctuations, thereby improving integration of local confidence across metacognitive hierarchies.

Methods

Participants

N = 162 participants post-inclusion criteria (75% female; 23% male) aged between 19 and 82 years (M = 54.38 ± 11.34) were recruited via the Neureka app from July 2022 to June 2025 (see Supplementary Table 18 for additional demographic information). We term these participants ‘citizen scientists’, as they received no compensation for their participation.

EMA assessment and validation

Participants took part in the ‘Brain Changer’ science challenge within the Neureka app, includes an EMA module tracking symptoms of depression and a perceptual decision-making task to measure metacognitive confidence. After an initial baseline assessment, participants were required to complete an 8-week EMA procedure tracking depressive mood symptoms and metacognitive confidence. At sign up, participants were required to select two 3-hour time windows (between 6:00 - 11:30 AM and 6:00 - 11:30 PM) to receive notifications for their assessment. Notifications were delivered at the start of each time window and participants were allowed to respond until the end of the selected time window. Each notification prompted participants to rate their experience on a sliding scale ranging from 0 (Not true at all) to 10 (Extremely true). The EMA items were adapted from the Quick Inventory of Depressive Symptomatology, self-report (QIDS-SR-16; Rush et al., 2016), a validated questionnaire to evaluate the Diagnostic and Statistical Manual of Mental Disorders (Fourth Edition, Text Revision, DSM-IV-TR) depression criterion symptom domains. The item on suicidality was not included. Three items pertaining to trouble sleeping, excessive sleeping, and appetite were only assessed in the morning and excluded from all analyses, which focused on afternoon assessments that coincided with the cognitive task.

To validate our EMA questionnaire, we correlated the EMA assessed in the morning across the second week with QIDS-SR-16 assessed on assessment day 14 and found it to be excellent (r = 0.71, 95% CI = [0.62, 0.79], p < .001), indicating good convergent validity (see Fig. 4a). Overall, depression scores were stable across each participant’s first 14 evenings (ICC = 0.888, 95% CI = [0.864, 0.910], p < .001; see Fig 4b and c), though this was perhaps exaggerated by low variance as depression was zero-inflated across participants throughout the study (see Fig. 4b). There was no systematic trend towards improvement or worsening over the 8 weeks of the study (β = -0.019, SE = 0.013, t = -1.416, p = .159).

Reliability and validity checks for depression EMA and confidence ratings.

(a) Bivariate correlations. Between-person Spearman’s rank correlation coefficients between non-detrended total depression scores in the morning (second week) and total QIDS-SR-16 scores assessed on assessment day 14. (b) Distribution of the sum of all depression EMA scores (i.e. total depression) across each participant’s first 14 evening assessment. Total depression scores show zero-inflation. (c) Average total depression over each participant’s first 14 evening assessments. There is minimal drift in average total depression scores in 162 participants across assessments. The error bars represent standard error. Internal consistency of total depression (d), local confidence (e), and global confidence (f) ratings for each participant’s first 14 assessments ordered by their means. ICC analyses were performed using evening EMA assessments of 162 participants with a total of 2268 observations.

Metacognitive confidence was assessed bi-daily in the evenings using a game in the app called ‘Meta-mind’ - a gamified version of a visuo-perceptual metacognitive task. ‘Metamind’ has been described and validated in a previous work (Fox et al., 2024). In this game, players are presented with two moving icons (left and right side) that differ in the number of dots contained within them and descend the screen. Players are instructed to navigate a spaceship to the icon containing more dots by touching the screen and allowing the ship to collide with the icon. After a choice is made, players are asked to rate their local confidence in the accuracy of their choice on a 1-6 scale ranging from ‘Low’ to ‘High’. Reaction times were measured starting from the appearance of the prompt, up until when the final confidence decision was made. After 20 trials, players evaluate their overall accuracy for that round on a 1-6 scale ranging from ‘Chance Level’ to ‘Perfect’. This block-by-block measure indexed global confidence. No feedback is provided to participants. The task uses a ‘two-down-one-up-staircase which maintains the accuracy within a range of 0.6 to 0.85. The initial staircasing was performed during a baseline testing session across 100 trials which has 20 burn-in trials before the staircase ‘settled’ within the desired accuracy range. The task difficulty level for each participant at the end of a session was always carried over to the next testing session to avoid restarting the staircase (and needing to discard more burn in trials). After the baseline assessment, participants completed 40 trials per day, and the staircase continue to operate throughout to account for any gradual improvements in objective performance that would confound confidence estimates.

To validate the staircasing procedure over the time series we fit a model comparing the accuracy from the 1st (assessment after baseline) and 15th assessment day and found no significant improvement in accuracy (β = 0.137, SE = 0.12, t = 1.144, p = .255). Furthermore, we did not find any significant associations between time and local (β = 0.013, SE = 0.016, t = 0.776, p = .439) or global confidence (β = -0.008, SE = 0.015, t = -0.497, p = .62). However, as expected, difficulty of the task increased over time (decreased log dot difference) as participants became more proficient (β = -0.042, SE = 0.02, t = -2.047, p = .042). The stability of the confidence data was very high for local confidence (Fig 4d; ICC = 0.871, 95% CI = [0.841, 0.896], p < .001), but only fair for global confidence (Fig 4e; ICC = 0.694, 95% CI = [0.644, 0.743], p < .001), although this can be attributed to the latter’s fewer datapoints (2 vs 40 per assessment). Overall, this suggests that the staircase was successful and this change in objective perceptual decision-making performance did not translate into changes in accuracy or confidence.

Data preparation

Behavioural outcomes and exclusions

N = 976 participants with a total of 6784 assessments (range = 1 – 28 assessments per person) were included for initial pre-processing. We began by excluding individual assessments (i.e. a single day of task data) where participants had selected the left or right option more than 95% of the time; 5 assessments across 5 participants (1 assessment each) failed this criterion across the study duration, resulting in 1 participant being excluded. Then, assessments that achieved accuracy rates < 60% or > 85% were excluded, suggesting the staircase had failed to converge for them; 182 assessments across 148 participants (range = 1 – 19 assessments per person) met this exclusion criterion, resulting in 11 participants being removed. Individual trials in an assessment were also excluded if participant’s choice or confidence reaction times were implausibly fast, i.e. < 150ms (81 trials across 33 participants, range = 1 – 18 trials per person) or slow, i.e. > 5s (560 trials across 204 participants, range = 1 – 26 trials per person). Following that, assessments with < 40 trials or 2 blocks of assessments for that day were also removed. This was true for 6597 assessments across 964 participants (range = 1 - 28 assessments per person), resulting in the removal of 55 participants. Following the trial and day level exclusions, participant-level data was excluded if their confidence reports had less than 0.17 standard deviation (1/confidence levels) across the whole time series (18 participants removed), and if they did not have at least 50% (14/28) of the ‘Meta-mind’ assessments and the corresponding EMA items at the same time point (727 participants removed). An additional 2 participants were removed due to missing demographic information.

After exclusions, we had a total of 7408 (6906) evening (morning) EMA assessments across 162 participants, of which 3306 coincided with metacognition assessments (median assessment per person = 20, range = 14 – 28). Mean bi-daily evening EMA responses and confidence ratings are provided in Supplementary Table 19, and data distribution are displayed in Supplementary Figure 1.

Detrending

The Kwiatkowski-Phillips-Schmidt-Shin (KPSS; Kwiatkowski et al., 1992) showed that on average, 75% (71% - 78%) of depression EMA items were stationary across participants. In contrast, 100% of local and global confidence data were stationary. However, total depression (sum of depression EMA items) was only stationary across 66% of participants. Given the accepted threshold of 70% (Bringmann et al., 2013), total depression was detrended to ensure stationarity. We regressed each participant’s total depression rating against a linear and exponential term (-0.2) of ‘assessment day’ (i.e., the number of days since baseline of the study). The residuals from this model were added to each participant’s mean score across the entire assessment period, and these detrended time series were then used in further analyses. Post detrending, KPSS showed that 95% of the depression items were stationary.

Statistical analysis

All participants provided informed consent in accordance with the European General Data Protection Regulation (GDPR). Data were pre-processed and analysed using R Statistical Software (v4.4.0) contained within a Docker container (Boettiger & Eddelbuettel, 2017; Nüst, Eddelbuettel, et al., 2020; Nüst, Sochat, et al., 2020).

Multilevel vector autoregressive models

We used a two-step multilevel vector autoregressive model using the ‘lme4’ package (Bates et al., 2015) to investigate within (temporal and contemporaneous) and between-person associations between depression and metacognitive confidence. Analysis was performed for local and global measures of metacognitive confidence, total depression (sum of all EMA items per day) and for each depression item. We adapted the methodology employed by Epskamp et al. (2018) to incorporate covariates as the mlVAR package does not allow for their inclusion. All models included a random intercept to account for individual differences in the dependent variable. Random slopes were not included in any of the analysis as the models failed to converge. Confidence and depression variables were scaled across the sample, and predictors were also within-person centred. Age (continuous in years), gender (factor; baseline = male), and education (binary: below-university level vs university) were entered as covariates and scaled. Below-university includes lower and upper secondary education, and no formal education, while university includes bachelor’s, master’s, and PhD degrees or equivalents. Gender and education were coded using a custom contrast such that the intercept would reflect the grand mean, while non-baseline levels were compared to variable’s respective baseline (i.e., female and no university level education).

Separate analyses were conducted for local and global confidence. We focused primarily on depression scores, i.e. the sum of all individual EMA items, but complemented those analyses with exploratory analyses of individual symptoms. The generic models used for these analyses were as follows (1,2):

Because metacognition was assessed bi-daily, the temporal associations were estimated at a two-day lag (t-2). Between-person partial correlations were estimated by averaging the regression coefficient (Depression_mean/Confidence_mean) in the forward and reverse analyses (Epskamp et al., 2018). Contemporaneous within-person associations were estimated by correlating the residuals of the forward regression and reverse regression. Because the temporal and trait associations have been partialed out, the residuals are assumed to represent a contemporaneous link (Epskamp et al., 2018). The models used for analyses is as follows:

P values for the exploratory symptom-level analyses were corrected for multiple comparisons using the Benjamini-Hochberg adaptive linear step-up procedure (Benjamini et al., 2006) with an alpha level of P < .05. This approach aims to reduce type-1 errors (false positives) whilst accounting for the positive correlations in the test-statistics of the outcomes, e.g. symptom level analyses where associations are expected to be similar to each other. Using the “AND” rule, between and within-person contemporaneous associations were only considered significant if both regression parameters used for averaging are PBH < .05. Bayesian models were not used to address the lack of random slopes in the chosen mlVAR framework as the approach required the ordinal outcomes to be scaled (Epskamp et al., 2018), which in turn, required the selection of likelihood families inconsistent with the data generating process.

Bayesian interaction models

We additionally ran Bayesian ordinal regressions for the interaction analyses using the brms package (Bürkner, 2017). This allowed us to fit random slopes which failed to converge using frequentist methods, while also allowing us to model the data generation process more accurately (Bürkner & Vuorre, 2019; Liddell & Kruschke, 2018). Unlike the frequentist models, global confidence as the dependent variable was not scaled. The Bayesian workflow recommended by Gelman et al. (2020) was used to validate the computations. With the exception of model 12, we could not conduct cross validation analyses due to compute limitations and kept the models maximal. Details on the weakly informative priors used can be found in Supplementary Note 1.

A cumulative probit family was used for the ordinal regression analysis with 8 chains, 2000 warmup samples and 6000 post warmup iterations. Divergent transitions were addressed by increasing adapt delta to 0.99 and tree depth to 12, remaining divergences were addressed by simplifying the model structure.

The models used for the interaction analyses is as follows:

The disc parameters were used to model variability in the latent thresholds used for the ratings. We omit the intercept from the disc model to ensure that the standard deviation of the latent scale at the intercept is set to one, which aids identifiability (Bürkner & Vuorre, 2019). For the Bayesian analyses, a one-sided test in the hypothesised direction was done. If 0 is located within the 95% credible interval for a one-sided hypothesis test, the parameter values are rejected as there is insufficient evidence in support of the interaction. We reported the 90% equal-tailed credible interval for symmetry. We adopt Stefan et al. (2019) Bayes factor thresholds to classify evidence strength.

Bayesian moderated mediation models

To investigate the mechanism underlying a significant interaction we observed between lagged local confidence and trait depression, we conducted a moderated mediation analysis (Hayes, 2015). Specifically, we tested whether trait depression would significantly impair the within-person autoregression of local confidence (Path a; Eq. 8) and subsequently impact the contemporaneous integration of local to global confidence, while accounting for the remaining direct effect of lagged local confidence (Path b and c’; Eq. 9). All predictor variables were scaled across the sample, but the predictors and outcome for path a (local confidence) were also within-person centred.

A Bayesian multilevel model including Path a and Path b and c’ was used. For the Path a model, unexplained variability in within-person confidence between participants was additionally modelled with parameter sigma. Pareto-Smoothed Importance Sampling (PSIS-LOO) cross-validation was used to evaluate model fit for Path a, and the model providing the best fit and parsimony was selected. See Supplementary Note 2 for details.

The index of moderated mediation was calculated using the product of the posteriors of Path a and Path b at ±1 SD of trait depression/symptoms. Similarly, a one-sided hypothesis test in the direction of the hypothesised interaction and moderated mediation findings was done to determine whether there was sufficient evidence supporting our hypothesis. If 0 is located within the 95% credible interval for a one-sided hypothesis test, we determined that there was insufficient evidence in support of the moderated mediation process. We reported the 90% equal-tailed credible interval for symmetry. We adopt Stefan et al., (2019) Bayes factor thresholds to classify evidence strength.

Computational model

To investigate distortions in the mechanisms underlying under-confidence in trait depression and its symptoms, we applied a recently developed computational drift-diffusion model (DDM) of confidence (Katyal & Fleming, 2026). The model builds on prior work modelling choice, decisional and confidence reaction times, and confidence ratings on a two-alternative forced choice task (Desender et al., 2022; Hellmann et al., 2023; Pleskac & Busemeyer, 2010), and was extended to account for distortions in confidence formation during post-decisional evidence accumulation (Katyal & Fleming, 2026). According to traditional DDM models, in each trial of a two-choice task, participants accumulate evidence (e) noisily over time at a drift-rate (v) until one of two decision boundaries is reached, triggering the corresponding response. After an initial non-decision time (Ter), evidence begins accumulating from a starting point (z), which is a proportion of the distance between the two decision boundaries (a). The reaction time taken for a trial is indexed by the sum of the nondecision time and decision time (Td) taken to reach the boundary. We can model the evidence accumulation process at each time step (Δt) as:

The noise (σ) is set to 1 for identifiability purposes (Desender et al., 2022). The At used was .001.

In post-decisional DDM, evidence continues to accumulate after the initial boundary has been reached, accounting for the correspondence between decision accuracy and confidence (Pleskac & Busemeyer, 2010). A ratio of the evidence accumulated after the decision (υpd) over those accumulated during the initial decision process (υ) can index a bias-free measure of metacognitive efficiency (υratio) that accounts for speed-accuracy tradeoff (Desender et al., 2022). This is analogous to Mratio which reflects the ratio of sensitivity of confidence to decision accuracy (meta-d’) vs first-order perceptual sensitivity (d’) within signal detection frameworks. Similar to the Mratio, a Vratio of 1 implies that the amount of evidence used to estimate confidence (post-decision drift rate) is equal to the evidence used to make the perceptual decision (drift rate), i.e. perfect metacognitive efficiency. However, unlike Mratio, a negative Vratio implies that evidence used to estimate confidence is in the opposite direction of correct evidence i.e., overconfidence in wrong choices and under-confidence in correct choices.

Katyal and Fleming (2026) extended this model by accounting for distortions in post-decisional drift-rate (υbias), which is meant to index a general bias for or against the decision made. We can model the post-decisional drift-rate for each decision in a trial (dn) as:

Furthermore, unlike Mratio the updated model fit allows Vratio to have positive or negative values. Positive values imply that post-decision evidence accumulates towards higher confidence when the decision is correct and towards lower confidence when the decision is incorrect; negative values imply that post-decision evidence accumulates towards lower confidence when the decision is correct and towards higher confidence when the decision is correct.

Following Katyal and Fleming (2026), we modelled the post-decisional evidence accumulation using a Wiener process that samples empirical confidence response times (rtconf) distributions for each participant:

The evidence accumulated (econf), which is an unbounded variable, is then subsequently transformed into a bounded probabilistic estimate of confidence. During this process, parameters indexing the threshold set for confidence estimates (abias) and the rate by which evidence integrates into confidence (mbias) are introduced:

Using differential evolution optimisation in R with the DEOptim package (Mullen et al., 2009), we estimated seven free parameters (v, a, Ter, z, vbias, abias, mbias, Vratio) by minimising the chi-squared error function as noted in Katyal and Fleming (2026). However, we further binned the responses at each quantile by their stimuli to improve recovery of parameter z:

oRTresp,stim,cl and sRTresp,stim,cl denotes the observed and simulated reaction times for each response chosen (resp; i.e., left or right), stimuli (stim; i.e., correct stimuli are on the left or right), and confidence level (cl; 1-6) respectively. oConffcl,ac and sConffcl,ac denotes confidence responses that were faster than their respective rtconf; oConfscl,ac and sConfscl,ac denotes confidence responses that were faster than their respective rtconf. The model was fit separately per participant by aggregating their metacognition assessments across all time points as a single dataset.

The retrieved parameters were z-scored across the whole sample before statistical analyses. We excluded parameter estimates that were larger than 3 standard deviations from the mean, reducing the number of samples to N = 154. To test whether between-person associations between depression and the relevant computational parameters were significant, the following model was used:

Parameter recovery was conducted using 100 sets of randomly sampled parameter values with 560 trials each and was found to be good to excellent (0.8 < r < 0.99, Ps < .05). See Supplementary Figure 2 for details.

Data availability

All supplementary information, processed data, code, and statistical output is openly accessible at https://osf.io/5m76h/.

Supplementary Materials

Mean EMA Responses and Confidence Ratings

Parameter recovery of simulated data generated by randomly drawn parameters

Between-person EMA Associations with Local Confidence

Between-person EMA Associations with Global Confidence

Contemporaneous EMA Associations with Local Confidence

Contemporaneous EMA Associations with Global Confidence

Lagged EMA Predicting Local Confidence (Forward)

Lagged EMA Predicting Global Confidence (Forward)

Lagged Local Confidence Predicting EMA (Reverse)

Lagged Global Confidence Predicting EMA (Reverse)

Same-level Contemporaneous Bayesian Interaction Predicting Global Confidence

Same-level Trait Bayesian Interaction Predicting Global Confidence

Same-level Lagged Bayesian Interaction Predicting Global Confidence

Cross-level Contemporaneous Bayesian Interaction Predicting Global Confidence

Cross-level Lagged Bayesian Interaction Predicting Global Confidence

Trait Depression Moderation on Local Confidence Autoregression

Index of Moderated Mediation on Global Confidence for Depression at ±1 SD

Direct Effect of Interaction on Global Confidence (Path c’)

Depression Associations with Computational Parameters

Demographic and descriptive statistics for the main samples

Mean Responses in EMA Questionnaires and Confidence

Weakly Informative Priors

PSIS-LOO Cross Validation of Mediation Model of Path a

Additional information

Funding

Science Foundation Ireland (SFI) (SFI-19/DP/7213)

  • Claire M Gillan

Science Foundation Ireland (SFI) (21/RC/10294)

  • Paveen Phon Amnuaisuk

  • Vanessa Teckentrup

  • Claire M Gillan